Information extraction method and system based on keywords

By constructing a logical relationship between keyword context association network and simulating real scenes, the problem of insufficient semantic association capture in the existing technology is solved, and accurate capture of deep semantic associations and improvement of content coherence is achieved.

CN120542418AActive Publication Date: 2025-08-26HANGZHOU ZHONGZHUO INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510622771.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Existing keyword-based information extraction methods are difficult to accurately capture semantic associations and context, resulting in deviations in core content refinement, and over-reliance on statistical features and ignoring in-depth analysis at the semantic level.

Method used

By building a keyword context association network, combining semantic connection strength and statistical characteristics, a preset blank context module framework is used to constrain keyword embedding according to context association relationships, simulate the logical relationship of the real scene, and ensure that the output content conforms to the actual application scenario.

Benefits of technology

Effectively capture deep semantic correlations, avoid deviations caused by semantic loss in traditional weight analysis, improve content coherence and logical consistency, and ensure that the output content conforms to actual application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542418A_ABST
    Figure CN120542418A_ABST
Patent Text Reader

Abstract

The invention relates to a keyword-based information extraction method and system, and relates to the field of information extraction, and the method comprises the following steps: receiving an information text; extracting keywords from the information text based on a preset extraction algorithm; determining an information field to which the information text belongs based on the keyword; obtaining a context category and a context association relationship of the keyword based on a preset algorithm; based on the context category and the context association relationship, filling the keyword into a preset blank context module to obtain a keyword context module; and performing scene simulation on the keyword context module to obtain key contents conforming to an actual scene, and outputting the key contents. The method has the advantages that deep semantic association is effectively captured by constructing the keyword context association network and combining semantic connection strength and statistical characteristics, and deviation caused by semantic deficiency in traditional weight analysis is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information extraction, and in particular to a keyword information extraction method and system. Background Art

[0002] With the rapid development of the internet and digital technologies, global data volumes are growing exponentially. The emergence of massive amounts of unstructured text data has led to an increasing problem of information overload. To meet the need for rapid access to core content, keyword-based information extraction methods have emerged. Their core goal is to use automated technology to identify user-focused keywords from vast amounts of text, quickly locate related content, and achieve efficient information filtering and aggregation.

[0003] In the existing automatic information extraction process, a weight analysis is performed based on the number of times a keyword appears, its position, and other attributes, and the relevance of the keyword to the information content is expressed through a weight coefficient. Existing technologies include setting a weight coefficient to determine the relevance, setting a dynamic weight coefficient combination for keywords, and calculating the product of the weights of keywords in sentence units between punctuation marks (M = Q1 ×… × Q γ ), and intercept the top 10 sentences with the highest product as the summary, so as to accurately remove redundant information and improve reading efficiency.

[0004] Regarding the above-mentioned related technologies, it is difficult to accurately capture semantic associations and contexts during the process of keyword weight analysis, which may lead to over-reliance on the statistical characteristics of keywords, such as weight products, while ignoring in-depth analysis at the semantic level, resulting in deviations in the extraction of core content. Summary of the Invention

[0005] In order to solve the problem that it is difficult to accurately capture semantic associations and contexts during keyword weight analysis, which may result in over-reliance on the statistical characteristics of keywords, such as weight products, while ignoring deep analysis at the semantic level, resulting in deviations in the extraction of core content, the present invention provides a keyword-based information extraction method and system.

[0006] In a first aspect, the present invention provides a keyword-based information extraction method, which adopts the following technical solution: A keyword-based information extraction method, comprising: Step 1: Receive information text; Step 2: extracting keywords from the information text based on a preset extraction algorithm; Step 3: Determine the information field to which the information text belongs based on the keywords; Step 4: Obtain the context category of the keyword based on a preset association algorithm; Step 5: Filling keywords into a preset blank context module based on the context category to obtain a keyword context module; Step 6: Build a scenario simulation with the keyword context module to obtain key content that conforms to the actual scenario and output it.

[0007] Optionally, the steps after step 5 and before step 6 further include: Step 50: searching for corresponding relevance between the keyword in any one of the keyword context modules and the other keywords from a preset keyword database; Step 51: constructing the keywords based on the correlation to form a contextual correlation network; Step 52: Analyze the number of connection edges and corresponding connection strengths of the keywords in the context association network; Step 53: Obtain the number of occurrences and position parameters of the keyword in the information text based on a preset traversal algorithm; Step 54: determining the keyword importance value of the keyword based on the number of connection edges, the connection strength, the number of occurrences, and the position parameter; Step 55: Filter the keywords whose keyword importance values ​​are lower than a preset importance threshold, discard them from the keyword context module to obtain a modified keyword context module, and define the remaining keywords in the modified keyword context module as modified keywords; Step 56: Outputting the modified keyword context module as the keyword context module; Step 57: When the number of the modified keywords is less than a preset quantity threshold, output a preset module alarm signal.

[0008] Optionally, when the number of the modified keywords is less than the critical number, the method of outputting the module alarm signal includes: Step 570: Lowering a keyword determination threshold preset in the extraction algorithm to form a modified extraction algorithm, wherein the keyword determination threshold is set in the extraction algorithm; Step 571: re-execute steps 2 to 55 based on the modified extraction algorithm; Step 5710: When the number of the modified keywords after re-execution of the step is greater than the threshold value, execute step 6; Step 5711: When the number of the modified keywords after re-execution of the step is still less than the quantity threshold, the module alarm signal is output.

[0009] Optionally, another method for extracting information when the number of the modified keywords after re-performing the step is still less than the critical number value includes: Step 57110: Searching for suspected information text on the Internet based on the keyword; Step 57111: Calculate a similarity value based on the predetermined judgment algorithm between the suspected information text and the information text; Step 57112: When the similarity value corresponding to the suspected information text is higher than a preset similarity threshold, the suspected information text is defined as a similar information text; Step 57113: Re-execute steps 2 to 4 based on the similar information text to obtain an alternative keyword context module and alternative keywords; Step 57114: Merge the replacement keyword context module and the modified keyword context module to form a mixed keyword context module; Step 57115: Determine a mixed quantity based on the replacement keyword and the revised keyword; Step 571150: When the mixed quantity is greater than the quantity threshold, outputting the mixed keyword context module as the keyword context module; Step 571151: When the mixed quantity is less than the quantity threshold, the blank context module corresponding to the mixed keyword context module is discarded.

[0010] Optionally, the method of constructing the keyword context module into a scenario simulation includes: Step 60: Determine module attributes based on the keyword context module; Step 61: searching for corresponding scenario simulation solutions, scenario simulation solution credibility, and scenario simulation solution credibility threshold from a preset scenario simulation library based on all the module attributes and the information domain; Step 610: If the credibility of the scenario simulation solution is greater than the scenario simulation solution credibility threshold, simulating the keyword corresponding to the keyword context module in the scenario simulation solution; Step 611: When the scenario simulation solution credibility is less than the scenario simulation solution credibility threshold, output a preset information missing alarm signal.

[0011] Optionally, the method of simulating the keyword corresponding to the keyword context module in the scenario simulation scheme includes: Step 6100: splicing the keyword context modules into a complete sentence based on the scenario simulation solution; Step 6101: Quantitatively analyze the logic of the complete statement to obtain the logic value of the complete statement; Step 6102: Sort the complete sentences based on the logical values ​​to obtain a credibility ranking of the complete sentences; Step 6103: Based on the credibility ranking, the complete sentence with the highest logical value is output as the key content, and a preset number of complete sentences are selected from the remaining complete sentences according to the credibility ranking to be output as supplementary content of the key content.

[0012] An optional method of sorting the complete statements based on the logical values ​​to obtain a credibility ranking of the complete statements includes: Step 61020: searching for a corresponding logic difference threshold from a preset logic database based on the information field; Step 61021: combining the complete statements based on the logic difference threshold to form a complete statement group, wherein the absolute difference between the logic values ​​of any two complete statements in the complete statement group is less than the logic difference threshold; Step 61022: Combining the keywords in the complete sentence group to obtain combined keywords; Step 61023: splicing the combined keywords into a combined complete sentence based on the scenario simulation solution; Step 61024: defining the largest logical value in the keyword context module in the complete sentence group as the maximum logical value; Step 61025: Quantitatively analyze the logic of the combined complete statement to obtain a combined logic value of the combined complete statement; Step 610250: If the combined logical value is greater than the maximum logical value, the combined complete statement is used to replace all complete statements forming the complete statement group and then sorted according to step 6102; Step 610251: If the combined logical value is less than the maximum logical value, no combination is performed, and the complete statements are directly sorted based on the logical value to obtain a credibility ranking of the complete statements.

[0013] Optionally, when the combined logical value is less than the maximum logical value, a method for determining whether to not perform the combination includes: Step 6102510: Analyze the complete sentence group to obtain logical relationships and credibility values; Step 6102511: Find the corresponding credibility value threshold from the preset credibility value database based on the logical relationship; Step 6102512: When the credibility value is higher than the credibility value threshold, replace all complete statements forming the complete statement group with the combined complete statement; Step 6102513: When the credibility value is lower than the credibility value threshold, no combination is performed.

[0014] Optionally, based on the credibility ranking, the method of outputting the complete sentence with the highest credibility as the key content, and selecting a preset number of complete sentences from the remaining complete sentences according to the credibility ranking and outputting them as supplementary content to the key content includes: Step 61030: Determine the content relevance of the information text based on the complete sentence; Step 61031: When the content relevance is higher than a preset reasonable relevance threshold, a preset number of the complete sentences with the highest credibility are output as the key content, and a preset number of the complete sentences are selected from the remaining complete sentences in order of credibility to be output as supplementary content to the key content; Step 61032: When the content relevance is lower than a preset reasonable relevance threshold, the information text is segmented based on the content relevance to obtain secondary information text; Step 61033: Classify the complete sentence based on the secondary information text to obtain a secondary complete sentence; Step 61034: Output a preset number of the secondary complete sentences with the highest credibility as the respective key contents, and select a preset number of the secondary complete sentences from the remaining secondary complete sentences according to the credibility sorting to output as supplementary content of the key contents.

[0015] In a second aspect, the present invention provides a keyword information extraction system, which adopts the following technical solution: A keyword information extraction system, comprising: A collection module, configured to collect the information text and the suspected information text; A memory for storing a program for controlling a keyword information extraction method as described above; The program in the processor memory can be loaded and executed by the processor to implement a keyword information extraction method.

[0016] In summary, this application includes at least one of the following beneficial technical effects: By constructing a keyword context association network and combining semantic connection strength with statistical characteristics, we can effectively capture deep semantic associations and avoid the deviation caused by semantic loss in traditional weight analysis.

[0017] By presetting the framework constraints of the blank context module, keywords are forced to be embedded in reasonable semantic positions according to contextual associations, avoiding logical breaks caused by keyword stacking and improving content coherence.

[0018] Through the dynamic combination of modular semantic units, the logical relationship of real scenarios is simulated to ensure that the output content conforms to the actual application scenario and avoid theoretical distortion. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flowchart of a keyword-based information extraction method and system in an embodiment of the present application.

[0020] Figure 2 Schematic diagram of a method for obtaining a keyword context module in an embodiment of the present application.

[0021] Figure 3 It is a schematic diagram of the context association network in an embodiment of the present application.

[0022] Figure 4 It is a flowchart of constructing a scenario simulation based on the keyword context module in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0024] The embodiment of the present application discloses a keyword-based information extraction method. Figure 1 , a keyword-based information extraction method includes: Step 1: Receive the information text.

[0025] Information text refers to text whose core purpose is to convey information or knowledge, such as news, industry reports, and popular science articles. Programs can receive input information text data through interfaces such as standard input, file reading, or network requests, and convert it into strings for processing.

[0026] Step 2: Extract keywords from the information text based on the preset extraction algorithm.

[0027] Extraction algorithms refer to computational methods that automatically identify candidate words from unstructured text and calculate their importance based on techniques such as statistics, semantic analysis, or rule matching, such as common word frequency statistics (TF-IDF), co-occurrence relationship analysis, or deep learning models (such as TextRank).

[0028] Keywords are the smallest semantic units that summarize the core content of a text, typically nouns or noun phrases. Keyword extraction typically involves breaking the text into word sequences using a word segmentation tool, then filtering out meaningless words using a stop word list. Finally, a target number of words are selected based on word frequency weighting and contextual relevance.

[0029] For example, based on semantic units, news text is divided into structural blocks such as title, lead, main body, and conclusion. A position weight gradient is established, such as title weight = 3, first paragraph = 2, main text = 1, and last paragraph = 1.5 (weight parameters can be adjusted according to actual conditions). Secondly, a part-of-speech filtering strategy is used to remove modifying words such as adjectives and adverbs, retaining nouns, verbs, and professional terms as the candidate set; finally, a dual scoring mechanism is established, combining word frequency statistics and position-weighted scores (for example, "artificial intelligence" appears once in the title and is scored 3 points, and appears three times in the main text, which is 3 points × 1 = 3 points, for a total score of 6 points). Core keywords are dynamically filtered according to the score threshold.

[0030] Step 3: Determine the information field to which the news text belongs based on the keywords.

[0031] The information domain refers to the industry, discipline, or subject matter that the text content involves or belongs to. For example, sub-directions such as "Technology - Artificial Intelligence," "Finance - Blockchain," and "Medical - Tumor Immunology" are included. Keywords are compared with the preset domain database, and the information domains corresponding to all keywords are extracted. The information domain with the most occurrences is the information domain to which the information text belongs. The domain database is obtained by staff through a large number of experiments. The process of the experiment is as follows: after keyword extraction from a large number of information texts in different fields, the extracted keywords are matched with the corresponding information domains, and the matching relationship is stored in the domain database.

[0032] Step 4: Obtain the context category of the keyword based on the preset association algorithm.

[0033] Context category refers to the semantic role it plays in the text, such as subject, event, time, etc.

[0034] The association algorithm uses a preset candidate keyword library to calculate the keyword's vocabulary matching degree to the candidate keyword library, and combines the contextual connection to calculate the keyword's contextual adaptability to each context type. The specific calculation method is as follows: context adaptability = vocabulary matching degree × context weight. The candidate keyword library is obtained by a large number of experiments conducted by the staff, and the candidate keyword library includes candidate keywords, context categories and attribution probabilities. The process of the experiment is: classify a large number of keywords of different context categories, and calculate the probability of the keyword appearing in the context, and use it as the attribution probability of the candidate keyword, and finally store the mapping relationship between the candidate keyword, context category and attribution probability in the candidate keyword library. For example: the candidate keyword library may include: Time context: day, month, future, 2025, etc. Subject context: company, team, laboratory, school, etc. Match the obtained keywords with the candidate keyword library. If the exact same keywords are found, the vocabulary matching degree is the candidate keyword's attribution probability value. If the exact same keywords are not found, the keyword with the most identical words is selected from the candidate keyword library and the attribution probability value is multiplied by the correction parameter. The correction parameter is calculated as follows: , the vocabulary matching degree of the keywords that are not completely identical is = correction parameter × attribution probability.

[0035] For example, if the keyword "last week" exists in the candidate keyword database and has a probability of 0.9 for belonging to the time context, then the module match for the keyword "last week" in the news text is 0.9. However, "high-temperature superconducting energy storage device" does not exist in the candidate keyword database, but "energy storage device" does exist, and has a probability of 0.8 for belonging to the object context. The correction parameters for "energy storage device" and "high-temperature superconducting energy storage device" are 4 / 8 = 1 / 2, so the vocabulary match for the keyword "high-temperature superconducting energy storage device" in the news text is 0.8 × 0.5 = 0.4.

[0036] The contextual association of keywords can be derived by analyzing the grammatical structure of the sentences to which the keywords belong. For example, the sentence in the news text containing the keyword "school" is "The school held a welcome party." Through grammatical structure analysis, it can be seen that "school" serves as the subject in this sentence, so "school" does not apply to the "location context," and therefore has a context weight of 0 as the "location context." However, "school" applies to the "subject context," so its context weight as the "subject context" is 1. Similarly, for example, the sentences in the news text containing the keyword "school" are "The school held a welcome party," "Held on the school playground," and "The school provided the venue for the party." In these sentences, "school" belongs to both the "subject context" and the "location context." Based on the frequency of occurrence in different contexts, the context weight of "school" as the "subject context" is 2 / 3, and the context weight of "school" as the "location context" is 1 / 3.

[0037] The contextual adaptability of the keywords to each context category is calculated, and the keywords are assigned to one or more context categories according to the context adaptability from high to low.

[0038] Step 5: Fill the keywords into the preset blank context module based on the context category to obtain a keyword context module.

[0039] A context module is a structured framework that can carry information and is pre-divided according to the context. Each module consists of keyword slots of a specific context type and is used to systematize keywords in information texts. Each blank context module will contain the corresponding context category, for example, the "time module (when)" will contain the time context type. Keywords with the same context category as the context module will be saved in the blank context module to obtain the keyword context module. For the specific operation process, please refer to Figure 2 .

[0040] Step 6: Build a scenario simulation with the keyword context module to obtain key content that conforms to the actual scenario and output it.

[0041] Scenario simulation refers to the restoration or construction of the actual application scenario described in the text through logical deduction and element reorganization based on keyword context modules, including but not limited to structured information such as Who, What, When, and How, combined with contextual associations and information fields, and generating core content expressions that conform to the scenario. For example, from the information text of "Company Releases New Product", modules such as Who (Company), What (Product), When (Time), and How (Technical Highlights) are extracted and combined into a complete sentence of "A certain company launches a certain technology product at a certain time". This complete sentence is the core content of the information text. This modular design allows scattered keywords to be dynamically combined according to logical relationships while maintaining the semantic independence between different modules, providing a structural foundation for subsequent information reorganization and scenario-based applications.

[0042] The steps after step 5 and before step 6 also include: Step 50: Find corresponding relevance between the keywords in any keyword context module and other keywords from a preset keyword database.

[0043] Relevance refers to the direct or indirect connection between two keywords at the semantic, contextual, or domain knowledge level. The system automatically matches the relevance strength of keywords using a pre-set keyword database.

[0044] The keyword database is set up by a number of technical personnel in this field. The staff first classifies the keywords into modules by analyzing a large amount of information text, and then judges the correlation between the keywords in the module and other keywords through manual experience: if the two keywords have a direct explanatory relationship in semantics, often appear together in similar texts, or belong to the same field knowledge system, the two keywords are marked as strongly correlated. If the correlation is weak or requires specific conditions to establish a connection, the two keywords are marked as weakly correlated. At the same time, the frequency of the two keywords appearing together in the information text is saved as the correlation strength. After review by multiple people, the correlation results are entered into the database in the format of "keyword A-keyword B-correlation strength" for easy subsequent calls. When a keyword in any keyword context module is received together with other keywords, the correlation between the two keywords is automatically found from the database and output.

[0045] Step 51: Construct keywords based on relevance to form a contextual association network.

[0046] A contextual association network refers to a structured topological system composed of keyword nodes and their association relationships. Nodes represent keywords, edges represent associations, such as semantic similarity, knowledge graph paths, or co-occurrence frequencies, and edge weights reflect the strength of associations.

[0047] Step 52: Analyze the number of connection edges and corresponding connection strengths of keywords in the context association network.

[0048] The number of connection edges refers to the number of direct connections between a keyword and other nodes in the contextual association network. The connection strength refers to the weight of each edge. Figure 3 In the figure, each keyword is a node, and the number on the line connecting each keyword and other keywords is the weight of the connection strength between the two keywords, and the number of lines is the number of connection edges. For example, the connection strength between "mobile phone" and "battery" is 0.9, and the number of connection edges of "mobile phone" is 6.

[0049] Step 53: Obtain the occurrence count and position parameters of the keyword in the information text based on a preset traversal algorithm.

[0050] A traversal algorithm systematically scans every character or semantic unit in a news article to comprehensively capture the distribution characteristics of keywords. The "occurrence count" refers to the cumulative frequency of a keyword's occurrence in the text, while the "position parameter" quantifies the spatial characteristics of the keyword within the text, such as whether it appears in key locations such as the title, first paragraph, or last paragraph. Traversal algorithms are common knowledge and will not be elaborated on here.

[0051] The text structure is parsed through the traversal algorithm, the number of occurrences of keywords is counted and their position parameters are recorded.

[0052] Step 54: Determine the keyword importance value of the keyword based on the number of connection edges, connection strength, number of occurrences and position parameters.

[0053] Keyword importance is a comprehensive assessment of a keyword's contribution to the content theme by quantifying its connection characteristics in the contextual network (such as the number and strength of connections), its frequency of occurrence in the text, and its positional distribution. First, the total number of connections and strength of the keywords in the network are calculated. This is then combined with the number of occurrences in the text (frequency weight) and positional parameters (such as the bonus given to key locations like the beginning and end of a paragraph and in the title) to calculate a comprehensive score using a weighted algorithm or empirical formula.

[0054] Substitute the corresponding parameters into the empirical formula: weight × ∑ (number of edges × connection strength) + weight × ∑ (number of occurrences × position parameter) = keyword importance.

[0055] Step 55: Select keywords whose importance values ​​are lower than a preset importance threshold, discard them from the keyword context module to obtain a modified keyword context module, and define the remaining keywords in the modified keyword context module as modified keywords.

[0056] Importance threshold refers to a numerical threshold preset according to domain needs or task objectives, which is used to measure the core degree of keywords in the context network. Keywords below this value are judged as non-core keywords. The modified keyword context module refers to the structured context framework remaining after removing low-importance keywords, retaining only keywords that contribute significantly to the content theme. The modified keywords refer to the remaining keywords that contribute significantly to the content theme retained in the modified keyword context module after removing non-core keywords with importance values ​​lower than the preset importance threshold. The importance threshold is obtained by experiments conducted by staff, through manual screening and exclusion of keywords with weaker contributions to the theme, and finally through multiple rounds of experiments and adjustments, repeatedly verifying the correlation between keywords and content themes, determining the critical keyword importance value that can reflect this correlation, and setting it as the importance threshold.

[0057] Step 56: Output the modified keyword context module as a keyword context module.

[0058] Step 57: When the number of the modified keywords is less than a preset quantity threshold, output a preset module alarm signal.

[0059] Module alerts indicate the current context module may be at risk of missing information or incomplete analysis. During information processing, if the number of corrected keywords after importance screening is insufficient, the semantic information of the information text is deemed too sparse to support effective scenario simulation or decision-making. In this case, an alert is output to warn of data quality deficiencies or algorithmic oversights.

[0060] In addition, when the number of the modified keywords is greater than a preset critical number, the modified keyword context module is output as a keyword context module.

[0061] The method for outputting a module alarm signal when the number of corrected keywords is less than a critical number includes: Step 570: Lowering the keyword determination threshold preset in the extraction algorithm to form a modified extraction algorithm.

[0062] The extraction algorithm is equipped with a keyword judgment threshold. Modifying the extraction algorithm means dynamically adjusting the judgment threshold in the original keyword extraction algorithm, relaxing the screening conditions to expand the range of candidate keywords, and thus adding more potential keywords to the optimized extraction algorithm. When the number of modified keywords is less than the critical value, it is necessary to query the keyword judgment threshold reduction range through the preset keyword judgment threshold table. The keyword judgment threshold table is obtained by the staff through a large number of experiments. The process of the experiment is: in different information fields, by gradually lowering the keyword judgment threshold, until words unrelated to the core of the information appear in the extracted keywords, and then the staff records the threshold reduction range and matches it with the corresponding information field and stores it in the keyword judgment threshold table. Subsequently, when the number of modified keywords is less than the critical value, the system automatically lowers the judgment threshold of the extraction algorithm according to the table.

[0063] For example, when the original algorithm has insufficient effective keywords due to a too high threshold, the keyword threshold parameter can be reduced from 0.8 to 0.6 to allow more keywords that were not originally included to enter the candidate pool and participate in subsequent steps.

[0064] Step 571: Re-execute steps 2 to 55 based on the modified extraction algorithm.

[0065] Step 5710: Execute step 6 when the number of modified keywords after re-execution of the step is greater than the quantity threshold.

[0066] After obtaining the modified keyword context module, it is still necessary to determine whether the keywords obtained after modifying the extraction algorithm can be applied after calculating the keyword importance value. If the number of modified keywords after re-executing the steps is greater than the critical value, it means that the newly extracted keywords meet the requirements, so that the corresponding modified keyword context module can be applied to the next step.

[0067] Step 5711: When the number of the modified keywords after re-execution of the step is still less than the critical number, a module alarm signal is output.

[0068] The number of corrected keywords after re-executing the steps is still less than the critical value, which means that the newly obtained keywords still cannot meet the minimum number required by the corrected keyword context module. Therefore, it means that the corresponding keywords are not sufficiently related to the core of the information content, and the information text is suspected to lack key information, so the output module alarm signal.

[0069] Another method for extracting information when the number of corrected keywords after re-performing the step is still less than the critical number is also included, the method comprising: Step 57110: Search the Internet for suspected information text based on keywords.

[0070] Suspected news text refers to a collection of candidate texts retrieved from the internet through keyword matching that may be related to the target news text. All keywords are entered into a search engine such as "Baidu," and a preset number of suspected news texts are downloaded and imported into the program in batches, sorted by publication time, from recent to recent. The preset number is determined by staff through experimentation.

[0071] Step 57111: Calculate the similarity value based on the preset judgment algorithm between the suspected information text and the information text.

[0072] A judgment algorithm refers to a mathematical model or rule used to calculate text similarity, such as the cosine similarity algorithm. The cosine similarity algorithm measures the directional similarity between two vectors by calculating the cosine of the angle between them in multidimensional space. Its core idea is that the closer the directions of the vectors, the more similar the content.

[0073] For example, consider Text 1: "Machine learning is the core of artificial intelligence" and Text 2: "Deep learning is a branch of machine learning." Extract the vocabulary from Text 1 and Text 2: {"machine learning," "artificial intelligence," "core," "deep learning," "branch"}. Use the TF-IDF method to construct vectors for these texts.

[0074] The vector of text 1 is {0.8, 0.4, 0.4, 0, 0}, and the vector of text 2 is {0.5, 0, 0, 0.7, 0.7}. , the similarity is obtained by the cosine similarity algorithm and then compared with the threshold to determine whether it is similar.

[0075] Similarly, the Jaccard index can also be used to determine whether two texts are similar. The above two methods are two methods for determining similarity, and specific practical methods include but are not limited to the above two methods.

[0076] Step 57112: When the similarity value corresponding to the suspected information text is higher than the preset similarity threshold, the suspected information text is defined as a similar information text.

[0077] The similarity threshold is a pre-set numerical standard used in determining information text similarity, used to measure the degree of similarity between two texts. Similarity thresholds for different information domains can be obtained using the cosine similarity method in step 57111. When the calculated similarity value of the suspected information text exceeds the similarity threshold, it indicates that the suspected information text and the information text are highly similar, and both reflect the same core information content.

[0078] This step objectively selects texts with true semantic associations by quantifying similarity indicators, avoiding errors caused by human subjective judgment and significantly improving the efficiency and accuracy of information deduplication or clustering.

[0079] Step 57113: Re-execute steps 2 to 4 based on similar information text to obtain alternative keyword context modules and alternative keywords.

[0080] The replacement keyword context module refers to the combination of all keywords extracted from similar news text and their corresponding context modules. Replacement keywords refer to all keywords extracted from similar news text. After obtaining similar news text, steps 2 to 4 are re-executed, using an association algorithm to extract keywords from the text and defining these keywords as replacement keywords.

[0081] Step 57114: Merge the replacement keyword context module and the modified keyword context module to form a hybrid keyword context module.

[0082] The hybrid keyword context module is a context module obtained by fusing the replacement keyword context module and the correction keyword context module, and has all the keywords of both.

[0083] For example, for information about solar panels, the keyword provided by the replacement keyword context module is "photovoltaic technology", and the keyword provided by the correction keyword context module is "solar panels". The "main module (who)" in the mixed keyword context module will contain both the keywords "photovoltaic technology" and "solar panels".

[0084] Step 57115: Determine the mixed quantity based on the replacement keyword and the revised keyword.

[0085] The mixed number refers to the total number of deduplicated words obtained by fusing alternative keywords with modified keywords. The keywords in the alternative keyword context module and the modified keyword context module are merged and then duplicates are removed to obtain a mixed keyword set. The number of words in the final statistical set is the mixed number.

[0086] Step 571150: When the mixed quantity is greater than the quantity threshold, the mixed keyword context module is output as a keyword context module.

[0087] When the mixed quantity is greater than the quantity critical value, it means that the number of alternative keywords extracted from the similar information text is sufficient to be simulated according to the scenario simulation plan and meets the requirements of step 620, and can be output as a keyword context module.

[0088] Step 571151: When the mixed quantity is less than the critical quantity, the blank context module corresponding to the mixed keyword context module is discarded.

[0089] When the mixed quantity is less than the critical value, it means that even if similar information texts are extracted, suitable keywords for the blank context module cannot be extracted, so the corresponding blank context module is discarded.

[0090] Reference Figure 4 ,The method of constructing a scenario simulation with the keyword context module includes: Step 60: Determine module attributes based on the keyword context module.

[0091] Module attributes refer to the properties of a keyword context module, representing the role that the keywords within the keyword context module represent in the scenario simulation. A keyword context module is created by populating a preset blank context module with keywords. Therefore, the keyword context module inherits the context category of the blank context module. For example, the keywords in the "subject module (who)" contain the module attributes of "who." Keywords within this module typically play the role of the subject in the scenario simulation. Therefore, module attributes can be determined by querying the attributes of the keyword context module.

[0092] Step 61: Based on all module attributes and information fields, the corresponding scenario simulation solution, scenario simulation solution credibility and scenario simulation solution credibility threshold are searched from the preset scenario simulation library.

[0093] The scenario simulation scheme is a standardized operational framework designed based on scenario templates. This framework is used to combine keywords within the keyword context module. The scenario simulation scheme's credibility is a matching metric, set based on historical data or expert experience. It is used to determine the degree of fit between module attributes and the pre-set scenario, preventing simulation bias caused by low-matching schemes. The scenario simulation library stores the mapping between module attributes, information domains, scenario simulation schemes, their credibility, and credibility thresholds.

[0094] The scenario simulation library was created by staff through analytical testing of existing information texts of different types and their core content. During the testing process, staff extracted the core content from a large number of information texts and compared the similarities in the presentation of similar information texts. They then organized the content into different modules and formed solutions based on the correlations between the modules. At the same time, staff will dynamically adjust the credibility of the solutions based on missing modules. They will also conduct a credibility assessment and finally store all data records in the scenario simulation library. After all module attributes and information fields are entered, the system searches the scenario simulation library for the scenario simulation solution composed of all the module attributes entered under that information field, and finds the corresponding scenario simulation solution credibility, as well as the scenario simulation solution credibility threshold under that information field.

[0095] Step 610: If the scenario simulation solution credibility is greater than the scenario simulation solution credibility threshold, the keyword corresponding to the keyword context module is simulated in the scenario simulation solution.

[0096] Simulation refers to substituting qualified keyword context modules into the selected scenario simulation plan to generate structured expressions or execute instructions.

[0097] When the credibility of the scenario simulation solution is greater than the preset scenario simulation solution credibility threshold, it means that the solution can be implemented, and the keywords corresponding to the keyword context module are simulated in the scenario simulation solution.

[0098] Step 611 : When the scenario simulation solution credibility is less than the scenario simulation solution credibility threshold, output a preset information missing alarm signal.

[0099] The missing information alarm signal is triggered automatically by the system when the credibility of a scenario simulation solution falls below a preset threshold. It alerts the user that the current keyword context module contains incomplete, conflicting, or noisy information and requires manual intervention or supplemental data. The output is textual, displaying "Missing Information Alert!" on the user interface.

[0100] The method of simulating the keywords corresponding to the keyword context module in the scenario simulation scheme includes: Step 6100: Splice the keyword context modules into complete sentences based on the scenario simulation solution.

[0101] A complete sentence is a semantically coherent, structurally complete expression formed by combining the logical connections and grammatical rules of keyword context modules (such as subject, event, time, and method). This statement must accurately convey the core information of the news and meet human language conventions. When simulating keywords corresponding to keyword context modules within a scenario simulation scheme, since fixed scenario simulation schemes cannot fully and accurately meet human language conventions, language models are required to correct them. The following illustrates how to use an NLP model to stitch and correct keyword context modules into complete sentences.

[0102] For example, the above improved simulation conditions based on scenario simulation scheme 2 result in the following simulation scenario: Scenario 1: "Recently, Company A released Car A, resulting in sales of 20,000 units." Scenario 2: "This spring, Company A released Car A, which resulted in sales of 20,000 units." (Note: This is the first part of the scenario simulation, and the accuracy and logic of the sentence have not yet been improved. For example, the phrase "resulted in sales of 20,000 units" in the example is a clear grammatical error.) The complete sentence in Scenario 1 is analyzed based on the NLP model, and fine-grained word segmentation is performed using the NLP model: {"recently / t", "company / n", "A / eng", "release / v", "released / ule", "car / n", "A / eng", "bring / v", "released / ule", "sold / v", "2w / m", "units / q"}. Based on fine-grained word segmentation, the key verb "bring / v" is abnormally paired with the verb "sell / v", resulting in "bringing and selling 2w units" being a clear grammatical error. Because "sold 2w units" is a keyword derived from the information text and is not modified, "bring / v" needs to be changed. Combined with the Chinese grammar rule library, the transitive verb "bring" must be followed by a result-oriented noun object, and the "sell + quantity" structure should be preceded by an adverb such as "already / accumulated". The sentence obtained in the scenario is replaced with "Recently, Company A released Car A and has sold 2w units."

[0103] Step 6101: Quantitatively analyze the logic of the complete statement to obtain the logical value of the complete statement.

[0104] The logical value is a quantitative result of whether a complete statement is logical, and is used to measure whether the statement conforms to the logical rigor of human cognition.

[0105] To obtain the logical value of a complete sentence, we analyze the strength of the relationship between the complete sentence and the information text, as well as factors such as the matching degree of keywords within each context module of the complete sentence, causal relationships, and temporal consistency. All these factors are then weighted and added together to obtain the logical value of the corresponding complete sentence.

[0106] For example, the logical value of a complete sentence is calculated based on three factors: context module matching, causal rationality, and information density. The context module matching can be calculated using the cosine similarity algorithm, and the remaining factors can also be calculated using corresponding algorithms. The relevant algorithms and factors are common knowledge and will not be elaborated here.

[0107] Step 6102: Sort the complete sentences based on the logical values ​​to obtain a credibility ranking of the complete sentences.

[0108] Credibility ranking refers to the order of credibility. This ranking prioritizes statements from highest to lowest logical value by quantitatively evaluating the logical value of each complete statement. This ranking allows for intuitive judgment of the credibility of complete statements and which complete statements best reflect the core content of the news article.

[0109] Step 6103: The complete sentence with the highest logical value is output as the key content, and a preset number of complete sentences are selected from the remaining complete sentences in order of credibility and output as supplementary content to the key content.

[0110] Staff determined the number of complete sentences to be supplemented by analyzing large amounts of information text, combining the characteristics and patterns of the information's field, and comprehensively considering user needs and information density. This is intended to ensure that the supplementary content provides sufficient and reasonable information reflecting the key content of the information text.

[0111] For example, the complete sentences derived from the news text are "In April, Company A sold 2,000 units of Car A." (credibility 0.9), "This month, Company A sold several units of Car A." (credibility 0.8), and "This month, Company A announced good news." (credibility 0.7). When output according to the rule, "In April, Company A sold 2,000 units of Car A." is directly output as the key content of the news text, while "This month, Company A sold several units of Car A." and "This month, Company A announced good news." are output as supplementary content below the key content and labeled "Supplementary Content."

[0112] Methods for sorting complete statements based on logical values ​​to obtain credibility rankings of the complete statements include: Step 61020: Find the corresponding logic difference threshold from the preset logic database based on the information field.

[0113] The logical database is a structured storage system used to store logical difference thresholds and other logical rule parameters corresponding to different information domains. The logical difference threshold is a preset numerical parameter that measures the maximum allowable difference between the logical values ​​of two complete statements. When the absolute difference between the logical values ​​of two statements is less than the threshold, the system considers their logical connection strong enough to be combined into a more complete statement. Conversely, if the difference exceeds the threshold, the logical connection is considered weak and the combination is not performed.

[0114] The logical database is based on extensive testing conducted by our staff. During this testing process, we extracted the core information from various news articles based on keywords, and then determined whether different complete sentences reflected the core of the article. Finally, we set different logical difference thresholds for each field. When the absolute difference between the logical values ​​of any two complete sentences is less than this threshold, it indicates that both complete sentences reflect the content of the news article to some extent. When the system receives an information field, it automatically searches the database for the logical difference threshold for the corresponding information field.

[0115] Step 61021: Combine complete statements based on the logical difference threshold to form a complete statement group.

[0116] The absolute difference between the logical values ​​of any two complete statements in a complete statement group is less than the logical difference threshold. A complete statement group is a set of semantically complete and logically consistent statements. This is done by traversing all statements, starting with the first statement, and selecting all statements whose logical value difference is less than the logical difference threshold to form a statement group. The group is then executed starting with the second statement.

[0117] Step 61022: Combine the keywords in the complete sentence group to obtain combined keywords.

[0118] Combined keywords are a collection of keywords extracted from complete sentence groups, formed through logical association and semantic integration. This aims to integrate core information from multiple sentences, covering the entire topic of the news while avoiding redundancy. For example, a complete sentence group containing "Company A sells mobile phone A." and "Company A sells mobile phone B." To fully reflect the core content of the news text, these keywords are extracted and combined, meaning the keywords "mobile phone A" and "mobile phone B" are combined to become "mobile phone A and mobile phone B."

[0119] Step 61023: Based on the scenario simulation plan, the combined keywords are spliced ​​into a combined complete sentence.

[0120] Combining complete sentences involves integrating the keywords and logical relationships within a complete sentence group to create a semantically complete, logically consistent, and core-topic complete sentence. For example, in step 61022, after obtaining the combined keywords, they are reassembled to create the combined complete sentence "Company A sells mobile phones A and B."

[0121] Step 61024: Define the largest logical value in the keyword context module within the complete sentence group as the maximum logical value.

[0122] Step 61025: Quantitatively analyze the logic of the combined complete statement to obtain the combined logic value of the combined complete statement.

[0123] The combined logical value is a numerical measure of the logical consistency of a complete sentence. It is used to quantify the overall rationality of the combined information of multiple sentences. Its core goal is to determine whether the combined result conforms to logical rules and is free of contradictions by analyzing the correlation between semantics, structure, time, and location elements between sentences.

[0124] The calculation method of the combinational logic value is the same as the logic calculation step of step 6101 and will not be repeated here.

[0125] Step 610250: If the combined logical value is greater than the maximum logical value, the combined complete statement is replaced with all complete statements forming the complete statement group and then sorted according to step 6102.

[0126] When the combined logical value is greater than the maximum logical value, it means that the combined complete sentence obtained after the combination is more consistent with the information text and can better reflect the core content. The combined complete sentence will replace all the complete sentences forming the complete sentence group and then be sorted according to step 6102.

[0127] For example, combine "Company A sells mobile phone A." with "Company A sells mobile phone B." to obtain the complete sentence "Company A sells mobile phone A and mobile phone B."

[0128] Step 610251: If the combined logical value is less than the maximum logical value, no combination is performed, and the complete statements are directly sorted based on the logical value to obtain a credibility ranking of the complete statements.

[0129] For example, "Company A sold mobile phone A" and "Company B sold mobile phone B" both have high logical values, reflecting the main content of the news text. However, when combined, the complete sentence is "Company A and Company B sold mobile phones A and B." This combination results in a lower logical value (Company A did not sell mobile phone B, and Company B did not sell mobile phone A). Therefore, the complete sentences are sorted directly based on their logical values ​​without combining them to determine their credibility.

[0130] The method further includes a method for determining whether to not perform the combination if the combination logic value is less than the maximum logic value, the method comprising: Step 6102510: Analyze the complete statement group to obtain logical relationships and credibility values.

[0131] Logical relationships refer to the semantic association patterns between statements within a complete statement group, describing how information is connected. Common types include causal relationships and parallel relationships. Credibility values ​​are quantitative assessments of the probability that a logical relationship holds true, used to measure the rationality and reliability of a statement group.

[0132] For example, the probability of causal relationships between complete sentences can be analyzed by weighted combination of three aspects: semantic similarity, event temporal sequence, and verb causal implication.

[0133] ① Semantic similarity (S): Calculate the cosine similarity of two complete sentences. The higher the cosine similarity, the more likely they are in a parallel relationship. If it is lower, it may be a causal relationship. The specific cosine similarity calculation method is mentioned in step 57111 and will not be repeated here.

[0134] ② Event temporality (T): Extract the "when" keyword from two complete sentences and determine the temporal context of the complete sentences based on the implicit clues of the time keyword. If event A clearly occurs before event B, a causal relationship is likely. Specific parameters for event temporality can be determined based on the domain of the information text.

[0135] ③ Verb Causality (C): Search the news text for words containing causal probabilities, such as "because," "so," and "result." Compute the cosine similarity between the two complete sentences involved in the causal relationship and the sentences preceding and following them. The higher the cosine similarity with the sentence following "because," the more likely the complete sentence is to be a "cause." Similarly, the higher the cosine similarity with the sentence following "so," the more likely the complete sentence is to be a "result." The value with the highest cosine similarity for each complete sentence is retained as the specific parameter representing verb causality.

[0136] The causal relationship probability is calculated by weighted combination of the above features: , where α, β, and γ are the preset weight parameters obtained by the staff after multiple experiments, S is the semantic similarity, T is the event temporality, and C is the verb causality.

[0137] Step 6102511: Find the corresponding credible value threshold from the preset credible value database based on the logical relationship.

[0138] The credibility threshold is a critical value used to determine whether two or more complete statements satisfy a specific logical relationship, based on pre-set logical rules and a credibility database. The credibility database is a structured knowledge base that stores the credibility value distributions of various logical relationships (such as causal relationships and parallel relationships) across a large number of experimental scenarios. Its core purpose is to quantify the probability of different logical relationships holding true in real applications using historical experimental data. Experimenters extract and verify logical relationships from a large amount of real-world information through manual annotation, statistical analysis, or machine learning. After calculating the probability using the example in steps 6102510, they compare the resulting probability with the actual causal relationship to determine the credibility value.

[0139] When two or more complete sentences are analyzed to obtain a suspected logical relationship, the corresponding credibility value threshold of the logical relationship is automatically found from the database.

[0140] Step 6102512: When the credibility value is higher than the credibility value threshold, the combined complete statement replaces all complete statements forming the complete statement group.

[0141] Combining complete sentences means combining two or more independent complete sentences into a logically coherent and semantically complete statement based on logical relationships through conjunctions, semantic integration or information supplementation.

[0142] Step 6102513: When the credibility value is lower than the credibility value threshold, no combination is performed.

[0143] When the credibility value is lower than the credibility value threshold, it means that there is no logical relationship between two or more independent complete sentences. In order to ensure the accuracy of the core content, the two or more independent complete sentences will not be combined.

[0144] The method of outputting the complete sentence with the highest credibility as key content based on credibility sorting and selecting a preset number of complete sentences from the remaining complete sentences according to credibility sorting as supplementary content for the key content includes: Step 61030: Determine the content relevance of the information text based on the complete sentence.

[0145] Content relevance refers to the property of whether multiple complete sentences belong to the same information text in terms of semantics, themes, entities, or event logic, and is used as a judgment criterion. Content relevance can be obtained by methods including but not limited to calculating text similarity or cosine similarity between two complete sentences.

[0146] Step 61031: When the content relevance is higher than a preset reasonable relevance threshold, a preset number of complete sentences with the highest credibility are output as key content, and a preset number of complete sentences are selected from the remaining complete sentences in order of credibility as supplementary content to the key content.

[0147] The reasonable relevance threshold is an indicator used to determine whether information content is highly relevant. This threshold can be determined by consulting a pre-set reasonable relevance threshold table. This table was developed through multiple tests by staff. By analyzing a large number of information texts from different fields, staff calculated the similarity between information texts reflecting the same content, and determined the minimum correlation value that represents the similarity between the two. This value was then set as the reasonable relevance threshold. When the content relevance is higher than the pre-set reasonable relevance threshold, it indicates that the information texts are highly relevant, and there is no risk of mistakenly inputting other information texts. Content output can proceed normally.

[0148] Step 61032: When the content relevance is lower than a preset reasonable relevance threshold, the information text is segmented based on the content relevance to obtain secondary information text.

[0149] Content-relevance-based segmentation refers to segmenting information text into multiple independent information texts by analyzing indicators such as semantic similarity, entity relevance, and event logic chains between sentences, ensuring that each independent information text only reflects a single core content.

[0150] Secondary information text is a separate unit created by segmenting the original information text based on its content relevance. The complete sentences within each unit are highly relevant, focusing on the same topic, entity, or event, but are less relevant to other secondary information texts. The presence of secondary information text indicates that multiple pieces of single, complete information exist within the input text, possibly due to the user mistakenly mixing multiple pieces of information into one.

[0151] Step 61033: Classify the complete sentence based on the secondary information text to obtain a secondary complete sentence.

[0152] A secondary complete sentence is a complete sentence that reflects the core content of a secondary news text after segmentation based on content relevance. Classification based on secondary news text refers to the process of further categorizing complete sentences within segmented secondary news texts based on semantics, function, or information type. Because users may mistakenly mix multiple pieces of information, different complete sentences may correspond to different secondary news texts after segmentation. Therefore, classification based on the core content reflected by the complete sentences is necessary to ensure the legitimacy of the information.

[0153] Step 61034: Output a preset number of the most credible secondary complete sentences as their respective key contents, and select a preset number of complete sentences from the remaining secondary complete sentences sorted by credibility as supplementary content to the key contents.

[0154] From the classification results of each secondary news text, a preset number of representative sentences (e.g., 2-5 sentences per category) are extracted and output as the key content of that category. Complete sentences with slightly lower logical values ​​are also classified as news texts, and the key content of the corresponding news texts is supplemented and output.

[0155] Based on the same inventive concept, an embodiment of the present invention provides a keyword-based information extraction system.

[0156] A keyword-based information extraction system, comprising: The acquisition module is configured to acquire information text and suspected information text. The memory is configured to store a program for controlling a keyword information extraction method. The processor is configured to load and execute the program stored in the memory.

Claims

1. A keyword-based information extraction method, characterized in that: include: Step 1: Receive information text; Step 2: extracting keywords from the information text based on a preset extraction algorithm; Step 3: Determine the information field to which the information text belongs based on the keywords; Step 4: Obtain the context category of the keyword based on a preset association algorithm; Step 5: Filling keywords into a preset blank context module based on the context category to obtain a keyword context module; Step 6: Build a scenario simulation with the keyword context module to obtain key content that conforms to the actual scenario and output it.

2. The keyword-based information extraction method according to claim 1, characterized in that: The steps after step 5 and before step 6 also include: Step 50: searching for corresponding relevance between the keyword in any one of the keyword context modules and the other keywords from a preset keyword database; Step 51: constructing the keywords based on the correlation to form a contextual correlation network; Step 52: Analyze the number of connection edges and corresponding connection strengths of the keywords in the context association network; Step 53: Obtain the number of occurrences and position parameters of the keyword in the information text based on a preset traversal algorithm; Step 54: determining the keyword importance value of the keyword based on the number of connection edges, the connection strength, the number of occurrences, and the position parameter; Step 55: Filter the keywords whose keyword importance values ​​are lower than a preset importance threshold, discard them from the keyword context module to obtain a modified keyword context module, and define the remaining keywords in the modified keyword context module as modified keywords; Step 56: Outputting the modified keyword context module as the keyword context module; Step 57: When the number of the modified keywords is less than a preset quantity threshold, output a preset module alarm signal.

3. The keyword-based information extraction method according to claim 2, characterized in that: When the number of the modified keywords is less than the critical number, the method of outputting the module alarm signal includes: Step 570: Lowering a keyword determination threshold preset in the extraction algorithm to form a modified extraction algorithm, wherein the keyword determination threshold is set in the extraction algorithm; Step 571: re-execute steps 2 to 55 based on the modified extraction algorithm; Step 5710: When the number of the modified keywords after re-execution of the step is greater than the threshold value, execute step 6; Step 5711: When the number of the modified keywords after re-execution of the step is still less than the quantity threshold, the module alarm signal is output.

4. The keyword-based information extraction method according to claim 3, characterized in that: Another method for extracting information when the number of the modified keywords after re-performing the step is still less than the critical number is also included, the method comprising: Step 57110: Searching for suspected information text on the Internet based on the keyword; Step 57111: Calculate a similarity value based on the predetermined judgment algorithm between the suspected information text and the information text; Step 57112: When the similarity value corresponding to the suspected information text is higher than a preset similarity threshold, the suspected information text is defined as a similar information text; Step 57113: Re-execute steps 2 to 4 based on the similar information text to obtain an alternative keyword context module and alternative keywords; Step 57114: Merge the replacement keyword context module and the modified keyword context module to form a mixed keyword context module; Step 57115: Determine a mixed quantity based on the replacement keyword and the revised keyword; Step 571150: When the mixed quantity is greater than the quantity threshold, outputting the mixed keyword context module as the keyword context module; Step 571151: When the mixed quantity is less than the quantity threshold, the blank context module corresponding to the mixed keyword context module is discarded.

5. The keyword-based information extraction method according to claim 1, characterized in that: The method for constructing the keyword context module into a scenario simulation includes: Step 60: Determine module attributes based on the keyword context module; Step 61: searching for corresponding scenario simulation solutions, scenario simulation solution credibility, and scenario simulation solution credibility threshold from a preset scenario simulation library based on all the module attributes and the information domain; Step 610: If the credibility of the scenario simulation solution is greater than the scenario simulation solution credibility threshold, simulating the keyword corresponding to the keyword context module in the scenario simulation solution; Step 611: When the scenario simulation solution credibility is less than the scenario simulation solution credibility threshold, output a preset information missing alarm signal.

6. The keyword-based information extraction method according to claim 5, characterized in that: The method of simulating the keyword corresponding to the keyword context module in the scenario simulation scheme includes: Step 6100: splicing the keyword context modules into a complete sentence based on the scenario simulation solution; Step 6101: Quantitatively analyze the logic of the complete statement to obtain the logic value of the complete statement; Step 6102: Sort the complete sentences based on the logical values ​​to obtain a credibility ranking of the complete sentences; Step 6103: Based on the credibility ranking, the complete sentence with the highest logical value is output as the key content, and a preset number of complete sentences are selected from the remaining complete sentences according to the credibility ranking to be output as supplementary content of the key content.

7. The keyword-based information extraction method according to claim 6, characterized in that: The method of sorting the complete sentences based on the logical values ​​to obtain a credibility sort of the complete sentences includes: Step 61020: searching for a corresponding logic difference threshold from a preset logic database based on the information field; Step 61021: combining the complete statements based on the logic difference threshold to form a complete statement group, wherein the absolute difference between the logic values ​​of any two complete statements in the complete statement group is less than the logic difference threshold; Step 61022: Combining the keywords in the complete sentence group to obtain combined keywords; Step 61023: splicing the combined keywords into a combined complete sentence based on the scenario simulation solution; Step 61024: defining the largest logical value in the keyword context module in the complete sentence group as the maximum logical value; Step 61025: Quantitatively analyze the logic of the combined complete statement to obtain a combined logic value of the combined complete statement; Step 610250: If the combined logical value is greater than the maximum logical value, the combined complete statement is used to replace all complete statements forming the complete statement group and then sorted according to step 6102; Step 610251: If the combined logical value is less than the maximum logical value, no combination is performed, and the complete statements are directly sorted based on the logical value to obtain a credibility ranking of the complete statements.

8. The keyword-based information extraction method according to claim 7, characterized in that: The method further includes a method for determining whether to not perform the combination if the combination logic value is less than the maximum logic value, the method comprising: Step 6102510: Analyze the complete sentence group to obtain logical relationships and credibility values; Step 6102511: Find the corresponding credibility value threshold from the preset credibility value database based on the logical relationship; Step 6102512: When the credibility value is higher than the credibility value threshold, replace all complete statements forming the complete statement group with the combined complete statement; Step 6102513: When the credibility value is lower than the credibility value threshold, no combination is performed.

9. The keyword-based information extraction method according to claim 6, characterized in that: The method of outputting the complete sentence with the highest credibility as the key content based on the credibility ranking, and selecting a preset number of complete sentences from the remaining complete sentences according to the credibility ranking as supplementary content to output the key content includes: Step 61030: Determine the content relevance of the information text based on the complete sentence; Step 61031: When the content relevance is higher than a preset reasonable relevance threshold, a preset number of the complete sentences with the highest credibility are output as the key content, and a preset number of the complete sentences are selected from the remaining complete sentences in order of credibility to be output as supplementary content to the key content; Step 61032: When the content relevance is lower than a preset reasonable relevance threshold, the information text is segmented based on the content relevance to obtain secondary information text; Step 61033: Classify the complete sentence based on the secondary information text to obtain a secondary complete sentence; Step 61034: Output a preset number of the secondary complete sentences with the highest credibility as the respective key contents, and select a preset number of the secondary complete sentences from the remaining secondary complete sentences according to the credibility sorting to output as supplementary content of the key contents.

10. A keyword information extraction system, characterized in that: include: A collection module, configured to collect the information text and the suspected information text; A memory for storing a program of a keyword information extraction method according to any one of claims 1 to 9; The processor loads and executes the program in the memory.

Citation Information

Patent Citations

  • Information abstract extraction method and system based on keywords

    CN110674296A

  • Automatic generation method for information theme

    CN117077632A

  • Financial information recommendation method and system based on multi-channel distribution

    CN119149827A

  • Power relay assembly

    KR1020240166209A

Cited By

  • Sensitive data-based information security data analysis early warning method and system

    CN121486000A