Technical economic risk event extraction method and system based on monitoring entity

By constructing a method for extracting technological and economic risk events based on monitored entities, the problem of the lack of a professional event type system in existing technologies is solved. This enables high-precision event identification and structured risk event extraction in the field of technological and economic security, supporting intelligent identification of risk situations and decision support.

CN120930752AActive Publication Date: 2025-11-11DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Patent Information

Application Number
CN202511056725.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing technologies lack a professional event type system and a monitoring entity-driven extraction mechanism in the field of technological and economic security. This makes it impossible to accurately extract event structures with industry characteristics and high-risk orientation, which affects the intelligent identification of risk status of key entities, trend early warning, and knowledge graph construction.

Method used

By employing a technology and economic risk event extraction method based on monitored entities, including entity identification, semantic fragment extraction, dependency parsing, semantic role labeling, and event structure construction, a standardized event structure is formed, and a multi-dimensional event chain is constructed to achieve structured risk event extraction from multi-source scientific and technological texts.

Benefits of technology

It improves the accuracy of identifying technical and economic security incidents, enhances the depth of incident understanding, and supports risk decision-making support capabilities in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930752A_ABST
    Figure CN120930752A_ABST
Patent Text Reader

Abstract

The invention provides a technical and economic risk event extraction method and system based on monitoring entities, and relates to the technical field of entity extraction, and the method comprises the steps: carrying out the entity recognition of a multi-source science and technology text, carrying out the entity comparison and screening of technical and economic monitoring entities, and outputting a candidate monitoring entity list; performing semantic fragment extraction by taking the candidate monitoring entity list as an anchor point, and performing event type judgment based on a trigger word matching rule; performing dependency syntactic analysis and semantic role labeling on the trigger word set, constructing a standardized event structure, extracting a semantic relationship, and outputting a risk event candidate structure set; identifying chapter-level master events, and aggregating paragraph-level associated events to form a master-slave event structure; and assembling the event chain according to the master-slave event structure, and performing standardized output. By means of the method and device, the technical problem that in the prior art, the reliability of extraction of economic risk events is poor can be solved, and the technical effect of improving the reliability of extraction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of entity extraction technology, and in particular to a method and system for extracting technical and economic risk events based on monitored entities. Background Technology

[0002] With the continuous development of science and technology intelligence systems and intelligent text processing technologies, higher demands are being placed on the automated extraction and structured analysis of important risk events in the field of technology and economy.

[0003] Currently, existing event extraction methods mainly rely on general natural language processing models and general event type systems. Although they have certain event recognition and structuring capabilities in open domains and news and public opinion scenarios, they are difficult to cover high-risk event types when facing technology, economic and security scenarios. Furthermore, they lack an extraction perspective centered on the monitored entity, and often fail to form a systematic tracking of entity risk dynamics.

[0004] In summary, existing technologies suffer from a lack of a professional event type system and entity-driven extraction mechanism for the field of technological and economic security. This results in the inability to accurately extract event structures with industry characteristics and high-risk orientation, further affecting the technical effectiveness and system reliability of applications such as intelligent identification, trend warning, and knowledge graph construction of key entity risk status in the technological and economic field. Summary of the Invention

[0005] The purpose of this application is to provide a method and system for extracting technical and economic risk events based on monitored entities. This is to address the technical problem in the existing technology that, due to the lack of a professional event type system and a monitoring entity-driven extraction mechanism for the field of technical and economic security, it is impossible to accurately extract event structures with industry characteristics and high risk orientation. This further affects the technical effectiveness and system reliability of applications such as intelligent identification, trend warning, and knowledge graph construction of key entity risk status in the field of technical and economics.

[0006] In view of the above problems, this application provides a method and system for extracting technical and economic risk events based on monitored entities.

[0007] Firstly, this application provides a method for extracting technical and economic risk events based on monitored entities, implemented through a system for extracting technical and economic risk events based on monitored entities. The method includes: performing entity recognition on multi-source scientific and technological texts, extracting technical and economic monitored entities, and comparing and filtering these entities based on a preset monitored entity library to output a candidate monitored entity list; extracting semantic segments using the candidate monitored entity list as anchors, identifying event trigger words, and determining event types based on trigger word matching rules; performing dependency parsing and semantic role labeling on the identified trigger word set to construct a standardized event structure, extracting semantic relationships, and outputting a risk event candidate structure set; identifying chapter-level main events and aggregating paragraph-level related events based on the risk event candidate structure set to form a master-slave event structure; and constructing an event chain according to the master-slave event structure and outputting it in a standardized manner.

[0008] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: acquiring historical technical and economic risk events and determining event type identifiers; listing trigger words based on the historical technical and economic risk events and obtaining a set of trigger word patterns; defining participating elements through the historical technical and economic risk events to obtain argument role templates; and performing semantic relationship association based on the event type identifiers, the set of trigger word patterns, and the argument role templates to obtain semantic relationship rules and construct a knowledge system for technical and economic risk events.

[0009] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: dividing the multi-source scientific and technological text into sentence units according to semantic boundaries to obtain sentence-level segmentation; performing lexical analysis on the sentence-level segmentation to obtain entity clues; performing dependency tree structure analysis on the sentence-level segmentation to obtain sentence semantic relations; correcting and deduplicating the entity clues and sentence semantic relations, and combining them to obtain multi-dimensional text information.

[0010] Preferably, the method for extracting technical and economic risk events based on monitoring entities further includes: performing literal matching of the multidimensional text information according to the preset monitoring entity library to obtain fuzzy matching entities; performing entity expansion recognition based on the fuzzy matching entities on the multidimensional text information to obtain expanded matching entities; identifying the expanded matching entities in the multidimensional text information, and extracting context windows from before and after the text with the entity identifier as the center, and outputting a candidate monitoring entity list.

[0011] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: performing dependency syntactic analysis on the trigger word set and outputting a dependency structure tree; performing semantic role labeling based on the dependency structure tree to obtain semantic roles; classifying and normalizing the semantic roles based on the argument role template to form the standardized event structure; constructing logical, causal, and temporal semantic relationships between multiple arguments of the trigger word set based on the standardized event structure, and outputting the risk event candidate structure set.

[0012] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: extracting semantic segments from the candidate monitored entity list and obtaining display tags; introducing natural language prompts to guide the input of the display tags to predict trigger words, obtaining trigger words, and obtaining the start and end positions of the trigger words; obtaining the events corresponding to the trigger words, performing bidirectional mapping judgment on the trigger word events based on the technical and economic risk event knowledge system, and outputting the trigger word set and event type judgment results.

[0013] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: archiving the risk event candidate structure set according to the monitored entity dimension and the paragraph position dimension to construct a chapter-level event candidate index table; performing a fusion scoring and sorting based on three types of strategies according to the chapter-level event candidate index table to select the main risk theme event; performing paragraph-level association event aggregation based on three types of association strategies according to the main risk theme event to obtain an association event set; and constructing a master-slave event structure output based on the main risk theme event and the association event set to obtain the master-slave event structure.

[0014] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: the three types of strategies include title semantic guidance strategy, event influence assessment strategy, and position priority strategy; the three types of association strategies include entity overlap association, event causal chain association, and semantic similarity association.

[0015] Preferably, the method for extracting technical and economic risk events based on monitored entities further includes: constructing the event chain of the master-slave event structure according to the chain construction dimension, wherein the chain construction dimension includes a time sequence chain, an entity-dominated chain, and a causal logic chain; and performing standardized representation based on the event chain.

[0016] Secondly, this application also provides a technology and economic risk event extraction system based on monitored entities, used to execute a technology and economic risk event extraction method based on monitored entities as described in the first aspect, including: an entity recognition module, used to perform entity recognition on multi-source scientific and technological texts, extract technology and economic monitoring entities, and perform entity comparison and screening on the technology and economic monitoring entities based on a preset monitoring entity library, outputting a candidate monitoring entity list; a semantic fragment extraction module, used to extract semantic fragments using the candidate monitoring entity list as anchor points, and perform event trigger word recognition, and determine the event type based on trigger word matching rules; a semantic role labeling module, used to perform dependency parsing and semantic role labeling on the identified trigger word set, construct a standardized event structure, extract semantic relationships, and output a risk event candidate structure set; a master-slave event structure formation module, used to identify chapter-level master events and aggregate paragraph-level related events based on the risk event candidate structure set, forming a master-slave event structure; and a standardized output module, used to construct an event chain based on the master-slave event structure and perform standardized output.

[0017] The technical solution provided in this application has at least the following technical effects or advantages: by achieving the technical goal of constructing a structured risk event set and establishing a multi-dimensional event chain centered on the monitored entity, it achieves the technical effects of improving the accuracy of technical and economic security event identification, enhancing the depth of event understanding, and supporting risk decision support capabilities in multiple scenarios.

[0018] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a method for extracting techno-economic risk events based on monitored entities, as proposed in this application.

[0021] Figure 2This is a schematic diagram of the structure of a technical and economic risk event extraction system based on monitored entities, as proposed in this application.

[0022] Figure labeling: Entity recognition module 11, semantic fragment extraction module 12, semantic role labeling module 13, master-slave event structure formation module 14, standardized output module 15. Detailed Implementation

[0023] This application provides a method and system for extracting technological and economic risk events based on monitored entities. It addresses the technical problem in existing technologies where the lack of a professional event type system and a monitoring entity-driven extraction mechanism for the technological and economic security field leads to the inability to accurately extract event structures with industry characteristics and high-risk orientations. This further impacts the technical effectiveness and system reliability of applications such as intelligent identification, trend warning, and knowledge graph construction of key entity risk situations in the technological and economic field. The application achieves the technical goal of constructing a structured risk event set centered on the monitored entity and establishing a multi-dimensional event chain for intelligent extraction. This results in improved accuracy in identifying technological and economic security events, enhanced depth of event understanding, and support for risk decision-making assistance in multiple scenarios.

[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.

[0025] Example 1, please refer to the appendix. Figure 1 This application provides a method for extracting technical and economic risk events based on monitored entities, applied to a system for extracting technical and economic risk events based on monitored entities, specifically including the following steps:

[0026] Entity recognition is performed on multi-source scientific and technological texts to extract technical and economic monitoring entities. Based on a preset monitoring entity library, the technical and economic monitoring entities are compared and screened to output a candidate monitoring entity list.

[0027] Specifically, from multiple sources of science and technology-related text data, key entities with independent semantics are identified using natural language processing (NLP) technology. These multi-source science and technology texts may include news reports, scientific papers, or corporate announcements, characterized by significant structural differences and diverse information distribution. Technology and economic monitoring entities refer to specific objects with risk concern value in the intersection of technology and economics, such as company names, university research institutions, industrial projects, and key technological equipment; these may be the core carriers or sources of potential risk events. The identified text entities are then subjected to fuzzy matching and semantic correction with an existing standardized entity set to remove ambiguity or redundant information; for example, "Company X" and "Technology Co., Ltd." are grouped under the same standardized entity, "Company Name." The screening process further incorporates factors such as the entity's risk labels, historical event frequency, and occurrence frequency to filter the matching results, retaining the most representative monitoring objects. The final output is a candidate monitoring entity list, which represents the set of entities with strong certainty and high risk relevance in the current text. Each entity is also labeled with its location in the text and surrounding semantics, providing precise anchors for subsequent event identification. For example, from a technology news article about the chip supply chain, the system may identify multiple entities such as "Entity A", "Entity B", and "Entity C". After comparison and screening, only "Entity A" and "Entity B" are retained as candidate monitoring entities, thereby focusing on the objects that truly have technological and economic risk value.

[0028] Semantic fragments are extracted using the candidate monitoring entity list as anchor points, event trigger words are identified, and event types are determined based on trigger word matching rules.

[0029] Specifically, the location of candidate monitored entities in the text is used as the core reference point. Several words are extended before and after the entity's context to extract sentences or phrases with complete semantic expression, capturing clues to potential events. The length of the semantic fragment is generally set according to the syntactic structure or a fixed window size; for example, thirty Chinese characters are extended before and after each entity to extract contextual content. Then, event trigger word identification is performed, identifying keywords in the semantic fragment that indicate the occurrence or change of state of an event, such as "release" or "termination of cooperation," reflecting important nodes in the monitored entity's participation in technological and economic activities. Next, the identified trigger words are matched with a pre-constructed set of trigger word patterns. Each trigger word pattern corresponds to a specific event type; for example, "Event D" and "Event E" match R-type events, and "Event F" and "Event G" match S-type events. The rule comparison mechanism clearly defines the event type represented by the trigger word. The overall process forms a closed-loop chain: locating the semantic range of candidate entities, triggering event identification based on the semantic range, and determining the event category through trigger word rules. This makes event extraction both targeted and structurally clear. For example, when "a certain company" is identified as the monitoring entity and the semantic fragment "due to a certain measure" is extracted, "a certain measure" can be identified as the trigger word, and the event can be determined as "a certain measure type" according to the matching rules, thereby effectively realizing the typological classification of risk events.

[0030] Dependency parsing and semantic role labeling are performed on the identified set of trigger words to construct a standardized event structure, extract semantic relationships, and output a set of candidate risk event structures.

[0031] Specifically, a dependency graph is established for each trigger word in the sentence to clarify the modification, dominance, and connection relationships between words in the sentence, such as the dependency paths between verbs and subjects, objects, and time adverbs, in order to determine the subject and object associated with the trigger word. Semantic role labeling then assigns semantic functions to words in the sentence, such as identifying which word represents the initiator, executor, affected object, time of occurrence, and location of occurrence, thereby achieving semantic structuring of event elements. Constructing a standardized event structure involves encoding and organizing the trigger word and its related semantic roles according to a pre-defined event template, forming a unified event triple or quintuple expression to ensure the comparability of the event structure and its usability for subsequent computational processing. Semantic relationship extraction then further identifies the logical, causal, and temporal relationships between different arguments from the standardized structure. Finally, a set of candidate risk event structures is output, which is a preliminary complete set of risk event units extracted from the original text in a structured manner. Each event unit contains a core entity, triggering behavior, and argument information, which can be used for subsequent tasks such as main event identification, chain construction, and graph reasoning.

[0032] Based on the aforementioned risk event candidate structure set, chapter-level master events are identified, paragraph-level related events are aggregated, and a master-slave event structure is formed.

[0033] Specifically, multiple standardized event structures extracted from the text are used as the input set. Each event structure contains monitored entities, trigger words, event types, and corresponding argument information, reflecting risk information in the text from different perspectives. Identifying the main event at the text level involves analyzing the semantic weights, positional distribution, and importance scores of all candidate events within a complete text to select the key event that best represents the core theme of the entire text. This main event has the highest sentence vector aggregation value or closely aligns with the semantic intent implied in the article's title. Aggregating paragraph-level related events involves further identifying other events from various paragraphs that are consistent with the main event in terms of entity, similar event type, or causal relationship, and aggregating them as subordinate events of the main event. For example, if the main event is "Event M," the related events could be "Event N," "Event L," etc. Forming a master-subordinate event structure involves organizing the main event and multiple related events hierarchically to construct a structured risk expression framework with the main event at its core and subordinate events as supplements, thereby enhancing the logical readability and risk analysis capabilities between events.

[0034] An event chain is constructed based on the master-slave event structure, and the output is standardized.

[0035] Specifically, the structure, which includes a main event and multiple related subordinate events, serves as the input basis. This structure clearly defines the core hierarchical relationships between events, as well as their trigger words, entities, and argument information. A component event chain refers to connecting these events in a logical order to form a sequence structure with temporal continuity, causal correlation, or entity dominance. Event chain construction methods include chronological chains, which arrange events according to the order in which they occurred; entity-dominated chains, which connect all related events with a key entity as the core; and causal logic chains, which emphasize the cause-and-effect relationship between events, such as "Event O" leading to "Event P," which in turn triggers "Event Q." Standardized output refers to converting the constructed event chain into a unified structured expression format, often presented as triples, JSON objects, or edge-node forms in a knowledge graph, facilitating subsequent storage, retrieval, visualization, and reasoning tasks for the computer system. For example, a set of events can be encoded as "Company A – Event H – Time B" and "Event H – Leads to – Project Suspension," with the output being a two-layer graph structure of time chain + causal chain, supporting multilingual annotation and risk label mapping. Ultimately, this standardized event chain not only improves the semantic organization capabilities of the event extraction system.

[0036] Furthermore, this application also includes: acquiring historical technological and economic risk events and determining event type identifiers; listing trigger words based on the historical technological and economic risk events and obtaining a set of trigger word patterns; defining participating elements through the historical technological and economic risk events and obtaining argument role templates; and performing semantic relationship association based on the event type identifiers, the set of trigger word patterns, and the argument role templates to obtain semantic relationship rules and construct a knowledge system of technological and economic risk events.

[0037] Specifically, acquiring historical technological and economic risk events and identifying event type identifiers refers to extracting typical cases from past events related to technological and economic security risks that have occurred and been recorded, in order to summarize and categorize the type labels of various risk events.

[0038] After identifying the event type, trigger words are listed based on historical technological and economic risk events to obtain a trigger word pattern set. Trigger words refer to the core verbs, nouns, or phrases in multi-source scientific and technological texts that indicate the occurrence of an event, acting as event trigger signals within the text. The trigger word pattern set summarizes the common combinations of trigger words in context.

[0039] Next, by defining participating elements through historical technological and economic risk events, an argument role template is obtained. Participating elements refer to various key entities or information units, such as the time, place, involved parties, mode of action, and outcome of the event. After categorizing the participating elements, the argument role template can be established.

[0040] Finally, semantic relationship associations are performed based on event type identifiers, trigger word pattern sets, and argument role templates to obtain semantic relationship rules and construct a knowledge system for technological and economic risk events. Semantic relationship association refers to identifying the logical relationships between trigger words and arguments within the event structure, such as "a certain action is performed by a certain subject" or "a certain cause leads to a certain result," forming a stable event expression structure. By continuously summarizing semantic rules, a complete knowledge system with event types, trigger words, argument templates, and semantic logic is constructed, thus providing fundamental support for subsequent risk event identification, classification, and early warning.

[0041] Furthermore, this application also includes: dividing the multi-source scientific and technological text into sentence units according to semantic boundaries to obtain sentence-level segmentation; performing lexical analysis on the sentence-level segmentation to obtain entity clues; performing dependency tree structure analysis on the sentence-level segmentation to obtain sentence semantic relations; correcting and deduplicating the entity clues and sentence semantic relations, and combining them to obtain multi-dimensional text information.

[0042] Specifically, preliminary processing is performed on science and technology text data from various sources (such as science and technology news, research reports, etc.). Long texts are divided into multiple sentences with complete meaning based on linguistic sentence segmentation rules and semantic completeness, achieving sentence-level segmentation. Semantic boundaries not only refer to the physical separation at point markers but also include semantically representing units that express a complete event or fact. For example, a 300-word announcement might be divided into ten independent sentence-level units, each constituting a sentence-level segmentation unit, providing basic units for subsequent analysis.

[0043] Each sentence undergoes part-of-speech tagging, word segmentation, and stemming to identify words that may constitute important entities, thus obtaining entity clues. Lexical analysis helps identify nouns, verbs, adjectives, and terms composed of consecutive words, such as "Company A" or "key materials." In scientific and technical texts, proper nouns often consist of multiple words, making lexical analysis crucial for initial entity extraction. The resulting entity clues form the basis for subsequent entity recognition and matching.

[0044] Within sentences segmented at the sentence level, a grammatical dependency structure is constructed between words, representing subject-verb, verb-object, and attributive-adverbial-complement relationships in a tree-like manner to obtain semantic relationships within the sentence. Dependency structures can reveal deep semantic relationships within a sentence. For example, dependency analysis of "the company closes its factory" shows that "closes" is the core verb, "company" is the subject, and "factory" is the object, thus clarifying the action, the executor, and the affected object. Structural analysis provides grammatical support for subsequent trigger word location and event role labeling.

[0045] After completing lexical and dependency structure analysis, the identified entity candidates and their corresponding syntactic relations are cross-validated to eliminate duplicate entities or recognition errors. Simultaneously, the entity representation format is standardized to improve overall recognition quality. Error correction operations include removing spelling errors and syntactic analysis biases, while deduplication refers to unifying the representation of the same entity under different expressions. For example, "Company X Technology Co., Ltd." and "Company X" are merged into a single entity. Furthermore, information such as entities, syntactic relations, and contextual positions are integrated to obtain a structured text representation containing multi-dimensional linguistic features, which is used for subsequent entity matching and event extraction.

[0046] Furthermore, this application also includes: performing entity literal matching on the multidimensional text information according to the preset monitoring entity library to obtain fuzzy matching entities; performing entity expansion recognition on the multidimensional text information based on the fuzzy matching entities to obtain expanded matching entities; identifying the expanded matching entities in the multidimensional text information, and extracting context windows from before and after the text with the entity identifier as the center, and outputting a candidate monitoring entity list.

[0047] Specifically, the entity names in the previously constructed monitoring entity database are compared with the multidimensional text information to be processed. Based on the literal similarity of words in the text, possible corresponding entities are identified. The preset monitoring entity database contains entries such as enterprises, institutions, and projects in the technology and economic fields that have risk monitoring value. Literal matching uses techniques such as string similarity, edit distance, and spelling variations to compare the descriptions in the text with the entity database, thereby identifying similar but not necessarily identical names.

[0048] Based on fuzzy entity matching, and combined with syntactic structure, part-of-speech information, and contextual semantic features, entities that are not directly matched but have actual referential relationships can be further identified. Then, through named entity recognition models, contextual coreference resolution algorithms, and other technologies, supplementary objects such as pronouns, abbreviations, and implicit entities can be identified.

[0049] All identified expanded matching entities are located within the original text, and their specific locations are recorded. Then, for each entity, a fixed length of text content is expanded both forward and backward for subsequent contextual analysis and event recognition. The context window contains twenty to fifty words to preserve the semantic environment surrounding the entity, ensuring complete contextual information is available when determining the event. Finally, all entities with location information and context windows are organized into a candidate monitoring entity list, providing input for the next step of trigger word recognition and event type determination.

[0050] Furthermore, this application also includes: performing dependency syntactic analysis on the trigger word set and outputting a dependency structure tree; performing semantic role labeling based on the dependency structure tree to obtain semantic roles; classifying and normalizing the semantic roles based on the argument role template to form the standardized event structure; constructing logical, causal, and temporal semantic relationships between multiple arguments of the trigger word set based on the standardized event structure, and outputting the risk event candidate structure set.

[0051] Specifically, based on the identified trigger words, natural language processing techniques are used to perform syntactic structure analysis on the sentences containing the trigger words, determine the grammatical dependencies between the trigger words, and express them in a tree structure. Dependency parsing is a method used to represent the subordinate relationships between words in a sentence, and it can identify grammatical roles such as subject-verb relations, verb-object relations, and modifiers.

[0052] After clarifying the grammatical structure of the sentence, the next step is to identify the semantic role of each word in the event, such as who is the performer of the action, who is the receiver of the action, when it happened, where it happened, and what caused it. Semantic role labeling is a method of parsing sentences at the semantic level. Common roles include agent, patient, time, place, cause, and manner. Taking "The company stopped production due to a shortage of raw materials" as an example, "the company" is the agent, "the shortage of raw materials" is the cause, and "stopping production" is the action.

[0053] Using pre-defined event argument role templates, the identified semantic roles are categorized and standardized. Role templates are structural patterns summarized based on different event types, used to standardize the naming and scope of various semantic roles. Aligning semantic roles with the templates allows diverse expressions to be standardized into a unified structure; for example, "the organization withdrew its investment in April" can be standardized as "Subject = the organization; Action = withdrawal; Time = April".

[0054] After obtaining the standardized event structure, we further analyze the semantic connections between the roles in the event to construct a logical chain. Logical relationships include parallel relationships, conditional relationships, and inverse relationships. Causal relationships reveal the motivation and result of the event, while temporal relationships describe the sequence of events.

[0055] Furthermore, this application also includes: extracting semantic segments from the candidate monitoring entity list and obtaining display markers; introducing natural language prompts for guidance, inputting the display markers to predict trigger words, obtaining trigger words, and obtaining the start and end positions of the trigger words; obtaining the events corresponding to the trigger words, performing bidirectional mapping judgment on the trigger word events based on the technical and economic risk event knowledge system, and outputting the trigger word set and event type judgment results.

[0056] Specifically, based on the identified candidate entities, the contextual semantic content of each entity's text location is extracted to form semantic fragments. Simultaneously, markers are generated for these fragments to provide clues. Each semantic fragment extends forward and backward from the identified entity, representing its semantic environment. The markers, on the other hand, are visual or model-based annotations of the entity, such as using special symbols, colors, or cue words to indicate its presence. This facilitates subsequent model recognition of key information within the context and focuses on the entity's relevant semantic regions.

[0057] By adding natural language prompts to a pre-trained language model, the model is guided to identify event-related verbs or expressions, i.e., trigger words, that are related to the monitored entity. Prompts can be provided by adding question templates such as "What happened to this company?" or "Does the following text contain any risk events?", thereby enhancing the model's ability to perceive event triggers. Once the pre-trained language model identifies a trigger word, it marks its specific position in the original semantic fragment, including the starting and ending character or word positions, ensuring that the trigger word can be used to locate and further construct the event structure.

[0058] The identified trigger words are compared with pre-built event type identifiers to determine their specific event type and verify the logical consistency of the mapping. Two-way mapping means mapping from trigger words to existing event types in the technical and economic risk event knowledge system (such as "change of controlling stake"), and also checking for typical usages of the trigger word within the event types to ensure the rationality and accuracy of the match. The final output includes a complete set of trigger words and their associated risk event type labels with the monitored entities, providing semantic anchors for subsequent argument extraction and event structure generation.

[0059] Furthermore, this application also includes: archiving the risk event candidate structure set according to the monitoring entity dimension and paragraph position dimension to construct a chapter-level event candidate index table; performing fusion scoring and sorting based on three types of strategies according to the chapter-level event candidate index table to select the risk theme master event; performing paragraph-level association event aggregation based on three types of association strategies according to the risk theme master event to obtain an association event set; and constructing a master-slave event structure output based on the risk theme master event and the association event set to obtain the master-slave event structure.

[0060] Specifically, the identified candidate risk events are organized and categorized according to the monitored entities and their paragraph positions in the text, and an index table is created. The monitored entity dimension means grouping events related to the same entity, such as a company, technology, or institution, while the paragraph position dimension refers to clearly marking the distribution of events in the entire text, such as appearing in the 2nd or 5th paragraph. This two-dimensional archiving method helps to grasp the context and entity relevance of events, providing an organizational basis for subsequent event focusing and screening.

[0061] The candidate events archived in the index table are comprehensively scored using three different strategies, and the most representative main event is selected based on the total score. The three strategies include: a title semantic guidance strategy, which considers whether the event aligns with the meaning of the title; an event influence assessment strategy, which measures the event's frequency, entity weight, and depth of reasoning within the text; and a position priority strategy, which considers whether the event is located near the beginning or end of the paragraph. The event with the highest combined score is identified as the main event of the chapter.

[0062] Around the main event, other events with semantic, logical, or entity-related connections are identified throughout the text and aggregated into a set. Three types of association strategies include entity overlap association, where the main event and other events share the same monitored entity; event causal chain association, where there is a clear causal or sequential influence relationship; and semantic similarity association, where different events are semantically similar or mutually supportive. This allows related events scattered across different paragraphs to be organized, enriching and supplementing the context behind the main event.

[0063] By taking the main event as the core and other related events as auxiliary content, a structured representation is formed according to a master-slave hierarchy. The master-slave structure not only reveals the full picture of the main event but also reflects its extension in a larger semantic context.

[0064] Furthermore, this application also includes: the three types of strategies include title semantic guidance strategy, event influence assessment strategy and position priority strategy; the three types of association strategies include entity overlap association, event causal chain association and semantic similarity association.

[0065] Specifically, the three strategies include title semantic guidance strategy, event influence assessment strategy, and position priority strategy. The title semantic guidance strategy refers to using the central meaning contained in the title of the article to guide the determination of the main theme of the event. The specific method is to match the semantic similarity of the candidate events with the title by sentence vectors or keywords. If an event is highly related to the title, it can be determined that the event is more dominant.

[0066] Event impact assessment strategy refers to evaluating candidate events based on their performance intensity and risk weight within the text. Factors considered in this strategy include whether the entities involved are key monitoring targets, whether the event is mentioned multiple times, and whether the event has a potential chain reaction effect. For example, an event about a "disruption of core material supply" that appears more than three times in the text and involves multiple industry segments is considered to have a significant impact.

[0067] The position-priority strategy refers to incorporating the location of events within the text's structure into the ordering criteria. Events located in the first and last paragraphs are generally considered more important because they often serve to guide and summarize the article. For example, in a six-paragraph press release, if an event is emphasized in both the first and last paragraphs, its priority is significantly higher than an event mentioned only in the middle paragraphs.

[0068] On the other hand, the three types of association strategies include entity overlap association, event causal chain association, and semantic similarity association. Entity overlap association refers to the presence of the same or highly related monitoring entities as the main event in the candidate event. For example, if a certain enterprise, research institution, or technology field in the main event reappears in other events, it can be regarded as having an entity association with the main event.

[0069] A causal chain of events refers to a clear causal relationship or logical sequence between two events. For example, if the main event is "Event W" and a related event is "the high-end equipment project is forced to terminate", then the latter is the result of the former. The two constitute a causal relationship and can establish a chain of events.

[0070] Semantic similarity association refers to events that have a certain degree of semantic similarity in terms of content, domain affiliation, or potential meaning. For example, if the main event is "supply chain disruption" and the event described in a certain paragraph is "unstable supply of key raw materials", although the two are not completely consistent, they are closely related in terms of the semantics of industrial risk and can be regarded as semantically similar association events.

[0071] Furthermore, this application also includes: constructing the event chain of the master-slave event structure according to the chain construction dimension, wherein the chain construction dimension includes a time sequence chain, an entity-dominated chain, and a causal logic chain; and performing a standardized representation based on the event chain.

[0072] Specifically, based on the various inherent relationships between a main event and its multiple subordinate events, these are connected into an ordered event chain through pre-defined dimensional standards. The chain construction dimensions refer to three logical methods for organizing the event chain: chronological chain, entity-driven chain, and causal chain. The chronological chain arranges events according to their chronological order in the original text; for example, an earlier "Event X" and a later "Event Y" form a sequential relationship. The entity-driven chain connects events based on the same or highly related monitoring entities involved; for example, two events, "Event I" and "Event J," involving the same research institution can form an entity-driven event development chain. The causal chain establishes connections based on causal reasoning between events; a clear causal logic exists between events, allowing for the construction of a complete causal chain.

[0073] After constructing the event chain, it needs to be standardized. Standardization refers to encoding the event chain in the original text using a unified format, structure, and semantic tags to facilitate storage, retrieval, and inference analysis. Standardized representations can use event triples, graph structures, or hierarchical JSON formats, where each event node includes a trigger word, time, argument information, and its relationship type with preceding and following nodes. For example, in a standardized chain, event A is connected to event B through a causal relationship, and event B is then connected to event C through a chronological relationship, forming a well-structured event chain representation that can be processed by algorithms.

[0074] In summary, the technical and economic risk event extraction method based on monitored entities provided in this application has the following technical effects: by realizing the technical goal of constructing a structured risk event set and establishing a multi-dimensional event chain with the monitored entity as the center, it achieves the technical effects of improving the accuracy of technical and economic security event identification, enhancing the depth of event understanding, and supporting risk decision-making assistance capabilities in multiple scenarios.

[0075] Example 2: Based on the same inventive concept as the method for extracting technical and economic risk events based on monitored entities in the foregoing examples, this application also provides a system for extracting technical and economic risk events based on monitored entities. Please refer to the appendix. Figure 2 The system includes: an entity recognition module 11, used to perform entity recognition on multi-source scientific and technological texts, extract technical and economic monitoring entities, and perform entity comparison and screening on the technical and economic monitoring entities based on a preset monitoring entity library, and output a candidate monitoring entity list; a semantic fragment extraction module 12, used to extract semantic fragments using the candidate monitoring entity list as anchor points, and perform event trigger word recognition, and determine the event type based on trigger word matching rules; a semantic role labeling module 13, used to perform dependency parsing and semantic role labeling on the identified trigger word set, construct a standardized event structure, extract semantic relationships, and output a risk event candidate structure set; a master-slave event structure formation module 14, used to identify chapter-level master events and aggregate paragraph-level related events based on the risk event candidate structure set, forming a master-slave event structure; and a standardized output module 15, used to build an event chain based on the master-slave event structure and perform standardized output.

[0076] Furthermore, the aforementioned technology and economic risk event extraction system based on monitored entities is also used for: acquiring historical technology and economic risk events and determining event type identifiers; listing trigger words based on the historical technology and economic risk events and obtaining a set of trigger word patterns; defining participating elements through the historical technology and economic risk events and obtaining argument role templates; and performing semantic relationship association based on the event type identifiers, the set of trigger word patterns, and the argument role templates to obtain semantic relationship rules and construct a knowledge system for technology and economic risk events.

[0077] Furthermore, the technology and economic risk event extraction system based on monitored entities is also used for: dividing the multi-source scientific and technological text into sentence units according to semantic boundaries to obtain sentence-level segmentation; performing lexical analysis on the sentence-level segmentation to obtain entity clues; performing dependency tree structure analysis on the sentence-level segmentation to obtain sentence semantic relations; correcting and deduplicating the entity clues and sentence semantic relations, and combining them to obtain multi-dimensional text information.

[0078] Furthermore, the aforementioned system for extracting technical and economic risk events based on monitored entities is also used for: performing literal matching of the multidimensional text information according to the preset monitoring entity library to obtain fuzzy matching entities; performing entity expansion recognition based on the fuzzy matching entities on the multidimensional text information to obtain expanded matching entities; identifying the expanded matching entities in the multidimensional text information, and extracting context windows from before and after the text with the entity identifier as the center, and outputting a list of candidate monitored entities.

[0079] Furthermore, the aforementioned system for extracting technical and economic risk events based on monitored entities is also used for: performing dependency syntactic analysis on the set of trigger words and outputting a dependency structure tree; performing semantic role labeling based on the dependency structure tree to obtain semantic roles; classifying and normalizing the semantic roles based on the argument role template to form the standardized event structure; constructing logical, causal, and temporal semantic relationships between multiple arguments of the set of trigger words based on the standardized event structure, and outputting the risk event candidate structure set.

[0080] Furthermore, the aforementioned technology and economic risk event extraction system based on monitored entities is also used for: extracting semantic segments from the candidate monitored entity list and obtaining display tags; introducing natural language prompts for guidance, inputting the display tags to predict trigger words, obtaining trigger words, and obtaining the start and end positions of the trigger words; obtaining the events corresponding to the trigger words, performing bidirectional mapping judgment on the trigger word events based on the technology and economic risk event knowledge system, and outputting the trigger word set and event type judgment results.

[0081] Furthermore, the aforementioned system for extracting technical and economic risk events based on monitored entities is also used for: archiving the candidate risk event structure set according to the monitored entity dimension and the paragraph position dimension to construct a chapter-level event candidate index table; performing a fusion scoring and sorting based on three types of strategies according to the chapter-level event candidate index table to select the main risk theme event; performing paragraph-level association event aggregation based on three types of association strategies according to the main risk theme event to obtain an association event set; and constructing a master-slave event structure output based on the main risk theme event and the association event set to obtain the master-slave event structure.

[0082] Furthermore, the aforementioned technical and economic risk event extraction system based on monitored entities is also used for: the three types of strategies include title semantic guidance strategy, event influence assessment strategy, and position priority strategy; the three types of association strategies include entity overlap association, event causal chain association, and semantic similarity association.

[0083] Furthermore, the aforementioned technology and economic risk event extraction system based on monitored entities is also used to: construct the event chain of the master-slave event structure according to the chain construction dimension, wherein the chain construction dimension includes a time sequence chain, an entity-dominated chain, and a causal logic chain; and perform standardized representation based on the event chain.

[0084] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The method and specific example for extracting technical and economic risk events based on monitored entities in the foregoing embodiment one are also applicable to the system for extracting technical and economic risk events based on monitored entities in this embodiment. Through the foregoing detailed description of the method for extracting technical and economic risk events based on monitored entities, those skilled in the art can clearly understand the system for extracting technical and economic risk events based on monitored entities in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0085] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0086] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for extracting techno-economic risk events based on monitored entities, characterized in that, include: Entity recognition is performed on multi-source scientific and technological texts to extract technical and economic monitoring entities. Based on a preset monitoring entity library, the technical and economic monitoring entities are compared and screened to output a candidate monitoring entity list. Semantic fragments are extracted using the candidate monitoring entity list as anchor points, event trigger words are identified, and event types are determined based on trigger word matching rules. Dependency parsing and semantic role labeling are performed on the identified set of trigger words to construct a standardized event structure, extract semantic relationships, and output a set of candidate risk event structures. Based on the aforementioned risk event candidate structure set, chapter-level master events are identified, paragraph-level related events are aggregated, and a master-slave event structure is formed. An event chain is constructed based on the master-slave event structure, and the output is standardized.

2. The method for extracting techno-economic risk events based on monitored entities as described in claim 1, characterized in that, The process of performing entity recognition on multi-source scientific and technological texts, extracting technical and economic monitoring entities, and comparing and filtering these entities based on a pre-set monitoring entity database to output a candidate monitoring entity list includes the following steps: Acquire historical technological and economic risk events and identify event type identifiers; Based on the historical technological and economic risk events, list the trigger words and obtain the trigger word pattern set; By defining the participating elements through the aforementioned historical technological and economic risk events, the argument role template is obtained; Based on the event type identifier, the trigger word pattern set, and the argument role template, semantic relationship associations are performed to obtain semantic relationship rules and construct a knowledge system for technological and economic risk events.

3. The method for extracting techno-economic risk events based on monitored entities as described in claim 2, characterized in that, The step of performing entity recognition on multi-source scientific and technological texts, extracting technical and economic monitoring entities, and comparing and filtering these entities based on a preset monitoring entity library to output a candidate monitoring entity list, also includes the following: The multi-source scientific and technological texts are divided into sentence units according to semantic boundaries to obtain sentence-level segmentation; Lexical analysis is performed on the sentence-level segmentation to obtain entity clues; Dependency tree structure analysis is performed on the sentence-level segmentation to obtain sentence semantic relations; The entity clues and sentence semantic relationships are corrected and deduplicated, and then combined to obtain multidimensional text information.

4. The method for extracting techno-economic risk events based on monitored entities as described in claim 3, characterized in that, The process involves entity recognition of multi-source scientific and technological texts, extraction of technical and economic monitoring entities, and comparison and screening of these entities based on a pre-defined monitoring entity database to output a candidate monitoring entity list, including: Based on the preset monitoring entity database, the multidimensional text information is subjected to entity literal matching to obtain fuzzy matching entities; The multidimensional text information is subjected to entity expansion recognition based on the fuzzy matching entity to obtain the expanded matching entity; The expanded matching entity is identified in the multidimensional text information, and a context window is extracted from the text before and after the entity identifier, and a candidate monitoring entity list is output.

5. The method for extracting techno-economic risk events based on monitored entities as described in claim 2, characterized in that, The process involves performing dependency parsing and semantic role labeling on the identified trigger word set, constructing a standardized event structure, extracting semantic relationships, and outputting a candidate structure set for risk events, including: Perform dependency parsing on the set of trigger words and output the dependency structure tree; Semantic roles are obtained by performing semantic role labeling based on the dependency structure tree; Based on the argument role template, the semantic roles are classified and normalized to form the standardized event structure; Based on the standardized event structure, construct the logical, causal, and temporal semantic relationships between multiple arguments of the trigger word set, and output the risk event candidate structure set.

6. The method for extracting techno-economic risk events based on monitored entities as described in claim 1, characterized in that, The step of extracting semantic segments using the candidate monitoring entity list as anchor points, identifying event trigger words, and determining event types based on trigger word matching rules includes: Semantic fragment extraction is performed on the candidate monitoring entity list, and display tags are obtained; Natural language prompts are introduced to guide the input of the displayed markers to predict the trigger words, obtain the trigger words, and acquire the start and end positions of the trigger words; Obtain the event corresponding to the trigger word, perform bidirectional mapping judgment on the trigger word event based on the aforementioned technical and economic risk event knowledge system, and output the trigger word set and event type judgment result.

7. The method for extracting techno-economic risk events based on monitored entities as described in claim 1, characterized in that, The process of identifying chapter-level master events and aggregating paragraph-level related events based on the risk event candidate structure set to form a master-detail event structure includes: The risk event candidate structure set is archived according to the monitoring entity dimension and paragraph position dimension to construct a chapter-level event candidate index table; Based on the chapter-level event candidate index table, a fusion scoring and ranking based on three strategies is performed to select the main risk theme event; Based on the main event of the risk theme, paragraph-level related events are aggregated using three types of association strategies to obtain a set of related events; The master-slave event structure is constructed by outputting the master-slave event structure based on the master event of the risk topic and the set of related events.

8. The method for extracting techno-economic risk events based on monitored entities as described in claim 7, characterized in that, The three types of strategies include title semantic guidance strategy, event influence assessment strategy, and position priority strategy; the three types of association strategies include entity overlap association, event causal chain association, and semantic similarity association.

9. The method for extracting techno-economic risk events based on monitored entities as described in claim 1, characterized in that, The process of constructing an event chain based on the master-slave event structure and performing standardized output includes: The event chain is constructed based on the chain construction dimension of the master-slave event structure, wherein the chain construction dimension includes a time sequence chain, an entity-dominated chain, and a causal logic chain; A standardized representation is performed based on the event chain.

10. A system for extracting techno-economic risk events based on monitored entities, characterized in that, The steps for implementing the method for extracting techno-economic risk events based on monitored entities as described in any one of claims 1 to 9 include: The entity recognition module is used to perform entity recognition on multi-source scientific and technological texts, extract technical and economic monitoring entities, and perform entity comparison and screening on the technical and economic monitoring entities based on a preset monitoring entity library, and output a candidate monitoring entity list. The semantic fragment extraction module is used to extract semantic fragments using the candidate monitoring entity list as anchor points, identify event trigger words, and determine the event type based on the trigger word matching rules. The semantic role labeling module is used to perform dependency parsing and semantic role labeling on the identified set of trigger words, construct a standardized event structure, extract semantic relationships, and output a set of candidate risk event structures. The master-slave event structure formation module is used to identify chapter-level master events and aggregate paragraph-level related events to form a master-slave event structure based on the risk event candidate structure set. The standardized output module is used to construct an event chain based on the master-slave event structure and perform standardized output.

Citation Information

Patent Citations

  • External risk event extraction method and device

    CN114528382A

  • Framework semantic mapping and type perception-based chapter event extraction method and system

    CN115168541A

  • Method, system and computer product for analyzing business risk using event information extracted from natural language sources

    US20050071217A1

Cited By

  • Enterprise risk identification and decision support method and system fused with natural language processing

    CN121810042A