Key event discrimination method and system based on hierarchical processing architecture
By building a hierarchical processing architecture and combining it with government affairs-specific semantic extraction and improved clustering methods, the problems of low efficiency and low accuracy in identifying key events in the government affairs field have been solved, efficient and accurate screening of key events has been achieved, and real-time response to urban governance has been supported.
Patent Information
- Application Number
- CN202510692617.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing technologies lack systematic solutions in the government affairs field, with low recognition efficiency and accuracy, making it difficult to screen key events efficiently and accurately, especially when faced with multi-channel and diverse event data, resulting in duplicate identification and misleading decisions.
A method based on a hierarchical processing architecture is adopted, including data preprocessing, semantic regularized agglomerative hierarchical clustering, and intra-cluster entity keyword extraction and comparison mechanism, to construct a three-level processing architecture of "classification-clustering-discrimination". Through government-specific semantic extraction, semantic regularized agglomerative hierarchical clustering, and intra-cluster entity keyword extraction technology, a full-process closed-loop processing from original text to key event judgment is achieved.
It significantly improves the system's processing efficiency and recognition accuracy, and can achieve real-time response and accurate duplication detection of large-scale event data in urban governance scenarios, reduce computational complexity, and improve recognition accuracy and system adaptability.
Smart Images

Figure CN120705313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a key event identification method and system based on a hierarchical processing architecture. Background Art
[0002] With the continuous development of the economy and society, urbanization rates are constantly increasing, and individual interests and demands are becoming more diverse. The channels and avenues for reporting issues across various sectors of society are also constantly expanding. As the main actors in social governance, relevant administrative departments receive a vast number of diverse inquiries daily. These departments need to screen and flag key incidents, prioritizing their handling or attention. For example, incidents such as water outages and heating pipe bursts have a wide impact, involving different groups within the same area reporting the same incident. Sometimes, the same incident is reported by the same person through different channels. If the issue is not addressed promptly or inappropriately, it can trigger mass incidents, seriously impacting social security and stability. Given the diverse reporting channels and the sheer volume of reports, the current solution still relies on manual screening or keyword matching. This is inaccurate, time-consuming, and labor-intensive, and can easily lead to omissions or misjudgment, resulting in the loss of important information. This can lead to decision-making errors and severely impact the effectiveness of grassroots governance. Therefore, it is imperative to leverage the intelligent capabilities of large-scale models to uniformly understand, analyze, screen, and flag incidents. This can provide early warnings based on the scope and severity of the incidents, improve work efficiency and screening accuracy, assist decision-making, and mitigate potential social security risks.
[0003] In practical applications, key event identification technology primarily focuses on accurately screening scenarios where multiple reports of a single event occur. This technology presents a pressing technical challenge in urban governance. The diverse and unpredictable nature of event types, the lack of standardized data formats across multiple channels, the potential for misjudgment caused by templated text, and the pressure of processing massive amounts of data all contribute to the difficulty of identification. If the system fails to accurately identify and consolidate key events, this can easily lead to duplicate statistical results, wasted resources, and misleading decision-making.
[0004] Specifically, technical issues in key event screening include but are not limited to:
[0005] 1. The amount of data reported daily is huge, resulting in low processing efficiency;
[0006] 2. Event sources are diverse and there is a lack of universal solutions;
[0007] 3. Event types and content descriptions are complex and changeable, which affects recognition accuracy.
[0008] Therefore, how to efficiently, universally and accurately screen key events among reported events has become a key technical challenge to improve system processing capabilities, decision support value and actual application effects.
[0009] In existing technologies, there are three main solutions for key event screening scenarios:
[0010] 1. Rule-based manual identification: This method uses preset keyword matching algorithms, geolocation matching rules, and named entity recognition technology to filter and identify events using keywords. This method typically involves professionals developing a series of criteria, such as including specific keywords in event descriptions, occurrences in the same or nearby areas, and time intervals within a specific range, to identify potential key events.
[0011] 2. Large language model screening method: Leveraging the powerful semantic understanding and generation capabilities of pre-trained large language models (such as Qwen), we can directly perform in-depth semantic extraction and analysis on all reported content. Specifically, all daily reported event content is input into the large language model, and its generalization capabilities are used to analyze the event content and identify key events that are repeatedly reported. Event content can also be screened and analyzed in multiple batches.
[0012] 3. Deep Learning Screening Methods: In recent years, semantic representation techniques based on deep learning have become a mainstream approach in natural language processing. These techniques effectively capture the semantic relationships between words and sentences by mapping text into a high-dimensional vector space. Specifically, a pre-trained embedding model is used to convert the content of reported events into a domain-specific, high-dimensional vector representation, thereby better capturing the deep semantic features of the text. Furthermore, by calculating the distance between vectors, the semantic similarity between texts can be quantified, and this semantic similarity can be used to match two texts to determine whether they meet the criteria for key events.
[0013] However, the main defects of the existing technology include but are not limited to the following aspects:
[0014] 1. Lack of systematic solutions: Existing technologies such as deep learning and large language model screening methods mostly use single technical means and fail to effectively integrate multiple methods such as classification, clustering, and semantic understanding. As a result, it is difficult to achieve a good balance between processing efficiency, recognition accuracy, and versatility, and cannot fully cope with complex and diverse data reporting scenarios.
[0015] 2. Inefficient recognition: For example, large language model screening methods directly use the model to perform pairwise comparisons across all event data, exponentially increasing computational complexity. This significantly reduces processing efficiency and increases response time when faced with the massive amounts of event data routinely collected by relevant departments, making it difficult to meet real-time or near-real-time processing requirements.
[0016] 3. Low recognition accuracy: Existing technologies generally lack effective preprocessing mechanisms tailored to the specific characteristics of government text. Reported data comes from diverse sources and formats, limiting model generalization. The semantic extraction process easily introduces a large amount of redundant or meaningless information, affecting the subsequent calculation of semantic similarity and reducing the accuracy of key event recognition.
[0017] To sum up, in the process of identifying key events in the government affairs field, existing technologies generally have key technical shortcomings such as a single system architecture, low recognition efficiency and low recognition accuracy, making it difficult to achieve efficient, accurate and universal screening of key events. Summary of the Invention
[0018] The purpose of the present invention is to provide a method and system for distinguishing key events based on a hierarchical processing architecture, aiming to solve the above-mentioned problems in the prior art.
[0019] An embodiment of the present invention provides a method for identifying key events based on a hierarchical processing architecture, comprising:
[0020] Acquire reported events from different channels, perform data preprocessing on the reported events, obtain processed event texts, and classify the event texts;
[0021] The classified texts are clustered using semantic regularized agglomerative hierarchical clustering method;
[0022] According to the clustering results, the repeated events in the text are identified by using the entity keyword extraction and comparison mechanism within the cluster to obtain the key event identification results.
[0023] An embodiment of the present invention provides a key event identification system based on a hierarchical processing architecture, comprising:
[0024] A classification module is used to obtain reported events from different channels, perform data preprocessing on the reported events, obtain processed event texts, and classify the event texts;
[0025] Clustering module, used to cluster the classified text using semantic regularized agglomerative hierarchical clustering method;
[0026] The discrimination module is used to discriminate repeated events in the text based on the clustering results using the entity keyword extraction and comparison mechanism within the cluster to obtain the key event discrimination results.
[0027] An embodiment of the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the above-mentioned method for identifying key events based on a hierarchical processing architecture are implemented.
[0028] An embodiment of the present invention further provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by a processor, the steps of the above-mentioned method for distinguishing key events based on a hierarchical processing architecture are implemented.
[0029] The use of the embodiments of the present invention can include the following beneficial effects: The embodiments of the present invention propose a key event identification method based on a hierarchical processing architecture, construct a hierarchical processing architecture of "classification-clustering-discrimination", and form a complete and systematic key event identification solution. This solution integrates a semantic extraction method dedicated to the government affairs field, a semantic regularized agglomerative hierarchical clustering method, and an intra-cluster entity keyword extraction technology. Through the multi-stage collaborative processing of semantic vectorization, clustering optimization, and key information verification, it significantly improves the system's processing efficiency and recognition accuracy, effectively meeting the actual needs of real-time response and accurate duplication judgment for large-scale event data in urban governance scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 is a flow chart of a method for distinguishing key events based on a hierarchical processing architecture according to an embodiment of the present invention;
[0032] Figure 2 is a flowchart of identifying key events according to an embodiment of the present invention;
[0033] Figure 3 This is an overall architecture diagram of the three-level processing framework of an embodiment of the present invention;
[0034] Figure 4 This is a flowchart of semantic extraction dedicated to the government affairs field according to an embodiment of the present invention;
[0035] Figure 5 2 is a schematic diagram of a key event identification system based on a hierarchical processing architecture according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.
[0037] Method Example
[0038] According to an embodiment of the present invention, a method for distinguishing key events based on a hierarchical processing architecture is provided. Figure 1 FIG. 1 is a flow chart of a method for distinguishing key events based on a hierarchical processing architecture according to an embodiment of the present invention. Figure 1 As shown, the key event identification method based on the hierarchical processing architecture according to an embodiment of the present invention specifically includes:
[0039] Step S101, obtaining reported events from different channels, performing data preprocessing on the reported events to obtain processed event texts, and classifying the event texts, specifically includes:
[0040] Perform data cleaning and filtering, text standardization, domain-specific processing, and validity verification on the reported events;
[0041] The data cleaning and filtering includes filtering empty text, removing non-string content and clearing interference symbols;
[0042] The text standardization processing includes constructing corresponding regular expressions according to different channels, extracting key content from the reported events through the corresponding regular expressions, and performing word segmentation processing on the text using a word segmentation tool;
[0043] The domain-specific processing is to use a pre-built stop word list specific to the government affairs field to screen the text and remove formatted words, single words and pure numeric fields that have no actual semantic meaning;
[0044] The validity verification is to check whether the processed text record is empty, and if so, mark and remove the text record;
[0045] Preliminarily classifying the event text according to event categories using a large language model; wherein the event categories include natural disasters, fire safety, traffic safety, public safety, social security, and conflicts and disputes;
[0046] Step S102: clustering the classified texts using a semantic regularized agglomerative hierarchical clustering method, specifically including:
[0047] A semantic embedding model is used to encode the elements within each type of text, generating a corresponding semantic embedding vector. Based on the semantic embedding vector, the cosine distance between each pair of texts is calculated using a cosine similarity function. A symmetric cosine distance matrix is constructed, and the cosine distance matrix is regularized. The regularized distance matrix is clustered using the average linkage method. After clustering, the clustering results are mapped back to the original data structure, and a corresponding cluster number is assigned to each valid record.
[0048] Step S103, based on the clustering results, uses the entity keyword extraction and comparison mechanism within the cluster to identify repeated events in the text, and obtains the key event identification results, which specifically include:
[0049] The clustering results are traversed, and isolated clusters containing only a single record in the clustering results are eliminated. Clusters containing multiple records are judged for duplicate content. If there are completely identical text fields, they are directly judged as duplicate events. If there are no completely identical text fields, the first k high-weight entity keywords of each record are extracted to obtain a keyword set for each record. The number of common keywords and the uniqueness ratio in the cluster are calculated based on the keyword set. If the number of common keywords is less than the preset minimum common keyword threshold and the uniqueness ratio is greater than the preset maximum uniqueness ratio upper limit, the cluster is eliminated and the remaining valid clusters are retained. The remaining valid clusters are input into the large language model for judgment to obtain the final key event judgment result.
[0050] The following is a specific example of a method for distinguishing key events based on a hierarchical processing architecture according to an embodiment of the present invention. Figure 2 As shown, the above technical solution of the embodiment of the present invention is described in detail.
[0051] To address the core issues of existing technologies, such as the difficulty in balancing accuracy and computational efficiency, and the limited system adaptability and generalization capabilities, the present invention proposes an innovative technical solution. This solution aims to build a new processing mechanism to achieve collaborative improvements at multiple levels, from data semantic understanding, feature optimization and extraction to intelligent clustering analysis. This solution significantly reduces computational overhead while improving judgment accuracy, thereby achieving more efficient, accurate, and transferable technical application effects. This solution specifically includes:
[0052] 1. Construct a three-level hierarchical architecture of "classification-clustering-discrimination" to form a structured key event screening system: The embodiment of the present invention proposes a key event identification method based on a hierarchical processing architecture, which is used to efficiently identify key events that are repeatedly reported through different channels in the field of urban governance. This method relies on the natural language rules that "the same event belongs to the same category" and "similar events have similar representations in the semantic space" to construct a hierarchical processing architecture consisting of three stages of classification, clustering and discrimination to form a systematic and scalable key event screening solution. By integrating special semantic extraction technology in the government field, semantic regularization agglomerative hierarchical clustering methods, and intra-cluster entity keyword extraction and comparison mechanisms, the embodiment of the present invention realizes the full-process closed-loop processing from raw text preprocessing, semantic vector representation, dynamic clustering analysis to the final key event determination, which significantly improves the accuracy and efficiency of the screening system.
[0053] 2. Semantic cleaning and normalization of government texts to improve effective semantic representation capabilities: The embodiment of the present invention proposes a text semantic extraction method for government scenarios, which aims to improve the accuracy of understanding event texts and the reliability of repeatability judgments. This method combines linguistic laws with government data characteristics to construct regular expressions and domain-specific stop word lists. On the basis of general stop words, it expands the common formatted expressions and redundant words without actual semantics in the reported text, significantly improving the effective proportion of semantic information after vectorization, and then extracts specific description content through different regular expressions according to the reporting channel. The preprocessing process includes: empty text filtering, interference symbol removal, word segmentation, elimination of government stop words, single-word words and pure digital content and other key steps to ensure that the final text only retains keywords with actual semantics and improves recognition accuracy.
[0054] 3. Introducing a semantic regularization mechanism to improve clustering efficiency and quality: The embodiment of the present invention makes targeted improvements to the traditional hierarchical clustering algorithm and proposes a semantic regularized hierarchical clustering method (SRHC). This method standardizes the semantic similarity matrix so that the data distribution approaches the standard normal distribution, thereby optimizing the similarity calculation and cluster merging strategy in the clustering process. Based on the principles of mathematical statistics, the algorithm significantly improves the computational efficiency while maintaining the stability and discrimination ability of the clustering results.
[0055] 4. Introduce a mechanism for extracting and comparing entity keywords within a cluster to improve the accuracy of judgment: The embodiment of the present invention proposes a technology for extracting entity keywords within a cluster. After completing the cluster analysis, the text content within each event cluster is deeply semantically analyzed to accurately identify and extract high-weight keywords related to key entities (such as places, names, etc.), and the top five keywords with the highest weight are selected for comparison. The entity keywords are then compared based on the uniqueness mechanism and the number of common keywords to screen out valid elements within the cluster. This mechanism not only formally ensures that the event cluster meets the judgment criteria of a "key event", but also implements strict verification at the content detail level, significantly improving the accuracy of the screening results.
[0056] The embodiment of the present invention is aimed at the key event screening scenario in the urban governance scenario, and proposes a key event identification method based on a hierarchical processing architecture. This method integrates the powerful semantic understanding ability of the large language model with the dynamic hierarchical clustering strategy driven by improved semantic embedding, and constructs a three-level processing framework of "classification-clustering-discrimination" to achieve high efficiency and accuracy in key event screening. Figure 3 The specific steps are as follows:
[0057] 1. Text semantic extraction methods for government affairs scenarios, such as Figure 4 As shown below: A standardized cleaning process is implemented on the original reported text, filtering out empty text, non-string content, and noise symbols. A word segmentation tool is used to segment the text, and regular expressions are used to extract information reported from different channels. Furthermore, a stop word list specific to the government sector is introduced to remove formatted words, single words, and purely numeric fields without actual meaning, ensuring that the words retained ultimately have actual semantic information. If a record is empty after cleaning, it is marked as invalid and removed.
[0058] 2. Intelligent classification based on a large language model categorizes reported content. Specifically, the system connects to a database to obtain the day's reported event data, and uses the large language model to intelligently classify and annotate it using prompt words and a structured output framework. The classification system covers six core categories: natural disasters, fire safety, traffic safety, public safety, social security, and conflicts and disputes. This process leverages the large model's deep understanding of the semantic boundaries of classification labels to accurately categorize event content, effectively narrow the scope of subsequent processing, and establish reasonable semantic boundaries for cluster analysis.
[0059] 3. Semantic vectorization and similarity modeling: A semantic embedding model is used to encode the elements within each category classified in the previous step, generating high-dimensional vectors with rich semantic representation capabilities. The cosine similarity function is used to calculate the semantic association between each pair of texts, constructing a symmetric semantic similarity matrix. This matrix is then regularized to make the data distribution more consistent with the standard normal distribution, thereby improving clustering efficiency and stability. Similar clusters are merged layer by layer using the average linkage method, and the distance threshold or upper limit of the number of clusters is manually set to control the clustering granularity. After clustering is complete, the results are mapped back to the original data structure, and a corresponding cluster number is assigned to each valid record.
[0060] 4. Entity keyword extraction and consistency verification: Traverse all clustering results and eliminate isolated clusters containing only a single record. For clusters containing multiple records, first determine whether there are completely repeated content fields. If so, directly determine it as a key event; if not, further extract high-weight entity keywords (such as locations, names, etc.) in each record to construct a keyword set within the cluster. Calculate the number of common keywords and the uniqueness ratio. Specifically, the number of common keywords refers to the number of keywords that appear in all event texts within the cluster. The formula is:
[0061] CKC=|common_keywords|≥MIN_COMMON_KEYWORDS (1);
[0062] Set a minimum threshold MIN_COMMON_KEYWORDS. If it is lower than this value, the keyword consistency is considered insufficient and the cluster is excluded. The uniqueness ratio measures the proportion of non-common parts in the keyword set within the cluster to avoid misjudgment due to low keyword overlap. The formula is:
[0063]
[0064] Among them, n is the number of records in the cluster, keywords i The set of keywords extracted from the i-th record. A maximum allowed value, MAX_UNIQUE_RATIO, is set. If the value exceeds this threshold, the keywords are considered too discrete and do not constitute the same event. This mechanism effectively overcomes the technical shortcomings of traditional methods in lacking semantic comparison at the key entity level, significantly improving screening accuracy.
[0065] 5. Verify semantic consistency within the cluster based on the large language model: Because there is still a small probability that non-key events may exist within the cluster after processing through the above steps, the large language model is called to perform semantic-level discrimination on the selected structured valid event clusters to confirm whether the events within the cluster truly belong to the same event. Finally, a specific description of the key event is output. The identified key events are marked and warned as events that require timely attention from management personnel, so that the corresponding personnel can use them as a reference for subsequent governance decisions.
[0066] From the above, it can be seen that the key points of the embodiments of the present invention are:
[0067] 1. Key event identification method based on hierarchical processing architecture: The embodiment of the present invention proposes a key event identification method based on the three-level architecture of "classification-clustering-discrimination"; this method integrates the government affairs-specific semantic extraction method, the semantic regularized agglomerative hierarchical clustering (SRHC) method and the intra-cluster entity keyword extraction technology, and significantly improves the system's processing efficiency and recognition accuracy through multi-stage collaborative processing of semantic vectorization, clustering optimization and key information verification.
[0068] 2. Semantic extraction method specifically for the government sector: First, regular expressions are constructed based on different reporting channels to filter out redundant content without actual semantics; then, a stop word list specifically for the government sector is constructed, and the general stop word library is expanded to include formatted words in government reports; a multi-stage standardized cleaning process is introduced, including empty text filtering, interference symbol removal, word segmentation processing, and terminology normalization; a similar keyword replacement mechanism is designed to unify and standardize terminology and abbreviations specific to the government sector to eliminate semantic ambiguity; the semantic validity of the text after vectorization is significantly improved, providing high-quality input for subsequent clustering analysis and improving recognition accuracy.
[0069] 3. Semantic Regularized Agglomerative Hierarchical Clustering Method: This method optimizes data distribution by standardizing the semantic similarity matrix; utilizes mathematical statistical laws to improve clustering efficiency and stability; and dynamically controls the cluster merging process to ensure that semantic similarity meets the preset threshold. This method significantly enhances the ability to quickly identify potential key events and improves recognition efficiency.
[0070] 4. Intra-cluster entity keyword extraction and consistency verification mechanism: Accurately extract high-weight entity keywords; define the number of common keywords (CKC) and uniqueness ratio (UR) as quantitative indicators, setting minimum thresholds and maximum limits; filter clusters with insufficient keyword overlap or overly discrete distribution to ensure semantic consistency; significantly improve the accuracy of key event screening in complex scenarios.
[0071] The detailed steps of the embodiment of the present invention are as follows:
[0072] S1: Technology Selection and Architecture Design
[0073] S1.1 adopts the embedding model paraphrase-multilingual-MiniLM-L12-v2 to support multilingual semantic representation and efficient text similarity calculation.
[0074] S1.2 uses Semantic Regularized Hierarchical Clustering (SRHC) to cluster data. This algorithm automatically divides the number of clusters by dynamically adjusting the distance threshold and calculates the distance between clusters using the average linkage method.
[0075] S1.3 uses textrank4keywords as the entity keyword extraction algorithm, which can effectively extract entity keywords such as names of people and places.
[0076] S2: Text Semantic Extraction Method for Government Affairs Scenarios
[0077] S2.1 reads the reported data, filters out empty text or non-string content from the data elements, cleans up special characters and unnecessary punctuation marks in the text, and then uses the Jieba word segmentation tool to segment the text.
[0078] S2.2 First, select different regular expressions based on the source channel to extract the reported content, and then filter out meaningless words, single-character phrases and pure digital content based on the stop word list exclusive to the government affairs field to ensure that the retained vocabulary has practical meaning.
[0079] S3: Initial classification of reported content
[0080] S3.1 uses a large language model to divide data into multiple categories according to predefined categories (natural disasters, fire safety, traffic safety, public safety, social security and conflicts and disputes), reducing the amount of input data for subsequent clustering algorithms, improving the overall process efficiency and robustness, and performing preliminary classification and labeling of the original data.
[0081] S4: Semantic Vectorization and Similarity Modeling
[0082] S4.1 encodes the elements in each category classified in the previous step using the semantic embedding Sentence-BERT model (paraphrase-multilingual-MiniLM-L12-v2), generates corresponding semantic embedding vectors with rich semantic representation capabilities, calculates the cosine distance between elements, and constructs a cosine distance matrix.
[0083] S4.2 then regularizes the matrix to make the data distribution more consistent with the standard normal distribution, thereby improving clustering efficiency and stability. Similar clusters are merged layer by layer through the average linkage method, and the distance threshold or the upper limit of the number of clusters is artificially set to control the clustering granularity. After the clustering is completed, the result is mapped back to the original data structure, and the corresponding cluster number is assigned to each valid record.
[0084] S5: Extract entity keywords within the cluster
[0085] S5.1 For the text content in each cluster, use the TextRank4Keyword algorithm to extract the top five entity keywords with the highest weights, including but not limited to entity keywords such as location information and the names of the parties.
[0086] S5.2 quantifies the semantic consistency within a cluster by calculating the number of common keywords (CKC) and the uniqueness ratio (UR). This mechanism sets a minimum common keyword threshold (MIN_COMMON_KEYWORDS) and a maximum uniqueness ratio (MAX_UNIQUE_RATIO) to effectively filter out clusters with insufficient keyword overlap or overly dispersed distribution.
[0087] S6: Intra-cluster semantic consistency verification based on large language models
[0088] S6.1 Traverse each cluster and check whether there is identical text content. If so, directly determine it as a "key event" cluster. For clusters without identical content: Combine the entity keyword extraction results to verify whether they meet the requirements for the number of common entity keywords and uniqueness ratio. Eliminate invalid clusters and retain valid clusters.
[0089] S6.2 passes the valid cluster content to the large language model for final review to confirm that the events in the cluster are different reported versions of the same event.
[0090] S6.3 Based on the audit results, a detailed description of the key event is ultimately output, and the identified key event is marked. At the same time, it is reported to the corresponding management personnel to inform them to give priority to handling or attention to the event for reference in subsequent governance decisions to ensure data consistency and accuracy.
[0091] Preferably, the vector embedding model of the embodiment of the present invention can be replaced with a bge or gte model; the clustering algorithm can be replaced with kmeans or DBSCAN; and the large language model can be replaced with other series models of qwen.
[0092] In summary, the embodiment of the present invention proposes a key event identification method based on a hierarchical processing architecture, which can be used to screen key events in the government field. The method adopts a hierarchical processing architecture of "classification-clustering-discrimination". Based on the natural laws that "the same event has the same category" and "similar events have similar characteristics in the semantic space", the method performs classification preprocessing on the event data and converts the full data processing into multiple small batch data processing, which significantly reduces the computational complexity. Theoretical analysis shows that the algorithm complexity is reduced from the original O(n 2 ) is reduced to O(m×k 2 ), where n is the total number of events, m is the number of event categories, and k is the average number of events of a single category, and the system processing efficiency is improved.
[0093] In response to the highly specialized and formatted nature of government data, an embodiment of the present invention proposes a government-specific semantic extraction method. By constructing a government-specific stop word list, regularized expressions, and redundant expression filtering rules, combined with a multi-stage standardized cleaning process (including empty text filtering, interference symbol removal, word segmentation, and terminology normalization), the effectiveness of text semantic information is significantly improved. Experiments have shown that this method improves the accuracy of text comprehension and recognition of specific terms, effectively reducing the rate of misjudgment due to semantic ambiguity, and providing a high-quality data foundation for subsequent semantic embedding and clustering analysis, thereby significantly enhancing the system's screening accuracy and efficiency for "key events."
[0094] At the same time, the embodiment of the present invention improves the traditional hierarchical clustering algorithm and proposes a semantic regularized agglomerative hierarchical clustering method. By regularizing (standardizing) the semantic similarity matrix, the data distribution conforms to the standard normal distribution, thereby optimizing the clustering process. This improvement utilizes the laws of mathematical statistics, allowing the system to significantly improve processing efficiency while maintaining clustering quality. It improves clustering processing speed and clustering quality, effectively supports subsequent key event judgment, and enables the system to identify potential key events more quickly and accurately.
[0095] The embodiment of the present invention also proposes a mechanism for extracting and verifying entity keywords within a cluster. After clustering is completed, the text content within each cluster is subjected to in-depth semantic analysis, and high-weight entity keywords related to key information such as location and name are accurately extracted. The semantic consistency within the cluster is quantified by calculating the number of common keywords (CKC) and the uniqueness ratio (UR). By setting the minimum common keyword threshold (MIN_COMMON_KEYWORDS) and the maximum uniqueness ratio upper limit (MAX_UNIQUE_RATIO), this mechanism effectively filters out clusters with insufficient keyword overlap or too discrete distribution, making up for the technical shortcomings of traditional methods in the lack of semantic comparison at the key entity level, and significantly improving the accuracy of key event screening in complex scenarios.
[0096] System Example
[0097] According to an embodiment of the present invention, a key event identification system based on a hierarchical processing architecture is provided. Figure 5 FIG. 1 is a schematic diagram of a key event identification system based on a hierarchical processing architecture according to an embodiment of the present invention. Figure 5 As shown, the key event identification system based on the hierarchical processing architecture according to an embodiment of the present invention specifically includes:
[0098] The classification module 50 is used to obtain reported events from different channels, perform data preprocessing on the reported events, obtain processed event texts, and classify the event texts;
[0099] The clustering module 52 is used to cluster the classified text using a semantic regularized agglomerative hierarchical clustering method, specifically for:
[0100] A semantic embedding model is used to encode the elements within each type of text, generating a corresponding semantic embedding vector. Based on the semantic embedding vector, the cosine distance between each pair of texts is calculated using a cosine similarity function. A symmetric cosine distance matrix is constructed, and the cosine distance matrix is regularized. The regularized distance matrix is clustered using the average linkage method. After clustering, the clustering results are mapped back to the original data structure, and a corresponding cluster number is assigned to each valid record.
[0101] The identification module 54 is used to identify repeated events in the text based on the clustering results using the entity keyword extraction and comparison mechanism within the cluster to obtain the key event identification results, specifically for:
[0102] The clustering results are traversed, and isolated clusters containing only a single record in the clustering results are eliminated. Clusters containing multiple records are judged for duplicate content. If there are completely identical text fields, they are directly judged as duplicate events. If there are no completely identical text fields, the first k high-weight entity keywords of each record are extracted to obtain a keyword set for each record. The number of common keywords and the uniqueness ratio in the cluster are calculated based on the keyword set. If the number of common keywords is less than the preset minimum common keyword threshold and the uniqueness ratio is greater than the preset maximum uniqueness ratio upper limit, the cluster is eliminated and the remaining valid clusters are retained. The remaining valid clusters are input into the large language model for judgment to obtain the final key event judgment result.
[0103] The embodiment of the present invention is a system embodiment corresponding to the above-mentioned method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.
[0104] In summary, compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0105] 1. More efficient processing flow: The embodiment of the present invention adopts a hierarchical processing architecture of "classification-clustering-discrimination". Based on the natural law that "the same event has the same category" and "the same type of events have similar characteristics in the semantic space", the full amount of data is decomposed into multiple small batches of data for processing through classification preprocessing, which significantly reduces the computational complexity. Theoretical analysis shows that the algorithm complexity is reduced from the original O(n 2 ) is reduced to O(m×k 2 ), where n is the total number of events, m is the number of event categories, and k is the average number of events within a single category. This effectively improves the system's processing efficiency, enabling real-time responses in urban governance scenarios and addressing the efficiency bottlenecks of traditional methods in large-scale data processing.
[0106] 2. Higher Accuracy: This embodiment of the present invention proposes a semantic extraction method specifically for the government sector. By constructing a government-specific stop word list, regularized expression rules, and a redundant expression filtering mechanism, combined with a multi-stage standardized cleaning process (such as empty text filtering, interference symbol removal, word segmentation, and terminology normalization), the effectiveness of text semantic information is significantly improved. Experimental results show that this method improves the accuracy of text comprehension and the recognition rate of specific terms, effectively reducing the misjudgment rate due to semantic ambiguity, providing a high-quality data foundation for subsequent semantic embedding and clustering analysis, and significantly enhancing the system's screening accuracy and reliability for key events.
[0107] 3. Faster Clustering Performance: This embodiment of the present invention improves upon traditional hierarchical clustering algorithms and proposes a semantically regularized agglomerative hierarchical clustering method. This method optimizes the clustering process by normalizing the semantic similarity matrix so that its distribution conforms to a standard normal distribution. This method leverages mathematical statistical principles to significantly improve processing efficiency while maintaining clustering quality. Compared to existing technologies, this embodiment of the present invention can more quickly and accurately identify potential key events.
[0108] 4. More rigorous result judgment: The embodiment of the present invention proposes a mechanism for extracting and verifying entity keywords within a cluster. After clustering is completed, the text content within each cluster is deeply semantically analyzed to accurately extract high-weight entity keywords related to key information such as location and name, and quantify the semantic consistency within the cluster by calculating the number of common keywords (CKC) and the uniqueness ratio (UR). This mechanism sets a minimum common keyword threshold (MIN_COMMON_KEYWORDS) and a maximum uniqueness ratio upper limit (MAX_UNIQUE_RATIO), effectively filtering out clusters with insufficient keyword overlap or overly discrete distribution, making up for the technical shortcomings of traditional methods in the lack of semantic comparison at the key entity level, and significantly improving the accuracy of key event screening in complex scenarios.
[0109] Device Example 1
[0110] An embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps described in the method embodiment when executed by the processor.
[0111] Device Example 2
[0112] An embodiment of the present invention provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by a processor, the steps described in the method embodiment are implemented.
[0113] The computer-readable storage medium in this embodiment includes, but is not limited to, ROM, RAM, magnetic disk, or optical disk.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying key events based on a hierarchical processing architecture, characterized in that include: Acquire reported events from different channels, perform data preprocessing on the reported events, obtain processed event texts, and classify the event texts; The classified texts are clustered using semantic regularized agglomerative hierarchical clustering method; According to the clustering results, the repeated events in the text are identified by using the entity keyword extraction and comparison mechanism within the cluster to obtain the key event identification results.
2. The method according to claim 1, characterized in that The data preprocessing of the reported event specifically includes: Perform data cleaning and filtering, text standardization, domain-specific processing, and validity verification on the reported events; The data cleaning and filtering includes filtering empty text, removing non-string content and clearing interference symbols; The text standardization processing includes constructing corresponding regular expressions according to different channels, extracting key content from the reported events through the corresponding regular expressions, and performing word segmentation processing on the text using a word segmentation tool; The domain-specific processing is to use a pre-built stop word list specific to the government affairs field to screen the text and remove formatted words, single words and pure numeric fields that have no actual semantic meaning; The validity verification is to check whether the processed text record is empty. If so, the text record is marked and removed.
3. The method according to claim 1, characterized in that Classifying the event text specifically includes: The event text is preliminarily classified according to event categories using a large language model; wherein the event categories include natural disasters, fire safety, traffic safety, public safety, social security, and conflicts and disputes.
4. The method according to claim 3, characterized in that The semantic regularized agglomerative hierarchical clustering method is used to cluster the classified texts, specifically including: A semantic embedding model is used to encode the elements in each type of text and generate the corresponding semantic embedding vector. The cosine distance between each pair of texts is calculated based on the semantic embedding vector using the cosine similarity function. A symmetric cosine distance matrix is constructed and regularized. The regularized distance matrix is clustered using the average link method. After clustering is completed, the clustering results are mapped back to the original data structure, and a corresponding cluster number is assigned to each valid record.
5. The method according to claim 4, characterized in that Based on the clustering results, the repeated events in the text are identified using the entity keyword extraction and comparison mechanism within the cluster. The key event identification results include: The clustering results are traversed, and isolated clusters containing only a single record in the clustering results are eliminated. Clusters containing multiple records are judged for duplicate content. If there are completely identical text fields, they are directly judged as duplicate events. If there are no completely identical text fields, the first k high-weight entity keywords of each record are extracted to obtain a keyword set for each record. The number of common keywords and the uniqueness ratio in the cluster are calculated based on the keyword set. If the number of common keywords is less than the preset minimum common keyword threshold and the uniqueness ratio is greater than the preset maximum uniqueness ratio upper limit, the cluster is eliminated and the remaining valid clusters are retained. The remaining valid clusters are input into the large language model for judgment to obtain the final key event judgment result.
6. A key event identification system based on a hierarchical processing architecture, characterized in that include: A classification module is used to obtain reported events from different channels, perform data preprocessing on the reported events, obtain processed event texts, and classify the event texts; Clustering module, used to cluster the classified text using semantic regularized agglomerative hierarchical clustering method; The discrimination module is used to discriminate repeated events in the text based on the clustering results using the entity keyword extraction and comparison mechanism within the cluster to obtain the key event discrimination results.
7. The system according to claim 6, characterized in that The clustering module is specifically used for: A semantic embedding model is used to encode the elements in each type of text and generate the corresponding semantic embedding vector. The cosine distance between each pair of texts is calculated based on the semantic embedding vector using the cosine similarity function. A symmetric cosine distance matrix is constructed and regularized. The regularized distance matrix is clustered using the average link method. After clustering is completed, the clustering results are mapped back to the original data structure, and a corresponding cluster number is assigned to each valid record.
8. The system according to claim 7, characterized in that The discrimination module is specifically used for: The clustering results are traversed, and isolated clusters containing only a single record in the clustering results are eliminated. Clusters containing multiple records are judged for duplicate content. If there are completely identical text fields, they are directly judged as duplicate events. If there are no completely identical text fields, the first k high-weight entity keywords of each record are extracted to obtain a keyword set for each record. The number of common keywords and the uniqueness ratio in the cluster are calculated based on the keyword set. If the number of common keywords is less than the preset minimum common keyword threshold and the uniqueness ratio is greater than the preset maximum uniqueness ratio upper limit, the cluster is eliminated and the remaining valid clusters are retained. The remaining valid clusters are input into the large language model for judgment to obtain the final key event judgment result.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the key event identification method based on a hierarchical processing architecture as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by the processor, the steps of the key event identification method based on the hierarchical processing architecture as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method for identifying disastrous meteorological hotspot events based on social signals
CN108595582A
Text clustering method and system based on keyword matching, storage medium and terminal
CN112732914A
Hotspot event identification method and system based on multilevel clustering
CN113064990A
Method and system for automatically generating network public opinion special report based on event aggregation
CN118364163A
Text hotspot clustering method based on large model
CN119474387A
Cited By
Event repetition judgment method and device, equipment and medium
CN121256399A
Policy consistency assessment method and system based on large model and iterative clustering
CN121980292A