A method, medium, and apparatus for regional event risk analysis
Patent Information
- Application Number
- CN202611236215.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-29
AI Technical Summary
但是上述方法仅关注纠纷事件等某一种类型的事件,未能综合冲突类、威胁类、行动类等多风险维度进行全面评估,导致风险刻画不完整;且未考虑区域人口基数差异,单纯比较事件频次导致人口密集区域的风险值虚高,而人口稀疏区域的风险值被低估,无法实现区域间的公平可比;同时依赖固定阈值或专家经验划分风险等级,缺乏基于全体区域统计分布的动态分级机制,导致分级结果主观性强、适应性差
[0011]本发明至少具有以下有益效果:通过对待处理事件文本进行多维度关键词匹配和文本占比,实现了风险的多维量化与自动化评估,避免了单一维度的片面性;通过对风险指数进行时间平滑处理,有效滤除了单日或短期随机波动,使得平滑风险值能够真实反映近期风险的发展趋势,提升了风险监测的稳定性和趋势捕捉能力;通过将各预设风险维度的标准化分值与区域人口分布数据进行加权融合,既综合了多维度风险信息,又引入了人口基数修正,避免了人口密集区域因事件绝对数量多而被高估、稀疏区域被低估的问题,提高了风险分级的准确性。
Smart Images

Figure CN122840698A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, medium, and device for regional event risk analysis. Background Technology
[0002] Regional event risk monitoring has significant application value in areas such as social governance, public safety, and emergency management. Traditional event risk monitoring methods mainly rely on manual collection and analysis, or simple statistical methods based on a single data source. For example, risk levels are determined solely based on whether specific keywords are present in the event text, or the number of events in each region is simply counted as a risk indicator. However, these methods only focus on one type of event, such as disputes, and fail to comprehensively assess multiple risk dimensions, including conflict, threat, and action-related events, resulting in incomplete risk characterization. Furthermore, they do not consider regional differences in population size; simply comparing event frequency leads to inflated risk values in densely populated areas and underestimated risk values in sparsely populated areas, failing to achieve fair comparisons between regions. Additionally, relying on fixed thresholds or expert experience to classify risk levels lacks a dynamic grading mechanism based on the statistical distribution of the entire region, resulting in highly subjective and poorly adaptable grading results.
[0003] Therefore, how to integrate multi-dimensional risk information to improve the accuracy of risk assessment has become an urgent technical problem to be solved. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention provides a regional event risk analysis method, which includes the following steps: S1. Perform keyword matching on several unprocessed event texts associated with each target area. Based on the matching results, obtain the risk index for each target area for several preset risk dimensions. The risk index is used to characterize the proportion of the frequency of event texts appearing in the corresponding target area under the corresponding preset risk dimension.
[0005] S2, perform time smoothing on the risk index of each target area under each preset risk dimension to obtain the smoothed risk value of each target area under each preset risk dimension, wherein the smoothed risk value is used to characterize the recent risk trend of the corresponding target area under the corresponding preset risk dimension.
[0006] S3, for any preset risk dimension, standardize the smoothed risk value of all target areas under the current preset risk dimension to obtain the standardized score of each target area under the current preset risk dimension.
[0007] S4, weighted and fused with the standardized scores of each target area under each preset risk dimension and the corresponding population distribution data of each target area, to obtain the comprehensive target score for each target area.
[0008] S5. Sort the comprehensive target scores corresponding to all target areas by numerical value, and determine the target risk level corresponding to each target area based on the percentile distribution after sorting.
[0009] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described regional event risk analysis method.
[0010] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0011] This invention has at least the following beneficial effects: By performing multi-dimensional keyword matching and text proportion analysis on the text of the event to be processed, it achieves multi-dimensional quantification and automated assessment of risk, avoiding the one-sidedness of a single dimension; by performing time smoothing processing on the risk index, it effectively filters out single-day or short-term random fluctuations, enabling the smoothed risk value to truly reflect the recent risk development trend, thus improving the stability and trend capture capability of risk monitoring; by weighting and integrating the standardized scores of each preset risk dimension with regional population distribution data, it not only integrates multi-dimensional risk information but also introduces population base correction, avoiding the problem of overestimation in densely populated areas due to the large absolute number of events and underestimation in sparsely populated areas, thereby improving the accuracy of risk classification. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a regional event risk analysis method provided in Embodiment 1 of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0016] Example 1 This first embodiment provides a method for regional event risk analysis, such as... Figure 1 As shown, the regional event risk analysis method includes the following steps: S1 performs keyword matching on several pending event texts associated with each target area, and obtains the risk index for each target area for several preset risk dimensions based on the matching results.
[0017] The target area is the geographical unit to be monitored for risks, such as provinces, cities, and sea areas. The event text to be processed consists of raw text data such as news reports and announcements after preliminary cleaning and deduplication. Preset risk dimensions include stress expression, action representation, regional conflict, security incident, and economic constraints, etc., to achieve a multi-dimensional characterization of risks and avoid the one-sidedness of single-dimensional assessment.
[0018] The risk index is used to characterize the proportion of event text occurrences in a corresponding target region under a preset risk dimension. Correspondingly, this embodiment uses automated data collection, keyword matching, and statistical calculation to measure the risk level of each region by the proportion of event texts involving various specific risk themes in each target region, thus transforming unstructured natural language text into structured risk quantification indicators.
[0019] In one specific embodiment, S1 includes the following steps: S111: For any target area, configure at least one data source and a set of data collection keywords corresponding to several preset risk dimensions.
[0020] S112, based on the data source and the set of data collection keywords, retrieve the initial text data corresponding to the current target area through a preset data interface.
[0021] S113, perform duplicate detection and filtering on the initial text data, and obtain several unprocessed event texts corresponding to the current target area.
[0022] Among them, the data source is the source of the raw text data, such as news media APIs, official public data interfaces, etc., to ensure the continuous acquisition and updating of data and support automated collection.
[0023] In one specific implementation, the data source includes approximately 1,200 mainstream news media outlets worldwide. Preferably, the data source automatically retrieves news data daily via a data API interface.
[0024] This embodiment pre-configures a set of representative keywords for each preset risk dimension and target region to retrieve relevant text from the data source. These keywords include regional keywords, basic category keywords, and risk keywords specific to each risk dimension. For each target region and each preset risk dimension, the regional keywords, basic category keywords, and risk keywords are logically combined, such as "regional keywords AND basic category keywords AND risk keywords" or "regional keywords AND (risk keyword 1 OR risk keyword 2)," forming the final query expression for data retrieval. Through API calls, initial text data related to specific regions and risk types is automatically and accurately retrieved from various data sources, improving the targeting and efficiency of data collection.
[0025] In one specific implementation, a set of keywords for the corresponding language is configured for each target region. For example, Chinese keywords are configured for target regions where Chinese is the primary language, and English keywords are configured for target regions where English is the primary language.
[0026] In one specific implementation, keyword matching supports lemmatization expansion. Specifically, for keywords configured with expansion tags, multiple lemmatization variants of the keyword are automatically generated to broaden the matching scope. For example, for English keywords, their plural forms can be automatically generated, such as by adding suffixes like 's' or 'es', to cover text expressions under different grammatical forms.
[0027] In this embodiment, regional keywords may include the Chinese name, English name, and common aliases of each target region; basic category keywords may include terms such as economy, trade, and defense, while irrelevant content such as entertainment, sports, and technology is excluded to ensure that the event text focuses on potential risk areas; risk keywords corresponding to the pressure expression dimension may include warnings, statements, deadlines, etc.; risk keywords corresponding to the action representation dimension may include deployment, mobilization, exercises, blockade, etc.; risk keywords corresponding to the regional conflict dimension may include standoffs, friction, military operations, etc.; risk keywords corresponding to the security event dimension may include explosions, attacks, etc.; and risk keywords corresponding to the economic restriction dimension may include restrictive measures, export controls, asset freezes, etc.
[0028] Because the same event may be reported repeatedly by multiple media outlets, or the same text may be retrieved multiple times due to crawling malfunctions, duplicate data can artificially inflate the number of texts, leading to an overestimation of the risk index's denominator or numerator, thus affecting the accuracy of the risk index. Therefore, this embodiment uses both unique identifiers and text content similarity for deduplication to ensure that each independent event is counted only once.
[0029] Specifically, the first round of deduplication is based on unique identifiers: First, extract the unique identifier of each initial text data, such as URL, news ID, hash value of title + publication time, etc., and build a hash set. For the current initial text data, if its unique identifier already exists in the hash set, the initial text data is discarded; otherwise, it is retained and the unique identifier is added to the hash set.
[0030] A second round of deduplication is performed based on text content similarity: For the initial text data that has passed the first round of deduplication, the text similarity of its main body is calculated to determine whether highly similar content exists. Correspondingly, by setting a similarity threshold, for any two initial text data, if the text similarity exceeds the similarity threshold, they are considered duplicates, and only one of them is retained. In this embodiment, initial text data with an earlier publication date, authoritative source, or more complete main body can be retained. The specific value of the similarity threshold can be set by the implementer according to the actual situation; for example, in this embodiment, the similarity threshold is 85%.
[0031] The remaining initial text data after two rounds of deduplication is the event text to be processed.
[0032] In one specific embodiment, S1 further includes the following steps: S121, For any target region, based on the region keyword set and basic category keyword set corresponding to the current target region, perform keyword matching on each pending event text corresponding to the current target region, and count the number of first texts corresponding to the current target region.
[0033] S122, for any preset risk dimension, based on the risk keyword set corresponding to the current preset risk dimension, perform keyword matching on each pending event text corresponding to the current target area, and count the number of second texts in the current target area under the current preset risk dimension.
[0034] S123, the ratio of the number of second texts to the number of first texts is determined as the risk index of the current target area for the current preset risk dimension.
[0035] In one specific embodiment, S121 includes the following steps: For any event text to be processed corresponding to the current target area, extract the keywords corresponding to the current event text to obtain a set of text keywords.
[0036] If the set of text keywords includes at least one keyword from the set of regional keywords and at least one keyword from the set of basic category keywords, then the current event text to be processed will be counted in the first text count.
[0037] In one specific embodiment, S122 includes the following steps: If the text keyword set includes at least one keyword from the regional keyword set and at least one keyword from the risk keyword set, then the current pending event text will be counted in the second text count.
[0038] In this process, keywords are extracted from each event text using word segmentation tools such as jieba and NLP word segmenters to obtain the corresponding set of text keywords.
[0039] The first text count represents the total number of basic event texts related to risk analysis in each target region. This embodiment requires the text to simultaneously contain regional keywords that ensure geographic attribution and basic category keywords that ensure content relevance, thus excluding irrelevant or noisy text. This allows the first text count to accurately reflect the effective event base of the target region.
[0040] The second text count represents the total number of event texts related to a specific risk dimension, further filtered after determining the region's affiliation. This embodiment ensures that the second text count reflects the number of valid events within that target region under a specific risk dimension by requiring the texts to contain both regional keywords and risk keywords.
[0041] In one specific implementation, keyword matching employs a word distance matching method. Specifically, for the current event text to be processed, assuming both regional keywords and risk keywords are detected, it is further determined whether the word distance between them in the current event text does not exceed a preset distance threshold. If the word distance does not exceed the preset distance threshold, the current event text to be processed is included in the second text count of the corresponding risk dimension. Here, word distance is represented by the number of word intervals between two keywords; for example, N / 2 indicates that there are no more than two interval words between two keywords.
[0042] Furthermore, the ratio of the number of second texts to the number of first texts is determined as the risk index of the current target region for the current preset risk dimension. This ratio calculation method eliminates the influence of differences in the total number of events between different target regions, making the risk levels of each target region comparable. Moreover, the higher the ratio, the greater the event density and the higher the risk level of the target region under a specific risk dimension, and the higher its risk index.
[0043] It should be noted that the same event text may simultaneously meet the risk keyword conditions of multiple preset risk dimensions, and will be counted separately in the second text count of each corresponding preset risk dimension.
[0044] The above-mentioned method accurately reflects the activity level of basic events truly relevant to risk analysis in each target area by using the first text quantity, effectively eliminating the interference of irrelevant text. The second text quantity accurately captures the event frequency of each target area under a specific risk dimension, improving the accuracy of risk identification and achieving horizontal comparability between target areas of different sizes, thus providing a data foundation for subsequent risk analysis.
[0045] In one specific embodiment, the regional event risk analysis method further includes the following steps after step S1 and before step S2: S10 processes each text of an event to be processed using a pre-trained sentiment analysis model to obtain the sentiment tendency information corresponding to each text of an event to be processed.
[0046] S20: For any target area and any preset risk dimension, count the number of third texts of pending event texts that have a negative emotional tendency under the current preset risk dimension in the current target area.
[0047] S30, the ratio of the number of third texts to the number of second texts in the current target area under the current preset risk dimension is used as the proportion of negative emotions in the current target area under the current preset risk dimension.
[0048] S40: Based on the proportion of negative emotions, the risk index of the current target area for the current preset risk dimension is weighted and corrected to obtain the corrected risk index, and the corrected risk index is used to replace the risk index for execution step S2.
[0049] The pre-trained sentiment analysis model is a deep learning model pre-trained on a large-scale labeled corpus, capable of automatically determining the sentiment tendency of text. Those skilled in the art will recognize that any existing sentiment analysis model falls within the scope of this invention, such as Bidirectional Encoder Representations from Transformers (BERT), RoBERTa, Long Short-Term Memory networks, GPT series, and other large language models.
[0050] This embodiment employs a BERT-based architecture, consisting of 12 stacked Transformer encoder layers with a hidden layer dimension of 768, 12 attention heads, and a total of approximately 110 million parameters. Each encoder layer contains two sub-layers: a multi-head self-attention mechanism and a feedforward neural network. Each sub-layer is followed by residual connections and layer normalization. The multi-head self-attention mechanism achieves dynamic feature extraction by calculating the correlation weights between different positions in the input sequence; the feedforward neural network performs non-linear transformations on the output of the self-attention layer to extract higher-level features. The input to the BERT model is represented as the sum of three embedding vectors corresponding to the event text to be processed: word embeddings, representing the basic semantic information of words; segment embeddings, used to distinguish different sentences in sentence pairs of the event text to be processed; and position embeddings, representing the positional order of words in the event text to be processed, using learnable positional encoding. A special label [CLS] is added to the beginning of the input sequence; the final hidden state corresponding to this label is used for downstream classification tasks. Sentences are separated by [SEP] separators, and when the sequence length is insufficient, [PAD] padding is used.
[0051] Specifically, the event text to be processed needs to undergo word segmentation, the addition of [CLS] and [SEP] tags, and the generation of attention masks to convert it into a numerical format acceptable to the BERT model. The final hidden state vector with the [CLS] tag is extracted from the output layer of the BERT model, and a fully connected layer is added above it as a classification head to map the [CLS] vector to the probability distributions of positive, negative, and neutral sentiment categories.
[0052] Among them, the emotional tendency information includes three categories: positive emotional tendency, negative emotional tendency, and neutral emotional tendency, which are used to distinguish the emotional attributes of event texts. Among them, negative emotional texts have a higher risk indication significance.
[0053] The third text quantity refers to the total number of pending event texts that simultaneously satisfy the region keywords, risk keywords, and have a negative sentiment tendency within a specific target area and risk dimension. This quantity reflects the frequency of high-risk emotional events. The ratio of the third text quantity to the second text quantity within the current target area under the current preset risk dimension is used as the negative sentiment proportion of the current target area under the current preset risk dimension, quantifying the concentration of negative sentiment events in risk-related texts. Correspondingly, a higher negative sentiment proportion indicates a more negative sentiment trend in the event within the target area and under the current preset risk dimension, signifying a higher risk.
[0054] The original risk index is based solely on the frequency of events, without considering the intensity of emotions. In this embodiment, the proportion of negative emotions is used as a moderating factor to either enhance or weaken the original risk index. Correspondingly, the higher the proportion of negative emotions, the greater the revised risk index, reflecting the escalation of risk due to worsening emotions. For example, the revised risk index = original risk index × (1 + proportion of negative emotions).
[0055] In one specific implementation, the pre-trained sentiment analysis model is constructed through the following steps: obtaining a training sample set, which includes multiple news texts labeled with sentiment tendency categories, including positive sentiment tendency, negative sentiment tendency, and neutral sentiment tendency; using the BERT-base architecture as the base model, inputting the training samples into the base model, and extracting the final hidden state vector labeled [CLS]; adding a fully connected layer as a classification head above the [CLS] vector, mapping the vector to the probability distributions of positive sentiment tendency, negative sentiment tendency, and neutral sentiment tendency; and employing the cross-entropy loss function and the AdamW optimizer, with the initial learning rate set to 2×10⁻⁶. -5 Up to 5×10 -5 Between these values, the batch size is 16 or 32, and the model parameters are adjusted through backpropagation until the model converges.
[0056] The above describes how, based on the original frequency-based risk index, sentiment analysis technology from natural language processing was introduced to construct a two-dimensional risk quantification mechanism that combines frequency and sentiment. This enhances the sensitivity and practicality of regional event risk monitoring methods and provides a higher-quality data foundation for subsequent tiered early warning systems.
[0057] S2, perform time smoothing on the risk index of each target area under each preset risk dimension to obtain the smoothed risk value of each target area under each preset risk dimension.
[0058] In this process, time smoothing can filter the time series data of the risk index, eliminating short-term random fluctuations, highlighting long-term trends, and avoiding misjudgments caused by abnormal daily fluctuations, thereby improving the stability of risk trend identification. Correspondingly, the smoothed risk value, as the risk index estimate obtained after smoothing, is used to characterize the recent risk trend of the corresponding target area under the corresponding preset risk dimension, that is, to reflect the recent development direction and intensity of the risk.
[0059] In one specific embodiment, S2 includes the following steps: S21. For any target area and any preset risk dimension, obtain the risk index of the current target area at the current time point and the previous L consecutive time points under the current preset risk dimension, where L is the preset time window length.
[0060] S22, perform a weighted average of L+1 risk indices to obtain the smoothed risk value of the current target area under the current preset risk dimension. The weight of the current time point is the preset benchmark weight, and the weight of the kth time point before the current time point is equal to the benchmark weight multiplied by the kth power of the preset decay factor. The decay factor is a value greater than 0 and less than 1, k=1, 2, ..., L.
[0061] In order to analyze the changing trend of risk, this embodiment obtains the risk index of the current time point and the L consecutive time points before it.
[0062] A weighted average is calculated for L+1 risk indices, with higher weights assigned to data closer to the current time point and exponentially decreasing weights for older data. This approach aims to more accurately capture recent risk trends while smoothing out short-term noise. In one specific implementation, the preset time window length L is set to 13, meaning a weighted average is calculated for the risk indices across 14 consecutive time points, including the current time point.
[0063] The closer the decay factor is to 1, the more long-term historical data is retained, and the smooth curve is more stable, but the response to changes is slower; in this embodiment, the decay factor is 0.7.
[0064] As described above, by obtaining the risk index of the current time point and the previous L consecutive time points, a time window for trend analysis is constructed, enabling the smoothing results to reflect the evolution of risk rather than isolated moments. By weighting and averaging the L+1 risk indices and using exponentially decaying weight allocation, the smoothed risk value can effectively filter out short-term random fluctuations and sensitively track the recent direction of risk changes. This provides more stable and reliable basic data for subsequent standardization processing and cross-regional comparisons, improving the accuracy and anti-interference ability of trend judgment during risk monitoring.
[0065] S3, for any preset risk dimension, standardize the smoothed risk value of all target areas under the current preset risk dimension to obtain the standardized score of each target area under the current preset risk dimension.
[0066] First, the overall average value μ and the overall standard deviation σ of the smoothed risk values of all target regions under the current risk dimension are calculated. Then, the standardized score of the i-th target region under the current preset risk dimension is: Z i =(Smoothed i -μ) / σ, where, Smoothed i Let be the smoothed risk value of the i-th target region under the current preset risk dimension, i=1,2,...,N, where N is the total number of target regions.
[0067] The above steps enable the conversion from raw smoothed risk values to comparable standardized scores, enhancing the analytical capabilities and interpretability of regional event risk monitoring methods.
[0068] S4, weighted and fused with the standardized scores of each target area under each preset risk dimension and the corresponding population distribution data of each target area, to obtain the comprehensive target score for each target area.
[0069] Among them, population distribution data reflects the population size or density of the target area, such as the number of permanent residents and population density (people / square kilometer), which is used to correct for differences in the population base in risk assessment and avoid overestimating the risk in densely populated areas due to the large absolute number of events.
[0070] In one specific embodiment, S4 includes the following steps: S41. Obtain population distribution data for each target area, and after normalizing the population distribution data, obtain the normalized population index corresponding to each target area.
[0071] S42, set the corresponding first fusion weight for each preset risk dimension, and set the corresponding second fusion weight for the normalized population indicator.
[0072] S43. For each target area, multiply the standardized scores of the current target area under each preset risk dimension by the corresponding first fusion weight and sum them to obtain the first weighted sum. Then, add the first weighted sum to the product of the normalized population index corresponding to the current target area multiplied by the second fusion weight to obtain the comprehensive target score corresponding to the current target area.
[0073] The population size or density of different target areas often varies greatly. If the original population values are directly used for weighted fusion, the overall score will be severely distorted, and the result will be dominated by the population of large areas. Therefore, the population distribution data is first normalized. For example, the maximum-minimum normalization method is used to map the population distribution data to the [0,1] interval to obtain the normalized population index corresponding to each target area.
[0074] Different preset risk dimensions contribute differently to the overall risk and can be adjusted through weighting; meanwhile, the influence of demographic factors can also be controlled through independent weighting. The specific values of the first and second fusion weights can be set by the implementer according to the actual situation.
[0075] The above-mentioned comprehensive target score, which reflects both the regional comprehensive risk level and the impact of population carrying capacity, is constructed by adding the weighted sum of the multi-dimensional standardized scores to the weighted terms of the normalized population index. This provides a unified, clear and interpretable numerical basis for subsequent ranking and classification.
[0076] S5. Sort the comprehensive target scores corresponding to all target areas by numerical value, and determine the target risk level corresponding to each target area based on the percentile distribution after sorting.
[0077] The comprehensive target scores of all target areas are sorted from largest to smallest or smallest to largest. This embodiment uses descending order as an example, with areas having higher scores indicating higher risk levels.
[0078] Based on the preset number of risk levels, determine the corresponding percentile cutoff points, and the specific percentile boundaries can be adjusted according to actual business needs. For example, if it is necessary to divide into 6 levels: extremely high risk, high risk, medium risk, low risk, extremely low risk, and no risk, the following percentile intervals can be set: extremely high risk corresponds to the interval ranking in the top 5% of the comprehensive target score; high risk corresponds to the interval ranking in ...
[0079] Further determine the target risk level for each target area, thereby providing reliable data support for differentiated risk response and resource optimization in scenarios such as social governance, public safety early warning, and emergency command.
[0080] This embodiment achieves multi-dimensional quantification and automated assessment of risk by performing multi-dimensional keyword matching and text proportion analysis on the text of the event to be processed, avoiding the one-sidedness of a single dimension. By performing time smoothing processing on the risk index, it effectively filters out single-day or short-term random fluctuations, enabling the smoothed risk value to truly reflect the recent risk development trend, thus improving the stability and trend capture capability of risk monitoring. By weighted fusion of the standardized scores of each preset risk dimension with regional population distribution data, it not only integrates multi-dimensional risk information but also introduces population base correction, avoiding the problem of overestimation in densely populated areas due to the large absolute number of events and underestimation in sparsely populated areas, thereby improving the accuracy of risk classification.
[0081] Example 2 Embodiment 2 of the present invention provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the regional event risk analysis method provided in the above embodiment.
[0082] Example 3 Embodiment 3 of the present invention provides an electronic device, which includes a processor and the non-transitory computer-readable storage medium of Embodiment 2 of the present invention.
[0083] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for regional event risk analysis, characterized in that, The method includes the following steps: S1, perform keyword matching on several unprocessed event texts associated with each target area, and obtain the risk index of each target area for several preset risk dimensions based on the matching results. The risk index is used to characterize the proportion of the frequency of event texts appearing in the corresponding target area under the corresponding preset risk dimension. S2, perform time smoothing on the risk index of each target area under each preset risk dimension to obtain the smoothed risk value of each target area under each preset risk dimension, wherein the smoothed risk value is used to characterize the recent risk trend of the corresponding target area under the corresponding preset risk dimension; S3, for any preset risk dimension, standardize the smoothed risk value of all target areas under the current preset risk dimension to obtain the standardized score of each target area under the current preset risk dimension; S4, weighted and fused the standardized scores of each target area under each preset risk dimension with the population distribution data corresponding to each target area to obtain the comprehensive target score for each target area; S5. Sort the comprehensive target scores corresponding to all target areas by numerical value, and determine the target risk level corresponding to each target area based on the percentile distribution after sorting.
2. The regional event risk analysis method according to claim 1, characterized in that, S1 includes the following steps: S111, for any target area, configure at least one data source and a set of data collection keywords corresponding to several preset risk dimensions respectively; S112, based on the data source and the set of data collection keywords, retrieve the initial text data corresponding to the current target area through a preset data interface; S113, perform duplicate detection and filtering on the initial text data to obtain several unprocessed event texts corresponding to the current target area.
3. The regional event risk analysis method according to claim 1, characterized in that, S1 also includes the following steps: S121, For any target area, based on the set of regional keywords and the set of basic category keywords corresponding to the current target area, perform keyword matching on each unprocessed event text corresponding to the current target area, and count the number of first texts corresponding to the current target area. S122, For any preset risk dimension, based on the risk keyword set corresponding to the current preset risk dimension, perform keyword matching on each pending event text corresponding to the current target area, and count the number of second texts in the current target area under the current preset risk dimension. S123, the ratio of the second text quantity to the first text quantity is determined as the risk index of the current target area for the current preset risk dimension.
4. The regional event risk analysis method according to claim 3, characterized in that, S121 includes the following steps: For any event text to be processed corresponding to the current target area, extract the keywords corresponding to the current event text to obtain a set of text keywords; If the set of text keywords includes at least one keyword from the set of regional keywords and at least one keyword from the set of basic category keywords, then the current event text to be processed will be counted in the first text count.
5. The regional event risk analysis method according to claim 4, characterized in that, S122 includes the following steps: If the set of text keywords includes at least one keyword from the set of regional keywords and at least one keyword from the set of risk keywords, then the text of the event to be processed will be counted in the second text count.
6. The regional event risk analysis method according to claim 3, characterized in that, The method, after step S1 and before step S2, further includes the following steps: S10, each text of the event to be processed is processed by a pre-trained sentiment analysis model to obtain the sentiment tendency information corresponding to each text of the event to be processed, wherein the sentiment tendency information includes positive sentiment tendency, negative sentiment tendency and neutral sentiment tendency. S20, for any target area and any preset risk dimension, count the number of third texts of the event texts to be processed in the current target area that have a negative emotional tendency under the current preset risk dimension; S30, the ratio of the number of third texts to the number of second texts in the current target area under the current preset risk dimension is used as the proportion of negative emotions in the current target area under the current preset risk dimension; S40, the risk index of the current target area for the current preset risk dimension is weighted and corrected according to the proportion of negative emotions to obtain the corrected risk index, and the corrected risk index is used to replace the risk index for executing step S2.
7. The regional event risk analysis method according to claim 1, characterized in that, S2 includes the following steps: S21. For any target area and any preset risk dimension, obtain the risk index of the current target area under the current preset risk dimension at the current time point and the previous L consecutive time points, where L is the preset time window length. S22, perform a weighted average of L+1 risk indices to obtain the smoothed risk value of the current target area under the current preset risk dimension. The weight of the current time point is the preset benchmark weight, and the weight of the kth time point before the current time point is equal to the benchmark weight multiplied by the kth power of the preset attenuation factor. The attenuation factor is a value greater than 0 and less than 1, k=1, 2, ..., L.
8. The regional event risk analysis method according to claim 1, characterized in that, S4 includes the following steps: S41, Obtain population distribution data for each target area, and after normalizing the population distribution data, obtain the normalized population index corresponding to each target area, wherein the population distribution data is the number of permanent residents or the population density. S42, set a corresponding first fusion weight for each preset risk dimension, and set a second fusion weight corresponding to the normalized population indicator; S43, for each target area, multiply the standardized scores of the current target area under each preset risk dimension by the corresponding first fusion weight and sum them to obtain the first weighted sum. Then, add the product of the first weighted sum and the normalized population index corresponding to the current target area multiplied by the second fusion weight to obtain the comprehensive target score corresponding to the current target area.
9. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the regional event risk analysis method as described in any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 9.