Risk personnel analysis method and system based on grassroots governance big data
By constructing a risk personnel analysis method based on big data for grassroots governance, collecting multi-source heterogeneous data and performing entity alignment and semantic standardization, and combining spatiotemporal co-occurrence and behavioral pattern similarity, high-risk groups are identified. This solves the problem of difficulty in discovering hidden gangs and risk transmission chains in traditional methods, and enables early intervention and precise prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 数尚(浙江)科技有限公司
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are insufficient for identifying hidden gangs, risk transmission chains, or clustered risk groups that lack a clear organizational structure but exhibit highly coordinated behavior in grassroots governance, resulting in delayed responses and an inability to detect and intervene in advance.
By constructing a risk personnel analysis method based on big data of grassroots governance, multi-source heterogeneous data is collected, entity alignment and semantic standardization are performed, a standardized personnel entity table is generated, and a weighted association graph is constructed by combining spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, dynamic risk scores are calculated, and high-risk groups are identified through risk propagation iteration.
It enables the automatic identification of hidden gangs, risk transmission chains, and clustered risk groups, improving the foresight and precision of grassroots governance. It allows for early intervention in the evolution of risks, significantly improving the timeliness and accuracy of risk prevention and control.
Smart Images

Figure CN121961247A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for risk personnel analysis based on big data in grassroots governance. Background Technology
[0002] In the field of grassroots social governance, the early identification and intervention of at-risk individuals is a crucial link in maintaining community safety and stability. Traditional methods of at-risk individual analysis mainly rely on data from single departments such as public security and community organizations, and statically assess individual behavior through preset rules and human experience. Typical technical solutions include: list comparison systems based on historical case and event records, early warning mechanisms triggered by abnormal behavior at a single point in time, and risk level classification methods relying on the subjective judgment of community grid workers. These technologies are effective in simple scenarios, but their limitations become apparent when facing the complex needs of social governance.
[0003] The core problem facing current technologies lies in the disconnect between individual risk assessment and group risk identification. They are unable to automatically identify potentially interconnected risk groups from discrete, anomalous individuals, particularly struggling to identify hidden gangs, risk transmission chains, or clustered risk groups that lack clear organizational structures but exhibit highly coordinated behavior. Existing systems often only identify known high-risk individuals or overtly related groups (such as relatives or cohabitants), lacking the ability to detect new types of risk groups that communicate through temporary communication tools, operate across regions, and exhibit covertly synchronized behavioral patterns. When multiple low-risk individuals form high-risk groups through complex relationship networks, traditional methods fail to capture this "whole is greater than the sum of its parts" risk aggregation effect, leading to delayed responses in grassroots governance and making it difficult to take effective intervention measures before the risk spreads. Summary of the Invention
[0004] This invention aims to provide a risk personnel analysis method and system based on big data in grassroots governance, enabling grassroots governance workers to identify risk groups that have not yet caused actual harm but have formed a clustering trend in advance, and to intervene in the early stage of risk evolution. This shifts the focus of social governance from post-event handling to pre-event prevention, significantly improving the foresight and accuracy of grassroots risk prevention and control.
[0005] To achieve the above objectives, the technical solution adopted by this invention is: a risk personnel analysis method based on big data of grassroots governance, comprising: Collect multi-source heterogeneous grassroots governance data, perform entity alignment and semantic standardization, and generate a standardized personnel entity table with attributes and initial screening labels; Based on the personnel attributes and initial screening tags in the standardized personnel entity table, and combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, a weighted association edge between personnel is constructed to form an initial risk personnel association graph; Based on the node connection relationships in the initial risk personnel association graph, and according to the multi-dimensional individual risk indicator system, combined with the time decay factor, the dynamic risk score of each person is calculated. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. Based on the node risk score and weighted association edges in the enhanced graph, the risk is propagated from high-risk nodes to neighboring nodes in multiple rounds until convergence. Connected subgraphs with risk scores higher than the threshold and a size greater than a preset number are identified as high-risk groups.
[0006] Preferably, the generation of the standardized personnel entity table with attributes and initial screening tags includes: Establish a list of data sources and access protocols for identifying at-risk individuals, and define a unified structured format for different data sources; Based on the data obtained from the access protocol, perform cross-source data alignment and deduplication based on the entity primary key, and construct a multi-dimensional primary key mapping table; Based on the alignment results of the multidimensional primary key mapping table, the structured fields are semantically standardized and initially screened for risk labels; Based on the results of semantic standardization and initial risk label screening, all aligned personnel entities and their attributes, along with the initial screening labels, are integrated into a standardized personnel entity table.
[0007] Preferably, the formation of the initial risk personnel association map includes: Trajectory data is extracted from the standardized personnel entity table to calculate the spatiotemporal co-occurrence intensity among personnel and to count pairs of personnel appearing in the same geographical area within the same time window. Based on the calculated spatiotemporal co-occurrence intensity, explicit social relationship fields are extracted from the standardized personnel entity table to generate relationship strength weights; By combining the spatiotemporal co-occurrence strength and relationship strength weights, behavioral attributes are extracted from the standardized personnel entity table to calculate the similarity of behavioral patterns among personnel; Based on the three-dimensional calculation results of spatiotemporal co-occurrence strength, relationship strength weight, and behavioral pattern similarity, weighted association edges are generated and an initial risk personnel association graph is constructed.
[0008] Preferably, the calculation of the dynamic risk score for each individual includes: Define basic risk indicators, behavioral abnormality indicators, social abnormality indicators, and environmental risk indicators to construct an individual risk indicator system; Normalize and assign weights to each risk indicator in the individual risk indicator system to eliminate differences in dimensionality; Based on the normalized and weighted indicators, the initial individual risk score is calculated and the contribution of historical events is adjusted by introducing a time decay factor. Based on the adjusted risk score, the dynamic risk score is embedded as a node attribute into the initial risk personnel association graph.
[0009] Preferably, the multi-round propagation iteration of execution risk from high-risk nodes to neighboring nodes includes: Define the rules for calculating the risk impact of a node on its neighbors. The risk impact is directly proportional to the edge weight and inversely proportional to the current risk level of the target node. Based on the risk impact calculation rules, and using the weighted association edges in the initial risk personnel association graph, the risk impact of all neighboring nodes on the current node is summarized. Based on the summarized risk impact, update the risk score of the current node. Based on the updated risk score, repeat the iteration until the risk change of all nodes is less than the preset convergence threshold.
[0010] Preferably, the connected subgraphs whose risk scores are higher than a threshold and whose size is greater than a preset number are identified as high-risk groups, including: Analyze the topological characteristics of high-risk connected subgraphs and calculate the average path length and clustering coefficient of the subgraphs; Based on the calculation results of the average path length of the subgraph and the clustering coefficient, when the average path length of the subgraph is large and the clustering coefficient is low, it is determined to be a propagation chain. Based on the calculation results of subgraph clustering coefficient and edge density, when the subgraph clustering coefficient is high and the edge density is large, it is identified as a hidden group. Based on the distribution of geographic coordinates of subgraph nodes, when nodes are concentrated in the same geographic area, it is determined to be a cluster risk.
[0011] Preferably, the method further includes: Based on the high-risk groups identified in the enhanced graph, a tracking window and a member activity monitoring mechanism are set up to continuously monitor the activity status of each member in the subgraph; Based on member activity monitoring results, detect structural changes in high-risk groups and identify events such as member increases, decreases, splits, or mergers; Based on the identified structural change events, combined with newly added grassroots governance data, the initial at-risk personnel association map is incrementally updated; Based on structural change events and incrementally updated map data, structural evolution and incremental data are integrated to generate dynamic risk map snapshots with evolution labels.
[0012] Preferably, the method further includes: Based on dynamic risk map snapshots, we define micro-scale early warning for sudden increases in individual risk, meso-scale early warning for expansion of subgraph size, and macro-scale early warning for abnormal risk density across the entire domain. Based on the defined three-level early warning scale, the early warning trigger conditions for each scale are calculated. When the trigger conditions are met, a structured early warning event is generated based on the original data in the dynamic risk map snapshot and associated with the original evidence chain supporting the early warning. Based on the structured early warning event and the original evidence chain, the output is a natural language interpretable summary containing risk type, behavioral characteristics and geographical location.
[0013] Preferably, the method further includes: Establish an intervention coding system and receive feedback from grassroots staff based on natural language interpretable summaries; Based on the feedback results, adjust the risk status of the corresponding nodes in the enhanced map; Optimize risk propagation parameters based on historical feedback data after adjusting risk status; Based on the system operation after optimizing the risk propagation parameters, an intervention effect evaluation report including the early warning verification rate and the handling rate is generated, and the overall risk baseline is updated.
[0014] On the other hand, this invention proposes a risk personnel analysis system based on big data of grassroots governance, including a data processing unit, a correlation graph construction unit, a risk scoring calculation unit, and a group risk reasoning unit; The data processing unit is used to collect multi-source heterogeneous grassroots governance data, perform entity alignment and semantic standardization, and generate a standardized personnel entity table with attributes and initial screening labels; The association graph construction unit is used to construct weighted association edges between personnel based on personnel attributes and initial screening tags in the standardized personnel entity table, combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, to form an initial risk personnel association graph; The risk scoring calculation unit is used to calculate the dynamic risk score of each person based on the node connection relationship in the initial risk personnel association graph, according to the multi-dimensional individual risk indicator system and combined with the time decay factor. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. The group risk reasoning unit is used to perform multiple rounds of risk propagation iteration from high-risk nodes to neighboring nodes based on the node risk scores and weighted association edges in the enhanced graph, until convergence, and to identify connected subgraphs with risk scores higher than the threshold and a size greater than a preset number as high-risk groups.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention achieves automatic expansion and identification of individual risks to group risks by constructing a three-element coupled association graph and risk propagation mechanism. By aligning entities and standardizing semantics of multi-source heterogeneous grassroots governance data, it breaks down data silos and lays the foundation for a comprehensive portrait of personnel. By integrating three dimensions—spatiotemporal co-occurrence, explicit social relationships, and behavioral pattern similarity—it constructs association edges, not only capturing explicit relationships but also discovering potential associations through implicit behavioral coupling. By performing risk propagation iterations on the association graph, it simulates the natural diffusion process of risk in the population network, allowing previously dispersed low-risk individuals to exhibit group risk characteristics under the aggregation effect, thereby automatically identifying connected subgraphs with high risk scores and significant scale. This enables grassroots governance workers to identify risk groups that have not yet caused actual harm but have formed a clustering trend in advance, allowing for intervention in the early stages of risk evolution. It shifts the focus of social governance from post-event handling to pre-event prevention, significantly improving the foresight and accuracy of grassroots risk prevention and control. Attached Figure Description
[0016] Figure 1 This is a flowchart of the risk personnel analysis method based on big data in grassroots governance according to the present invention; Figure 2 This is a block diagram of the risk personnel analysis system based on big data in grassroots governance according to the present invention. Detailed Implementation
[0017] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0018] This embodiment proposes, as follows: Figure 1 The method shown is a risk personnel analysis method based on big data of grassroots governance. This method expands the identification from individual risk to group risk, thereby discovering hidden groups, transmission chains, or cluster risks. Specifically, it includes: Collect multi-source heterogeneous grassroots governance data, and perform entity alignment and semantic standardization to generate a standardized personnel entity table with attributes and initial screening labels. Specifically, this includes: establishing a list of data sources and access protocols for identifying at-risk personnel, and defining a unified structured format for different data sources; performing cross-source data alignment and deduplication based on entity primary keys based on the data obtained from the access protocols, and constructing a multi-dimensional primary key mapping table; performing semantic standardization and initial screening of risk labels on the structured fields based on the alignment results of the multi-dimensional primary key mapping table; and integrating all aligned personnel entities and their attributes and initial screening labels into a standardized personnel entity table based on the results of semantic standardization and initial screening of risk labels.
[0019] It effectively breaks down data barriers among multiple departments such as public security, civil affairs, grid management, and health in grassroots governance, eliminating the problems of "same person, different name" or "different people, same name" caused by inconsistent data formats and chaotic identification systems, and significantly improving the accuracy and completeness of personnel entity identification. At the same time, through the initial screening of risk labels, it enables rapid filtering and risk focusing on massive amounts of grassroots data, providing a high-quality and highly consistent data foundation for subsequent correlation modeling and group risk identification, and avoiding misjudgments and waste of computing resources caused by data noise or redundancy.
[0020] Based on the personnel attributes and initial screening tags in the standardized personnel entity table, and combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relationships, and behavioral pattern similarity, a weighted association edge is constructed between personnel to form an initial risk personnel association graph; specifically including: Trajectory data is extracted from the standardized personnel entity table to calculate the spatiotemporal co-occurrence intensity among personnel, and pairs of personnel appearing in the same geographical area within the same time window are counted. Based on the calculated spatiotemporal co-occurrence intensity, explicit social relationship fields are extracted from the standardized personnel entity table to generate relationship strength weights. Combining spatiotemporal co-occurrence intensity and relationship strength weights, behavioral attributes are extracted from the standardized personnel entity table to calculate the similarity of behavioral patterns among personnel. Based on the three-dimensional calculation results of spatiotemporal co-occurrence intensity, relationship strength weights, and behavioral pattern similarity, weighted association edges are generated and an initial risk personnel association graph is constructed.
[0021] Breaking away from the limitations of traditional networks built solely on explicit relationships (such as kinship and cohabitation), this approach significantly enhances the ability to perceive implicit connections by integrating three dimensions: spatiotemporal behavior, social relationships, and behavioral patterns. It can not only identify risk co-occurrences in known relationships but also discover potential risk groups with no direct social connections but highly coordinated behaviors, effectively improving the coverage and semantic depth of the association graph. This lays a structural foundation for the subsequent accurate identification of complex group forms such as hidden gangs and risk transmission chains.
[0022] Based on the node connection relationships in the initial risk personnel association graph, and according to the multi-dimensional individual risk indicator system, combined with the time decay factor, the dynamic risk score of each person is calculated. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. The calculation of each individual's dynamic risk score includes: defining basic risk indicators, behavioral abnormality indicators, social abnormality indicators, and environmental risk indicators to construct an individual risk indicator system; normalizing and weighting each risk indicator in the individual risk indicator system to eliminate dimensional differences; calculating the initial individual risk score based on the normalized and weighted indicators and introducing a time decay factor to adjust the contribution of historical events; and embedding the dynamic risk score as a node attribute into the initial risk personnel association graph based on the adjusted risk score.
[0023] By integrating multi-dimensional indicators and using a time decay mechanism, the risk score can reflect both long-term accumulated risk and highlight the sensitivity of recent abnormal behavior, significantly improving the timeliness and discriminativeness of individual risk assessment; at the same time, the dynamic score is embedded in the graph as a node attribute.
[0024] Based on the node risk score and weighted association edges in the enhanced graph, the risk is propagated from high-risk nodes to neighboring nodes in multiple rounds until convergence. Connected subgraphs with risk scores higher than the threshold and a size greater than a preset number are identified as high-risk groups.
[0025] Specifically, the multi-round propagation iteration of execution risk from high-risk nodes to neighboring nodes includes: defining the calculation rules for the risk impact of a node on its neighbors, where the risk impact is directly proportional to the edge weight and inversely proportional to the current risk level of the target node; based on the risk impact calculation rules and the weighted association edges in the initial risk personnel association graph, summarizing the risk impact of all neighboring nodes on the current node; updating the risk score of the current node based on the summarized risk impact; and repeating the iteration based on the updated risk score until the risk change of all nodes is less than a preset convergence threshold.
[0026] It effectively simulates the real diffusion and aggregation process of risk in population networks, breaking through the limitations of relying solely on static individual scores to identify risks. It can aggregate scattered low-to-medium risk individuals into groups with overall high-risk characteristics through association structures. By introducing a dynamic influence mechanism that is positively correlated with edge weights and negatively correlated with the current risk level of target nodes, it not only strengthens the radiation effect of high-risk core nodes but also avoids the excessive accumulation of risks in already high-risk areas, making the propagation results more consistent with the evolution of grassroots risks. Ultimately, it achieves the automatic discovery of hidden, loosely organized but behaviorally coordinated risk groups, significantly improving the sensitivity and accuracy of group risk identification.
[0027] Identifying connected subgraphs with risk scores exceeding a threshold and a size greater than a preset number as high-risk groups includes: analyzing the topological characteristics of high-risk connected subgraphs and calculating the average path length and clustering coefficient of the subgraphs; based on the calculation results of the average path length and clustering coefficient of the subgraphs, when the average path length of the subgraphs is large and the clustering coefficient is low, it is determined to be a propagation chain; based on the calculation results of the clustering coefficient and edge density of the subgraphs, when the clustering coefficient of the subgraphs is high and the edge density is large, it is determined to be a hidden group; based on the geographical coordinate distribution of the subgraph nodes, when the nodes are concentrated in the same geographical area, it is determined to be a cluster risk.
[0028] It enables refined classification and identification of high-risk groups, not only determining whether "risk groups exist" but also accurately distinguishing their organizational forms and behavioral characteristics. Through joint analysis of topological structure and geographical distribution, it can automatically identify risk types with different governance needs—such as linearly spreading transmission chains, tightly coupled hidden groups, or regionally concentrated cluster risks—thereby providing differentiated and precise handling suggestions for grassroots staff, significantly improving the pertinence of risk response strategies and the efficiency of governance resource allocation.
[0029] In one embodiment, the method further includes: setting a tracking window and a member activity monitoring mechanism based on the high-risk groups identified in the enhanced graph, and continuously monitoring the activity status of each member in the subgraph; detecting structural changes in the high-risk group and identifying member additions, subtractions, splits, or mergers based on the member activity monitoring results; performing incremental updates on the initial risk personnel association graph based on the identified structural change events and combined with newly added grassroots governance data; and generating a dynamic risk graph snapshot with evolution labels by fusing structural evolution and incremental data based on the structural change events and the incrementally updated graph data.
[0030] This system enables continuous tracking and dynamic perception of the lifecycle of at-risk groups, effectively overcoming the drawbacks of traditional static maps that are "built once and become invalid in the long term." By monitoring activity and identifying structural changes, it can promptly capture the evolutionary trends of at-risk groups, such as expansion, disintegration, or reorganization, effectively avoiding misjudgments or missed controls due to information lag. At the same time, relying on the incremental update mechanism, it maintains the timeliness and integrity of the map while ensuring system response efficiency, upgrading grassroots governance from "snapshot-style assessment" to "continuous monitoring," significantly improving the adaptability to dynamic risks and the accuracy of intervention timing.
[0031] In one embodiment, the method further includes: defining micro-scale early warning for sudden increases in individual risk, meso-scale early warning for expansion of subgraph size, and macro-scale early warning for abnormal risk density across the entire region, based on dynamic risk map snapshots; calculating early warning triggering conditions for each scale according to the defined three-level early warning scales; when the triggering conditions are met, generating structured early warning events based on the original data in the dynamic risk map snapshots and associating them with original evidence chains supporting the early warnings; and outputting a natural language interpretable summary containing risk type, behavioral characteristics, and geographical location based on the structured early warning events and the original evidence chains.
[0032] A multi-scale risk early warning system covering individuals, groups, and regions has been constructed, enabling governance entities at different levels to respond to risk signals of corresponding granularity as needed, avoiding information overload or underreporting of key risks. By automatically linking early warnings with the original evidence chain and generating semantically clear and complete natural language summaries, the understanding threshold for grassroots personnel without technical backgrounds has been significantly reduced, and the credibility and operability of early warning results have been improved.
[0033] In one embodiment, the method further includes: establishing an intervention measure coding system and receiving feedback results from grassroots staff based on natural language interpretable summaries; adjusting the risk status of corresponding nodes in the enhanced map according to the content of the feedback results; optimizing risk propagation parameters based on historical feedback data after adjusting the risk status; and generating an intervention effect evaluation report including the early warning verification rate and the handling rate based on the system operation after optimizing the risk propagation parameters, and updating the overall risk baseline.
[0034] By feeding back the results of human intervention into the map status and propagation parameters, false alarms and false negatives in the system are effectively suppressed, and the accuracy and adaptability of risk assessment are improved. At the same time, based on quantifiable intervention effect evaluation, dynamic updates of the risk baseline across the entire domain are achieved.
[0035] On the other hand, this invention proposes a risk personnel analysis system based on big data in grassroots governance, such as... Figure 2 As shown, it includes a data processing unit, a correlation graph construction unit, a risk score calculation unit, a group risk reasoning unit, a dynamic evolution tracking unit, an early warning generation unit, and a feedback processing unit; The data processing unit is used to collect multi-source heterogeneous grassroots governance data, perform entity alignment and semantic standardization, and generate a standardized personnel entity table with attributes and initial screening labels; The association graph construction unit is used to construct weighted association edges between personnel based on personnel attributes and initial screening tags in the standardized personnel entity table, combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, to form an initial risk personnel association graph; The risk scoring calculation unit is used to calculate the dynamic risk score of each person based on the node connection relationship in the initial risk personnel association graph, according to the multi-dimensional individual risk indicator system and combined with the time decay factor. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. The group risk reasoning unit is used to perform multiple rounds of risk propagation iteration from high-risk nodes to neighboring nodes based on node risk scores and weighted association edges in the enhanced graph, until convergence, and identify connected subgraphs with risk scores higher than the threshold and a size greater than a preset number as high-risk groups. The dynamic evolution tracking unit is used to identify high-risk groups in the enhanced map, set up tracking windows and member activity monitoring mechanisms, detect structural changes in high-risk groups, perform incremental map updates based on newly added grassroots governance data, and generate dynamic risk map snapshots with evolution labels. The early warning generation unit is used to define three levels of early warning scales based on dynamic risk map snapshots, calculate the early warning triggering conditions at each scale, generate structured early warning events and associate them with the original evidence chain, and output natural language interpretable summaries. The feedback processing unit is used to establish an intervention measure coding system, receive feedback from the grassroots level, analyze the feedback content to adjust the risk status of nodes, optimize risk propagation parameters based on historical feedback, generate an intervention effect evaluation report, and update the risk baseline.
[0036] In addition, each unit in the above system is also used to implement other steps of the risk personnel analysis method based on grassroots governance big data during execution, as follows: Step 1: Unified Collection and Structured Processing of Multi-Source Heterogeneous Grassroots Governance Data In grassroots governance scenarios, the data sources involved are diverse, including but not limited to public security population registration, community grid worker visit records, entry and exit registration of key locations, communication base station trajectories, water, electricity, and gas usage records, social security and medical insurance reimbursement data, online public opinion information, and emergency reporting records. These data vary significantly in format, granularity, update frequency, and semantic expression; directly using them for risk analysis will lead to information fragmentation or misjudgment. Therefore, it is essential to first collect and structure this multi-source heterogeneous data in a unified manner to lay the data foundation for subsequent association mapping.
[0037] Step 1.1: Establish a list of data sources and access protocols for identifying at-risk individuals. In this step, based on grassroots governance scenarios, eight core data sources closely related to personnel behavior, activity trajectories, and social relationships are identified, and standardized access protocols are defined for each data source. For example, for visit records reported by community grid workers, JSON format is used to define fields including personnel ID, address, family members, description of recent abnormal behavior, and reporting time; for communication base station trajectory data, CSV format is used, including fields such as device identifier, base station ID, timestamp, and signal strength. Through the pre-defined access protocols, it is ensured that data from different sources has a unified syntax structure before entering the system, avoiding subsequent processing failures due to formatting issues. This step provides structured input for subsequent data cleaning and entity alignment.
[0038] Step 1.2: Perform cross-source data alignment and deduplication based on entity primary keys. Following the structured data defined in step 1.1, this step focuses on aligning the same person's information from different data sources. Since different data sources identify people differently (e.g., ID card number, mobile phone number, device ID, grid number, etc.), a multidimensional primary key mapping table needs to be constructed. For example, if a person is identified by ID card number A in the public security system, appears as mobile phone number B in communication data, and is registered as household number C in water and electricity data, then a mapping relationship of A, B, and C is established through cross-validation (e.g., name + date of birth + address). During the alignment process, Jaccard similarity is used to measure the degree of matching of text fields (e.g., address description): ; in Given the word set after word segmentation of two strings. ( When a preset threshold (e.g., 0.75) is reached, the records are considered to be the same entity. After alignment, duplicate records are merged, retaining the latest or most complete field value. This step ensures that each real-world individual corresponds to only one unique identifier in the system, avoiding redundant nodes in the graph.
[0039] Step 1.3: Perform semantic normalization and initial risk labeling on structured fields. After entity alignment is completed in step 1.2, the attribute fields of each entity are semantically standardized. For example, the "address" field is uniformly converted to the standard administrative division code (such as GB / T2260); the "occupation" field is mapped to the National Occupational Classification Code; and the "abnormal behavior description" is automatically labeled with preliminary risk tags through keyword matching (such as "gathering in crowds," "entering and leaving late at night," and "frequent changes of residence"). This process introduces a rule engine, setting up a tag rule base based on grassroots governance experience. For example, if a person changes their address more than 3 times within 30 days and fails to report each time to the public security system, they are labeled with "abnormal mobility." Although these preliminary tags are not directly used for the final risk assessment, they provide a basis for subsequent graph edge weight calculations. This step improves the semantic consistency of the data, allowing attributes from different sources to participate in calculations within the same semantic space.
[0040] Step 1.4: Generate a standardized personnel entity table with attributes and initial screening tags. Based on the semantic standardization results of step 1.3, this step integrates all aligned personnel entities, their attributes, and initial screening labels into a standardized personnel entity table. Each row in the table corresponds to a unique personnel ID, and the columns include: basic attributes (name, gender, age, standard address code), behavioral attributes (number of active base stations in the past 30 days, water and electricity consumption fluctuation coefficient), relationship attributes (list of family member IDs, list of cohabitant IDs), and a set of initial screening risk labels (such as {"abnormal mobility", "frequent nighttime activity"}). This table serves as a node pool for subsequent association graph construction, ensuring that each node has rich contextual information. This step completes the transformation from raw heterogeneous data to structured, semantically consistent, and initially labeled entity representations, providing high-quality node input for the construction of graph edges.
[0041] Step 2: Constructing association edges based on the spatiotemporal-relational-behavior ternary coupling After obtaining the standardized personnel entity table, it is necessary to construct the association edges between personnel to form an initial association graph. Traditional methods rely only on explicit relationships (such as relatives and cohabitation), making it difficult to discover implicit connections. This embodiment proposes a three-element coupling association edge construction method, which integrates three dimensions—spatiotemporal co-occurrence, social relations, and behavioral similarity—to quantify the strength of potential associations between personnel.
[0042] Step 2.1: Extract spatiotemporal co-occurrence events and calculate co-occurrence intensity Following the personnel entity table generated in step 1.4, this step extracts spatiotemporal co-occurrence events from trajectory-based data (such as base station positioning and location QR code scanning records). Spatiotemporal co-occurrence is defined as: two individuals appearing in the same geographical area (e.g., within a radius of 200 meters) within the same time window (e.g., ±15 minutes). For each pair of individuals... Count the number of times it co-occurs during the observation period. And calculate the co-occurrence intensity : ; in For personnel The total number of activity records. This formula normalizes co-occurrence frequency to avoid false high-frequency co-occurrences from highly active individuals. For example, if 10 activities per day If the activity occurs twice a day, and both occur twice in total, then... ,reflect A high proportion of the activities were related to Overlap. This step transforms the original trajectory data into quantified co-occurrence relationships, providing a spatiotemporal dimension for the associated edges.
[0043] Step 2.2: Integrate explicit social relationships to generate relationship strength weights Building upon the spatiotemporal co-occurrence intensity obtained in step 2.1, this step introduces explicit social relationships (from the fields of family members, cohabitants, emergency contacts, etc. in step 1.4) as a supplement. For each pair of individuals with explicit relationships... Assign a base value to the relationship strength If there is no explicit relationship, then Furthermore, if the two appear simultaneously in multiple relationship scenarios (such as being both relatives and cohabitants), then... The weighting can be increased to 1.2 or 1.5 to reflect the tightness of the relationship. This step transforms structured relationship data into numerical weights, which complement spatiotemporal co-occurrence: explicit relationships provide high-confidence connections, while spatiotemporal co-occurrence captures potential interactions.
[0044] Step 2.3: Calculate behavioral pattern similarity as the third dimension of association. Following the relationship weighting in step 2.2, this step further explores associations from the perspective of behavioral attributes. Behavioral patterns include: activity time distribution (e.g., percentage of nighttime activities), venue type preferences (e.g., frequent visits to internet cafes and bars), and resource usage patterns (e.g., sudden increases in water and electricity consumption). For personnel... Represent its behavior vector as ,in For the first Normalized values of behavioral indicators. Cosine similarity is used to calculate behavioral similarity: ; This value ranges from 0 to 1, with higher values indicating more similar behavioral patterns. For example, if two people frequently appear at the same entertainment venue between 2 AM and 5 AM, and their water and electricity consumption both show periodic sharp drops, then... Close to 1.
[0045] Step 2.4: Integrate the three dimensions to generate weighted relational edges and construct the initial graph. Based on the three types of strength values obtained in steps 2.1 to 2.3, this step generates the final associated edge weights through weighted fusion. : ; in For the preset weighting coefficients, satisfy In grassroots governance scenarios, [the following is set up] To balance the contributions of time, space, relationships, and behavior. Only when ( An edge is added to the graph only when a threshold (e.g., 0.25) is met. Finally, using the personnel entities from step 1.4 as nodes and the weighted edges generated in this step as connections, an initial risk personnel association graph is constructed. ,in For a set of nodes, Edge set, This is an edge weight mapping. This graph not only contains explicit relationships but also integrates spatiotemporal and behavioral clues, providing a structural foundation for subsequent risk propagation reasoning.
[0046] Step 3: Dynamic Calculation of Individual Risk Scores and Enhancement of Node Attributes While the initial graph contains rich associations, the nodes themselves lack quantifiable risk levels. This step aims to calculate a dynamic risk score for each individual node and embed this score as a node attribute into the graph, providing input for subsequent group risk inference.
[0047] Step 3.1: Define a multi-dimensional individual risk indicator system Following the initial risk map constructed in step 2.4, this step first establishes an individual risk indicator system. This system includes four categories of indicators: (1) basic risk indicators, such as whether the individual has a criminal record or is a key target for control; (2) behavioral abnormality indicators, such as the frequency of nighttime activities and the rate of location switching; (3) social abnormality indicators, such as the number of high-risk individuals associated with the individual and the average weight of the associated edges; and (4) environmental risk indicators, such as the density of recent events in the individual's community and the number of high-risk locations in the surrounding area. Each category of indicator has several specific measurement items, for example, "location switching rate" is defined as the number of times an individual enters or exits different types of locations per unit time. This indicator system covers both internal and external risk factors for individuals, ensuring the comprehensiveness of the scoring.
[0048] Step 3.2: Normalize and assign weights to each risk indicator. After defining the indicators in step 3.1, this step normalizes the original indicator values to eliminate dimensional differences. For positive indicators (higher values indicate higher risk), Min-Max normalization is used: ; For negative indicators (such as stable length of residence), the values are inverted and then normalized. Subsequently, based on grassroots governance experience, weights are assigned to the four types of indicators. ,satisfy For example, suppose This emphasizes the importance of behavioral and social abnormalities. This step ensures that different indicators are weighted and integrated under a unified scale, preventing any one type of indicator from dominating the score due to its large numerical range.
[0049] Step 3.3: Calculate the initial individual risk score and introduce the time decay factor. Based on the normalized indicators and weights from step 3.2, the personnel... Calculate the initial risk score : ; in for In the Normalized scores on similar indicators. To further reflect the timeliness of risk, a time decay factor is introduced. ,in The number of days since the event. This is the attenuation coefficient (e.g., 0.1). For historical event indicators (e.g., criminal records, records of abnormal behavior), their contribution is calculated as follows: Attenuation. For example, aberrant behavior from 30 days ago contributes to [a certain percentage]. The contribution three days ago was This step allows the risk score to dynamically reflect recent behavior, avoiding the long-term dominance of historical records.
[0050] Step 3.4: Embed the dynamic risk score as a node attribute into the association graph. Obtain the dynamic risk score in step 3.3 Then, this step attaches it as a core attribute to the graph node. Above. Simultaneously, original attributes (such as address and occupation) and initial screening tags are retained to form an enhanced node representation. The updated atlas is denoted as ,in To enhance the node set, this graph not only includes structural connections but also quantifies node risk, providing necessary input for subsequent risk propagation calculations based on graph structures.
[0051] Step 4: Group Risk Reasoning and Subgraph Recognition Based on Risk Propagation Dynamics After obtaining the enhanced graph with risk scores, it is necessary to infer group-level risks from individual risks. This step simulates the propagation process of risk in the interconnected network, identifies high-risk subgraphs, and corresponds to hidden groups, propagation chains, or clustered risk groups.
[0052] Step 4.1: Define the risk propagation rule and the neighborhood influence function The enhanced map following step 3.4 This step establishes risk propagation rules. It assumes that risk can spread from high-risk nodes to neighboring nodes via associated edges, with the propagation strength influenced by both edge weights and the node's current risk level. Define the nodes. To the neighbors Risk impact for: ; in For edge weights, Risk to the source node This indicates the risk-receiving capacity of the target node that is not yet saturated. The formula reflects the intuition that "high risk spreads to low-risk areas through strong connections" and avoids risk values exceeding 1.
[0053] Step 4.2: Perform multiple rounds of risk propagation iterations until convergence. Based on the propagation rules in step 4.1, this step synchronously updates the risk scores for all nodes in the graph. Let the... After round of iterations, the node The risk is Then the first The cycle is updated to: ; in for The neighborhood group, Set a learning rate (e.g., 0.1) to control the propagation speed. Iterate continuously until the risk change at all nodes is less than the convergence threshold. (e.g., 0.001). This process simulates the natural spread of risk in social networks, increasing the risk value of individuals who were originally low-risk but closely connected to high-risk groups, thus revealing potentially infected individuals.
[0054] Step 4.3: Identify high-risk connected subgraphs based on the converged risk distribution After the risk propagation converges in step 4.2, the steady-state risk score of each node is obtained. This step sets a risk threshold. (e.g., 0.6), all Nodes are marked as high-risk nodes. Then, the largest connected subgraph consisting of these high-risk nodes and their connecting edges is extracted from the graph. If the subgraph size (number of nodes) exceeds a preset minimum size... (e.g., 3 people) are considered as a candidate risk group.
[0055] Step 4.4: Determine the type and summarize the risk characteristics of high-risk subgraphs. Following the high-risk subgraphs identified in step 4.3, this step further determines their risk type. If the subgraph exhibits a chain-like structure (long average path length, low clustering coefficient), it is identified as a propagation chain; if it exhibits a cluster-like structure (high clustering coefficient, high edge density), it is identified as a hidden group; if the subgraph nodes are concentrated in the same geographical area (e.g., the first 6 digits of the standard address code are the same), it is identified as a clustered risk. Simultaneously, a risk characteristic summary is generated, including: core high-risk nodes, main correlation dimensions (e.g., predominantly spatiotemporal co-occurrence), and dominant risk behavior (e.g., nighttime gatherings). This summary provides interpretable judgment criteria for grassroots staff, supporting the formulation of subsequent intervention measures.
[0056] Step 5: Tracking the dynamic evolution of at-risk groups and incrementally updating the map After identifying high-risk sub-graphs, it's crucial to understand that at-risk groups in grassroots governance scenarios are not static; their members, structures, and behavioral patterns continuously evolve over time. Relying solely on a single point in time for analysis can easily lead to overlooking emerging groups or misjudging disbanded communities. Therefore, this step focuses on dynamically tracking identified at-risk groups and incrementally updating the associated graphs based on new data to ensure the timeliness and continuity of risk identification.
[0057] Step 5.1: Set up a risk group tracking window and a member activity monitoring mechanism Following the high-risk sub-graphs and their risk characteristic summaries generated in step 4.4, this step first assigns a unique tracking identifier to each sub-graph and sets a dynamic monitoring window (e.g., 7 days as a tracking cycle). Within this window, the activity status of each member in the sub-graph is continuously monitored, including: whether there are still trajectory reports, whether new abnormal behaviors have occurred, and whether new connections have been established with other high-risk nodes. Activity level is assessed using a comprehensive indicator. measure: ; in For personnel In the current window The number of newly added data records, Its historical average number of records, For indicator functions, when The current risk score is still above the threshold. The value is 1 if the condition is met, otherwise it is 0. If a member has two consecutive windows If the activity level is low, it is marked as "low activity" and may have become isolated from the group.
[0058] Step 5.2: Detect subgraph structure changes and identify member addition / removal events. Based on the activity monitoring results in step 5.1, this step further analyzes the dynamic changes in the internal structure of the subgraph. Specifically, this includes: (1) Member exit: The original member is continuously inactive and there are no new edges connecting them; (2) Member addition: A new node is added through a strongly associated edge ( (2) Connecting to at least two existing members in the subgraph; (3) Subgraph splitting: The original connected subgraph is broken into two or more connected components due to the exit of a key node; (4) Subgraph merging: Multiple strong edges connect two independent high-risk subgraphs. Such events are identified by connected component detection and edge density change rate in graph theory. For example, subgraph In the window The edge density is: ; like If the value is 0.15, a structural anomaly alarm will be triggered.
[0059] Step 5.3: Perform incremental map construction based on newly added grassroots data While identifying structural changes in step 5.2, the system continuously receives new data from grassroots governance channels (such as newly reported visit records, new base station trajectories, and QR code scanning information for new locations). This step performs lightweight processing on this incremental data: entity alignment and edge weight calculation are only performed on the parts involving existing nodes or potential new nodes, avoiding a full reconstruction. Specifically, if the new data includes personnel... and For co-occurrence records, only update Or create a new edge (if) First time exceeding the threshold );like For new personnel, a node is created for them and linked to the existing graph (e.g., by matching by address or mobile phone number). This incremental update strategy significantly reduces computational overhead and adapts to the reality of limited resources in grassroots systems.
[0060] Step 5.4: Generate a dynamic risk map snapshot by fusing structural evolution and incremental data. Based on the structural change detection results in step 5.2 and the incremental map update in step 5.3, this step generates a dynamic risk map snapshot at the end of each tracking window. This snapshot not only includes all current nodes and edges, but also adds evolutionary tags, such as stable groups, expanding propagation chains, and disbanding clusters. It also preserves a sequence of historical snapshots. This dynamic map snapshot is used for subsequent multi-time series analysis. It provides frontline staff with a "lifecycle view of at-risk groups," enabling them to determine whether a group is emerging, growing, splitting, or disappearing, thereby developing differentiated response measures.
[0061] Step Six: Generation of Multi-Scale Risk Warnings and Output of Interpretable Summary While dynamic risk maps can reflect population evolution, frontline staff require intuitive and actionable early warning information. This step aims to transform the complex structure in the map into multi-scale early warning signals and generate semantically interpretable summaries to facilitate understanding and decision-making by non-experts.
[0062] Step 6.1: Define a three-level scale system for risk warning. Following the snapshot of the dynamic risk map generated in step 5.4, this step establishes a three-level early warning scale: (1) Micro scale – for individuals, warning of a sudden increase in their risk score or their first entry into a high-risk submap; (2) Meso scale – for submaps, warning of their expansion in scale, structural changes, or cross-regional migration; (3) Macro scale – for the entire region, warning of a surge in the total number of high-risk submaps, the concentrated appearance of specific types of gangs, or an abnormal increase in the risk density of a certain community. Each scale corresponds to a different response level: micro scale is verified by grid members, meso scale is intervened by the street comprehensive management center, and macro scale is coordinated by the district-level command center.
[0063] Step 6.2: Calculate the early warning triggering conditions at each scale based on map features. Within the scale framework of step 6.1, this step sets quantitative trigger conditions for each type of warning. For example, the micro-level warning trigger condition is: an individual's risk score increases by more than 0.4 within 7 days, or the individual is included in a high-risk submap for the first time; the meso-level warning conditions include: a week-on-week increase of ≥50% in the number of submap nodes, or a shift of more than 2 kilometers in the submap's centroid (weighted average of node geographic coordinates); the macro-level warning conditions include: a week-on-week increase of more than 20% in the number of high-risk submaps in the entire region, or a proportion of "hidden gangs" in a certain street exceeding 60% of the total number of submaps. These conditions are all based on computable features in the map, ensuring that the warnings are objective and reproducible.
[0064] Step 6.3: Generate structured early warning events and associate them with the original chain of evidence. When the triggering conditions in step 6.2 are met, this step automatically generates a structured early warning event object, including: early warning level, involved personnel / subgraph ID, triggering reason, time window, and geographical location range. More importantly, it automatically associates the original evidence chain supporting the early warning, such as "Personnel A and B co-occurred 8 times in the past 5 days (Base Station ID: X123, Y456)" and "Subgraph S added member C, whose behavior vector has a similarity of 0.82 with the core members." The evidence chain directly references the original records processed in steps 1 to 5, ensuring that the early warning is traceable and verifiable, avoiding misjudgment disputes caused by "black box judgment."
[0065] Step 6.4: Output a natural language interpretable summary for human evaluation. Based on the structured early warning events in step 6.3, this step further generates natural language summaries. For example: "A three-person gang located in XX Street has recently expanded to five members. The new members frequently visit the same internet cafe in the early morning and exhibit high-density spatiotemporal co-occurrence with the original members, suggesting a suspected new type of cluster risk." The summary integrates risk type, behavioral characteristics, evolutionary trends, and geographical location, using concise language and clear logic. This summary is directly pushed to the grassroots governance platform, allowing grid workers, community police officers, and other frontline personnel to quickly grasp the situation and decide whether to conduct on-site verification, deployment, or reporting. This step bridges the gap between data analysis and human decision-making, transforming complex graph reasoning results into executable instructions.
[0066] Step 7: Feedback on Intervention Measures and Closed-Loop Correction of Risk Status After an early warning is issued, grassroots staff will take intervention measures such as on-site verification, interviews and education, and key monitoring. The effectiveness of these measures needs to be fed back to the system to correct the risk score and graph structure, forming a closed loop of analysis, early warning, intervention, and feedback to avoid continuous false alarms or omissions in the system.
[0067] Step 7.1: Establish an intervention coding system and feedback interface Following the early warning summary in step 6.4, this step designs an intervention measure coding system covering common grassroots governance actions, such as "I01 - No abnormalities found during on-site verification," "I02 - Included in the key monitoring list," "I03 - Joint interview with the police," and "I04 - Risk resolved." Simultaneously, a feedback interface is set up on grassroots work terminals (such as mobile apps). After completing the intervention, staff select the corresponding code and fill in a brief result (e.g., "The person involved has moved out, no abnormal behavior"). This feedback data is transmitted back to the analysis system through a secure channel as a basis for risk status correction.
[0068] Step 7.2: Parse the feedback content and map it to the status adjustment of the graph nodes. After receiving the feedback data from step 7.1, this step parses its semantics and converts it into graph operation instructions. For example, if the feedback is "I04 - Risk Relief", the risk score of the corresponding personnel node will be forcibly set to 0.1, and its risk propagation capability for the next 7 days will be frozen (i.e., the risk propagation capability will be set in the propagation formula). If the feedback is "I02 - Included in key monitoring", then its basic risk weight will be increased. The threshold is reduced to 0.4, and the tracking window is shortened to 3 days. For group interventions (such as the entire gang being dismantled), the entire subgraph is marked as "disposed of" and its evolution tracking is suspended.
[0069] Step 7.3: Optimize risk propagation parameters based on historical feedback In addition to real-time correction, this step also utilizes historical feedback data to optimize internal system parameters. For example, it analyzes cases where "no abnormalities were reported after an early warning" to identify common characteristics (such as frequent occurrences at night without behavioral anomalies), and adjusts the edge weight coefficients for such scenarios accordingly. Conversely, if a certain type of behavior (such as frequently changing SIM cards) is repeatedly confirmed as high-risk in the feedback, its weight in the behavior vector is increased. Parameter optimization is achieved through online learning, eliminating the need for full retraining.
[0070] Step 7.4: Generate an intervention effectiveness assessment report and update the risk baseline. After completing state correction and parameter optimization in steps 7.2 and 7.3, this step generates an intervention effectiveness evaluation report, which includes: the total number of warnings in this period, the verification rate, the false alarm rate, and the handling rate of high-risk subgraphs. Simultaneously, based on successful intervention cases, the overall risk baseline is updated—for example, if a certain type of gang does not relapse after multiple interventions, its characteristic pattern is added to the "low-risk template library," and similar structures will no longer trigger warnings in the future. This report is used by higher-level departments to evaluate the effectiveness of grassroots governance and provides more accurate prior knowledge for the next round of risk analysis.
[0071] Step 8: Integrating System Deployment Architecture with Grassroots Governance Business Processes This step describes how to deploy this method as a working system and seamlessly integrate it with existing business processes to ensure that technical capabilities translate into effective governance.
[0072] Step 8.1: Design a layered distributed system architecture to adapt to the underlying IT environment. Building upon the closed-loop capability of step 7.4, this step constructs a layered system architecture: (1) Edge layer – deployed on street or community servers, responsible for local data collection, preliminary alignment, and real-time early warning push, meeting the security requirement that data does not leave the domain; (2) Regional layer – deployed in district-level data centers, performing map construction, risk propagation calculation, and cross-community group identification; (3) Central layer – located on the city-level platform, summarizing the risk situation of each district, generating macro-level early warnings, and issuing parameter update packages. Incremental data and feedback results are synchronized between layers through encrypted message queues to avoid full transmission. This architecture takes into account data security, computing efficiency, and governance levels, and adapts to the differences in grassroots IT infrastructure.
[0073] Step 8.2: Integrate with existing grassroots governance information systems to achieve automatic data transfer. Building upon the architecture established in step 8.1, this step integrates the system with existing business systems such as the public security population database, the grid management platform, and the key location monitoring system. Data such as personnel registration, event reporting, and location scanning are automatically retrieved via standard APIs or database views, eliminating the need for manual data entry. Simultaneously, early warning summaries and intervention feedback are automatically pushed to the corresponding personnel's to-do lists through the workflow engine, embedding themselves into their daily operation interfaces. For example, when a grid worker logs into the work app, the homepage displays "3 individuals at risk need to be verified," and clicking on it directly displays the evidence chain and handling suggestions.
[0074] Step 8.3: Configure role permissions and risk information tiered disclosure strategy Given the sensitivity of risk information, this step configures permissions based on grassroots governance roles: grid workers can only see micro-level warnings within their assigned grid; street-level comprehensive management cadres can view meso-level sub-maps and evolution trends; and district-level commanders can access the overall macro-level situation. Simultaneously, highly sensitive information (such as communication trajectory details) is anonymized and only temporarily decrypted upon approval when necessary. This tiered disclosure strategy ensures effective information utilization while complying with personal information protection regulations, enhancing system compliance and acceptability.
[0075] Step 8.4: Establish a routine operation mechanism and personnel operation training system Finally, this step drives the system from technical deployment to routine operation. It involves developing operating procedures for the risk personnel analysis system, clarifying requirements such as data update frequency, early warning response time limits, and feedback entry standards; organizing tiered training to ensure grid workers master the early warning verification process and technical personnel understand the parameter adjustment logic; and establishing an operations and maintenance team responsible for system monitoring, troubleshooting, and version iteration.
[0076] In summary, by collecting multi-source heterogeneous grassroots governance data, performing entity alignment and semantic standardization, a standardized personnel entity table is generated. Based on personnel attributes and tags, and combining the three-element coupling of spatiotemporal co-occurrence, explicit social relationships, and behavioral pattern similarity, weighted association edges between personnel are constructed to form an initial risk personnel association graph. Based on the graph node connection relationships and multi-dimensional individual risk indicators, combined with the time decay factor, a dynamic risk score is calculated and embedded into the graph. Through multiple rounds of risk propagation iteration from high-risk nodes to neighboring nodes, high-risk connected subgraphs are identified.
[0077] This invention enables automatic expansion and identification from individual risk to group risk, allowing for early detection of hidden groups, transmission chains, or cluster risks, and intervention in the early stages of risk evolution.
[0078] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A risk personnel analysis method based on big data of grassroots governance, characterized in that, include: Collect multi-source heterogeneous grassroots governance data, perform entity alignment and semantic standardization, and generate a standardized personnel entity table with attributes and initial screening labels; Based on the personnel attributes and initial screening tags in the standardized personnel entity table, and combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, a weighted association edge between personnel is constructed to form an initial risk personnel association graph; Based on the node connection relationships in the initial risk personnel association graph, and according to the multi-dimensional individual risk indicator system, combined with the time decay factor, the dynamic risk score of each person is calculated. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. Based on the node risk score and weighted association edges in the enhanced graph, the risk is propagated from high-risk nodes to neighboring nodes in multiple rounds until convergence. Connected subgraphs with risk scores higher than the threshold and a size greater than a preset number are identified as high-risk groups.
2. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The generated standardized personnel entity table with attributes and initial screening labels includes: Establish a list of data sources and access protocols for identifying at-risk individuals, and define a unified structured format for different data sources; Based on the data obtained from the access protocol, perform cross-source data alignment and deduplication based on the entity primary key, and construct a multi-dimensional primary key mapping table; Based on the alignment results of the multidimensional primary key mapping table, the structured fields are semantically standardized and initially screened for risk labels; Based on the results of semantic standardization and initial screening of risk labels, all aligned personnel entities and their attributes, along with the initial screening labels, are integrated into a standardized personnel entity table.
3. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The formation of the initial risk personnel association map includes: Trajectory data is extracted from the standardized personnel entity table to calculate the spatiotemporal co-occurrence intensity among personnel and to count pairs of personnel appearing in the same geographical area within the same time window. Based on the calculated spatiotemporal co-occurrence intensity, explicit social relationship fields are extracted from the standardized personnel entity table to generate relationship strength weights; By combining the spatiotemporal co-occurrence strength and relationship strength weights, behavioral attributes are extracted from the standardized personnel entity table to calculate the similarity of behavioral patterns among personnel; Based on the calculation results of the three-dimensional dimensions of spatiotemporal co-occurrence strength, relationship strength weight, and behavioral pattern similarity, weighted association edges are generated and an initial risk personnel association graph is constructed.
4. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The calculation of each individual's dynamic risk score includes: Define basic risk indicators, behavioral abnormality indicators, social abnormality indicators, and environmental risk indicators to construct an individual risk indicator system; Normalize and assign weights to each risk indicator in the individual risk indicator system to eliminate differences in dimensionality; Based on the normalized and weighted indicators, the initial individual risk score is calculated and the contribution of historical events is adjusted by introducing a time decay factor. Based on the adjusted risk score, the dynamic risk score is embedded as a node attribute into the initial risk personnel association graph.
5. The risk personnel analysis method based on grassroots governance big data according to claim 1, characterized in that, The multi-round propagation iteration of execution risk from high-risk nodes to neighboring nodes includes: Define the rules for calculating the risk impact of a node on its neighbors. The risk impact is directly proportional to the edge weight and inversely proportional to the current risk level of the target node. Based on the risk impact calculation rules, and using the weighted association edges in the initial risk personnel association graph, the risk impact of all neighboring nodes on the current node is summarized. Based on the summarized risk impact, update the risk score of the current node. Based on the updated risk score, repeat the iteration until the risk change of all nodes is less than the preset convergence threshold.
6. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The connected subgraphs whose risk scores are higher than a threshold and whose size exceeds a preset number are identified as high-risk groups, including: Analyze the topological characteristics of high-risk connected subgraphs and calculate the average path length and clustering coefficient of the subgraphs; Based on the calculation results of the average path length of the subgraph and the clustering coefficient, when the average path length of the subgraph is large and the clustering coefficient is low, it is determined to be a propagation chain. Based on the calculation results of subgraph clustering coefficient and edge density, when the subgraph clustering coefficient is high and the edge density is large, it is identified as a hidden group. Based on the distribution of geographic coordinates of subgraph nodes, when nodes are concentrated in the same geographic area, it is determined to be a cluster risk.
7. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The method further includes: Based on the high-risk groups identified in the enhanced graph, a tracking window and a member activity monitoring mechanism are set up to continuously monitor the activity status of each member in the subgraph; Based on member activity monitoring results, detect structural changes in high-risk groups and identify events such as member increases, decreases, splits, or mergers; Based on the identified structural change events, combined with newly added grassroots governance data, the initial at-risk personnel association map is incrementally updated; Based on structural change events and incrementally updated map data, structural evolution and incremental data are integrated to generate dynamic risk map snapshots with evolution labels.
8. The risk personnel analysis method based on grassroots governance big data according to claim 1, characterized in that, The method further includes: Based on dynamic risk map snapshots, we define micro-scale early warning for sudden increases in individual risk, meso-scale early warning for expansion of subgraph size, and macro-scale early warning for abnormal risk density across the entire domain. Based on the defined three-level early warning scale, the early warning trigger conditions for each scale are calculated. When the trigger conditions are met, a structured early warning event is generated based on the original data in the dynamic risk map snapshot and associated with the original evidence chain supporting the early warning. Based on the structured early warning event and the original evidence chain, the output is a natural language interpretable summary containing risk type, behavioral characteristics and geographical location.
9. The risk personnel analysis method based on big data of grassroots governance according to claim 1, characterized in that, The method further includes: Establish an intervention coding system and receive feedback from grassroots staff based on natural language interpretable summaries; Based on the feedback results, adjust the risk status of the corresponding nodes in the enhanced map; Optimize risk propagation parameters based on historical feedback data after adjusting risk status; Based on the system operation after optimizing the risk propagation parameters, an intervention effect evaluation report including the early warning verification rate and the handling rate is generated, and the overall risk baseline is updated.
10. A risk personnel analysis system based on grassroots governance big data for implementing the method as described in any one of claims 1-9, characterized in that, It includes a data processing unit, a correlation graph construction unit, a risk scoring calculation unit, and a group risk reasoning unit; The data processing unit is used to collect multi-source heterogeneous grassroots governance data, perform entity alignment and semantic standardization, and generate a standardized personnel entity table with attributes and initial screening labels; The association graph construction unit is used to construct weighted association edges between personnel based on personnel attributes and initial screening tags in the standardized personnel entity table, combined with the three-element coupling of spatiotemporal co-occurrence, explicit social relations and behavioral pattern similarity, to form an initial risk personnel association graph; The risk scoring calculation unit is used to calculate the dynamic risk score of each person based on the node connection relationship in the initial risk personnel association graph, according to the multi-dimensional individual risk indicator system and combined with the time decay factor. The dynamic risk score is then embedded as a node attribute into the initial risk personnel association graph to obtain an enhanced graph. The group risk reasoning unit is used to perform multiple rounds of risk propagation iteration from high-risk nodes to neighboring nodes based on the node risk scores and weighted association edges in the enhanced graph, until convergence, and to identify connected subgraphs with risk scores higher than the threshold and a size greater than a preset number as high-risk groups.