Method and system for generating dynamic portraits of residential communities based on spatiotemporal data mining
By performing layered processing, timeliness compensation, and scenario adaptation on community data, we have resolved the obstacles to multi-source data collaboration and delayed responses, achieved rapid response to emergencies and accurate group behavior analysis, and improved the level of refined community management.
Patent Information
- Application Number
- CN202510975883.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-16
AI Technical Summary
The existing technology for generating dynamic portraits of residential communities suffers from the difficulty in effectively coordinating multi-source data due to differences in collection frequency, resulting in spatiotemporal discontinuities. It also lacks the ability to adapt to sudden scenarios such as holiday peaks and extreme weather, and fails to effectively distinguish the behavioral characteristics of different groups, leading to inaccurate allocation of public service resources.
By performing layered processing on access control, garage, and environmental sensor data, performing timeliness compensation and missing value filling, using scene classifiers to identify special event types, generating scene-adaptive weight combination data, and performing correlation fusion and weighted integration of spatiotemporal features, a targeted weight matrix for multi-source fusion is constructed.
It realizes a seamless data base that is continuous in time and space, has minute-level response capabilities, accurately distinguishes the activity characteristics of different groups, supports on-demand and targeted allocation of public service resources, provides a basis for real-time decision-making, and improves the level of refinement of smart community management.
Smart Images

Figure CN120492985B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cell management technology, and in particular to a method and system for generating a dynamic cell portrait based on spatiotemporal data mining. Background Art
[0002] Refined community management has become a crucial component of modern urban governance. Residential dynamic profiling, a core tool for intelligent management, builds a digital twin reflecting the real-time operational status of a residential community through the integration and analysis of multi-dimensional data such as access control, garage turnover, and environmental monitoring. This approach aims to transform dispersed IoT sensor data into a visual basis for management decisions, providing data support for security deployment, resource scheduling, and service optimization.
[0003] However, existing technologies for generating dynamic community portraits have significant limitations: First, due to differences in the frequency of data collection, it is difficult to effectively coordinate multiple sources, resulting in temporal and spatial discontinuities in the resulting portraits. Second, they lack the ability to adapt to unexpected scenarios such as holiday peaks and extreme weather, causing management strategies to lag behind real-time changes. Third, they fail to effectively distinguish between the behavioral characteristics of different groups, leading to misallocation of public service resources. These shortcomings make it difficult for the generated community portraits to accurately reflect the distribution of pedestrian flows and the status of facility usage, making them incapable of supporting refined community governance. Summary of the Invention
[0004] Based on this, the purpose of the present invention is to provide a method and system for generating dynamic portraits of a community based on spatiotemporal data mining, which can integrate multi-frequency data in real time, dynamically respond to scene changes, and accurately portray group behavior.
[0005] The purpose of the present invention is achieved by the following scheme:
[0006] In one aspect, the present invention provides a method for generating a dynamic profile of a community based on spatiotemporal data mining, comprising the following steps:
[0007] S1: Acquire and preprocess the access control data, garage data, and environmental sensor data of the cell to be tested, and generate a hierarchical data set containing high-frequency access control data, medium-frequency environmental data, and low-frequency garage data;
[0008] S2: Perform timeliness compensation and missing value filling on the stratified data set, perform behavioral fluctuation detection and grouping processing, and generate initial weight distribution data and grouped behavioral feature weight data;
[0009] S3: Perform data distribution anomaly detection based on the initial weight distribution data, call the scene classifier to identify special event types in the distribution characteristics, and generate a scene classification result;
[0010] S4: Perform parameter mapping and weight update processing on the scene classification judgment result and the initial weight distribution data to generate scene adaptation weight combination data;
[0011] S5: Perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, build a targeted weight matrix for multi-source fusion, and perform weighted integration processing of spatiotemporal features based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the community. The dynamic portrait of the community is used to indicate the distribution characteristics of the population and the resource usage status in the community.
[0012] On the other hand, the present invention provides a system for generating dynamic community portraits based on spatiotemporal data mining, which is configured with the following modules:
[0013] The data acquisition and preprocessing module is used to acquire and preprocess the access control data, garage data and environmental sensor data of the tested community to generate a hierarchical data set containing high-frequency access control data, medium-frequency environmental data and low-frequency garage data;
[0014] The behavioral feature processing module is used to compensate for timeliness and fill missing values in the layered data set, detect behavioral fluctuations and perform grouping processing to generate initial weight distribution data and grouped behavioral feature weight data;
[0015] The abnormal event recognition module is used to detect abnormalities in data distribution patterns based on the initial weight distribution data, call the scene classifier to identify special event types in the distribution characteristics, and generate scene classification judgment results;
[0016] The scene weight update module is used to perform parameter mapping and weight update processing on the scene classification judgment results and the initial weight distribution data to generate scene adaptation weight combination data;
[0017] The dynamic portrait construction module is used to perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, build a targeted weight matrix for multi-source fusion, and perform weighted integration processing of spatiotemporal features based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the community. The dynamic portrait of the community is used to indicate the distribution characteristics of the population and the resource usage status within the community.
[0018] To sum up, the method for generating dynamic community portraits based on spatiotemporal data mining provided in this application can effectively overcome the data gaps, response lags and group recognition inaccuracies of existing community portrait technologies through a multi-source data fusion mechanism and a dynamic scene adaptation strategy. In the data preprocessing stage, access control, garage, and environmental sensor data are layered to eliminate the synergy barriers between high-frequency and low-frequency data, thereby achieving a seamless data base with spatiotemporal continuity. In the dynamic weight generation stage, timeliness compensation and missing value filling mechanisms are used to ensure the effective connection between historical data and real-time data, thereby resolving the decision-making lag caused by data update delays in traditional methods. In the scene perception link, the collaborative operation of anomaly detection and classifiers can achieve minute-level response capabilities to emergencies such as holiday peaks and extreme weather, overcoming the passive response defects of existing technologies. In the group behavior analysis dimension, fluctuation detection and pattern grouping processing can accurately distinguish the activity characteristics of different groups, thereby achieving on-demand and targeted allocation of public service resources. Finally, in the portrait generation stage, correlation fusion and spatiotemporal weighted integration technologies are used to construct a visual map that simultaneously reflects the thermal characteristics of crowd distribution and the usage status of facilities, providing real-time decision-making basis for community security deployment, commercial operation optimization, and environmental resource scheduling, and comprehensively improving the refinement and response efficiency of smart community management.
[0019] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A flowchart of a method for generating a dynamic cell portrait based on spatiotemporal data mining provided in an embodiment of the present application;
[0021] Figure 2 A schematic diagram of a process for generating group behavior feature weight data provided in an embodiment of the present application;
[0022] Figure 3 A schematic diagram of the process of generating a dynamic cell portrait provided in an embodiment of the present application;
[0023] Figure 4 A structural diagram of a community dynamic portrait generation system based on spatiotemporal data mining provided in another embodiment of the present application. DETAILED DESCRIPTION
[0024] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0026] In one embodiment, Figure 1 As shown, a method for generating a dynamic cell portrait based on spatiotemporal data mining is provided. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0027] S1: Acquire and preprocess the access control data, garage data and environmental sensor data of the tested community to generate a hierarchical data set containing high-frequency access control data, medium-frequency environmental data and low-frequency garage data.
[0028] Specifically, access control data covers card swiping records, facial recognition records, passage time and other information of the access control systems at various entrances and exits of the community, unit doors, elevators, etc. The data collection frequency is relatively high, usually in seconds or minutes, so it is classified as high-frequency data; garage data mainly includes records of vehicles entering and exiting the garage, parking space occupancy, etc. The collection frequency is between access control data and environmental sensor data, and is defined as medium-frequency data; environmental sensor data involves temperature and humidity sensors, noise sensors, air quality sensors, etc. in the community. The collection frequency is relatively low, usually in hours or longer periods, so it is classified as low-frequency data.
[0029] After acquiring the data, the system will preprocess the acquired access control data, garage data and environmental sensor data, including data cleaning, format normalization and preliminary feature extraction. For example, it will extract features such as personnel identity, passage time, and passage direction from access control records, extract features such as vehicle entry and exit time and parking space occupancy time from garage data, and extract environmental parameters such as temperature, humidity, and light intensity from environmental sensor data. Based on the inherent characteristics and acquisition frequency of the data, the processed data will be divided into three levels: high-frequency access control data, medium-frequency environmental data and low-frequency garage data, to construct a hierarchical data set. The high-frequency access control data is used to reflect the flow of people in real time, the medium-frequency environmental data is used to dynamically monitor changes in the community environment, and the low-frequency garage data is used for long-term parking space usage analysis.
[0030] It should be noted that the access control data used is only for community management purposes, and its scope of application is strictly limited to improving the efficiency of refined community management and does not involve other activities beyond community management needs. Through the rational use of access control data, the goal is to optimize community resource scheduling, strengthen security control, and improve service optimization. For example, using access control access records to analyze changes in community traffic density at different times to provide decision support for the reasonable arrangement of security personnel patrol frequency and routes; and optimize the operating hours of public facilities based on residents' traffic patterns. Throughout the data processing process, relevant data protection laws and regulations are always followed to fully protect the privacy rights of residents, ensure that access control data is used exclusively for community management, and effectively maintain the normal life order and legal rights of community residents.
[0031] S2: Perform timeliness compensation and missing value filling on the stratified data set, perform behavioral fluctuation detection and grouping processing, and generate initial weight distribution data and grouped behavioral feature weight data.
[0032] Specifically, due to differences in physical properties and application scenarios, access control, garage, and environmental sensor data have different collection frequencies, leading to spatiotemporal gaps in the time series, affecting the accuracy of subsequent overall analysis. To address this issue, the system performs time-sensitive processing on the stratified datasets. By constructing a temporal interpolation model, low- and medium-frequency data are interpolated along the time axis based on the historical trends of each data type, ensuring that the sampling intervals are relatively consistent with the high-frequency access control data in the temporal dimension, thus compensating for the spatiotemporal gaps. For missing values in the data, the system uses a machine learning-based missing value imputation method, training corresponding imputation models based on the characteristics of different data types. For access control data, imputation considers the traffic patterns of people and the data characteristics of adjacent time points; for garage data, it incorporates information such as vehicle entry and exit frequency and parking duration; and for environmental sensor data, imputation is based on the temporal continuity and spatial correlation of environmental parameters to ensure data integrity.
[0033] After completing timeliness compensation and missing value filling, the system performs behavioral fluctuation detection. By calculating the statistical characteristic values of each data type within different time windows and comparing and analyzing them with historical data from the same period, it identifies time periods and behavioral characteristics where data fluctuations exceed normal ranges. For example, an abnormally high frequency of traffic records for a unit door within a specific time period in access control data may indicate a special event or a gathering of people; large fluctuations in temperature and humidity within a short period of time in environmental sensor data may indicate a malfunction in environmental monitoring equipment or a sudden change in the external environment.
[0034] Based on the behavioral fluctuation detection results, the system groups the data according to different behavioral characteristics. For example, access control data is divided into groups such as weekday commuting and non-commuting hours, and resident and visitor traffic. Garage data is divided into groups such as weekday and holiday peak parking hours, and parking space usage by vehicle type. Environmental sensor data is grouped according to seasons, diurnal variations, and other factors. For each data group, its relative importance within the overall dataset is calculated to form an initial weight distribution. Simultaneously, the corresponding weight is calculated for each group's behavioral characteristics to generate group behavioral characteristic weights.
[0035] S3: Perform data distribution law anomaly detection based on the initial weight distribution data, call the scene classifier to identify special event types in the distribution characteristics, and generate scene classification judgment results.
[0036] Specifically, the system uses a multidimensional data distribution law detection algorithm based on the initial weight distribution data, such as the DBSCAN algorithm based on cluster analysis combined with Mahalanobis distance discrimination, to analyze the distribution characteristics of the data and identify abnormal points that deviate from the normal distribution pattern. These abnormal points may indicate errors in the data collection process, equipment failures, or data changes when special events occur.
[0037] At the same time, the system constructs a scene classifier based on the Convolutional Neural Network (CNN) architecture in deep learning. It uses various scene types marked in historical data as training samples. The training samples include normal working day scenes, holiday peak scenes, extreme weather scenes, etc. The system extracts key spatiotemporal features such as the time series pattern of access control access, the flow fluctuation pattern of vehicles entering and exiting the garage, and the spatial change pattern of environmental parameters from the data distribution characteristics as input feature vectors. After multi-layer convolution, pooling operations and training and learning of fully connected layers, it can achieve accurate identification of special event types in the current data distribution characteristics and output scene classification judgment results, such as "holiday peak scene" or "rainstorm weather scene".
[0038] S4: Perform parameter mapping and weight update processing on the scene classification judgment results and the initial weight distribution data to generate scene adaptation weight combination data.
[0039] Specifically, the system uses the scenario classification results to call the corresponding parameter mapping model, and adapts different data weight adjustment strategies and parameter mapping relationships to different special event types. For example, during peak holiday seasons, the weight of visitor access in access control data will be increased, while the weight of temporary parking space usage in garage data will increase. During extreme weather events, the weight of environmental sensor data closely related to meteorological conditions will be significantly increased, while the weight of access to indoor public areas in access control data may also increase.
[0040] Specifically, the system combines the initial weight distribution data with the parameter mapping model and dynamically updates the initial weights through mathematical mapping relationships and weight adjustment algorithms. Specifically, data types and features related to the current special event are assigned a larger weight coefficient in the weight calculation formula, while data with lower relevance are appropriately weighted. After parameter mapping and weight update processing, scenario-adaptive weight combination data is generated. This data reflects the importance of each data type and behavioral feature in a specific scenario, allowing data weights to adapt to changes in different scenarios.
[0041] S5: Perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, build a targeted weight matrix for multi-source fusion, and perform weighted integration processing of spatiotemporal features based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the community. The dynamic portrait of the community is used to indicate the distribution characteristics of the population and the resource usage status in the community.
[0042] Specifically, the system constructs a multi-source data fusion model, comprehensively considering factors such as data type, behavioral characteristics, and scenario adaptability to establish correlation weight relationships between data. For example, it correlates visitor access weights in access control data with temporary parking space usage weights in garage data to further refine the weight calculation method; it also combines air quality weights in environmental sensor data with outdoor area access weights in access control data to account for the impact of environmental factors on human behavior. Through this correlation fusion process, a targeted multi-source fusion weight matrix is constructed, which comprehensively and accurately reflects the comprehensive weight distribution of various data types and behavioral characteristics in specific scenarios.
[0043] Next, based on the environmental sensor data in the hierarchical dataset, the system performs a weighted integration process on the targeted weight matrix using spatiotemporal characteristics. Using geographic information technology (GIS) and spatiotemporal data analysis methods, the environmental sensor data is located and mapped on the spatial layout of the community, and its spatiotemporal distribution characteristics are analyzed. For example, considering the impact of differences in environmental parameters in different regions on population distribution and resource utilization within the community, areas with more favorable environmental conditions are given higher weights when integrating weights, while areas with less favorable environmental conditions are given appropriate weights. At the same time, combined with the time dimension, the system analyzes the impact of changes in environmental parameters in different time periods on the dynamic state of the community, such as the impact of temperature differences between morning and evening on the use of outdoor activity areas. Through weighted integration of spatiotemporal characteristics, the targeted weight matrix is deeply integrated with the spatiotemporal characteristics of environmental sensor data, ultimately generating a dynamic portrait of the community. The generated dynamic portrait of the community is presented in an intuitive visual form, including but not limited to a variety of chart combinations such as heat maps, bar charts, and line charts. Heat maps are used to display the density of people in a residential community, visually reflecting the degree of concentration of people in each area through different shades of color; bar charts can show changes in the utilization rate of resources such as parking spaces and public facilities; and line charts are used to depict the dynamic changes in various key indicators of the community, such as access control frequency and environmental parameters in different time periods.
[0044] The dynamic portrait of a community can reflect the distribution characteristics of the population and resource usage status within the community, providing real-time and comprehensive decision-making basis for community managers, helping to achieve accurate deployment of community security and control, reasonable optimization of resource scheduling, and scientific decision-making for service optimization, thereby improving the level of refined management of the community.
[0045] To sum up, the method for generating dynamic community portraits based on spatiotemporal data mining provided in this application can effectively overcome the data gaps, response lags and group recognition inaccuracies of existing community portrait technologies through a multi-source data fusion mechanism and a dynamic scene adaptation strategy. In the data preprocessing stage, access control, garage, and environmental sensor data are layered to eliminate the synergy barriers between high-frequency and low-frequency data, thereby achieving a seamless data base with spatiotemporal continuity. In the dynamic weight generation stage, timeliness compensation and missing value filling mechanisms are used to ensure the effective connection between historical data and real-time data, thereby resolving the decision-making lag caused by data update delays in traditional methods. In the scene perception link, the collaborative operation of anomaly detection and classifiers can achieve minute-level response capabilities to emergencies such as holiday peaks and extreme weather, overcoming the passive response defects of existing technologies. In the group behavior analysis dimension, fluctuation detection and pattern grouping processing can accurately distinguish the activity characteristics of different groups, thereby achieving on-demand and targeted allocation of public service resources. Finally, in the portrait generation stage, correlation fusion and spatiotemporal weighted integration technologies are used to construct a visual map that simultaneously reflects the thermal characteristics of crowd distribution and the usage status of facilities, providing real-time decision-making basis for community security deployment, commercial operation optimization, and environmental resource scheduling, and comprehensively improving the refinement and response efficiency of smart community management.
[0046] In one embodiment, S1 of a method for generating a dynamic cell portrait based on spatiotemporal data mining provided by the present invention specifically includes the following steps:
[0047] S11: Based on the IoT sensing device interface, the access control data, garage data and environmental sensor data of the tested community are collected in real time. The data verification rules are called to filter out abnormal records in the original data stream and fill in missing fields to generate a standard multi-source data set.
[0048] Specifically, this embodiment relies on various IoT sensing devices deployed within the community under test, including access control systems, garage management systems, and environmental monitoring sensors, to achieve real-time collection of community access control data, garage data, and environmental sensor data through their device interfaces. Access control data primarily covers personnel entry and exit records, including key fields such as access card number, personal identity information, entry timestamp, and direction of travel; garage data covers vehicle entry and exit information, including license plate number, entry and exit time, and parking space number; and environmental sensor data includes real-time measurements of environmental parameters such as temperature and humidity, light intensity, noise decibels, and air quality index, along with corresponding timestamps.
[0049] During the data collection process, the system calls the preset data validation rules to perform preliminary cleaning of the raw data stream. The data validation rules cover data format validation, data integrity validation, and data logic validation. Among them, the system directly filters out abnormal records that do not comply with the validation rules. For missing fields in some records, the system adopts corresponding data completion strategies based on the device type, data type, and historical data characteristics. For example, for missing traffic direction fields in access control data, they can be inferred and completed based on the traffic direction patterns recorded at the same access control point at adjacent times. Occasional missing temperature and humidity values in environmental sensor data can be interpolated and estimated based on data from adjacent time points and the temperature and humidity change trends in the area where the sensor is located, thereby completing the missing fields and generating a standard multi-source dataset with a unified format, complete content, and quality requirements.
[0050] S12: Extract the collection frequency features of the standard multi-source data set, calculate the time interval distribution statistics of each data source, identify the categories of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data, and generate a frequency classification label set.
[0051] Specifically, the system performs time series analysis on access control data, environmental sensor data, and garage data, calculating statistics on their time interval distribution, including key metrics such as the mean time interval, standard deviation of the time interval, and median time interval, to quantitatively describe the frequency characteristics of data collection. For example, by analyzing the time difference between two consecutive entry and exit records in access control data, it is calculated that the average time interval is short, indicating that access control data has high-frequency collection characteristics; while the time intervals between vehicle entry and exit records in garage data are relatively large, belonging to the medium-frequency collection category; and the collection intervals of environmental sensor data are often even longer, classifying them as low-frequency data.
[0052] By setting a reasonable frequency threshold interval, the system determines the frequency category of each data source based on the calculated time interval distribution statistics, identifies high-frequency access control data, medium-frequency environmental data, and low-frequency garage data, and assigns corresponding frequency classification labels to them respectively, generating a frequency classification label set that marks the collection frequency characteristics of each type of data in the time dimension.
[0053] S13: Divide the frequency classification label set into time windows, configure a second-level sliding window for high-frequency access control data, a minute-level fixed window for medium-frequency environmental data, and an hour-level fixed window for low-frequency garage data, to generate a hierarchical dataset structure.
[0054] Specifically, the system configures adaptive time windows for each type of data based on the frequency of collection, as revealed by the frequency classification label set, to achieve refined, hierarchical organization of the data. For high-frequency access control data, a second-level sliding window is configured. The window length can be set from a few seconds to tens of seconds depending on the specific application scenario, and it slides along the time axis at a fixed time step. This ensures that the dynamic characteristics of access control traffic flow within a very short period of time can be captured. For example, during peak commuting hours, a second-level sliding window can accurately reflect the high-frequency and high-density fluctuations in traffic flow. For medium-frequency environmental data, a minute-level fixed window is configured, with a window length generally ranging from a few minutes to tens of minutes. This fixed window effectively integrates the changing trends of environmental data on a medium time scale, balancing the preservation of data detail with computational complexity. For example, it can be used to analyze minute-level fluctuations in light intensity within a residential area under different weather conditions. For low-frequency garage data, an hour-level fixed window is configured. A window length of several hours is sufficient to capture the cyclical changes in vehicle entry and exit in garages, making it suitable for capturing the macro-distribution of parking behavior across the diurnal cycle and different time periods, such as the significant difference in parking space utilization during daytime and nighttime on weekdays.
[0055] S14: Match and integrate the standard multi-source dataset and the hierarchical dataset structure, aggregate the time series data features by window type and add hierarchical identifiers to generate a hierarchical dataset. The hierarchical dataset is used to indicate the temporal hierarchical structure of the multi-source spatiotemporal data.
[0056] Specifically, the system matches and integrates standard multi-source datasets and hierarchical dataset structures. Time dimension matching uses an algorithm to map the timestamp of each data record to the time interval of the corresponding window. Spatial dimension matching associates data with coordinates to the corresponding physical area window based on spatial indexes. When a record meets the conditions of multiple windows simultaneously, it is assigned to the window with the later end time based on the principle of temporal proximity, and the "cross-window record" is marked in the metadata.
[0057] The system aggregates time-series data features by window type. High-frequency access control windows aggregate and calculate statistical features such as pass frequency, deduplication of personnel signature codes, two-way traffic ratio, and average interval between passes, with individual counts for abnormal passes. Medium-frequency environmental windows aggregate and calculate metrics such as mean, maximum, minimum, and standard deviation for numerical parameters, and record the duration of valid states for state-based parameters. Low-frequency garage windows aggregate and calculate statistics such as total vehicle ingress and egress, parking space turnover, average dwell time, and peak-hour traffic volume, correlating these with parking space saturation warning thresholds for the corresponding area. The computer system adds hierarchical identifiers using a multidimensional encoding system encompassing frequency level, window type, area code, and timestamp. An algorithm generates a unique hash value as the primary key, generating a hierarchical dataset. The hierarchical dataset is stored in separate tables by frequency level, with composite indexes established to support queries. Daily integrity checks are automatically performed, and data re-collection is triggered when deviations exceed a set value. The hierarchical dataset indicates the hierarchical structure of the timeliness of multi-source spatiotemporal data.
[0058] In one embodiment, Figure 2 As shown, S2 of a method for generating a dynamic cell portrait based on spatiotemporal data mining provided by the present invention specifically includes the following steps:
[0059] S21: Processing the layered data set based on the exponential decay function, calculating the timeliness compensation weight of the high-frequency access control data, and generating weighted access control data.
[0060] Specifically, the system extracts high-frequency access control data from the hierarchical data set, obtains the generation time and current time of each data, and calculates the time difference between the two Preferably, the calculation formula for the timeliness compensation weight is:
[0061] ;
[0062] in, is the timeliness compensation weight, is the weight of the original access control data, is the attenuation coefficient, Generate the time difference between the current time and the high-frequency access control data, is a natural constant. The system calculates the weight of high-frequency access control data one by one. During the calculation process, the system automatically associates the time window information corresponding to the high-frequency access control data to ensure that the time difference The calculation accuracy is consistent with the window division granularity. After completing the timeliness compensation weight calculation for all high-frequency access control data, the system associates and integrates the compensation weight with the original access control data to generate weighted access control data. This data retains all field information of the original access control data and adds a timeliness compensation weight field for subsequent data fusion processing.
[0063] S22: The layered data set is processed based on the linear interpolation algorithm to calculate the missing value filling parameters of the low-frequency garage data, and a complete data sequence is constructed through the data of adjacent time periods to generate smoothed garage data.
[0064] Specifically, the linear interpolation algorithm is used to effectively fill the missing value problem in the low-frequency garage data. The basic principle of the linear interpolation algorithm is to construct a linear function based on the data values of two adjacent valid time points of the missing data to estimate the data value of the missing time point. Specifically, assuming that in the time series, the garage data at time point and There are complete records, and the time point (lie in and If data is missing, then the missing value It can be calculated by the following formula:
[0065] ;
[0066] in, and Time points and To ensure data smoothness and consistency, the data series was constructed using complete data from adjacent time periods. After filling in missing values, the garage data was processed again using smoothing techniques such as moving averages to reduce fluctuations caused by the randomness of data collection, resulting in a more stable trend in the time series.
[0067] S23: Perform feature merging processing on the weighted access control data, smoothed garage data, and medium-frequency environmental data in the layered data set, assign initial weight coefficients according to the data source type, construct a unified data structure, and generate initial weight distribution data.
[0068] Specifically, the system extracts core feature fields from weighted access control data, smoothed garage data, and hierarchical datasets. Weighted access control data extracts features such as pass frequency and compensation weights; smoothed garage data extracts features such as vehicle entry and exit volume and dwell time; and medium-frequency environmental data extracts features such as the measured values of various environmental parameters. The system assigns initial weight coefficients to each of the three data types based on their data source type. These weight coefficients are determined based on the importance of the data type in the overall analysis and stored in a coefficient configuration table for subsequent adjustment. The system constructs a unified data structure consisting of a timestamp, spatial identifier, feature value array, data type identifier, and weight coefficient fields. The feature fields from the three data types are populated into this structure using a unified format. During the populating process, the system automatically verifies the data format for consistency and performs standardization conversion on records that do not conform to the format requirements. After feature merging, the system aggregates the data in the unified data structure by time window to generate initial weight distribution data. This data retains the original features of each data source and the assigned initial weight coefficients for subsequent anomaly detection.
[0069] S24: Based on the access control data and time window information in the hierarchical data set, behavioral fluctuation detection and processing are performed, the speed variance of the movement trajectory is analyzed for abnormal time periods, and the behavioral pattern groups of different groups are divided, and grouped behavioral feature weight data are generated. The grouped behavioral feature weight data is used to indicate the spatiotemporal activity characteristics of different resident groups.
[0070] Specifically, the system detects behavioral fluctuations based on access control data and time window information from a hierarchical dataset. The system extracts individual signature codes and corresponding time window information from the access control data, constructing a trajectory sequence for each signature code. The trajectory sequence contains location coordinates within different time windows. The system calculates the variance of movement speed within the trajectory sequence and, by analyzing changes in the variance, identifies periods of abnormal speed variance. These periods correspond to turning points in the individual's behavior patterns. Based on the abnormal periods, the system divides the trajectory sequence into distinct behavioral pattern segments and calculates characteristic parameters such as average movement speed, range of activity, and time window distribution for each segment. The system performs cluster analysis on the behavioral pattern segments of all signature codes, grouping segments with similar characteristic parameters into the same group and thus defining different behavioral pattern groups. The system calculates a feature weight for each group based on the group's proportion in the overall sample and the discriminatory power of the feature parameters. This generates grouped behavioral feature weight data, which includes group identifiers, feature parameters, and weight values. This data, linked to the corresponding signature code by the group identifier, indicates the spatiotemporal activity characteristics of different resident groups, providing a basis for subsequent resource allocation analysis. Preferably, the initial weight distribution data is generated by the following steps:
[0071] S241: Desensitize the access control movement record data in the hierarchical dataset, replace the specific location with area code mapping and remove personal identification information to generate an anonymous spatiotemporal feature dataset.
[0072] Specifically, the system desensitizes the access control movement record data in the hierarchical data set, extracts the location coordinate information from the access control movement record data, and maps the specific location coordinates to the corresponding area code using the preset area coding rules. The area code is divided according to the physical division of the cell, and each code corresponds to a specific geographical area. The system scans the personal identification information fields in the access control movement record data, including personnel feature codes, name-related fields, etc., and permanently removes these fields from the data records to ensure that the processed data cannot be reversely associated with a specific individual. During the processing process, the system retains non-identifying fields such as the timestamp, area code, and direction of travel of the access control movement record to maintain the temporal and spatial correlation of the data.
[0073] After the desensitization process is completed, the system performs an integrity check on the data to ensure that all records have completed regional code mapping and no residual personal identification information is left, generating an anonymous spatiotemporal feature dataset, which is used for subsequent regional activity analysis and complies with data privacy protection requirements.
[0074] S242: Jointly analyze and process the anonymous spatiotemporal feature dataset and the hierarchical time window to calculate the change rate of pedestrian flow density and the distribution of length of stay in each area during the same period of time, and generate a regional activity intensity matrix.
[0075] Specifically, the system performs a joint analysis and processing on the anonymous spatiotemporal feature dataset and the hierarchical time window, and associates the records in the anonymous spatiotemporal feature dataset to the corresponding hierarchical time window by timestamp, ensuring that each record matches a unique time window identifier. The system groups the data by the combination of area code and time window, counts the number of records under each combination, and calculates the pedestrian density of each area in the same time period. The pedestrian density is the ratio of the number of records in the area in the corresponding time window to the area of the area. The system calculates the ratio of the change in pedestrian density in the continuous time window to the initial density to obtain the pedestrian density change rate. At the same time, the system extracts the timestamp information recorded under each area code, calculates the time difference between different records in the same area, obtains the length of stay, and counts the distribution of the length of stay, including the proportion of records in different time intervals.
[0076] For example, within each time window, the number of people passing through each area is counted, and the rate of change of pedestrian density between adjacent time windows is calculated. For example, during the morning rush hour from 7:00 to 9:00 on weekdays, the number of people passing through the entrance area of a residential building is calculated every 15 minutes with a 15-minute time window. The difference in the number of people in adjacent windows is compared to obtain the rate of change of pedestrian density. If the number of people in a certain window increases by 50% compared to the previous window, the rate of change is 50%, which reflects the rapid gathering of residents during commuting hours. The system integrates the rate of change of pedestrian density and the distribution of dwell time according to the dimensions of regional code and time window, constructing a two-dimensional matrix with the row dimension being the regional code and the column dimension being the time window. The matrix elements contain the parameters of the pedestrian density change rate and dwell time distribution for the corresponding region and time period, generating a regional activity intensity matrix, which is used for subsequent clustering modeling.
[0077] S243: Perform cluster modeling on the regional activity intensity matrix, divide it into life service type, commuting type and leisure type behavior pattern groups according to the peak characteristics of the pedestrian density, and generate group behavior feature weight data.
[0078] Specifically, the system performs cluster modeling on the regional activity intensity matrix, extracts the pedestrian density change rate data of each area in the regional activity intensity matrix, identifies the peak pedestrian density of each area in all time windows, and records the time window where the peak occurs, the density value corresponding to the peak, and the number of time windows in which the peak lasts.
[0079] Preferably, the system can use a clustering algorithm to group areas, classifying areas where the peak of pedestrian density mainly occurs in specific time windows in the morning and evening and has a high correlation with the location of life service facilities as life service groups; areas where the peak of pedestrian density is concentrated in fixed time periods in the morning and evening on weekdays and has a high correlation with the entrance and exit of the community are classified as commuting groups; areas where the peak of pedestrian density is distributed on non-working days and weekday evenings and has a high correlation with public activity areas are classified as leisure groups. The system calculates feature weights for each group. The weights are determined based on the proportion of the number of areas included in the group to the total number of areas and the significance of the peak of pedestrian density. The weight values are associated with the feature parameters of the corresponding group and stored to generate group behavior feature weight data. This data is used to indicate the spatiotemporal distribution characteristics of different behavior pattern groups, providing a basis for the allocation of public service resources.
[0080] In one embodiment, S3 of a method for generating a dynamic cell portrait based on spatiotemporal data mining provided by the present invention specifically includes the following steps:
[0081] S31: Analyze the distribution pattern of the initial weight distribution data, calculate the degree to which the variance of the weight distribution of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data deviates from the historical benchmark value, and generate a data anomaly indicator.
[0082] Specifically, for high-frequency access control data, medium-frequency environmental data, and low-frequency garage data, the system calculates the degree to which the variance of their weight distribution deviates from the historical benchmark value to quantify the degree of abnormality of the data. Specifically, the system analyzes the distribution rules of the initial weight distribution data, extracts the weight distribution sequences of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data in the initial weight distribution data, and simultaneously retrieves the weight distribution benchmark sequences of the three types of data stored in the same period in history. The system calculates the variance of the current weight distribution sequence and the historical benchmark sequence, and quantifies the degree of deviation of the current distribution from the historical benchmark through the variance difference. The calculation of the degree of deviation covers all time windows to ensure that the distribution changes in different time periods are captured. The system integrates the degree of deviation of the three types of data according to the data type to form an indicator set containing the data type identifier, the corresponding weight distribution variance, and the quantitative value of the degree of deviation, and generates a data anomaly indicator, which is associated with the corresponding records in the initial weight distribution data through the time window.
[0083] S32: Perform scenario classification processing on data anomaly indicators, combine public holiday calendar information with meteorological disaster warning signals to identify three special scenario types: holiday peak, extreme weather response and emergency response, and generate event type labels.
[0084] Specifically, the system classifies data anomaly indicators into scenarios, associates them with built-in public holiday calendar information, extracts data anomaly indicators during holiday periods, and analyzes the changing characteristics of the weight distribution of high-frequency access control data during these periods. When the changing characteristics match the preset peak pattern, the scenario is marked as a candidate for a holiday peak scenario. The system connects to the external meteorological disaster warning signal interface to obtain real-time and forecasted meteorological disaster information, matches it with the changes in the weight distribution of medium-frequency environmental data in the data anomaly indicators, and marks the scenario as a candidate for an extreme weather response scenario when the abnormal weight distribution of environmental data is synchronized with the meteorological disaster warning signal in time.
[0085] The system simultaneously monitors sudden and drastic changes in the weight distribution of three types of data in the data anomaly indicators. When the amplitude of the change exceeds the normal fluctuation range and has no obvious periodicity, it is marked as a candidate for emergency response scenario. At the same time, the system performs feature matching on the three types of candidate scenarios, and assigns corresponding event type labels to the scenarios that meet the matching conditions, generating event type labels that include time windows, data anomaly indicator related items, and scenario types.
[0086] S33: Conflict check is performed on the event type tags to eliminate false alarm signals caused by equipment failure and confirm the dominant event characteristics, and generate scene classification judgment results. The scene classification judgment results are used to guide the community security resource allocation strategy.
[0087] Specifically, the system performs a conflict check on event type tags, extracts the original data collection records corresponding to the event type tags, and checks whether the records contain device status abnormality indicators. If so, the corresponding event type tag is marked as a suspected false alarm. The system also performs a secondary verification on suspected false alarm tags, comparing the data of adjacent devices in the same area over the same period to determine whether the abnormal indicators are caused by a single device failure, thereby eliminating false alarm signals caused by device failure.
[0088] Preferably, the system performs feature analysis on the remaining valid event type tags, identifying multiple tags appearing within the same time window. By analyzing the impact range and duration of the abnormal indicators corresponding to each type of tag, the system determines the dominant event characteristics. The system associates the dominant event characteristics with the corresponding scenario type, integrating them into a final judgment result that includes the time window, dominant event type, associated abnormal indicators, and impact range. This generates a scenario classification judgment result, which is used to guide the adjustment and implementation of the community security resource allocation strategy by associating the corresponding response strategy template with the event type.
[0089] In one embodiment, S4 of a method for generating a dynamic cell portrait based on spatiotemporal data mining provided by the present invention specifically includes the following steps:
[0090] S41: Perform parameter mapping processing on the scene classification judgment result, query the preset scene-weight correction coefficient mapping relationship table, and generate a weight correction parameter set. The scene-weight correction coefficient mapping relationship table is used to indicate the adjustment ratio of data weights for different event types.
[0091] Specifically, the system extracts the dominant event type, impact range, and time window information from the scene classification results, uses this information as search keywords, and queries the preset scene-weight correction coefficient mapping relationship table. Among them, the preset scene-weight correction coefficient mapping relationship table is divided into sub-tables according to event type. The sub-tables contain weight correction coefficients corresponding to different impact ranges. The correction coefficients are expressed in percentage form and are used to indicate the adjustment ratio of the weights of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data under this event type. For example, in the holiday peak scenario, the weight correction coefficient of access control data is set to +0.2, environmental data is +0.1, and garage data is -0.1; in the extreme weather response scenario, the weight correction coefficient of access control data is set to +0.1, environmental data is +0.3, and garage data is +0.1; in the emergency response scenario, the weight correction coefficient of access control data is set to +0.3, environmental data is +0.2, and garage data is +0.2.
[0092] Specifically, the system extracts the corresponding data type correction coefficient based on the matched subtable entry and generates a weight correction parameter set based on the time window information. The parameter set includes the event type identifier, data type identifier, correction coefficient value, and valid time window field. The valid time window is consistent with the time window in the scene classification result. During the generation process, the system performs an integrity check on the mapping results to ensure that each data type has a corresponding correction coefficient. If any are missing, the default correction coefficient is used to fill in the gaps. The generated weight correction parameter set is used for subsequent weight adjustment processing.
[0093] S42: Perform coefficient superposition processing on the initial weight distribution data and the weight correction parameter set, adjust the initial weight distribution by superimposing the correction coefficient of the weight correction parameter set according to the data category, and generate updated weight data.
[0094] Specifically, the system assigns correction coefficients from the weight correction parameter set to the high-frequency access control data, medium-frequency environmental data, and low-frequency garage data in the initial weight distribution data, based on data type. For example, in the initial weight distribution data, access control data has a weight of 0.4, environmental data has a weight of 0.3, and garage data has a weight of 0.3. For peak holiday traffic, the system assigns +0.2 from the weight correction parameter set to access control data, +0.1 to environmental data, and -0.1 to garage data.
[0095] Furthermore, the system performs a superposition update calculation on the initial weight based on the assigned correction coefficient. The following formula is used:
[0096] ;
[0097] in, represents the updated weight, is the initial weight, is the weight correction coefficient. Calculated using the above example data, the updated weights for access control data are 0.4+0.2=0.6, for environmental data 0.3+0.1=0.4, and for garage data 0.3+(-0.1)=0.2. The system then performs a range check on the adjusted weight values. If they exceed the preset weight threshold range, they are truncated to the threshold boundary value and a truncation flag is recorded. After completing the weight adjustment for all data types, the system re-associates the adjusted weight values with the original data features to generate updated weight data. This data retains the structure of the initial weight distribution data and only replaces the weight value fields.
[0098] S43: Perform structured integration processing on the updated weight data, uniformly encode the weight values of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data into a multi-dimensional vector, and generate scene adaptation weight combination data. The scene adaptation weight combination data is used to guide the real-time data priority allocation strategy.
[0099] Specifically, the system performs structured integration processing on the updated weight data, grouping the updated weight data by time window. Each group contains the adjusted weight values for high-frequency access control data, medium-frequency environmental data, and low-frequency garage data within that time window. The system assigns a fixed dimension index to each data type, where high-frequency access control data corresponds to the first dimension, medium-frequency environmental data corresponds to the second dimension, and low-frequency garage data corresponds to the third dimension. The system arranges the three weight values in each time window group in the order of the dimension index and encodes them into a three-dimensional vector, where the element value in the vector is the corresponding weight value.
[0100] During the encoding process, the system standardizes the weight values and converts them into numerical values in the range of 0-1. The standardization is based on the maximum weight value of the data type in the same historical period. After the encoding is completed, the system arranges the three-dimensional vectors of all time windows in chronological order to form a vector sequence. The sequence contains the time window identifier and the corresponding three-dimensional vector. The system performs an integrity check on the vector sequence to ensure that each time window has a corresponding vector record, and the missing vectors are supplemented by interpolation of the previous and next adjacent vectors. The generated scene adaptation weight combination data is stored in the form of a vector sequence to guide the priority allocation strategy in the real-time data transmission and processing process, where the data type corresponding to the dimension with a higher value in the vector is given a higher processing priority.
[0101] In one embodiment, Figure 3 As shown, S5 of the method for generating a dynamic cell portrait based on spatiotemporal data mining provided by the present invention specifically includes the following steps:
[0102] S51: Perform similarity fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, analyze the scene-behavior matching degree between the group behavior and the current scene feature, and generate a fusion weight coefficient.
[0103] Specifically, the system performs similarity fusion on the scene adaptation weight combination data and the group behavior feature weight data, extracts the three-dimensional vector sequence in the scene adaptation weight combination data, and the group feature parameters and weight values in the group behavior feature weight data. By calculating the cosine similarity between the scene feature vector and the feature parameters of each behavior group, the degree of matching between the group behavior and the current scene characteristics is quantified, that is, the scene-behavior matching degree. The matching degree calculation covers all time windows and behavior groups to ensure that the correlation between each group behavior in different scenarios is captured. Preferably, the calculation formula for scene-behavior matching degree is:
[0104] ;
[0105] in, is the scene-behavior matching degree, For the Class scene weight value, For the Class behavior weight value, is the total number of weight categories, is the modulus of the scene weight vector, The system calculates the weighted sum of the scene adaptation weight and the weight of the grouped behavioral characteristics based on the matching degree. The weight distribution is determined by the matching degree; the higher the matching degree, the greater the weight of the corresponding behavioral group. During the summation process, the system maintains the consistency of the data dimensions and integrates the results into a unified coefficient sequence to generate a fusion weight coefficient. This coefficient is associated with both the scene identifier and the behavioral group identifier for subsequent matrix conversion processing.
[0106] S52: Perform spatial matrix conversion on the fusion weight coefficients, construct a three-dimensional weight distribution structure according to the data source type dimension and the behavior group dimension, and generate a targeted weight matrix.
[0107] Specifically, the system performs a spatial matrix conversion on the fusion weight coefficients. The first dimension is divided by data source type, encompassing high-frequency access control data, medium-frequency environmental data, and low-frequency garage data. The second dimension is divided by behavioral group, encompassing already defined groups such as life services, commuting, and leisure. The third dimension, using the time window, maintains the same granularity as previously described. The system populates the fusion weight coefficients according to this three-dimensional structure, with each matrix element corresponding to the weight of a specific data source type and behavioral group within a specific time window. During the filling process, the system uses a mean-filling method for missing element values, calculating the mean of the weights of adjacent elements in the same dimension and filling in the missing values. After filling, the system normalizes the matrix to ensure that the sum of the weights of all elements within each time window is a fixed value, maintaining the integrity of the weight distribution. The resulting targeted weight matrix retains three-dimensional index information to guide the weighted processing of real-time data streams.
[0108] S53: Perform spatiotemporal feature weighting processing on the real-time data stream in the layered data set, apply a targeted weight matrix to fuse the crowd density characteristics and facility utilization rate characteristics, and generate a dynamic portrait of the community including a thermal distribution map.
[0109] Specifically, the system extracts spatiotemporal features of real-time data streams from hierarchical datasets. Crowd density features are derived by calculating the rate of change in the number of people in each area over different time periods. Facility utilization features are calculated based on data such as garage parking occupancy and public facility usage records. For example, during peak hours in the morning and evening, the crowd density feature at the entrance of a residential building is higher, while the garage parking space utilization feature reaches its peak during the day on weekdays.
[0110] Preferably, the system substitutes the extracted spatiotemporal feature data into a targeted weight matrix for weighted fusion calculation. For example, for the life service behavior group, when calculating its corresponding thermal distribution, the weights of the access control data and environmental data under the behavior group are mainly referred to, and the crowd density characteristics and facility utilization rate characteristics are weighted summed. The formula is:
[0111] ;
[0112] in, represents the weighted fusion feature value, 、 、 Respectively represent the weights of access control data, environmental data, and garage data in the current behavior group. 、 、 They represent the characteristics of pedestrian density, facility utilization rate, and garage data respectively.
[0113] Based on the weighted fusion of characteristic values and the neighborhood's geographic information, the system creates a heat map. The heat map uses different color gradients to visually represent the activity intensity of different areas within the neighborhood. For example, red areas indicate high-intensity activity, such as building entrances during peak hours in the morning and evening; green areas indicate low-intensity activity, such as leisure plazas late at night. Furthermore, key indicators of crowd density and facility utilization are overlaid on the heat map to enhance the information richness and intuitiveness of the dynamic portrait. This dynamic neighborhood portrait can reflect the distribution characteristics of the community and the status of resource utilization in real time, providing decision-making support for neighborhood managers.
[0114] To sum up, the method for generating dynamic community portraits based on spatiotemporal data mining provided in this application can systematically solve the core defects of existing community portrait technologies such as poor scene adaptability, disconnection of group characteristics and insufficient visualization accuracy through multi-dimensional feature fusion and dynamic weighting mechanism. At the scene-behavior collaboration level, dynamic matching of scene adaptation weights and group behavior weights based on a similarity fusion algorithm can achieve real-time quantification of the correlation strength between behavioral patterns and emergencies, overcoming the strategy failure problem caused by the rigidification of scene features in traditional methods. In the weight structured conversion stage, a three-dimensional weight distribution structure constructed through spatial matrix processing can achieve a refined coupled expression of multi-feature weights, providing a mathematical basis for subsequent precise weighting. In the dynamic portrait generation stage, the spatiotemporal characteristics of real-time data streams are integrated using a targeted weight matrix, achieving collaborative visualization of crowd density characteristics and facility utilization characteristics under a unified spatial benchmark, thereby eliminating portrait distortion caused by feature fragmentation. The resulting thermal distribution map, by integrating the dual characteristics of scene responsiveness and group activity patterns, can construct a visual decision-making model that simultaneously reflects instantaneous crowd hotspots and long-term resource occupancy patterns. This provides full-dimensional data support for the targeted deployment of community security resources, dynamic scheduling of public service facilities, and optimization of commercial network layout, significantly improving the situational awareness accuracy and resource allocation efficiency of smart community governance.
[0115] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0116] Based on the same inventive concept, the embodiments of the present application also provide a system for generating dynamic cell portraits based on spatiotemporal data mining for implementing the method for generating dynamic cell portraits based on spatiotemporal data mining mentioned above. The implementation solution provided by this system is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more embodiments of the system for generating dynamic cell portraits based on spatiotemporal data mining provided below can be found in the above limitations of the method for generating dynamic cell portraits based on spatiotemporal data mining, and will not be repeated here.
[0117] In another aspect, the present invention provides a system 600 for generating a dynamic cell portrait based on spatiotemporal data mining. The system is configured with the following modules:
[0118] The data acquisition and preprocessing module 610 is used to acquire and preprocess the access control data, garage data and environmental sensor data of the cell to be tested, and generate a hierarchical data set containing high-frequency access control data, medium-frequency environmental data and low-frequency garage data;
[0119] Behavior feature processing module 620 is used to perform timeliness compensation and missing value filling on the layered data set, and perform behavior fluctuation detection and grouping processing to generate initial weight distribution data and grouped behavior feature weight data;
[0120] Abnormal event identification module 630, used to perform data distribution law anomaly detection based on the initial weight distribution data, call the scene classifier to identify special event types in the distribution characteristics, and generate a scene classification judgment result;
[0121] The scene weight update module 640 is used to perform parameter mapping and weight update processing on the scene classification determination result and the initial weight distribution data to generate scene adaptation weight combination data;
[0122] The dynamic portrait construction module 650 is used to perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, construct a targeted weight matrix of multi-source fusion, and perform weighted integration processing of spatiotemporal features based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the community. The dynamic portrait of the community is used to indicate the distribution characteristics of the population and the resource usage status in the community.
[0123] Preferably, the data acquisition preprocessing module 610 provided in this application is configured with the following units:
[0124] The multi-source data acquisition unit is used to collect access control data, garage data, and environmental sensor data from the test area in real time based on the IoT sensing device interface. It uses data verification rules to filter out abnormal records in the original data stream and fill in missing fields to generate a standard multi-source data set.
[0125] The frequency feature classification unit is used to extract the collection frequency features of the standard multi-source data set, calculate the time interval distribution statistics of each data source, and identify the categories of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data, and generate a frequency classification label set;
[0126] The time window configuration unit is used to divide the frequency classification label set into time windows, configuring a second-level sliding window for high-frequency access control data, a minute-level fixed window for medium-frequency environmental data, and an hour-level fixed window for low-frequency garage data, thereby generating a hierarchical dataset structure.
[0127] The hierarchical data integration unit is used to match and integrate standard multi-source datasets and hierarchical dataset structures, aggregate time series data features by window type and add hierarchical identifiers to generate hierarchical datasets. Hierarchical datasets are used to indicate the temporal hierarchical structure of multi-source spatiotemporal data.
[0128] Preferably, the behavior feature processing module 620 provided in this application is configured with the following units:
[0129] The access control timeliness weighting unit is used to process the layered data set based on the exponential decay function, calculate the timeliness compensation weight of the high-frequency access control data, and generate weighted access control data;
[0130] The garage data smoothing unit is used to process the hierarchical data set based on the linear interpolation algorithm, calculate the missing value filling parameters of the low-frequency garage data, and construct a complete data sequence through the data of adjacent time periods to generate smoothed garage data;
[0131] The initial weight construction unit is used to perform feature merging processing on weighted access control data, smoothed garage data, and medium-frequency environmental data in the layered data set, assign initial weight coefficients according to data source type, construct a unified data structure, and generate initial weight distribution data;
[0132] The behavior grouping feature unit is used to detect and process behavior fluctuations based on the access control data and time window information in the hierarchical data set, analyze the speed variance of the movement trajectory during abnormal time periods, and divide the behavior pattern groups of different groups, and generate group behavior feature weight data. The group behavior feature weight data is used to indicate the spatiotemporal activity characteristics of different resident groups.
[0133] Preferably, the behavior grouping feature unit provided in the present application includes a spatiotemporal data desensitization subunit, an activity intensity analysis subunit and a behavior pattern clustering subunit, wherein the spatiotemporal data desensitization subunit is used to desensitize the access control mobile record data in the hierarchical data set, replace the specific location through area code mapping and remove personal identification information to generate an anonymous spatiotemporal feature data set; the activity intensity analysis subunit is used to jointly analyze and process the anonymous spatiotemporal feature data set and the hierarchical time window, calculate the change rate of pedestrian density and the distribution of stay time in each area during the same period, and generate a regional activity intensity matrix; the behavior pattern clustering subunit is used to perform cluster modeling processing on the regional activity intensity matrix, divide it into life service type, commuting type and leisure type behavior pattern groups according to the peak characteristics of pedestrian density, and generate grouped behavior feature weight data.
[0134] Preferably, the abnormal event identification module 630 provided in this application is configured with the following units:
[0135] The distribution anomaly analysis unit is used to analyze the distribution regularity of the initial weight distribution data, calculate the degree to which the variance of the weight distribution of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data deviates from the historical benchmark value, and generate a data anomaly indicator;
[0136] The special scenario recognition unit is used to classify data anomaly indicators and identify three special scenario types: holiday peak, extreme weather response, and emergency response, by combining public holiday calendar information and meteorological disaster warning signals, and generate event type labels;
[0137] The scene determination and verification unit is used to perform conflict verification on event type tags, eliminate false alarm signals caused by equipment failures, confirm the characteristics of the dominant event, and generate scene classification determination results. The scene classification determination results are used to guide the community security resource allocation strategy.
[0138] Preferably, the scene weight updating module 640 provided in this application is configured with the following units:
[0139] A weight mapping processing unit is used to perform parameter mapping processing on the scene classification determination results, query a preset scene-weight correction coefficient mapping relationship table, and generate a weight correction parameter set. The scene-weight correction coefficient mapping relationship table is used to indicate the adjustment ratio of data weights for different event types;
[0140] The weight update adjustment unit is used to perform coefficient superposition processing on the initial weight distribution data and the weight correction parameter set, adjust the initial weight distribution by superimposing the correction coefficient of the weight correction parameter set according to the data category, and generate updated weight data;
[0141] The weight vector integration unit is used to perform structured integration processing on the updated weight data, uniformly encode the weight values of high-frequency access control data, medium-frequency environmental data and low-frequency garage data into a multi-dimensional vector, and generate scene-adapted weight combination data. The scene-adapted weight combination data is used to guide the real-time data priority allocation strategy.
[0142] Preferably, the dynamic portrait construction module 650 provided in this application is configured with the following units:
[0143] A weight fusion calculation unit is used to perform similarity fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, analyze the scene-behavior matching degree between the group behavior and the current scene feature, and generate a fusion weight coefficient;
[0144] The matrix structure conversion unit is used to perform spatial matrix conversion processing on the fusion weight coefficient, construct a three-dimensional weight distribution structure according to the data source type dimension and the behavior group dimension, and generate a targeted weight matrix;
[0145] The dynamic portrait generation unit is used to perform spatiotemporal feature weighting processing on the real-time data stream in the layered data set, and to fuse the pedestrian density characteristics and facility utilization rate characteristics using a targeted weight matrix to generate a dynamic portrait of the community including a thermal distribution map.
[0146] In one embodiment, the present application also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method for generating dynamic portraits of a community based on spatiotemporal data mining is implemented.
[0147] In one embodiment, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for generating dynamic portraits of a community based on spatiotemporal data mining is implemented.
[0148] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0149] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0150] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for generating a dynamic portrait of a community based on spatiotemporal data mining, characterized in that: The following steps are involved: S1: Acquire and preprocess the access control data, garage data, and environmental sensor data of the cell to be tested, and generate a hierarchical dataset containing high-frequency access control data, medium-frequency environmental data, and low-frequency garage data; S2: performing timeliness compensation and missing value filling on the stratified data set, and performing behavioral fluctuation detection and grouping processing to generate initial weight distribution data and grouped behavioral feature weight data; S3: Performing data distribution law anomaly detection based on the initial weight distribution data, calling a scene classifier to identify special event types in the distribution characteristics, and generating a scene classification determination result; S4: performing parameter mapping and weight updating processing on the scene classification determination result and the initial weight distribution data to generate scene adaptation weight combination data; S5: Perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, construct a targeted weight matrix of multi-source fusion, and perform spatiotemporal feature weighted integration processing based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the cell, which is used to indicate the distribution characteristics of the population and the resource usage status in the cell.
2. The method according to claim 1, characterized in that Said S1 comprises: S11: Based on the IoT sensing device interface, access control data, garage data, and environmental sensor data from the tested community are collected in real time. Data validation rules are used to filter out abnormal records in the original data stream and fill in missing fields to generate a standard multi-source data set. S12: extracting the collection frequency features of the standard multi-source data set, calculating the time interval distribution statistics of each data source, identifying the categories of high-frequency access control data, medium-frequency environmental data, and low-frequency garage data, and generating a frequency classification label set; S13: Divide the frequency classification label set into time windows, configure a second-level sliding window for high-frequency access control data, a minute-level fixed window for medium-frequency environmental data, and an hour-level fixed window for low-frequency garage data, to generate a hierarchical data set structure; S14: Matching and integrating the standard multi-source dataset and the hierarchical dataset structure, aggregating time series data features by window type and adding hierarchical identifiers to generate a hierarchical dataset, where the hierarchical dataset is used to indicate the temporal hierarchical structure of the multi-source spatiotemporal data.
3. The method according to claim 1, characterized in that The S2 includes: S21: Processing the layered data set based on an exponential decay function, calculating a timeliness compensation weight of the high-frequency access control data, and generating weighted access control data; S22: Processing the layered data set based on a linear interpolation algorithm, calculating missing value filling parameters for low-frequency garage data, and constructing a complete data sequence through adjacent time period data to generate smoothed garage data; S23: performing feature merging processing on the weighted access control data, the smoothed garage data, and the medium-frequency environmental data in the layered data set, assigning initial weight coefficients according to data source types, and constructing a unified data structure to generate initial weight distribution data; S24: Based on the access control data and time window information in the hierarchical data set, behavioral fluctuation detection processing is performed, the speed variance of the movement trajectory is analyzed for abnormal time periods, and the behavioral pattern groups of different groups are divided, and grouped behavioral feature weight data is generated. The grouped behavioral feature weight data is used to indicate the spatiotemporal activity characteristics of different resident groups.
4. The method according to claim 3, characterized in that The calculation formula of the timeliness compensation weight is: ; in, is the timeliness compensation weight, is the weight of the original access control data, is the attenuation coefficient, Generate the time difference between the current time and the high-frequency access control data, is a natural constant.
5. The method according to claim 3, characterized in that The S24 includes: S241: Desensitizing the access control movement record data in the hierarchical dataset, replacing specific locations with area code mapping and removing personal identification information to generate an anonymous spatiotemporal feature dataset; S242: Jointly analyzing and processing the anonymous spatiotemporal feature dataset and the hierarchical time window, calculating the change rate of pedestrian flow density and the distribution of length of stay in each area during the same period, and generating a regional activity intensity matrix; S243: Perform cluster modeling processing on the regional activity intensity matrix, divide it into life service type, commuting type and leisure type behavior pattern groups according to the peak characteristics of the pedestrian flow density, and generate group behavior feature weight data.
6. The method according to claim 1, characterized in that The S3 includes: S31: Analyzing the distribution regularity of the initial weight distribution data, calculating the degree to which the variance of the weight distribution of the high-frequency access control data, the medium-frequency environmental data, and the low-frequency garage data deviates from the historical benchmark value, and generating a data anomaly indicator; S32: Scenario classification is performed on the data anomaly indicators, and three special scenario types, namely holiday peak, extreme weather response, and emergency response, are identified by combining public holiday calendar information and meteorological disaster warning signals, and event type labels are generated; S33: Perform conflict verification on the event type tag to eliminate false alarm signals caused by equipment failure and confirm the dominant event characteristics, and generate a scene classification judgment result. The scene classification judgment result is used to guide the community security resource configuration strategy.
7. The method according to claim 1, characterized in that The S4 includes: S41: performing parameter mapping processing on the scene classification determination result, querying a preset scene-weight correction coefficient mapping relationship table to generate a weight correction parameter set, wherein the scene-weight correction coefficient mapping relationship table is used to indicate the adjustment ratio of data weights for different event types; S42: performing coefficient superposition processing on the initial weight distribution data and the weight correction parameter set, superimposing the correction coefficients of the weight correction parameter set according to data categories to adjust the initial weight distribution and generate updated weight data; S43: Perform structured integration processing on the updated weight data, uniformly encode the weight values of high-frequency access control data, medium-frequency environmental data and low-frequency garage data into a multi-dimensional vector, and generate scene adaptation weight combination data. The scene adaptation weight combination data is used to guide the real-time data priority allocation strategy.
8. The method according to any one of claims 1 to 7, characterized in that The S5 includes: S51: performing similarity fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, analyzing the scene-behavior matching degree between the group behavior and the current scene feature, and generating a fusion weight coefficient; S52: performing spatial matrix conversion processing on the fusion weight coefficient, constructing a three-dimensional weight distribution structure according to the data source type dimension and the behavior group dimension, and generating a targeted weight matrix; S53: Perform spatiotemporal feature weighting processing on the real-time data stream in the layered data set, apply a targeted weight matrix to fuse the crowd density feature and the facility utilization rate feature, and generate a dynamic portrait of the community including a thermal distribution map.
9. The method according to claim 8, characterized in that The calculation formula of the scene-behavior matching degree is: ; in, is the scene-behavior matching degree, For the Class scene weight value, For the Class behavior weight value, is the total number of weight categories, is the modulus of the scene weight vector, is the scene adjustment factor.
10. A system for generating dynamic community portraits based on spatiotemporal data mining, characterized in that: The system comprises: The data acquisition and preprocessing module is used to acquire and preprocess the access control data, garage data and environmental sensor data of the tested community to generate a hierarchical data set containing high-frequency access control data, medium-frequency environmental data and low-frequency garage data; A behavior feature processing module is used to perform timeliness compensation and missing value filling on the layered data set, and to perform behavior fluctuation detection and grouping processing to generate initial weight distribution data and grouped behavior feature weight data; An abnormal event identification module is used to perform data distribution law abnormality detection based on the initial weight distribution data, call a scene classifier to identify special event types in the distribution characteristics, and generate a scene classification judgment result; A scene weight updating module is used to perform parameter mapping and weight updating processing on the scene classification determination result and the initial weight distribution data to generate scene adaptation weight combination data; The dynamic portrait construction module is used to perform correlation fusion processing on the scene adaptation weight combination data and the group behavior feature weight data, construct a targeted weight matrix of multi-source fusion and perform weighted integration processing of spatiotemporal features based on the environmental sensor data in the hierarchical data set to generate a dynamic portrait of the cell. The dynamic portrait of the cell is used to indicate the distribution characteristics of the population and the resource usage status in the cell.
Citation Information
Patent Citations
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Perceptual data-based spatio-temporal portrait analysis method for personnel and places
CN119476429A