A regional economic vitality intelligent monitoring method based on big data analysis

By constructing a dynamic and evolvable economic vitality monitoring indicator system and integrating multi-level features, and by accessing high-frequency data in real time for anomaly detection, the problems of adaptability and accuracy in regional economic vitality monitoring have been solved, enabling dynamic, precise, and real-time monitoring and anomaly identification of regional economic vitality.

CN122491967APending Publication Date: 2026-07-31JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-05-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for monitoring regional economic vitality suffer from problems such as the inability of indicator systems to evolve adaptively, insufficient real-time data utilization, inadequate spatial accuracy, and insufficient fusion of multi-source data, resulting in lagging and inaccurate monitoring results.

Method used

Construct a dynamic and evolvable economic vitality monitoring indicator system, access multi-source heterogeneous high-frequency dynamic data in real time, and conduct anomaly assessment and early warning through spatial gridding and multi-level feature fusion, combined with a distributed anomaly detection model.

Benefits of technology

It enables dynamic, precise, and real-time monitoring of regional economic vitality, and can promptly identify abnormalities in economic activity and provide accurate policy recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491967A_ABST
    Figure CN122491967A_ABST
Patent Text Reader

Abstract

This invention relates to the field of regional economic monitoring and big data analysis technology, and discloses an intelligent monitoring method for regional economic vitality based on big data analysis. This method uses an information gain ratio algorithm to periodically screen indicators to construct a dynamically evolving indicator system; it accesses multi-source high-frequency data streams in real time, and generates multi-dimensional temporal feature vectors through streaming preprocessing and spatial gridding aggregation; it uses a two-level fusion mechanism of intra-unit attention and spatial graph attention to generate spatial context-aware fusion feature vectors; it uses an isolated forest model for multi-scale anomaly detection, identifies spatial anomaly units and their deviation types, determines anomaly clusters through comprehensive distance clustering, extracts backbone anomaly paths, and finally generates intelligent early warning information. This invention solves the problems of fixed indicator systems, poor data timeliness, coarse spatial granularity, and shallow fusion, achieving dynamic, refined, and real-time monitoring and early warning of regional economic vitality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of regional economic monitoring and big data analysis technology, specifically to an intelligent monitoring method for regional economic vitality based on big data analysis. Background Technology

[0002] Regional economic vitality is a core indicator for measuring a region's overall competitiveness and development potential. Accurate and timely monitoring of regional economic vitality is crucial for formulating economic policies, optimizing resource allocation, and guiding industrial layout. However, traditional methods for monitoring regional economic vitality still have many shortcomings in terms of indicator system construction, data utilization, analysis timeliness, spatial accuracy, and multi-source data fusion, making it difficult to meet the urgent needs of current economic governance for monitoring accuracy and response speed.

[0003] Regarding indicator systems, existing technologies generally employ pre-determined fixed sets of indicators and fixed weights for evaluation. Once set, these systems remain unchanged, failing to adapt to the dynamic evolution of regional economic structures. When regional industrial structures undergo transformation, existing indicators may lack measurement items for emerging economic activities, leading to systematic deviations from reality in evaluation results. Similarly, fixed weights cannot reflect changes in the relative importance of factors across different economic cycle stages.

[0004] In terms of data utilization, existing technologies mostly rely on low-frequency updated data such as corporate annual reports and industry planning data, lacking effective utilization of real-time behavioral data such as logistics activity, nighttime economy intensity, and frequency of personnel movement. This makes it difficult to capture high-frequency fluctuations and subtle signals in economic vitality, resulting in obvious monitoring blind spots.

[0005] In terms of the timeliness of analysis, existing technologies generally adopt a batch processing mode, which often takes several days or even weeks from data generation to the output of evaluation results. The monitoring results are essentially a retrospective description of historical states. When early signs of sudden changes in the economy appear, the delayed response may cause policy intervention to miss the best window of opportunity.

[0006] In terms of spatial accuracy, existing monitoring technologies are mostly limited to the macro scale of provincial and municipal administrative divisions. Some schemes use nighttime light remote sensing data as a proxy indicator, but nighttime light data has drawbacks such as limited activity coverage, low collection frequency, and unclear mapping with economic activity types. Using administrative boundaries as the unit of analysis masks the spatial heterogeneity within the region and fails to reveal the micro-level differences in vitality between different streets and business districts.

[0007] In terms of multi-source data fusion, existing technologies mostly adopt simple weighted summation or feature splicing methods, ignoring the nonlinear interaction relationship and synergistic effect between data of different dimensions, making it difficult to characterize the deep coupling mechanism between factors such as enterprise registration growth and logistics activity improvement.

[0008] In summary, existing technologies have significant shortcomings in areas such as adaptive evolution of indicator systems, utilization of high-frequency real-time data, timeliness of analysis, fine-grained spatial analysis, and nonlinear fusion of multi-source data. There is an urgent need for an intelligent monitoring method that can systematically solve the above problems. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides an intelligent monitoring method for regional economic vitality based on big data analysis, thereby resolving the problems mentioned in the background.

[0010] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent monitoring of regional economic vitality based on big data analysis, comprising the following steps: S1. Construct a dynamic and evolvable economic vitality monitoring indicator system. The indicator system includes a core indicator layer and multiple extended indicator layers. The core indicator layer includes a unified basic vitality indicator. Each extended indicator layer corresponds to a type of regional economic characteristic. The indicator items in each extended indicator layer are dynamically increased or decreased as the regional economic structure changes. Using an information gain rate algorithm, effective indicators with an information gain rate exceeding a preset threshold in the current monitoring period are selected from the mutual information between each indicator item and a preset economic vitality benchmark label. The selected effective indicators are dynamically combined to form the monitoring indicator set for the current period. Preferably, the basic vitality indicators of the core indicator layer in step S1 include the rate of change in the number of registered enterprises, the density of newly added market entities, the activity level of factor flows, and the public service load index; the extended indicator layer includes at least three of the following: the manufacturing vitality extended indicator layer, the service industry vitality extended indicator layer, the emerging industry vitality extended indicator layer, and the open economy vitality extended indicator layer; the candidate indicator items in each of the extended indicator layers are periodically screened by the information gain rate algorithm, so that different indicator combinations can be used for vitality monitoring in different monitoring periods.

[0011] Preferably, the preset economic vitality benchmark label in step S1 is generated by dynamically weighting and fusing the regional GDP growth rate and tax revenue growth rate. The dynamic weights are updated adaptively based on the contribution of the core indicator layer and the extended indicator layer in the previous monitoring period, forming a closed-loop self-feedback mechanism of indicator screening, vitality assessment, and weight adaptation. The preset threshold is automatically adjusted based on adaptive kernel density estimation. The ability of each candidate indicator to distinguish economic vitality is measured by calculating the information gain rate between each candidate indicator and the economic vitality benchmark label. When the information gain rate of any candidate indicator is lower than the preset threshold in the current period, it is removed from the monitoring indicator set of the current period and re-evaluated in the next period.

[0012] S2. Real-time access to multi-source heterogeneous high-frequency dynamic data, including enterprise business registration data stream, logistics transportation trajectory data stream, commercial consumption transaction data stream, talent flow data stream, and public service activity data stream; and perform streaming preprocessing on the multi-source heterogeneous high-frequency dynamic data to obtain a standardized real-time data stream in a unified format. Preferably, in step S2, the access of multi-source heterogeneous high-frequency dynamic data adopts the message middleware Kafka to build a distributed data bus. Each data source pushes data in real time in a streaming manner through the data bus. Dynamic semantic recognition and conflict resolution are performed at the data access edge. Fast field normalization is compatible with semi-structured and unstructured information. The streaming preprocessing includes data format standardization, timestamp alignment, missing value imputation and outlier detection processing. The access of the logistics transportation trajectory data stream includes accessing freight vehicle trajectory data, logistics order flow data and warehouse inbound and outbound flow data.

[0013] S3. Divide the target area to be monitored into multiple spatial grid units, and use each spatial grid unit as the statistical granularity to perform spatial aggregation statistics on various types of data in the standardized real-time data stream to generate a multi-dimensional temporal feature vector of each spatial grid unit in each monitoring time window. Preferably, the spatial grid unit division method in step S3 is as follows: the target area is divided into regular hexagonal grids with a preset side length, and each hexagonal grid serves as a basic monitoring unit covering the entire target area; when performing spatial aggregation statistics on various types of real-time data, the data is allocated to the corresponding hexagonal grid according to its latitude and longitude coordinates, and the temporal characteristic indicators of various types of data in each hexagonal grid are statistically analyzed with a preset time window as the statistical period, forming a multidimensional temporal characteristic vector sequence of each hexagonal grid in a continuous time window.

[0014] S4. Perform multi-level feature fusion on the multi-dimensional temporal feature vectors of each spatial grid unit, including: the first level performs attention-weighted fusion on the temporal features of each dimension within each spatial grid unit to generate a unit-level fused feature vector; the second level performs spatial context fusion on the unit-level fused feature vectors based on the spatial adjacency relationship between each spatial grid unit using a spatial self-attention mechanism to generate a spatial context-aware fused feature vector for each spatial grid unit. Preferably, the first-level attention-weighted fusion method in step S4 is as follows: learnable attention weight coefficients are assigned to the temporal features of each dimension within each spatial grid unit, and the temporal features of each dimension are aggregated into a unit-level fusion feature vector by weighted summation; the second-level spatial self-attention fusion method is as follows: a weighted spatial graph based on road network and functional area is constructed, with each spatial grid unit as a node in the graph and the connection weight in the weighted spatial graph as a quantitative expression of spatial adjacency relationship. The graph attention network is used to perform message passing and feature updating of each node feature using the weighted spatial graph as a mask to generate a spatial context-aware fusion feature vector for each spatial grid unit.

[0015] S5. Input the spatial context-aware fusion feature vector of each spatial grid unit into the pre-constructed distributed economic vitality anomaly detection model. The distributed economic vitality anomaly detection model scores the vitality anomaly of each spatial grid unit on a unit-by-unit basis, identifies spatial anomaly units whose vitality status deviates from the historical normal pattern and the deviation type, the deviation type including rise, decline or structural inactivation; and determines anomaly clusters based on the spatial aggregation characteristics of the spatial anomaly units, and extracts economic vitality signals within the anomaly clusters. Preferably, the distributed economic vitality anomaly detection model in step S5 is constructed using the isolated forest algorithm. During model training, the spatial context-aware fusion feature vectors of each spatial grid unit during historical normal periods are used as training samples. The isolated forest model supports feature dimension alignment and incremental retraining as the monitoring index set changes dynamically. When scoring anomalies, the average path length of the fusion feature vectors of each spatial grid unit in the isolated forest during the current monitoring period is calculated, and the normalized value of the average path length is used as the vitality anomaly score. When the vitality anomaly score of any spatial grid unit exceeds the anomaly judgment threshold, the spatial grid unit is marked as a spatial anomaly unit.

[0016] Preferably, the distributed economic activity anomaly detection model employs a multi-scale time detection mechanism, including a short-term detection window, a medium-term detection window, and a long-term detection window, each corresponding to a historical normal pattern baseline over a different time span. The short-term detection window is used to detect weekly-level economic activity fluctuation anomalies, the medium-term detection window is used to detect monthly-level economic activity trend changes, and the long-term detection window is used to detect quarterly-level economic activity structural changes. The anomaly scores from each detection window are weighted and combined to obtain a comprehensive economic activity anomaly score.

[0017] Preferably, the method for determining the anomalous clusters in step S5 is as follows: based on the comprehensive distance matrix that integrates spatial adjacency distance and temporal feature similarity, the DBSCAN clustering algorithm is used to perform spatial clustering on the spatial grid cells marked as spatial anomalous units, and the continuous anomalous regions formed by the clustering are determined as anomalous clusters. The backbone anomalous paths of each anomalous cluster are extracted to track the dynamic propagation direction of the anomalous signal in space. The backbone anomalous paths are extracted in the following way: taking each spatial grid cell in the anomalous cluster as a node, and using the reciprocal of the absolute value of the difference in anomalous scores between units as the edge weight, a directed weighted graph is constructed. The shortest path algorithm or minimum spanning tree is used to extract the propagation path with the largest anomalous score gradient, and the unit with the earliest occurrence of an anomalous time on the path is taken as the candidate for the anomalous source. Extracting the economic vitality signal in the anomalous cluster includes: statistically analyzing the economic vitality anomalous scores of each spatial grid cell in the anomalous cluster to obtain the area, anomalous intensity, and anomalous duration of the anomalous cluster, and backtracking the distribution of the attention weight coefficients in each dimension in step S4 to identify the temporal feature dimension that contributes the most to the anomalous score as the basis for tracing the source of the anomalous signal.

[0018] S6. Based on the economic vitality signal within the abnormal cluster, generate intelligent monitoring and early warning information that includes abnormal spatial location information, abnormality level, and abnormal signal source analysis results.

[0019] Preferably, the generation of intelligent monitoring and early warning information in step S6 further includes: mapping the spatial boundary of the anomaly cluster to the actual geographical area, associating the enterprise directory and industry distribution information in the corresponding area, generating a comprehensive early warning report containing the affected area range, the severity level of the anomaly, the key driving factors of the anomaly, and the list of associated key enterprises, and sending the comprehensive early warning report to the terminal in real time through a message push interface.

[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention constructs a dynamically evolving economic vitality monitoring indicator system, introduces an information gain rate algorithm to periodically screen indicator items, and combines a dynamic weight adaptive update mechanism for economic vitality benchmark labels to form a closed-loop self-feedback of indicator screening, vitality assessment, and weight adaptation. This solves the problem that traditional fixed indicator systems cannot adapt to the dynamic evolution of regional economic structures and ensures that monitoring indicators always remain adapted to regional economic characteristics.

[0021] 2. This invention overcomes the limitations of traditional methods that rely on low-frequency statistical data by real-time access to multi-source heterogeneous high-frequency dynamic data such as enterprise registration, logistics and transportation trajectories, commercial consumption transactions, talent mobility and public service activities, and performs dynamic semantic recognition and conflict resolution at the data access edge. It can capture high-frequency fluctuation characteristics and subtle signals in economic vitality, significantly improving the timeliness and sensitivity of monitoring.

[0022] 3. This invention divides the target area into regular spatial grid units, performs fine-grained spatial aggregation statistics at the grid level, and combines multi-level attention feature fusion with a weighted spatial map based on road networks and functional zones to achieve a precise depiction of the micro-spatial pattern of economic vitality. This solves the problem of traditional methods using administrative divisions as units to mask the spatial heterogeneity within a region.

[0023] 4. This invention constructs a distributed economic vitality anomaly detection model by employing the isolated forest algorithm and introduces a multi-scale time detection mechanism to comprehensively evaluate anomalies at the weekly, monthly, and quarterly levels. It can accurately identify anomalous units that deviate from historical normal patterns and their deviation types, trace the dynamic propagation path of anomalous signals in space, and provide timely and reliable decision-making basis for precise intervention in economic policies. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall process of the method of the present invention; Figure 2 This is a schematic diagram illustrating the construction and screening process of the dynamic and evolvable index system of this invention; Figure 3 This is a schematic diagram of the intelligent monitoring and early warning information generation and push process of the present invention; Figure 4 This is a schematic diagram of the multi-level spatiotemporal feature fusion process of the present invention; Figure 5 This is a schematic diagram of the anomaly detection and anomaly cluster analysis process of the present invention. Detailed Implementation

[0025] Please see Figures 1-5 This embodiment provides a method for intelligent monitoring of regional economic vitality based on big data analysis. The method consists of six core steps sequentially linked to form a complete processing chain, covering the entire process from dynamic construction of a monitoring indicator system, real-time access and spatial grid aggregation of multi-source high-frequency data, multi-level spatiotemporal feature fusion, to distributed anomaly detection and anomaly cluster analysis, and finally the generation and delivery of intelligent monitoring and early warning information. Each step works in concert to achieve dynamic, precise, and real-time monitoring of regional economic vitality, as detailed below: Step S1: Construct a dynamic and evolvable economic vitality monitoring indicator system Step S1 establishes a monitoring indicator system that adapts to changes in the regional economic structure.

[0026] 1. Establish a hierarchical structure for the indicator system. The indicator system is structured in two layers: a core indicator layer and an extended indicator layer. The core indicator layer includes four basic vitality indicators: the rate of change in the number of registered enterprises (the ratio of the difference between the number of registered enterprises at the end of the current period and the end of the previous period to the number of registered enterprises at the end of the previous period), the density of newly registered market entities (the number of newly registered market entities per unit area), the activity level of factor mobility, and the public service load index. The core indicator layer remains stable throughout the monitoring period, forming the basic framework for evaluation.

[0027] Above the core indicator layer, at least three extended indicator layers are established, categorized into four types: manufacturing vitality, service sector vitality, emerging industry vitality, and open economy vitality. Each category contains several candidate indicator items. For example, the manufacturing vitality extended indicator layer may include industrial electricity consumption and industrial land transfer area; the service sector vitality extended indicator layer may include total retail sales of consumer goods and catering revenue; the emerging industry vitality extended indicator layer may include the number of newly added high-tech enterprises and R&D investment intensity; and the open economy vitality extended indicator layer may include total import and export volume and actual utilization of foreign capital. The indicator items in each extended indicator layer are dynamically added or removed according to changes in the regional economic structure, and are added or deleted based on adjustments to regional industrial planning or changes in statistical methods, ensuring that the indicator system always matches the regional economic characteristics.

[0028] 2. Perform periodic indicator screening. At the beginning of each monitoring period, the information gain ratio algorithm is used to screen candidate indicator items in the extended indicator layer, forming the monitoring indicator set for the current period together with the core indicator layer. The specific screening method is the information gain ratio algorithm. To perform the information gain ratio calculation, an economic vitality benchmark label needs to be determined. In this embodiment, the economic vitality benchmark label is generated by dynamically weighting and fusing the regional GDP growth rate and tax revenue growth rate. Let the economic vitality benchmark label be... The regional GDP growth rate was Tax revenue growth rate was The dynamic weights are respectively and Then we have: ,in, and satisfy ,and , The initial values ​​of these two dynamic weights can each be set to 0.5, and then automatically updated according to the feedback mechanism, which will be detailed later.

[0029] Calculate the relationship between each candidate indicator and the benchmark label. The information gain ratio between them. Let a candidate index be a discretized feature variable. Its possible values ​​are First, calculate the baseline label. Information entropy : ,in, Indicates the reference label Values The probability is obtained by frequency statistics of historical sample data.

[0030] Then, given the known values ​​of the feature variable X, the benchmark label is calculated. Conditional information entropy : ,in, This indicates that the characteristic variable X takes the value of The probability, Indicates in The information entropy of the benchmark label Y on a subset of samples.

[0031] Information gain Defined as the difference between information entropy and conditional information entropy: ; Information gain ratio Introducing SplitInfo based on Information Gain As a penalty factor, to avoid the algorithm favoring features with too many values: ; when When the value is 0, the candidate indicator is directly excluded.

[0032] Candidate indicators whose information gain ratio exceeds a preset threshold are included in the current period's monitoring indicator set, while those below the threshold are removed but retained in the candidate pool for re-evaluation in the next period, allowing different combinations of indicators to be used in different monitoring periods.

[0033] 3. Operation of the closed-loop self-feedback mechanism. The dynamic weights in the aforementioned economic vitality benchmark labels. and The preset threshold for information gain ratio screening is adjusted through an adaptive mechanism, forming a complete closed loop of "indicator screening - vitality assessment - weight adaptation".

[0034] Dynamic weight updates: At the end of each monitoring period, review the contribution of each indicator to the vitality assessment results within that period (quantified by the attention weight coefficient in step S4 or the importance score in anomaly detection in step S5). If the contribution of indicators highly correlated with GDP growth is high, while the discriminative power of indicators related to tax revenue growth declines due to one-off policy interference, then the weights will be increased in the next period. Lower Let the weight of the t-th monitoring period be... The feedback signal of the indicator contribution calculated after this round of evaluation is: The weights will then be updated as follows for the next cycle: ,in, The learning rate determines the step size for weight updates, typically a small positive value such as 0.01 to 0.05 to ensure smooth weight changes. After the update... and Perform normalization to ensure that the sum of the two is 1.

[0035] For adaptive adjustment of the preset threshold, an adaptive kernel density estimation method is used. Let the information gain ratio values ​​of all candidate indicator items recorded over the past M periods constitute the sample set. , , …, The kernel density estimation function is: ,in, For the kernel function, this embodiment uses the Gaussian kernel function; The bandwidth parameter controls the smoothness of the density estimate and can be determined using the Silverman rule. Based on the estimated probability density distribution, the lower bound of the probability density function is taken. quantiles (e.g.) The 20th percentile (i.e., the 20th percentile) is used as the preset threshold for the next cycle, so that the threshold is automatically adjusted according to the actual distribution of information gain rate, and always maintains a moderately strict screening intensity for candidate indicators.

[0036] Through the above three steps, step S1 establishes a dynamic indicator system with adaptive capabilities, providing a monitoring dimension framework that matches the current economic environment for subsequent steps.

[0037] Step S2: Real-time access and streaming preprocessing of multi-source heterogeneous high-frequency dynamic data This step enables real-time access and standardized processing of multi-source heterogeneous high-frequency dynamic data, providing a unified data input format for subsequent steps.

[0038] 1. A distributed data bus is built using the message middleware Kafka. Data from each data source is pushed to the corresponding topic in real time in a streaming manner: business registration data is pushed to "business_register", logistics transportation trajectory data is pushed to "logistics_track", commercial consumption transaction data is pushed to "commerce_trans", talent flow data is pushed to "talent_flow", and public service activity data is pushed to "public_service". Parallel read and write operations for each topic partition ensure high throughput and low latency.

[0039] The five data streams cover: business registration data stream (enterprise establishment, alteration, and cancellation registration records, including unified social credit code, enterprise name, registered capital, and registered address latitude and longitude); logistics transportation trajectory data stream (freight vehicle GPS trajectory, logistics order origin and destination and completion time, warehouse inbound and outbound volume); commercial consumer transaction data stream (offline POS and online payment transaction amount, time, merchant latitude and longitude); talent mobility data stream (job postings and resume submissions, changes in social security payment location, and employment destinations of college graduates); and public service activity data stream (public transportation card swipes, public utility payments, and government service queuing volume).

[0040] 2. Perform dynamic semantic recognition and conflict resolution at the data access edge. Before data enters the Kafka bus, deploy an edge processing module close to the data source to perform the following processing: Dynamic semantic recognition: The edge processing module maintains a dynamically updatable semantic mapping table, recording the correspondence between identical semantic fields from different source systems (e.g., "enterprise registration address" is "reg_address" in data source A and "qy_zcdz" in data source B). This table is updated periodically through synchronization or manual maintenance. When data arrives at the edge, the module looks up the mapping rules based on the data source identifier and converts the original fields into standardized field names.

[0041] Conflict Resolution: When there are field value conflicts between different versions of the same data from different sources, the edge processing module resolves them based on a preset priority strategy and data credibility rating. For example, if a company's registered address differs between business registration data and social security registration data, the module prioritizes the address information in the business registration data based on the preset rule that "business registration data has higher credibility in the registered address field," while retaining a conflict marker for subsequent manual verification. For the timestamp field, if the time deviation of multiple source records of the same business event is within a preset tolerance range (e.g., 5 minutes), it is aligned to the earliest recorded timestamp; if the deviation exceeds the tolerance range, it is retained as an independent event to avoid information loss due to simple alignment.

[0042] After edge semantic recognition and conflict resolution, heterogeneous data from various sources has been converted into preliminary standardized data that follows a unified field naming convention and eliminates field semantic conflicts.

[0043] 3. Streaming Preprocessing. The initially standardized data output from the edge is further processed by data format standardization, timestamp alignment, missing value imputation, and outlier detection to form a final standardized real-time data stream that can be directly used in step S3.

[0044] Data format standardization: All numeric fields are uniformly converted to double-precision floating-point format, timestamp fields are uniformly converted to Unix timestamp millisecond format, category fields are uniformly converted to predefined enumeration type encoding, and text fields are uniformly converted to UTF-8 encoding format.

[0045] Timestamp alignment: For timestamps of different records in the same data stream, due to data transmission delays, system clock deviations, etc., timestamps may not be strictly arranged in the order of event occurrence. An event sequence sorting method based on event time is adopted, maintaining a sliding window of a preset duration (e.g., 30 seconds). Within the window, arriving data records are reordered according to event occurrence time, and the sorted record sequence is output.

[0046] Missing value imputation: For cases where individual fields in a record are missing, imputation is performed based on the field type and context information. Missing values ​​for numeric fields are filled with the mean of the field in the previous time window; missing values ​​for categorical fields are filled with the mode of the field in the previous time window; missing values ​​for geographic coordinate fields are not imputed, and the record is marked as a missing coordinate record. In step S3, it is temporarily stored separately during spatial aggregation and will be reallocated after the coordinate information is completed.

[0047] Outlier detection and handling: Outlier detection is performed on numeric fields using a method based on the interquartile range. Let the 25th percentile of a numeric field within the sliding window be... The 75th percentile is Interquartile range If a certain value is less than or greater than If a value is found to be outlier, it is considered an outlier. Outliers are processed using the same mean imputation strategy as missing values, and outlier markers are recorded for data quality assessment.

[0048] After completing the above streaming preprocessing, the resulting standardized real-time data stream is consumed by step S3 in a unified format. Each data record contains standardized field names, standardized data types, aligned event timestamps, and latitude and longitude coordinates (records with missing coordinates are marked separately).

[0049] Step S3: Spatial gridding aggregation and multidimensional temporal feature generation This step aggregates the standardized real-time data stream in both spatial and temporal dimensions, transforming it into a multi-dimensional temporal feature vector with grid cells as the granularity and time windows as the period.

[0050] 1. Spatial grid cell division. The target region is divided into regular hexagonal grids with preset side lengths. Hexagonal grids have the optimal boundary-to-area ratio under the condition of equal area, and the center distance between adjacent cells is equal everywhere, which is beneficial for subsequent spatial self-attention mechanism modeling of adjacency relationships.

[0051] The preset side length is selected based on spatial accuracy requirements (e.g., 500 meters in urban core areas, 1000 to 2000 meters in urban-rural fringe areas and rural areas). Each hexagonal grid is assigned a unique identifier, and each unit seamlessly and completely covers the entire target area.

[0052] 2. Extract latitude and longitude coordinates from each record in the standardized real-time data stream, and map them to the corresponding hexagonal grid cells using a spatial coordinate transformation algorithm (mapping the point to the cell belonging to the nearest hexagonal center point based on the latitude and longitude of the grid center point, side length, and rotation angle). Records with missing coordinates are stored in a buffer, and the coordinates are filled in before backtracking and allocation.

[0053] 3. Set a fixed time window as the statistical period (such as one hour or one day). At the end of each statistical period, statistically analyze the various types of data accumulated in the past time window for each grid cell and generate multi-dimensional time series characteristic indicators for that grid cell in that time window.

[0054] Taking a single hexagonal grid cell u within a time window t as an example, the statistically generated feature dimensions include: business registration (number of newly registered enterprises, number of deregistered enterprises, and number of registered enterprises), logistics activities (density of freight vehicle trajectory points, number of logistics orders, and net value of warehouse inflows and outflows), consumption activities (number of transactions, average transaction amount, and proportion of transactions during nighttime), talent mobility (number of job postings, number of resumes submitted, and net inflow of social security personnel), and public services (number of public transportation card swipes, number of water and electricity bill payments, and number of queuing numbers for government services). All of these indicators are collectively referred to as dimensional features. For a single grid cell within a time window, a total of d feature values ​​are generated, denoted as the feature vector. ,in, Represents grid cells In the time window Inner Statistical values ​​of each dimension feature This represents the total number of dimensions.

[0055] As time progresses, each grid cell In continuous A multidimensional temporal feature vector sequence was formed over the time window: ,in, This represents the total number of statistical time windows to date. This sequence serves as the input data for step S4, where multi-level feature fusion is performed.

[0056] Step S4: Multi-level spatiotemporal feature fusion This step performs two-level feature fusion on the multidimensional temporal feature vectors of each grid cell, compressing the high-dimensional original features into a low-dimensional embedded representation that incorporates the spatial context.

[0057] First level: Intra-cell attention weighted fusion. For each grid cell... In each monitoring time window ,That 3D original feature vector This includes simultaneous window readings of multiple economic vitality indicators. These indicators vary in their ability to reflect the state of economic vitality and their importance may change over time and across different regions. This layer employs an attention mechanism, using learnable weighting coefficients to weight and fuse features across different dimensions.

[0058] Specifically, it refers to spatial grid units. of Assign a set of learnable attention weights to the dimensional features: ,in, and The weight coefficients are learned through a lightweight attention network. The input to the attention network is grid cells. The time-series feature statistics (including the mean and variance of each dimension of features over the time window) over several consecutive time windows are output as a normalized d-dimensional weight vector.

[0059] After obtaining the attention weights, for We perform weighted aggregation across all dimensions to generate a unit-level fused feature vector: ,in, for Dimensional unit-level fusion of feature vectors (when using multiple attention heads) It can be greater than 1 when using a single attention head and directly outputting a scalar. In this embodiment, each attention head outputs a fusion scalar value, and the outputs of multiple attention heads are concatenated to form an intra-unit fusion feature embedding. The process of generating unit-level fusion feature vectors is a dimensionality reduction and information condensation of the original high-dimensional features, so that the feature dimensions with the most economically viable representation capabilities receive greater information retention weights.

[0060] Second level: Spatial self-attention fusion based on weighted spatial graphs. The unit-level fusion feature vector generated by the first-level fusion is... This only includes information about the grid cell itself, without considering its association with surrounding grid cells. Economic activities exhibit significant spatial spillover effects and radiation characteristics; the economic vitality of a grid cell is often influenced by its neighboring or functionally related cells, and may also radiate outwards. This level constructs a weighted spatial map based on road networks and functional zones, and utilizes graph attention networks to fuse the spatial context of each cell's features.

[0061] The specific method for constructing a weighted spatial map is as follows. Each hexagonal grid cell is considered a node in the map. The existence and weight of connecting edges between nodes are not simply determined by the Euclidean distance between the cell centers, but rather by a comprehensive consideration of road network connectivity and functional area correlation between cells. Road network connectivity is calculated using GIS road network data: if two grid cells are directly connected by a graded road (national highway, provincial highway, urban arterial road, or higher), and the actual travel distance along the road does not exceed a preset threshold (e.g., 3 kilometers), then a road connection is considered to exist between them. Functional area correlation is determined using point-of-interest (POI) data and land use type data: if two grid cells are designated as the same type of functional area (e.g., both are commercial areas, both are industrial areas), or have complementary functions (e.g., a commercial area and a nearby residential area), then a functional correlation exists between them. For node pairs that simultaneously satisfy both road connectivity and functional area correlation, a higher connecting edge weight is assigned; for node pairs that satisfy only one of these conditions, a medium edge weight is assigned; for node pairs that satisfy neither condition, no connecting edge is assigned.

[0062] The adjacency weight matrix, composed of connection weights, serves as a mask for message passing in the subsequent graph attention network. Let the total number of nodes in the graph be... The adjacency weight matrix is ,in For nodes To the node The weight of the connecting edges; if and If there is no boundary between them, then .

[0063] In a graph attention network, each node The Layer input features are First, for the nodes... With all its neighboring nodes Attention coefficient between Perform the calculation: ,in, The shared linear transformation weight matrix maps node features to a new feature space; Wa is the attention coefficient calculation vector; [·‖·] represents the vector concatenation operation; This is the activation function. Next, the attention coefficients mentioned above are combined with the connection weights of the weighted spatial graph. Element-wise fusion is performed to obtain the normalized attention score masked by the graph. : ,when When (i.e., the two nodes are not connected by an edge). The neighboring node does not participate in message aggregation. This mechanism ensures that only neighboring units that are determined to have a real economic and geographical relationship through the road network and functional areas will affect the feature update of this unit, thus avoiding the erroneous inclusion of spatially adjacent but geographically isolated unrelated units into information fusion.

[0064] node The updated features are attention-weighted aggregations of the features of its neighboring nodes: ,in, Transform the weight matrix to its value; This is a non-linear activation function. After message passing and feature updating through a multi-layer graph attention network, the output feature vector of each spatial grid cell is its spatial context-aware fusion feature vector: ,in, The total number of layers in the graph attention network, That is, a grid cell In the time window Spatial context-aware fusion feature vectors.

[0065] The second level of integration constructs a weighted graph by introducing prior knowledge of road networks and functional zones, and uses the graph as a mask to guide the spread of attention, thus integrating the integration features into the spatial structural constraints of the economy.

[0066] Step S5: Distributed Economic Vitality Anomaly Detection and Anomaly Cluster Analysis This paper proposes anomaly detection and anomaly cluster analysis based on spatial context-aware fusion feature vectors, which consists of four parts: anomaly detection modeling, multi-scale temporal detection, spatial anomaly unit labeling, anomaly cluster determination, and backbone path extraction.

[0067] The first step: Anomaly detection modeling. The Isolation Forest algorithm is employed. This algorithm utilizes the characteristic that outliers are sparsely distributed in the feature space and easily isolated by a small number of random partitions. It does not rely on explicit assumptions about the distribution of normal data, effectively addressing the multimodal characteristics of normal patterns in economic vitality data, and supports efficient training and inference on large-scale data.

[0068] Model training phase: Select a historical period confirmed as a period of normal economic operation (such as a period without major economic shocks in the same period of the previous year), and use the spatial context-aware fusion feature vectors of all spatial grid units within this period for each time window as the training sample set. The total number of samples in the training set is denoted as Ntrain. By randomly sampling subsamples from the training sample set, randomly selecting feature dimensions, and randomly selecting splitting thresholds, several isolated binary trees are recursively constructed. The set of all isolated binary trees constitutes an isolated forest.

[0069] Let the number of isolated binary trees in an isolated forest be . The sample size for each tree is For the construction of each tree, trees are randomly drawn from the training set without replacement. A sample is placed into the root node of the tree. At each node, a feature dimension is randomly selected, and within the range of values ​​for that dimension, a split point is randomly chosen to divide the current node's samples into left and right child nodes: samples with feature values ​​less than the split point go to the left child node, and those with feature values ​​greater than the split point go to the right child node. This partitioning process is recursively executed until the number of samples in a child node is 1 (the sample is completely isolated) or the tree depth reaches a preset maximum depth limit (usually set to 1). (the rounded-up value).

[0070] Model inference phase: For the current monitoring time window Spatial context-aware fusion feature vector of any spatial grid cell u Let the path be traversed through each isolated binary tree in the isolated forest, recording the number of edges (i.e., path length) traversed from the root node to the final leaf node in each tree. The path length in the tree is Then the average path length of this sample in the entire isolated forest is: ; To convert path length into standardized anomaly scores, the average path length needs to be normalized. For a sample size of... For an isolated binary tree, the theoretical value of its expected average path length is: ,in, For harmonic numbers, (Euler's constant). Its function is to standardize the path length, making the scores comparable under different sample sizes.

[0071] Then the grid cell In the time window The vitality abnormality score is: The score ranges from (0, 1). When the average path length is short (i.e., samples are easily isolated), A value close to 1 indicates that the sample is highly likely to be an anomaly; when the average path length is close to the expected value... hour, A value close to 0.5 indicates that the sample has no obvious abnormal characteristics; when the average path length is longer than the expected value, A value close to 0 indicates that the sample is more normal than the normal sample.

[0072] Incremental retraining mechanism: Since the monitoring indicator set in step S1 changes dynamically over time, the feature dimensions input into the isolated forest may change. To address this issue, the isolated forest model supports feature dimension alignment and incremental retraining as the monitoring indicator set dynamically changes. When the monitoring indicator set for a new period differs from the previous period, the system identifies the feature dimensions corresponding to newly added and deleted indicators. For deleted feature dimensions, the corresponding dimension is marked as "no operation" in the retained isolated binary trees, meaning that when the tree is partitioned to that dimension, it uniformly returns to continue random partitioning downwards without affecting the existing tree structure. For newly added feature dimensions, the system incrementally constructs a new supplementary isolated binary tree group based on recently accumulated sample data containing the new dimension, and the supplementary tree participates in subsequent anomaly scoring calculations. This mechanism ensures that the anomaly detection model can smoothly adapt to the dynamic evolution of the indicator system without needing to retrain from scratch for each indicator change.

[0073] The second step: multi-scale time detection. Considering that abnormal economic activity may manifest in different forms at different time scales—in the short term it may be sporadic fluctuations, in the medium term it may be a trend reversal signal, and in the long term it may be a structural decline in activity—this embodiment introduces a multi-scale time detection mechanism based on the isolated forest anomaly detection framework, establishing three independent time detection windows.

[0074] Short-term detection window: The time span is 7 days (one week), used to detect abnormal fluctuations in economic activity at the weekly level. This window constructs a baseline of historical normal patterns using sample data from the most recent 7 days, and can keenly capture sudden fluctuations beyond the weekend effect.

[0075] The China Time Detection Window: Spanning 30 days (approximately one month), this window is used to detect monthly shifts in economic activity trends. It covers a relatively complete monthly economic cycle, smoothing out weekly fluctuations and demonstrating good ability to identify trend changes.

[0076] Long-term monitoring window: The time span is 90 days (approximately one quarter), used to detect structural changes in economic activity at the quarterly level. This window can filter out short-term random disturbances and reveal fundamental structural changes in economic activity over a longer period, such as structural inactivation in a region that shifts from a growth trajectory to sustained contraction.

[0077] Each detection window maintains its own independent isolated forest model instance. Within each monitoring time window, each model instance performs anomaly scoring on the current sample, obtaining anomaly scores across three dimensions: and The final comprehensive vitality abnormality score is obtained by weighting and combining these three scores: ,in, , and The weighting coefficients for the short-time, medium-time, and long-time detection window scores are respectively, satisfying... The weighting coefficient can be configured according to the actual early warning needs: if the focus is on rapid early warning, it can be appropriately increased. If the focus is on identifying structural risks, the level can be appropriately increased. .

[0078] The third step: Spatial anomaly unit marking and deviation type identification. Setting anomaly detection thresholds. When the comprehensive vitality anomaly score of any spatial grid cell exceeds this threshold, it is marked as a spatial anomaly cell: ,in, Represents a set of spatially anomalous units.

[0079] For grid cells marked as spatial anomaly cells, their deviation types are further determined. The deviation type determination is based on the direction of change and synergistic relationship of multi-scale anomaly scores. If the comprehensive anomaly score is mainly driven by a short-term window ( Higher than and If the score is high and the normalized value of the core indicator is higher than the historical average, it is judged as rising; if the short-term and medium-term scores both rise and the core indicator continues to decline, it is judged as declining; if the long-term score remains high for more than three cycles and the core indicator hovers at a low level, it is judged as structural inactivation.

[0080] The fourth step: Identification of anomaly clusters and extraction of backbone anomaly paths. Spatial anomaly units typically exhibit spatial clustering and propagation. This step reveals the spatial structured information of anomalies through clustering and path tracing.

[0081] To perform anomalous clustering, a comprehensive distance metric needs to be constructed that considers both spatial proximity and temporal similarity. For any two grid cells labeled as spatial anomalous units... and The combined distance is defined as: ,in, The geographical distance (in kilometers) between the center points of the two units. The Euclidean distance between the multidimensional temporal feature vectors of two units within the current time window is... ; The weighting coefficients for the temporal distance are used to adjust the importance of temporal features in the overall distance. Based on this overall distance matrix, the DBSCAN clustering algorithm is used to perform spatial clustering on the set of spatial anomaly cells. The DBSCAN algorithm requires setting two parameters: neighborhood radius. And the minimum number of samples, minPts. Neighborhood radius. Based on the statistical distribution setting of the comprehensive distance, the kth percentile of the comprehensive distance between all pairs of anomalous units is usually taken (the value of k is set according to the prior anomalous cluster scale); minPts is set according to the minimum number of units in a meaningful cluster, for example, it is set to 3, that is, at least 3 spatial anomalous units must be clustered together to form an anomalous cluster.

[0082] After clustering, each cluster of consecutively aggregated spatial anomaly units is identified as an anomaly cluster. For each anomaly cluster, its backbone anomaly path is extracted to track the dynamic propagation direction of the anomalous signal in space. The specific method for extracting the backbone anomaly path is as follows: using each spatial grid unit within the anomaly cluster as a node, a directed weighted graph is constructed using the reciprocal of the absolute value of the difference in anomaly scores between units as the edge weight. Let the comprehensive anomaly score of node i be... ,node The overall abnormality score is Then the edge weight is defined as: ,in, A small positive constant is used to prevent the denominator from being zero. The shortest path algorithm or minimum spanning tree algorithm is used to extract the propagation path with the largest anomaly score gradient—under the natural assumption that anomaly scores propagate from high to low, the node with the highest score is usually the source of the anomaly, and the gradual decrease in score along the propagation path reflects the attenuation of the anomaly impact during spatial propagation. The unit on the path where the anomaly first appears is marked as a candidate anomaly source.

[0083] Extract economic activity signals within the abnormal clusters: statistically analyze the coverage area of ​​the abnormal clusters, the mean and maximum values ​​of the abnormal scores, and the duration of the abnormality; backtrack the distribution of the first-level attention weight coefficients in step S4 across all dimensions, identify the time-series feature dimension that contributes the most to the abnormal score (i.e., the dimension whose attention weight in the abnormal unit deviates the most significantly from the average weight of the normal unit), and the economic activity type corresponding to this dimension is the key driving factor of the abnormal signal.

[0084] This step realizes the complete link from anomaly detection to anomaly understanding, covering anomaly type identification, spatial pattern clustering, propagation path tracing, and driving factor identification.

[0085] Step S6: Intelligent monitoring and early warning information generation and push This step transforms the abnormal cluster analysis results into a comprehensive early warning report and pushes it to the terminal in real time.

[0086] First, the actual geographic mapping of the spatial boundaries of the anomaly clusters. The anomaly clusters identified in step S5 consist of a set of spatial grid cells. The center coordinates of these grid cells are connected to construct a convex hull or polygonal boundary, which is then overlaid onto the actual electronic map. Through the reverse geocoding service in the geographic information system, the spatial range covered by the anomaly clusters is mapped to an identifiable actual geographic area description, such as "XX Street and surrounding area in XX District, XX City" or "XX business district and three adjacent blocks". This mapping allows recipients of the warning information to intuitively understand the specific geographic location of the anomaly.

[0087] Second, information on enterprises and industries within the associated region. Based on the actual geographical scope mapped from the anomaly clusters, the enterprise registration database accessed in step S2 is used to query all registered enterprises with their registered addresses within this range, extracting information such as enterprise directory (enterprise name, unified social credit code), registered capital, industry classification, and number of employees. Simultaneously, combined with industry distribution data, the proportion of enterprises in each industry category and their scale within the region are statistically analyzed to identify key affected industries. Enterprises above a certain size, leading enterprises, and core nodes in the industrial chain within the region are marked as key associated enterprises.

[0088] Third, generate a comprehensive early warning report. The above information is summarized and integrated to generate a structured comprehensive early warning report. The report's content modules include: Anomaly spatial location information—providing the geographical location of the anomaly in the form of text description and coordinate range; Anomaly severity level—classifying the anomaly level into three levels: warning, attention, and alert, based on the statistical value of the comprehensive anomaly score within the anomaly cluster (e.g., a score mean of 0.65-0.75 is "warning," 0.75-0.85 is "attention," and above 0.85 is "alert"); Anomaly deviation type, providing the rising, declining, or structurally inactive category determined in step S5; Key anomaly driving factors, listing the time-series feature dimensions that contribute most to the anomaly score obtained from step S5 backtracking, and adding the change amount and rate of change from one period to the current period to each driving factor; List of key enterprises associated with the anomaly, listing key enterprises in the affected area; Recommended attention items, automatically generating targeted recommendations based on the driving factors and deviation types. For example, if the driving factor is logistics activity and the deviation type is declining, the recommendation is "Recommend paying attention to the operational status of logistics infrastructure and the operational status of logistics enterprises in the area."

[0089] Fourth, real-time push notifications of early warning information. After the comprehensive early warning report is generated, it is sent to relevant terminals in real time via a pre-defined message push interface. The message push interface supports multiple communication protocols, including but not limited to HTTP RESTful API, MQTT protocol, and email SMTP protocol, ensuring compatibility with different types of receiving terminals (economic monitoring dashboard systems, work terminals of economic management departments, and mobile terminals of decision-makers). The pushed message body contains the complete text content of the comprehensive early warning report and a structured JSON data payload, facilitating secondary processing and visualization at the receiving end.

[0090] This step transforms the quantitative analysis results into geographically and industrially readable early warning information for decision-makers.

Claims

1. A regional economic vitality intelligent monitoring method based on big data analysis, characterized in that, Includes the following steps: S1. Construct a dynamic and evolvable economic vitality monitoring indicator system. The indicator system includes a core indicator layer and multiple extended indicator layers. The core indicator layer includes a unified basic vitality indicator. Each extended indicator layer corresponds to a type of regional economic characteristic. The indicator items in each extended indicator layer are dynamically increased or decreased as the regional economic structure changes. Using an information gain rate algorithm, effective indicators with an information gain rate exceeding a preset threshold in the current monitoring period are selected from the mutual information between each indicator item and a preset economic vitality benchmark label. The selected effective indicators are dynamically combined to form the monitoring indicator set for the current period. S2. Real-time access to multi-source heterogeneous high-frequency dynamic data, including enterprise business registration data stream, logistics transportation trajectory data stream, commercial consumption transaction data stream, talent flow data stream, and public service activity data stream; and perform streaming preprocessing on the multi-source heterogeneous high-frequency dynamic data to obtain a standardized real-time data stream in a unified format. S3. Divide the target area to be monitored into multiple spatial grid units, and use each spatial grid unit as the statistical granularity to perform spatial aggregation statistics on various types of data in the standardized real-time data stream to generate a multi-dimensional temporal feature vector of each spatial grid unit in each monitoring time window. S4. Perform multi-level feature fusion on the multi-dimensional temporal feature vectors of each spatial grid unit, including: the first level performs attention-weighted fusion on the temporal features of each dimension within each spatial grid unit to generate a unit-level fused feature vector; the second level performs spatial context fusion on the unit-level fused feature vectors based on the spatial adjacency relationship between each spatial grid unit using a spatial self-attention mechanism to generate a spatial context-aware fused feature vector for each spatial grid unit. S5. Input the spatial context-aware fusion feature vector of each spatial grid unit into the pre-constructed distributed economic vitality anomaly detection model. The distributed economic vitality anomaly detection model scores the vitality anomaly of each spatial grid unit on a unit-by-unit basis, identifies spatial anomaly units whose vitality status deviates from the historical normal pattern and the deviation type, the deviation type including rise, decline or structural inactivation; and determines anomaly clusters based on the spatial aggregation characteristics of the spatial anomaly units, and extracts economic vitality signals within the anomaly clusters. S6. Based on the economic vitality signal within the abnormal cluster, generate intelligent monitoring and early warning information that includes abnormal spatial location information, abnormality level, and abnormal signal source analysis results. 2.The regional economic vitality intelligent monitoring method based on big data analysis of claim 1, wherein, The basic vitality indicators of the core indicator layer mentioned in step S1 include the rate of change in the number of registered enterprises, the density of newly added market entities, the activity of factor mobility, and the public service load index; the extended indicator layer includes at least three of the following: the extended indicator layer for manufacturing vitality, the extended indicator layer for service industry vitality, the extended indicator layer for emerging industry vitality, and the extended indicator layer for open economy vitality; the candidate indicator items in each of the extended indicator layers are periodically screened through the information gain rate algorithm, so that different combinations of indicators can be used for vitality monitoring in different monitoring periods. 3.The regional economic vitality intelligent monitoring method based on big data analysis of claim 2, wherein, In step S1, the preset economic vitality benchmark label is generated by dynamically weighting and fusing the regional GDP growth rate and tax revenue growth rate. The dynamic weights are updated adaptively based on the contribution of the core indicator layer and the extended indicator layer in the previous monitoring period, forming a closed-loop self-feedback mechanism of indicator screening, vitality assessment, and weight adaptation. The preset threshold is automatically adjusted based on adaptive kernel density estimation. The ability of each candidate indicator to distinguish economic vitality is measured by calculating the information gain rate between each candidate indicator and the economic vitality benchmark label. When the information gain rate of any candidate indicator is lower than the preset threshold in the current period, it is removed from the monitoring indicator set of the current period and re-evaluated in the next period. 4.The regional economic vitality intelligent monitoring method based on big data analysis of claim 1, wherein, The access to multi-source heterogeneous high-frequency dynamic data in step S2 uses the message middleware Kafka to build a distributed data bus. Each data source pushes data in real time in a streaming manner through the data bus. Dynamic semantic recognition and conflict resolution are performed at the data access edge. It is compatible with the rapid field normalization of semi-structured and unstructured information. The streaming preprocessing includes data format standardization, timestamp alignment, missing value imputation and outlier detection processing. The access to the logistics transportation trajectory data stream includes access to freight vehicle trajectory data, logistics order flow data, and warehouse inbound and outbound flow data. 5.The regional economic vitality intelligent monitoring method based on big data analysis of claim 1, wherein, The spatial grid unit division method in step S3 is as follows: the target area is divided into regular hexagonal grids with a preset side length, and each hexagonal grid serves as a basic monitoring unit covering the entire target area; when performing spatial aggregation statistics on various types of real-time data, the data is assigned to the corresponding hexagonal grid according to its latitude and longitude coordinates, and the temporal characteristic indicators of various types of data in each hexagonal grid are statistically analyzed with a preset time window as the statistical period, forming a multidimensional temporal characteristic vector sequence of each hexagonal grid in a continuous time window. 6.The regional economic vitality intelligent monitoring method based on big data analysis of claim 1, wherein, The first-level attention weighted fusion method in step S4 is as follows: learnable attention weight coefficients are assigned to the temporal features of each dimension within each spatial grid cell, and the temporal features of each dimension are aggregated into a cell-level fusion feature vector by weighted summation. The second-level spatial self-attention fusion method is as follows: a weighted spatial graph based on road network and functional area is constructed, with each spatial grid unit as a node in the graph and the connection weight in the weighted spatial graph as a quantitative expression of spatial adjacency relationship. The graph attention network is used to perform message passing and feature updating of each node feature with the weighted spatial graph as a mask, and spatial context-aware fusion feature vector of each spatial grid unit is generated.

7. The intelligent monitoring method for regional economic vitality based on big data analysis according to claim 1, characterized in that, The distributed economic vitality anomaly detection model described in step S5 is constructed using the isolated forest algorithm. During model training, the spatial context-aware fusion feature vectors of each spatial grid unit during historical normal periods are used as training samples. The isolated forest model supports feature dimension alignment and incremental retraining as the monitoring index set changes dynamically. When scoring anomalies, the average path length of the fusion feature vectors of each spatial grid unit in the isolated forest for the current monitoring period is calculated, and the normalized value of the average path length is used as the vitality anomaly score. When the vitality anomaly score of any spatial grid unit exceeds the anomaly judgment threshold, the spatial grid unit is marked as a spatial anomaly unit.

8. The intelligent monitoring method for regional economic vitality based on big data analysis according to claim 7, characterized in that, The distributed economic activity anomaly detection model employs a multi-scale time detection mechanism, including short-term, medium-term, and long-term detection windows. Each detection window corresponds to a historical normal pattern baseline over a different time span. The short-term detection window is used to detect weekly-level economic activity fluctuation anomalies, the medium-term detection window is used to detect monthly-level economic activity trend changes, and the long-term detection window is used to detect quarterly-level economic activity structural changes. The anomaly scores from each detection window are weighted and combined to obtain a comprehensive economic activity anomaly score.

9. The intelligent monitoring method for regional economic vitality based on big data analysis according to claim 1, characterized in that, The method for determining the anomalous clusters in step S5 is as follows: Based on the comprehensive distance matrix that integrates spatial adjacency distance and temporal feature similarity, the DBSCAN clustering algorithm is used to perform spatial clustering on the spatial grid cells marked as spatial anomalous units. The continuous anomalous regions formed by the clustering are determined as anomalous clusters, and the backbone anomalous paths of each anomalous cluster are extracted to track the dynamic propagation direction of the anomalous signal in space. The backbone anomaly path is extracted in the following way: taking each spatial grid cell in the anomaly cluster as a node, constructing a directed weighted graph with the reciprocal of the absolute value of the difference between anomaly scores between cells as the edge weight, using the shortest path algorithm or minimum spanning tree to extract the propagation path with the largest anomaly score gradient, and taking the cell with the earliest anomaly time on the path as the candidate anomaly source. Extracting economic vitality signals within anomaly clusters includes: statistically analyzing the economic vitality anomaly scores of each spatial grid unit within the anomaly cluster to obtain the area, anomaly intensity, and anomaly duration of the anomaly cluster; and retrospectively analyzing the distribution of attention weight coefficients in each dimension in step S4 to identify the temporal feature dimension that contributes the most to the anomaly score as the basis for tracing the source of the anomaly signal.

10. The intelligent monitoring method for regional economic vitality based on big data analysis according to claim 1, characterized in that, The generation of intelligent monitoring and early warning information in step S6 further includes: mapping the spatial boundary of the anomaly cluster to the actual geographical area, associating the enterprise directory and industry distribution information in the corresponding area, generating a comprehensive early warning report containing the affected area range, the severity level of the anomaly, the key driving factors of the anomaly, and the list of associated key enterprises, and sending the comprehensive early warning report to the terminal in real time through the message push interface.