Multi-scale feature extraction and correlation analysis method
By constructing a multi-level correlation graph system and introducing a hierarchical attention mechanism, the feature definition and conduction problems in multi-scale spatial correlation analysis are solved, and unified representation of data at different spatial scales and accurate analysis of cross-scale correlations is achieved.
Patent Information
- Application Number
- CN202510820128.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing technology has problems such as ignoring single-dimensional feature, lack of cross-scale analysis perspective, unified feature definition, and lack of correlation conduction mechanism in multi-scale spatial correlation analysis, resulting in the inability to accurately identify and quantify multi-scale spatial correlation.
Build a multi-level correlation diagram system, design differentiated multi-dimensional feature definition and extraction strategies, extract features through statistical aggregation and pattern aggregation, introduce a hierarchical attention mechanism to achieve cross-scale correlation dissemination, and integrate correlation information at different levels.
It realizes a unified representation of data at different spatial scales, accurately extracts periodic, trend and burst characteristics, solves cross-scale information transmission, and improves the accuracy and adaptability of multi-scale correlation analysis.
Smart Images

Figure CN120336828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-scale feature extraction and correlation analysis method, belonging to the field of data analysis and modeling. Background Art
[0002] In the field of data science and complex system analysis, spatial correlation analysis refers to studying whether there are statistically interdependent or correlated relationships between different spatial units.
[0003] When the attributes or behaviors of a spatial unit are affected by other spatial units, the phenomenon of spatial correlation appears. This correlation is not only reflected in the geographical space (such as the economic association between cities), but also exists in the abstract conceptual space (such as the association between upstream and downstream enterprises in the industrial chain). With the advent of the big data era, accurately identifying and quantifying this multi-scale spatial correlation has become a key challenge in many fields, and it has important guiding significance for regional coordinated development, optimized layout of the industrial chain, efficient allocation of resources, etc.
[0004] At present, the research in the field of multi-scale spatial correlation analysis at home and abroad mainly has the following deficiencies:
[0005] Firstly, when analyzing the correlation between entities, existing research often uses single-dimensional or limited-dimensional features, ignoring the comprehensive impact of multi-dimensional features on correlation judgment.
[0006] Secondly, traditional correlation analysis methods usually operate at a single spatial scale and lack a cross-scale analysis perspective. However, the correlation in actual systems often exhibits obvious scale-dependent characteristics - entities that are uncorrelated or weakly correlated at a certain spatial scale may show strong correlation at another spatial scale. For example, two seemingly unrelated micro-individuals may have close connections at a higher-level industrial cluster or regional economic scale.
[0007] Thirdly, when dealing with feature extraction and characterization of different spatial scales, existing technologies usually adopt a unified feature definition and extraction method, ignoring the significant differences in feature emphases of entities at different scales. For example, micro-individuals may be more concerned about short-term fluctuation characteristics, medium-sized groups are more concerned about periodic changes, and macro-systems pay more attention to long-term trends.
[0008] Fourthly, existing research lacks the exploration of the correlation conduction mechanism between spatial scales and it is difficult to achieve effective integration of cross-scale correlations. For example, the correlation between micro-individuals will affect the correlation between medium-sized groups through aggregation, and further affect the overall correlation of the macro-system. The lack of a conduction mechanism makes it impossible to achieve cross-scale comprehensive correlation analysis. Summary of the Invention
[0009] The present invention provides a multi-scale feature extraction and correlation analysis method to solve the problems existing in the above-mentioned prior art. First, the present invention designs differential multi-dimensional feature definitions and extraction strategies for different spatial scales to ensure that the feature representation matches the entity scale characteristics. Second, by establishing a cross-scale analysis framework, the identification and comparison of correlations at different spatial scales are realized, and scale-dependent features are effectively captured. Third, the present invention constructs a cross-scale conduction mechanism of correlations, revealing how micro-correlations affect the evolution paths of meso- and macro-correlations through hierarchical aggregation. Finally, when defining correlations, the present invention comprehensively considers the direct correlations at the current spatial scale and the conduction correlations from other scales, so as to obtain a more comprehensive and accurate multi-scale spatial correlation evaluation result.
[0010] The technical solutions adopted by the present invention are as follows:
[0011] A multi-scale feature extraction and correlation analysis method, comprising the following steps:
[0012] S1: Construct a spatial multi-scale correlation graph system with multiple hierarchical correlation graphs, each correlation graph corresponding to a specific spatial scale, and there is a hierarchical aggregation relationship between different hierarchical correlation graphs;
[0013] S2: For each node of the bottom-layer correlation graph, extract its time-series index data, and decompose the time-series index data into periodic, trend, and sudden components;
[0014] S3: For the nodes of non-bottom-layer correlation graphs, extract aggregation features in two ways: statistical aggregation and pattern aggregation;
[0015] S4: Based on the time-series index data extracted in S2 and the aggregation features extracted in S3, calculate the attribute correlation and graph structure correlation of node pairs within each layer, and fuse the two to obtain the intra-layer correlation index corresponding to the layer;
[0016] S5: Define a hierarchical aggregation correlation index according to the intra-layer correlation index to realize inter-layer correlation propagation;
[0017] S6: Introduce a hierarchical attention mechanism, fuse the correlation information of different layers, and calculate the spatial multi-scale correlation.
[0018] Furthermore, S1 specifically includes:
[0019] (1.1) Construct a set of spatial multi-scale correlation graphs:
[0020] , , ,
[0021] Wherein: Represents the association graph of the i-th layer, corresponding to a specific spatial scale; Represents the set of nodes corresponding to the spatial scale of the i-th layer, representing the unit individuals at this scale Set; Represents the set of association relationships between unit individuals; Represents the set of node attributes, depicting the multi-dimensional characteristics of each node;
[0022] (1.2) Establish the hierarchical aggregation relationship between different hierarchical association graphs, so that the nodes in the node set V i of the i-th layer are formed by the aggregation of several nodes in the node set V i−1 of the (i - 1)-th layer, reflecting the hierarchical progression relationship of spatial scales.
[0023] Furthermore, S2 specifically includes:
[0024] (2.1) For each node of the bottom-layer association graph , extract its time-series index data:
[0025] ,
[0026] where, represents the index value of the j-th node at time t, and T represents the length of the time series;
[0027] (2.2) Decompose the time-series index data into three-dimensional features:
[0028] ,
[0029] where: represents the periodic component, reflecting the periodic law of the index data; represents the trend component, capturing the long-term change trend of the index data; represents the sudden component, characterizing the sudden events and abnormal patterns in the index data;
[0030] (2.3) Achieve the three-dimensional feature decomposition through the following optimization model:
[0031] ,
[0032] where, is the reconstruction loss function, used to measure the difference between the original time-series data and the reconstructed data after decomposition;
[0033] is the periodic constraint, used to ensure the periodic characteristics of the periodic component;
[0034] is a smoothness constraint used to ensure the smoothness of the trend component;
[0035] is a sparsity constraint used to ensure the sparsity of the burst component;
[0036] 、 、 are trade-off parameters used to balance the influence of different constraint terms on the total loss.
[0037] Furthermore, S3 is specifically as follows:
[0038] For the nodes of layer , i≥1, assume is aggregated and formed by the node set of the lower layer, then calculate its statistical aggregation feature and pattern aggregation feature ;
[0039] Calculation of statistical aggregation feature:
[0040] ,
[0041] where , , , , , respectively represent the mean, standard deviation, lower quartile, upper quartile, skewness and kurtosis; When correspond to respectively, so as to calculate the load statistical aggregation feature from three dimensions of periodicity, trend and burst;
[0042] Calculation of pattern aggregation feature:
[0043] ,
[0044] where represents the assortativity feature, which measures the similarity degree of the internal members of the aggregated node in key attributes; is the modularity feature, which measures the degree of functional or attribute clustering inside.
[0045] For the assortativity feature , multi-dimensional attributes such as spatial distance decay relationship, attribute similarity, and temporal behavior pattern similarity can be comprehensively considered; for the modularity representation , attributes such as community structure compactness, functional clustering degree, resource distribution density and heterogeneity can be comprehensively considered.
[0046] Further, S4 specifically includes:
[0047] (4.1) For any node pair in the th layer, calculate the attribute correlation :
[0048] ,
[0049] ,
[0050] ,
[0051] where
[0052] , are obtained from the time series index data extracted by S2;
[0053] , are obtained from the statistical aggregation features extracted by S3;
[0054] , are obtained from the pattern aggregation features extracted in step S3; is a multivariate correlation metric function, and the Pearson correlation coefficient or Spearman rank correlation coefficient, etc. can be selected.
[0055] (4.2) For any node pair in the th layer, calculate the graph structure correlation :
[0056] ,
[0057] where represents the neighborhood subgraph of node , represents the graph kernel method, and the Weisfeiler-Lehman kernel or random walk kernel, etc. can be selected.
[0058] (4.3) Fuse the attribute correlation obtained in (4.1) and the graph structure correlation obtained in (4.2) to obtain the intra-layer correlation index:
[0059] ,
[0060] where and are weight parameters;
[0061] Further, in S5, the hierarchical aggregation correlation index is:
[0062] ,
[0063] Among them, is the set of lower-level nodes converging into the node . is a correlation aggregation function, responsible for refining the key features of the relationships between lower-level nodes, and can be achieved through methods such as maximum value, average value, and attention mechanism;
[0064] Furthermore, S6 specifically includes:
[0065] (6.1) Introduce a hierarchical attention mechanism to assign dynamic weights to each layer:
[0066] ,
[0067] Among them is a multi-layer perceptron, used to extract the features of each hierarchical graph and generate weights.
[0068] (6.2) Calculate the spatial multi-scale correlation based on intra-layer correlation and inter-layer correlation:
[0069] ,
[0070] Among them, , are respectively , at the layer corresponding nodes.
[0071] The present invention has the following beneficial effects:
[0072] (1) By constructing a hierarchical spatial multi-scale association graph system, a unified representation of the association relationships of data at different spatial scales is achieved;
[0073] (2) Through the multi-dimensional feature decomposition method, accurate extraction of the periodic, trend, and sudden characteristics of data is achieved;
[0074] (3) By means of statistical aggregation and pattern aggregation, the problem of cross-scale information transmission is solved, and information abstraction and knowledge refinement from micro to macro are achieved;
[0075] (4) Through intra-layer correlation analysis and inter-layer correlation propagation, quantitative characterization of the association relationships within the same scale and between different scales is achieved;
[0076] (5) Through the hierarchical attention mechanism, dynamic fusion of multi-scale correlations is achieved, improving the accuracy and adaptability of correlation analysis. Description of the Drawings
[0077] Figure 1 This is the flowchart of the present invention. Specific embodiments
[0078] The present invention will be further described below with reference to the accompanying drawings.
[0079] As Figure 1 , this embodiment is the application of a multi-scale feature extraction and correlation analysis method of the present invention in the power load analysis scenario. By constructing a four-level graph structure from individual users to the overall city, multi-dimensional extraction, cross-scale aggregation, and correlation propagation of power load characteristics are realized. By implementing the present invention, the internal correlation patterns of power loads at different spatial scales can be obtained, providing data support and decision-making basis for the accurate prediction, efficient scheduling, and scientific planning of power systems.
[0080] (I) Scenario description:
[0081] In power load analysis, there are data at different spatial scales, including electricity consumption data of individual users, electricity consumption data of industrial parks or communities, and electricity consumption data of urban areas, etc. These data show different characteristics and correlations at different spatial scales. For example, the power load of an individual user may be affected by factors such as their living habits and the operation cycle of production equipment, with obvious periodic and sudden characteristics; while the power load of an industrial park or community is comprehensively affected by the electricity consumption behaviors of multiple users, showing more complex trend and homogeneity characteristics; the power load of an urban area further synthesizes the electricity consumption situations of multiple industrial parks or communities, reflecting the power demand pattern of the entire city.
[0082] (II) Method application:
[0083] Step 1: Construct a power load spatial multi-scale correlation graph system.
[0084] In this embodiment, a four-level spatial multi-scale correlation graph system is constructed , , , , corresponding to the following spatial scales respectively:
[0085] : The individual power user level, including industrial users, commercial users, and residential users.
[0086] : The industrial park / community level, composed of multiple adjacent users.
[0087] : The urban area level, composed of multiple industrial parks / communities.
[0088] : The overall city level, composed of multiple urban areas.
[0089] When implementing specifically, taking the electricity users in a certain city as an example:
[0090] At the level, select 1000 typical electricity users in the city as nodes , including 500 industrial users, 300 commercial users and 200 residential users.
[0091] Establish an edge set according to the similarity of user geographical location and electricity consumption type . When the geographical distance between two users is less than 500 meters and the similarity of electricity consumption type is greater than 0.7, establish a connection between them.
[0092] At the level, converge the users in according to the affiliated park / community to form 50 nodes , representing 50 industrial parks / communities.
[0093] At the level, converge the parks / communities in according to the affiliated urban area to form 10 nodes , representing 10 urban areas.
[0094] At the level, converge the urban areas in into 1 node , representing the entire city.
[0095] Step 2: Multi-dimensional feature extraction and representation of electric load.
[0096] For each electric power user node at the level , collect its 15-minute granularity electric load data for one year to form time series data .
[0097] Decompose the time series index data into three-dimensional features:
[0098] ,
[0099] Realize the three-dimensional feature decomposition through the following optimization model:
[0100] ,
[0101] Among them, the loss function adopts the L2 norm, = 0.1, = 0.05, = 0.2.
[0102] Periodic constraint Defined as:
[0103] ,
[0104] where T1 = 96 (representing the daily cycle) and T2 = 672 (representing the weekly cycle).
[0105] Smoothness constraint Defined as:
[0106] ,
[0107] where represents the second-order difference operator.
[0108] Sparsity constraint Defined as:
[0109] ,
[0110] The L1 norm is adopted to promote sparsity.
[0111] Taking a certain industrial user as an example, its power load is decomposed into:
[0112] Periodic component Capturing the load differences between weekdays and rest days, as well as the periodic changes in production shifts.
[0113] Trend component Capturing the long-term trends caused by capacity expansion or seasonal changes.
[0114] Sudden component Capturing abnormal events such as equipment maintenance and sudden failures.
[0115] Step 3: Cross-scale feature aggregation of power load.
[0116] For the industrial park nodes at the level, assuming it is aggregated by 20 user nodes at the level, the feature aggregation is carried out in the following way:
[0117] Statistical aggregation feature calculation:
[0118] ,
[0119] ,
[0120] ,
[0121] where, and , are respectively the sets of periodic, trend, and burst components of all nodes in
[0122] Mode convergence feature calculation:
[0123] ,
[0124] Among them, represents the assortativity feature, which measures the similarity of the internal members of the converging nodes in key attributes; is the modularity feature, which measures the degree of functional or attribute clustering within
[0125] Assortativity feature r: Calculate the homogeneity of the distribution of electricity consumption types and voltage levels of users in the park:
[0126] ,
[0127] Among them, is the information entropy of the electricity consumption type, is the information entropy of the voltage level, N is the number of electricity consumption type categories, and M is the number of voltage level categories.
[0128] Modularity feature q: Calculate the tightness of the connection of nodes within the park:
[0129] - ,
[0130] Taking a high-tech industrial park as an example, by statistical convergence, its mean periodic load characteristic = 5.2 MW, and the standard deviation = 1.8 MW, indicating that the overall load of the park has obvious working / non-working period characteristics; through mode convergence, its assortativity feature r = 0.78 and modularity feature q = 0.65, indicating that the internal power users of the park have high homogeneity and close functional associations.
[0131] Step 4: Intra-layer correlation analysis of the power load layer.
[0132] For any two industrial park nodes at the and levels, perform intra-layer correlation analysis according to the following steps:
[0133] Calculate the attribute correlation:
[0134] ,
[0135] Among them, = 0.6, = 0.4, represents the Pearson correlation coefficient, represents the cosine similarity.
[0136] Calculate the graph structure correlation:
[0137] ,
[0138] where WL represents the Weisfeiler-Lehman graph kernel, represents the neighborhood subgraph of represents the number of iterations.
[0139] Fuse the attribute and structure correlations:
[0140] ,
[0141] Taking two industrial parks (high-tech park) and (traditional manufacturing park) as an example: The attribute correlation = 0.42, indicating that there are certain differences in the load characteristics between the two parks; the graph structure correlation = 0.35, indicating that the network structures of the two parks are quite different; the intra-layer correlation after fusion = 0.4, indicating that the overall correlation of the power loads of the two parks is relatively low.
[0142] Step Five: Propagate the inter-layer correlation of power loads.
[0143] For the industrial park nodes at the level, the calculation of its level aggregation correlation is as follows:
[0144] Calculate the average correlation between the lower-layer nodes:
[0145] ,
[0146] where represents the correlation coefficient between any two nodes and at the level, represents the set the number of nodes in, so N represents the number of all possible different node pairs in the set. This formula is used to characterize the correlation between and its lower-layer nodes (i.e., the nodes in
[0147] Step Six: Multi-scale Correlation Fusion of Power Loads
[0148] Finally, fuse the correlation information at different levels and calculate the spatial multi-scale correlation:
[0149] Calculate the attention weights at each level through a multi-layer perceptron:
[0150] ,
[0151] The input of the MLP is the feature vectors of the graphs at each level, including the number of nodes, edge density, average degree, etc.
[0152] Suppose the obtained attention weights are α = [0.15, 0.35, 0.3, 0.2], representing the degree of emphasis on each level from G0 to G3.
[0153] Calculate the spatial multi-scale correlation between any two nodes:
[0154] ,
[0155] where , are respectively , the corresponding nodes at the m-th level.
[0156] Taking the evaluation of the multi-scale correlation between a high-tech industrial park and a traditional manufacturing industrial park as an example: At the level, their correlation is = 0.4; at the level, the correlation of the corresponding urban area nodes is = 0.55; at the level, for the corresponding nodes in the same city, the correlation is = 1.0; the aggregated correlations at their respective levels are , ; the final multi-scale correlation calculation is:
[0157]
[0158] ,
[0159] This result shows that although the direct correlation between the two parks at the level is low, considering the higher-level associations and the internal structure of the levels, their comprehensive correlation reaches a medium level.
[0160] Through the above specific implementation steps, the method of the present invention can comprehensively capture the characteristics and correlations of power loads at different spatial scales, providing a basis for power load forecasting, power system dispatching, power system planning, etc. Compared with traditional methods, the present invention can more comprehensively and accurately capture the multi-scale characteristics of power loads, improve the operation efficiency and reliability of the power system, and reduce the operation cost.
[0161] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.
Claims
1. A multi-scale feature extraction and correlation analysis method, characterized in that: It includes the following steps: S1: Construct a spatial multi-scale association graph system with multiple hierarchical association graphs. Each association graph corresponds to a specific spatial scale, and there is a hierarchical convergence relationship between different hierarchical association graphs; S2: For each node of the bottommost association graph, extract its time series index data and decompose the time series index data into periodic, trend, and burst components; S3: For the nodes of non-bottommost association graphs, extract convergence features in two ways: statistical convergence and pattern convergence; S4: Based on the time series index data extracted in S2 and the convergence features extracted in S3, calculate the attribute correlation and graph structure correlation of node pairs within each level, and fuse the two to obtain the intra-level correlation index corresponding to that level; S5: Define the hierarchical convergence correlation index according to the intra-level correlation index to achieve inter-level correlation propagation; S6: Introduce a hierarchical attention mechanism, fuse the correlation information of different levels, and calculate the spatial multi-scale correlation.
2. The multi-scale feature extraction and correlation analysis method according to claim 1, wherein: S1 specifically includes: (1.1) Construct a set of spatial multi-scale association graphs: , , , Wherein: represents the i-th layer association graph, corresponding to a specific spatial scale; represents the set of nodes corresponding to the i-th layer spatial scale, representing the unit individuals at this scale set; represents the set of association relationships between unit individuals; represents the set of node attributes, depicting the multi-dimensional characteristics of each node; (1.2) Establish a hierarchical aggregation relationship between different hierarchical association graphs, so that the node set V in the i-th layer i is formed by the aggregation of several nodes in the node set V of the (i-1)-th layer i−1 , reflecting the hierarchical progressive relationship of spatial scales.
3. The multi-scale feature extraction and correlation analysis method according to claim 1, wherein: S2 specifically includes: (2.1) For each node of the bottom - layer association graph extract its time - series index data: , Among them, represents the index value of the j-th node at time t, and T represents the length of the time series; (2.2) Decompose the time series index data into three-dimensional features: , Wherein: represents the periodic component, reflecting the periodic law of the index data; represents the trend component, capturing the long-term change trend of the index data; represents the sudden component, characterizing the sudden events and abnormal patterns in the index data; (2.3) Achieve the three-dimensional feature decomposition through the following optimization model: , Among them, is the reconstruction loss function; , , are the periodicity constraint, smoothness constraint, and sparsity constraint respectively, , , are the trade-off parameters.
4. The multi-scale feature extraction and correlation analysis method according to claim 3, characterized in that: S3 specifically is: For the nodes of layer , i≥1, assume that is formed by converging from the node set of the lower layer , and calculate its statistical convergence feature , pattern convergence feature respectively; Statistical convergence feature calculation: , Among them, , , , , , respectively represent the mean, standard deviation, lower quartile, upper quartile, skewness and kurtosis; when correspond to respectively to calculate the load statistical aggregation characteristics from three dimensions of periodicity, trend and suddenness; Pattern convergence feature calculation: , Among them, represents the assortativity feature, measuring the similarity degree of the internal members of the aggregation node in terms of key attributes; is the modularity feature, measuring the degree of functional or attribute clustering within.
5. The multi-scale feature extraction and correlation analysis method according to claim 4, wherein: S4 specifically includes: (4.1) For any node pair in the layer, calculate the attribute correlation : , , , Among them, is a multi-variable correlation metric function; (4.2) For any node pair in the layer, calculate the graph structure correlation : , Among them, represents the neighborhood subgraph of the node , and represents the graph kernel method; (4.3) Fuse the attribute correlation obtained in (4.1) and the graph structure correlation obtained in (4.2) to obtain the intra-level correlation index: , Among them and are weight parameters.
6. The multi-scale feature extraction and correlation analysis method according to claim 4, wherein: In S5, the hierarchical convergence correlation index is: , Among them, is the set of lower-layer nodes converging into the node, is the correlation aggregation function.
7. The multi-scale feature extraction and correlation analysis method according to claim 4, wherein: S6 specifically includes: (6.1) Introduce a hierarchical attention mechanism and assign dynamic weights to each level: , Among them is a multi-layer perceptron; (6.2) Calculate the spatial multi-scale correlation based on the intra-level correlation and inter-level correlation: , Among them, , are respectively , at the layer corresponding nodes.
Citation Information
Patent Citations
Distributed time series data analysis processing method and system
CN118484566A
Power load prediction method based on adaptive multi-scale hypergraph neural network
CN119168113A
Aero-engine missing value filling method based on multi-scale time sequence and attribute feature extraction
CN119807632A
Harbor flexible load multi-level interaction control method oriented to ship-port-network cooperation
CN119944713A