Quick association query system for multi-modal data of electric power station area
Through data association extraction, cleaning and indexing modules, combined with distributed databases, the heterogeneity and security problems of multimodal data in the power station area are solved, efficient and real-time data association and analysis are achieved, and multi-dimensional business applications are supported.
Patent Information
- Application Number
- CN202510439058.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
AI Technical Summary
There are data islands, heterogeneity, query efficiency bottlenecks, semantic gaps, security and privacy issues in traditional power data management, making it difficult to achieve efficient, real-time correlation and analysis of multimodal data.
The data association extraction module, data fusion cleaning module, distributed index module and data storage management module are adopted, and the association relationship of multimodal data is established by combining common attribute matching, machine learning algorithms and distributed databases. The data accuracy and security are achieved through global and local index models.
It realizes efficient correlation and analysis of multi-modal power station data, improves query efficiency, supports cross-domain data fusion, meets real-time requirements, ensures data security, optimizes resource allocation and strategic decision-making.
Smart Images

Figure CN120353830A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database security, and particularly to a multi-modal data rapid association query system for power distribution areas. Background Art
[0002] With the advancement of the dual-carbon goal and the construction of a new power system, power enterprises need to improve operation efficiency and decision-making accuracy through digital means. However, the traditional power data management has the following pain points:
[0003] Data islands: Financial data and power consumption data belong to different systems, lacking a unified association framework.
[0004] Heterogeneity problem: There are significant differences between the two types of data in the time dimension, space dimension, and theme dimension.
[0005] Query efficiency bottleneck: The data scale of power distribution areas is huge, and traditional centralized indexing technologies are difficult to meet the millisecond-level real-time query requirements.
[0006] Power system data presents multi-source heterogeneous characteristics, including structured financial statements, semi-structured equipment ledgers, unstructured inspection images, etc. For example, the Guangming Power Model released by the State Grid integrates multi-modal data such as text, images, and videos to achieve the intelligence of power grid planning and operation and maintenance. However, multi-modal data fusion faces the following technical challenges:
[0007] Semantic gap: The semantic expression methods of different modal data vary greatly, and traditional association methods are difficult to establish effective mappings.
[0008] Real-time requirement: Power consumption data needs to be monitored in real time, while the update cycle of financial data is relatively long. Cross-modal association needs to balance the timeliness and historical analysis requirements.
[0009] Security and privacy: Financial data involves the core secrets of enterprises, and power consumption data contains user privacy. The fusion process needs to meet the requirements of data security and compliance.
[0010] The existing technologies have the following problems
[0011] Traditional methods rely on manual rules: For example, matching financial and power consumption data through the distribution area ID. However, in actual scenarios, the ID coding rules are not unified, resulting in a high association failure rate.
[0012] Lack of intelligent mining ability: Traditional association rule mining cannot discover non-linear association patterns.
[0013] Poor scalability of centralized indexing: Traditional B-trees or hash indexes face single-point performance bottlenecks when processing data in the hundreds of millions, and cannot support distributed storage architectures.
[0014] Missing multimodal index: Existing indexing technologies mainly target unimodal spatial data and are difficult to handle joint queries of multiple types of data such as time series, text, and images simultaneously.
[0015] Difficulty in normalizing heterogeneous data: The currency unit of financial data and the physical unit of power consumption data need to be uniformly converted, and traditional methods are prone to introducing errors.
[0016] Therefore, to solve the problems existing in the prior art, a fast multi-modal data correlation query system for power distribution areas is proposed. Summary of the Invention
[0017] The present invention solves the problems of the prior art through the following technical solutions. The present invention includes a data correlation extraction module, a data fusion and cleaning module, a distributed indexing module, a data storage and management module, and a data integration and processing module;
[0018] The data correlation extraction module is used to identify common attributes or key indicators in financial data and power consumption data of the distribution area and establish an association relationship between the two types of data;
[0019] The data fusion and cleaning module is used to process errors, missing values, and inconsistencies in the original data and perform normalization processing on financial data and power consumption data;
[0020] The distributed indexing module includes a global indexing model and a local indexing model, and is used to achieve fast retrieval and pruning of multi-modal data;
[0021] The data storage and management module is used to design a data model to store the associated data set and support efficient query and analysis;
[0022] The data integration and processing module is used to implement data extraction, transformation, cleaning, and storage processes to ensure data accuracy and integrity.
[0023] Furthermore, the data correlation extraction module further includes a common attribute matching unit and an association pattern mining unit;
[0024] The common attribute matching unit is used to match financial data and power consumption data through common attributes such as distribution area ID and timestamp;
[0025] The association pattern mining unit is used to discover potential association patterns in the data by using statistical methods or machine learning algorithms such as cluster analysis and association rule mining.
[0026] Furthermore, the data storage and management module stores data using a distributed database or a NoSQL database, processes large-scale data sets in combination with a distributed computing framework, and supports SQL language for query and operation.
[0027] Furthermore, the global index model organizes and manages the storage partition boundaries through a hash structure, which is used to prune the trajectories in a specific partition area without calling a task to calculate this partition;
[0028] The local index model adopts a variant R-tree structure, and each internal node contains the complete trajectory IDs of all segments of the subtree, which is used to prune the trajectories that do not need to traverse the child nodes.
[0029] Furthermore, the content of the normalization process of the data fusion and cleaning module includes: aligning the dimensions and unifying the units of the electricity consumption, load curve, and electricity consumption pattern indicators of the electricity consumption data in the substation area with the income, cost, assets, liabilities, and other indicators of the financial data, and eliminating data heterogeneity.
[0030] Furthermore, the data integration and processing module includes a data extraction unit, a data conversion and cleaning unit, and a security protection unit;
[0031] The data extraction unit is used to extract the required data from the original data source;
[0032] The data conversion and cleaning unit is used to perform format conversion, outlier processing, and missing value filling operations;
[0033] The security protection unit is used to implement data privacy protection measures to prevent the leakage or tampering of sensitive information.
[0034] Furthermore, in the variant R-tree structure of the local index model, the TID set of each node u contains all the trajectory IDs involved in the subtree rooted at u.
[0035] A method for fast multi-modal data correlation query in a power substation area, the method includes the following steps:
[0036] Including the following steps:
[0037] Step 1: Establish the correlation relationship between the financial data and the electricity consumption data in the substation area through common attribute matching or association rule mining;
[0038] Step 2: Clean and normalize the associated data to eliminate errors, missing values, and inconsistencies;
[0039] Step 3: Use the distributed index model to index the processed data to achieve fast pruning and retrieval of trajectory data;
[0040] Step 4: Perform efficient data query and analysis through the data storage and management module to support business applications such as investment benefit analysis in the substation area
[0041] Furthermore, in step two, an outlier detection is performed on the load curve of the power consumption data and the cost data of the financial data by using a data cleaning algorithm, and the missing values are filled by interpolation or a machine learning model to ensure the data quality.
[0042] Furthermore, in step three, the global index prunes specific partition trajectories based on a hash structure, and the local index prunes trajectories that do not require traversing child nodes through a TID set. The combination of the two reduces the data retrieval range.
[0043] The present invention has the following advantages compared with the prior art: The multi-modal data fast correlation query system for power distribution areas realizes the efficient correlation of financial data and power consumption data through common attribute matching and machine learning algorithms, solves the heterogeneity problems of the two types of data in terms of time, space, and theme, and provides a general framework for cross-domain data fusion.
[0044] The distributed index model combines a hash structure and a variant of the R-tree, supporting the fast pruning and retrieval of trajectory data: The global index prunes invalid regions through hash partition boundaries, reducing the invocation of computing tasks;
[0045] The local index directly filters child nodes that do not need to be traversed through a TID set, significantly improving the query efficiency of large-scale data and reducing the retrieval time-consuming.
[0046] Aiming at the problems of errors, missing values, and inconsistencies in the original data, the heterogeneity of financial data and power consumption data is eliminated through data cleaning and normalization processing, laying a high-quality data foundation for subsequent analysis.
[0047] A distributed database, a NoSQL database, and a distributed computing framework are adopted to support the storage, query, and analysis of massive multi-modal data, meet the business requirements of large data volume and complex dimensions in power distribution areas, and have good scalability and flexibility.
[0048] The data integration process realizes automated processing, ensures data accuracy and integrity, and at the same time guarantees data security through security protection measures.
[0049] Break the isolated state of financial data and power consumption data, establish a unified data model, support multi-dimensional correlation analysis, and provide cross-domain data support for business scenarios.
[0050] Through the joint modeling of power consumption and financial data in the power distribution area, the effectiveness analysis of the power distribution area assets is realized, which helps the management to accurately evaluate the investment benefits, optimize the resource allocation, and improve the refined management level of the power distribution area.
[0051] Provide data-driven support for strategic decision-making. For example, through the correlation analysis of power consumption patterns and financial costs, identify inefficient power distribution areas and formulate targeted optimization strategies to enhance the competitiveness of the enterprise.
[0052] The data-driven management mode reduces manual intervention and improves operational efficiency, meeting the industry trends of smart grid and digital transformation. Brief Description of the Drawings
[0053] Figure 1 is the system block diagram of the present invention;
[0054] Figure 2 is the distributed index model diagram of the present invention;
[0055] Figure 3 is the global index model diagram of the four-layer tree structure of the present invention;
[0056] Figure 4 is the local index model diagram of SR numbers of the present invention;
[0057] Figure 5 is the local index connection example diagram of the present invention. Detailed Embodiment
[0058] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0059] As Figures 1 to 5 shown, this embodiment provides a technical solution: a multi-modal data fast correlation query system for power distribution areas, including a data correlation extraction module, a data fusion and cleaning module, a distributed index module, a data storage and management module, and a data integration and processing module;
[0060] The data correlation extraction module is used to identify the common attributes or key indicators in financial data and power distribution area power consumption data, and establish the correlation relationship between the two types of data;
[0061] The data fusion and cleaning module is used to process errors, missing values, and inconsistencies in the original data, and perform normalization processing on financial data and power consumption data;
[0062] The distributed index module includes a global index model and a local index model, and is used to achieve fast retrieval and pruning of multi-modal data;
[0063] The data storage and management module is used to design a data model to store the associated data set, and support efficient query and analysis;
[0064] The data integration and processing module is used to implement the data extraction, transformation, cleaning, and storage processes to ensure data accuracy and integrity.
[0065] Furthermore, the data correlation extraction module further includes a common attribute matching unit and an association pattern mining unit;
[0066] It includes a common attribute matching unit for matching financial data and power consumption data through common attributes such as substation area ID and timestamp;
[0067] An association pattern mining unit, which is used to discover potential association patterns in data by using statistical methods or machine learning algorithms such as cluster analysis and association rule mining.
[0068] The data storage and management module stores data using a distributed database or a NoSQL database, processes large-scale data sets in combination with a distributed computing framework, and supports SQL language for querying and operation.
[0069] The global index model organizes and manages the storage partition boundaries through a hash structure, which is used to prune the trajectories in a specific partition area without calling tasks to calculate this partition;
[0070] The local index model adopts a variant R-tree structure, and each internal node contains the complete trajectory IDs of all segments of the subtree, which is used to prune the trajectories that do not need to traverse the child nodes.
[0071] The content of the normalization process of the data fusion and cleaning module includes: aligning the dimensions and unifying the units of the power consumption, load curve, and power consumption pattern indicators of the substation area power consumption data with the income, cost, assets, liabilities, etc. indicators of the financial data to eliminate data heterogeneity.
[0072] The data integration and processing module includes a data extraction unit, a data transformation and cleaning unit, and a security protection unit;
[0073] The data extraction unit is used to extract the required data from the original data source;
[0074] The data transformation and cleaning unit is used to perform format conversion, outlier processing, and missing value filling operations;
[0075] The security protection unit is used to implement data privacy protection measures to prevent the leakage or tampering of sensitive information.
[0076] Furthermore, in the variant R-tree structure of the local index model, the TID set of each node u contains all the trajectory IDs involved in the subtree with u as the root node.
[0077] A method for fast association query of multi-modal data in a power substation area, the method includes the following steps:
[0078] It includes the following steps:
[0079] Step 1: Establish an association relationship between financial data and power consumption data of the substation area through common attribute matching or association rule mining;
[0080] Step 2: Clean and normalize the associated data to eliminate errors, missing values, and inconsistencies;
[0081] Step 3: Use a distributed indexing model to index the processed data to achieve fast pruning and retrieval of trajectory data;
[0082] Step 4: Perform efficient data querying and analysis through the data storage and management module to support business applications such as investment benefit analysis in the substation area
[0083] Furthermore, in Step 2, an outlier detection is performed on the load curve of the electricity consumption data and the cost data of the financial data using a data cleaning algorithm, and the missing values are filled by interpolation or machine learning models to ensure data quality.
[0084] Furthermore, in Step 3, the global index prunes specific partition trajectories based on a hash structure, and the local index prunes trajectories that do not require traversing child nodes through a TID set. The combination of the two reduces the data retrieval range.
[0085] Cleaning and Fusion of Substation Area Electricity Consumption Data and Financial Data
[0086] (1) Research on the Technology of Extracting and Associating Financial Data and Substation Area Electricity Consumption Data
[0087] Comprehensively understand the characteristics and structures of substation area financial data and electricity consumption data. Financial data usually includes information such as revenue, cost, assets, and liabilities, while electricity consumption data includes information such as the electricity consumption, load curve, and electricity consumption pattern of users. These two types of data may have differences in terms of time, space, and theme, so data cleaning and preprocessing are required to ensure data consistency and accuracy.
[0088] Determine the association method between financial data and electricity consumption data. This can be achieved by identifying common attributes or key indicators in the two types of data. For example, we can use attributes such as substation area ID and timestamp to match financial data and electricity consumption data, thereby establishing an association relationship between them. In addition, we can also consider using some statistical methods or machine learning algorithms, such as clustering analysis and association rule mining, to discover potential association patterns in the data.
[0089] Design a suitable data model to store and manage the associated data set. This data model should be able to support efficient querying and analysis operations and have good scalability and flexibility. A commonly used method is to use a distributed database to store and manage data and use SQL language for querying and operations. In addition, some advanced big data technologies, such as NoSQL databases and distributed computing frameworks, can also be considered to process large-scale and complex data sets.
[0090] Implement an effective data extraction and integration process to ensure data accuracy and integrity. This includes extracting the required data from the original data sources, performing necessary transformation and cleaning operations, establishing associations between data, and storing the integrated data into the target data model. During this process, we need to pay attention to data privacy and security issues and take appropriate measures to protect sensitive information from being leaked or tampered with.
[0091] (2) Research on the Fusion Cleaning and Normalization Technology of Financial and Electricity Consumption Data
[0092] In the analysis of the asset effectiveness of the substation area, the fusion analysis of financial data and electricity consumption data is of great significance for improving operational efficiency, optimizing resource allocation, and making strategic decisions. However, the original data often contains problems such as errors, missing values, and inconsistencies, which pose challenges to data fusion. This project focuses on the handling of errors and missing values in the electricity consumption data of the substation area, as well as the normalization technology of financial data and electricity consumption data, aiming to improve data quality and lay a solid foundation for subsequent fusion analysis;
[0093] The technical solution is carried out in the way of distributed indexing:
[0094] The distributed index model consists of two parts: the global index model and the local index model. As shown in the figure, each tree node at each layer maintains a hash structure for storing partition boundaries to organize and manage data. The global index can prune the trajectories passing through specific partition regions without calling any tasks to calculate that partition. The local index model is a data structure that is a variant of the R-tree. The main difference from the traditional R-tree is that each internal node u in the tree contains the complete trajectory IDs of all segments contained in the subtree rooted at u (represented as the TID set of u). The local index can prune all trajectories that pass through the nodes in the local index without traversing their child nodes.
[0095] The global index model has a four-layer tree structure. First, the root node of the tree structure saves a hash map for mapping each text item to the ID of the first-layer grouping, and its mapping relationship is expressed as: w → g i .
[0096] Secondly, each text item in the first layer corresponds to a grouping, and each text item saves a hash map for mapping each time slice to the ID of the second-layer sub-grouping, and its mapping relationship is expressed as: s → g i,j .
[0097] Thirdly, each start time slice in the second layer corresponds to a grouping, and each time slice also saves a hash map for mapping each time slice to the D of the third-layer sub-grouping, and its mapping relationship is expressed as: s′ → g i,j,k .
[0098] Then, each end time slice of the third layer corresponds to a group, and each time slice also stores a hash map for mapping each trajectory segment to D in the sub - groups of the fourth layer. The mapping relationship is expressed as:
[0099] Finally, each leaf node of the fourth layer corresponds to a partition of the trajectory data. For each partition P i,j,k,l , save the IDs of all trajectories in this partition and denote them as TID.
[0100] Local Index Model
[0101] To efficiently organize and manage the data within the partition, construct an SR - tree for each partition as a local index. The structure of the local index model is as Figure 4 shown. The main difference between the SR - tree and the traditional R - tree is that each internal node u in the tree contains the complete trajectory IDs (denoted as the TID set of u) of all segments contained in the subtree rooted at u2.
[0102] For example, the tids of the three child nodes of the internal node u are 3, 1, 3, and all are included in the TID set of u2.
[0103] The number of segments contained in an SR - tree subtree is usually more than the number of trajectories passing through this subtree. Therefore, retaining the TID set can filter the trajectories of a node without traversing its child nodes. To index the segments of trajectory T , in addition to retaining the geographical text , also assign a tuple (tid, sid), where tid is the ID of the trajectory and sid represents the position in the trajectory.
[0104] The construction process is as follows:
[0105] Design a variant structure of the R - tree. The internal nodes in the tree retain all the trajectory IDs contained in the subtrees to efficiently organize the data within the partition; design a global index with a four - layer tree structure using a hash map to achieve distributed management of the partitions. The pseudo - code of the distributed index construction algorithm based on MDP is shown in Algorithm 1.
[0106] Algorithm 1: Distributed Index Construction Algorithm Based on MDP
[0107] Input: Data partition P = {P i,j,k,l};
[0108] Output: Distributed index I G ;
[0109] Foreach: P = {P i,j,k,l} ∈ {P i,j,k,l} do;
[0110] I L = Build an R-tree index for partition P i,j,k,l ;
[0111] I G ← Collect the boundary information of I L ;
[0112] Return: I G ;
[0113] The trajectory similarity search algorithm based on the distributed index model The pseudocode of the trajectory similarity search algorithm based on the distributed index model is shown in Algorithm 2. Given a query trajectory q, a similarity threshold r, and a trajectory dataset P, first sample the trajectory data from P and use the inverted index algorithm, time slicing algorithm
[23] , and SFTR algorithm
[113] to construct a partitioner; then, divide P into several partitions and construct local SR-tree indexes for each partition respectively; next, collect statistical data from each partition to build a global index; finally, filter out the candidate partitions that may be similar to the query trajectory through the global index, and find all similar trajectory data through the local index in these candidate partitions. Since each partition is independent, the search process using the local index can run in parallel on each partition, and finally the query results of each partition are aggregated and returned. The entire trajectory similarity search can be divided into three parts: global search, local search, and result refinement, which will be introduced in detail in turn below.
[0114] Global Search This section designs a global search method as shown in Figure GlobalPruning.
[0115] The steps are as follows: a) Obtain the text summary of trajectory q;
[0116] b) Traverse the start time slices of each item in the item set to obtain a set of start time slices Traverse the end time slices of each slice in to obtain a set of end time slices
[0117] c) Traverse the tree nodes of the end time slices in to obtain a set of tree nodes S A , and query the corresponding partitions through the global index as the candidate set;
[0118] d) Traverse each candidate partition and calculate the start time distance and the trajectory segment trajectory space distance through ;
[0119] e) Further verify the candidate partitions according to Theorem 2.1 and Theorem 1. Specifically, for candidate partition P i,j,k,l , if it shares at least one text item with q and and sd(q, v a ) ≤ ε, then P i ,j,k,l may contain trajectories similar to q, otherwise P i,j,k,l will be filtered out.
[0120] Local Search After obtaining candidate partitions through global search, the FilteringAndVerify function in Algorithm 2 is executed in parallel on each candidate partition to obtain the local search results. Specifically, first use the local index to find all tree nodes whose distance from q is no greater than ε. Then, obtain the trajectory segments i ∈ T, and obtain the text summary and spatial summary of T from segment i. Further verify the trajectory segments using text summary pruning, start and end point pruning of the spatial summary, and signature pruning of the spatial summary. Finally, return the TIDs of all unqualified trajectories.
[0121] Result Refinement Reconstruct the trajectories that are not pruned during the global and local query processes (i.e., the trajectories whose IDs are not in A'), and execute the Refine function on each complete trajectory to calculate the upper bound T UB of the time similarity. If the trajectories are definitely not similar, thus obtaining the final search results.
[0122] Algorithm 2: Trajectory Similarity Search Algorithm Based on Distributed Index Model
[0123] Input: Global index I G , local index I L , query trajectory q, similarity threshold r, weight values λ1 and λ2
[0124] Output: Search result A
[0125] Obtain the set C of partitions where similar trajectories may exist = GlobalPruning(q, τ, λ1, λ2, I L );
[0126] foreach P i,j,k,l in C do / / This loop is executed in parallel
[0127] I L = Obtain the local index in partition P i,j,k,l ;
[0128] Obtain the set A' of trajectory segments that may be similar = FilteringAndVerify(q, τ, λ1, λ2, I L );
[0129] P' = The complete trajectory set obtained by reorganizing the trajectories through the ID of the trajectory segments in A'.
[0130] For each T in P′, do A ← Refine(T, q, τ, λ1, λ2);
[0131] Return A;.
[0132] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0133] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0134] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-modal data fast correlation query system for a power distribution area, characterized in that, It includes a data association extraction module, a data fusion and cleaning module, a distributed index module, a data storage and management module, and a data integration and processing module; The data association extraction module is used to identify the common attributes or key indicators in financial data and power consumption data of the substation area, and establish the association relationship between the two types of data; The data fusion and cleaning module is used to process errors, missing values, and inconsistencies in the original data, and perform normalization processing on financial data and power consumption data; The distributed index module includes a global index model and a local index model, and is used to achieve fast retrieval and pruning of multi-modal data; The data storage and management module is used to design a data model to store the associated data set, and support efficient query and analysis; The data integration and processing module is used to implement the data extraction, transformation, cleaning, and storage processes to ensure data accuracy and integrity.
2. According to claim 1, a fast association query system for multi-modal data of power substations is characterized by: The data association extraction module further includes a common attribute matching unit and an association pattern mining unit; The common attribute matching unit is used to match financial data and power consumption data through the common attributes of the substation area ID and timestamp; The association pattern mining unit is used to discover potential association patterns in the data by using statistical methods or machine learning algorithms such as clustering analysis and association rule mining.
3. The multimodal data fast correlation query system for a power distribution area according to claim 1, wherein: The data storage and management module stores data using a distributed database or a NoSQL database, combines a distributed computing framework to process large-scale data sets, and supports SQL language for query and operation.
4. A multimodal data rapid correlation query system for a power distribution area according to claim 1, characterized in that: The global index model organizes and manages the storage partition boundaries through a hash structure, and is used to prune the trajectories in a specific partition area without calling tasks to calculate this partition; The local index model adopts an R-tree variant structure, and each internal node contains the complete trajectory IDs of all segments of the subtree, and is used to prune the trajectories that do not need to traverse the child nodes.
5. The multimodal data rapid correlation query system for a power distribution area according to claim 1, wherein: The content of the normalization processing of the data fusion and cleaning module includes: aligning the dimensions and unifying the units of the power consumption, load curve, and power consumption pattern indicators of the substation area power consumption data with the income, cost, assets, and liability indicators of the financial data, and eliminating data heterogeneity.
6. The multimodal data fast correlation query system for a power distribution area according to claim 1, characterized in that: The data integration and processing module includes a data extraction unit, a data transformation and cleaning unit, and a security protection unit; The data extraction unit is used to extract the required data from the original data source; The data transformation and cleaning unit is used to perform format conversion, outlier processing, and missing value filling operations; The security protection unit is used to implement data privacy protection measures to prevent the leakage or tampering of sensitive information.
7. The multimodal data fast correlation query system for a power distribution area according to claim 1, wherein: In the R-tree variant structure of the local index model, the TID set of each node u contains all the trajectory IDs involved in the subtree rooted at u.
8. A method for quickly associating and querying multimodal data in a power distribution area, the method being based on the system according to any one of claims 1-7, characterized in that, The method includes the following steps: includes the following steps: Step 1: Establish the association relationship between financial data and power consumption data of the substation area through common attribute matching or association rule mining; Step 2: Clean and normalize the associated data to eliminate errors, missing values, and inconsistencies; Step 3: Use the distributed index model to index the processed data to achieve fast pruning and retrieval of trajectory data; Step 4: Perform efficient data query and analysis through the data storage and management module to support the business application of substation area investment benefit analysis.
9. A method for quickly associating and querying multimodal data in a power distribution area according to claim 8, characterized in that: In step 2, an outlier detection is performed on the load curve of the electricity consumption data and the cost data of the financial data by using a data cleaning algorithm, and the missing values are filled by the interpolation method or a machine learning model.
10. A method for rapid association query of multi-modal data of power substations according to claim 8, characterized in that: In step 3, the global index prunes specific partition trajectories based on the hash structure, and the local index prunes the trajectories that do not need to traverse child nodes through the TID set. The combination of the two reduces the data retrieval range.