Multi-source heterogeneous data fusion processing method and system

By calculating the impedance identification and dynamic scheduling mode of multi-source heterogeneous data fusion paths, the problem of decreased fusion accuracy caused by incompatible data sources in existing technologies is solved, achieving high-precision data fusion and fast response, adapting to the processing needs of different scenarios.

CN121389034BActive Publication Date: 2026-04-10ZHEJIANG SIJI TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing multi-source heterogeneous data fusion processing methods fail to identify and eliminate incompatible data sources, resulting in an increase in data sources leading to a decrease in fusion accuracy, and failing to balance response speed and processing accuracy.

Method used

By calculating the impedance identifiers of the fusion paths between each data source node, a data fusion topology is formed. Low-impedance fusion paths are identified and connected subsets are formed. The connected subset with the smallest sum of impedance identifiers is selected as the data fusion topology. The scheduling mode is dynamically switched at the processing anchor point deployment status probe. The streaming or batch processing mode is determined according to the data fluctuation amplitude. Data labeling is optimized by combining content sensitivity and scene data labeling.

Benefits of technology

Selective fusion was achieved, which improved the accuracy of data fusion, took into account the response speed and processing accuracy in different scenarios, avoided interference from incompatible data sources on the fusion results, and reduced the risk of data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121389034B_ABST
    Figure CN121389034B_ABST
Patent Text Reader

Abstract

The application provides a multi-source heterogeneous data fusion processing method and system, relates to the field of data processing, and comprises the following steps: acquiring a plurality of data source nodes, calculating impedance identifiers of fusion paths between the data source nodes, and forming a data fusion topology according to the impedance identifiers; deploying a state probe of a processing anchor to monitor fusion data of the data fusion topology, determining a scheduling mode of the processing anchor according to a monitoring result; obtaining a fusion result based on the scheduling mode; determining a static data mark according to a content sensitivity of the fusion result, determining a dynamic data mark according to an aggregation number of the data source nodes, determining a scene data mark based on a current processing scene, and taking the highest mark in the static data mark, the dynamic data mark and the scene data mark as a data mark. The application filters low-impedance data source nodes to form a data fusion topology by calculating impedance identifiers of fusion paths, improves data fusion accuracy, and takes into account the response speed of a sudden scene and the processing accuracy of a stable scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a multi-source heterogeneous data fusion processing method and system. BACKGROUND

[0002] With the development of Internet of Things, cloud computing and big data technology, a large amount of multi-source heterogeneous data has been generated in the fields of energy, transportation, medical treatment and the like. These data are widely sourced from different types of sensors, monitoring devices, business systems and the like, and have characteristics such as various data formats, inconsistent time and space scales, and uneven data quality. In order to fully tap the value of data, it is necessary to effectively fuse and process multi-source heterogeneous data.

[0003] The existing multi-source heterogeneous data fusion processing method usually adopts a full-amount fusion strategy, that is, the data of all data sources are directly aggregated and then uniformly processed. In the fusion process, the connection relationship between each data source is regarded as identical, and the time synchronization, spatial matching degree and data distribution consistency between data sources are not evaluated. When there is a time offset, spatial coverage range mismatch or data distribution characteristic difference between some data sources, these data sources are still included in the fusion processing, resulting in an increase in the number of data sources and a decrease in fusion accuracy.

[0004] Therefore, there is a problem in the prior art that incompatible data sources cannot be identified and removed. SUMMARY

[0005] The present application provides a multi-source heterogeneous data fusion processing method and system to solve the problems mentioned in the prior art.

[0006] In a first aspect of the present application, a multi-source heterogeneous data fusion processing method is provided, comprising: acquiring a plurality of data source nodes, calculating impedance identifiers of fusion paths between the data source nodes, and forming a data fusion topology according to the impedance identifiers;

[0007] processing fusion data of the data fusion topology by an anchor point deployment state probe, and determining a scheduling mode of the processing anchor point according to a monitoring result;

[0008] obtaining a fusion result based on the scheduling mode;

[0009] determining a static data marker according to a content sensitivity of the fusion result, determining a dynamic data marker according to an aggregated number of the data source nodes, determining a scene data marker based on a current processing scene, and taking the highest marker among the static data marker, the dynamic data marker and the scene data marker as a data marker.

[0010] Optionally, in a possible implementation manner of the first aspect, the calculating of the impedance identifiers of the fusion paths between the data source nodes comprises:

[0011] obtaining sampling time intervals of each data source node, and calculating a ratio of the sampling time intervals of any two data source nodes;

[0012] obtaining coverage area coordinates of each data source node, and calculating an overlapping area of the coverage areas of any two data source nodes;

[0013] extracting data distribution features of each data source node, and calculating a data distribution feature distance of any two data source nodes;

[0014] determining impedance identifiers of the fusion paths based on the ratio of the sampling time intervals, the overlapping area of the coverage areas, and the data distribution feature distance.

[0015] Optionally, in a possible implementation manner of the first aspect, the forming the data fusion topology according to the impedance identifiers comprises:

[0016] marking a fusion path with an impedance identifier less than a preset impedance threshold as a low-impedance fusion path;

[0017] identifying data source nodes connected by the low-impedance fusion path, and grouping data source nodes in mutual communication into a same connected subset;

[0018] calculating a sum of the impedance identifiers of the fusion paths in each connected subset, and selecting a connected subset with a smallest sum of the impedance identifiers as the data fusion topology.

[0019] Optionally, in a possible implementation manner of the first aspect, after the connected subset with the smallest sum of the impedance identifiers is selected as the data fusion topology, the method further comprises:

[0020] identifying isolated data source nodes outside the data fusion topology;

[0021] calculating impedance identifiers of fusion paths between each isolated data source node and each data source node in the data fusion topology;

[0022] determining that the corresponding isolated data source node is included in the data fusion topology when there is a fusion path with an impedance identifier less than a preset impedance threshold.

[0023] Optionally, in a possible implementation manner of the first aspect, the determining the impedance identifiers of the fusion paths based on the ratio of the sampling time intervals, the overlapping area of the coverage areas, and the data distribution feature distance comprises:

[0024] taking a logarithm of the ratio of the sampling time intervals to obtain a time deviation amount;

[0025] calculating a ratio of the overlapping area of the coverage areas to a union area of the coverage areas of each data source node to obtain a space matching amount;

[0026] The data distribution feature distance is normalized to obtain a distribution deviation;

[0027] When any one of the time deviation, the space matching amount, and the distribution deviation exceeds the corresponding threshold, the impedance identifier of the fusion path is marked as high impedance;

[0028] When none of the time deviation, the space matching amount, and the distribution deviation exceeds the corresponding threshold, the time deviation, the space matching amount, and the distribution deviation are combined to obtain the impedance identifier of the fusion path.

[0029] Optionally, in a possible implementation manner of the first aspect, the monitoring, by the processing anchor point deployment state probe, of the fusion data of the data fusion topology and the determination of the scheduling mode of the processing anchor point according to a monitoring result, comprise:

[0030] The processing anchor point deployment state probe collects data fluctuation amplitudes of data source nodes in the data fusion topology;

[0031] When the data fluctuation amplitude exceeds a first fluctuation threshold, the scheduling mode of the processing anchor point is determined as a stream processing mode;

[0032] When the data fluctuation amplitude does not exceed the first fluctuation threshold, the scheduling mode of the processing anchor point is determined as a batch processing mode.

[0033] Optionally, in a possible implementation manner of the first aspect, the obtaining of the fusion result based on the scheduling mode comprises:

[0034] When the scheduling mode is the stream processing mode, a fusion operation is performed on current data of the data source nodes in the data fusion topology to obtain the fusion result;

[0035] When the scheduling mode is the batch processing mode, a fusion operation is performed on accumulated data of the data source nodes in the data fusion topology within a first time window to obtain the fusion result.

[0036] Optionally, in a possible implementation manner of the first aspect, the determination of the static data label according to the content sensitivity of the fusion result comprises:

[0037] A data type identifier in the fusion result is extracted;

[0038] A sensitivity level corresponding to the data type identifier is queried based on a sensitivity mapping table;

[0039] The sensitivity level is taken as the static data label.

[0040] Optionally, in a possible implementation manner of the first aspect, the determination of the dynamic data label according to the aggregated number of the data source nodes comprises:

[0041] counting a number of data source nodes participating in generating the fusion result, to obtain an aggregation number;

[0042] when the aggregation number exceeds a first aggregation threshold, upgrading the static data mark to a dynamic data mark by a first level;

[0043] when the aggregation number does not exceed the first aggregation threshold, taking the static data mark as the dynamic data mark.

[0044] Optionally, in a possible implementation manner of the first aspect, when the aggregation number exceeds the first aggregation threshold, the static data mark is upgraded to the dynamic data mark by the first level, including:

[0045] calculating a number ratio of the aggregation number to the first aggregation threshold;

[0046] determining a number of upgrade levels according to the number ratio;

[0047] upgrading the static data mark to a dynamic data mark by a level corresponding to the number of upgrade levels.

[0048] Optionally, in a possible implementation manner of the first aspect, after the static data mark is upgraded to the dynamic data mark by the level corresponding to the number of upgrade levels, the method further includes:

[0049] monitoring a usage state of the fusion result;

[0050] when an aggregation operation of the fusion result ends and the fusion result returns to single-point data corresponding to a single data source node, downgrading the dynamic data mark to a static data mark.

[0051] Optionally, in a possible implementation manner of the first aspect, the method further includes:

[0052] receiving a scene type identifier carried in the processing request;

[0053] when the scene type identifier indicates an abnormal processing scene, upgrading the dynamic data mark to the scene data mark by a second level;

[0054] when the scene type identifier indicates a normal processing scene, taking the dynamic data mark as the scene data mark.

[0055] In a second aspect of the present application, a multi-source heterogeneous data fusion processing system is provided, including:

[0056] a topology construction module, configured to acquire a plurality of data source nodes, calculate impedance identifiers of fusion paths between the data source nodes, and form a data fusion topology according to the impedance identifiers;

[0057] a state monitoring module configured to monitor fusion data of the data fusion topology by a processing anchor point deployment state probe, and determine a scheduling mode of the processing anchor point according to a monitoring result;

[0058] a fusion execution module configured to obtain a fusion result based on the scheduling mode;

[0059] a label determination module configured to determine a static data label according to a content sensitivity of the fusion result, determine a dynamic data label according to an aggregation quantity of the data source node, determine a scene data label based on a current processing scene, and take a highest label among the static data label, the dynamic data label and the scene data label as a data label.

[0060] In a third aspect, the present application provides an electronic device, comprising a memory, a processor and a computer program, wherein the computer program is stored in the memory, and the processor executes the computer program to implement the method of the first aspect and various possible aspects related to the first aspect.

[0061] The multi-source heterogeneous data fusion processing method provided by the present application has the following beneficial effects:

[0062] 1. The present application identifies the impedance labels of fusion paths between data source nodes, comprehensively considers three dimensions of a sampling time interval ratio, an overlapping area of coverage regions and a data distribution characteristic distance, and quantifies the compatibility degree between data source nodes. The fusion paths with impedance labels less than a preset impedance threshold are marked as low-impedance fusion paths, the data source nodes connected by the low-impedance fusion paths are identified and classified into connected subsets, and the connected subset with the smallest total impedance label is selected as a data fusion topology. The data source nodes with time asynchronization, spatial mismatch or abnormal data distribution are identified and removed through the impedance labels, the interference of incompatible data sources on the fusion result is avoided, the problem that the increase of data sources leads to the decrease of fusion accuracy is solved, selective fusion is realized, and the accuracy of data fusion is improved.

[0063] 2. The present application collects data fluctuation amplitudes of each data source node in the data fusion topology by a processing anchor point deployment state probe, and determines a scheduling mode of the processing anchor point according to whether the data fluctuation amplitude exceeds a first fluctuation threshold. When the data fluctuation amplitude exceeds the first fluctuation threshold, the scheduling mode is determined as a stream processing mode, and a fusion operation is performed on current data to quickly respond to sudden changes. When the data fluctuation amplitude does not exceed the first fluctuation threshold, the scheduling mode is determined as a batch processing mode, and a fusion operation is performed on accumulated data in a first time window to improve processing accuracy. The scheduling mode is dynamically switched according to the data fluctuation amplitude, the stream processing mode is adopted in a burst scenario to guarantee response speed, and the batch processing mode is adopted in a smooth scenario to guarantee processing accuracy, so that different needs for response speed and processing accuracy in different scenarios are taken into account. Attached Figure Description

[0064] Figure 1 This is a flowchart illustrating the multi-source heterogeneous data fusion processing method provided in the embodiments of this application;

[0065] Figure 2 This is an application scenario diagram provided in the embodiments of this application;

[0066] Figure 3 This is a schematic diagram of the structure of the multi-source heterogeneous data fusion processing system provided in the embodiments of this application;

[0067] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0069] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0070] In related technologies, multi-source heterogeneous data fusion processing methods employ a full-scale fusion strategy, directly aggregating data from all data source nodes for unified processing. During the fusion process, this method treats the connections between various data source nodes as equivalent, without evaluating time synchronization, spatial matching, or data distribution consistency. When some data source nodes have time offsets, mismatched spatial coverage, or differences in data distribution characteristics, this method still includes these data source nodes in the fusion process, leading to an increase in data source nodes and a decrease in fusion accuracy. Furthermore, this method uses a fixed stream-batch integrated architecture, applying the same processing method regardless of data state changes, making it impossible to balance response speed and processing accuracy. In addition, this method uses static labeling for data sensitivity, determining the sensitivity level only based on the content attributes of the data, ignoring the impact of data aggregation and usage scenarios on sensitivity, thus posing a risk of data leakage.

[0071] To solve the problem that the increase of data sources leads to the decrease of fusion accuracy because incompatible data source nodes cannot be identified and removed in the prior art, a multi-source heterogeneous data fusion processing method is proposed. The impedance identifiers of fusion paths between data source nodes are calculated, the time synchronization, spatial matching degree and data distribution consistency between data source nodes are comprehensively evaluated, low impedance fusion paths are screened to form connected subsets, and the connected subset with the smallest impedance identifier sum instead of the connected subset with the most data source nodes is selected as the data fusion topology, which reflects the fusion principle that less and precise is better than more and miscellaneous, and realizes selective fusion of data source nodes. Meanwhile, state probes are deployed at processing anchors to monitor data fluctuation amplitudes, and the scheduling mode is dynamically switched according to the data fluctuation amplitudes, the stream processing mode is used in a burst scenario to ensure response speed, and the batch processing mode is used in a smooth scenario to ensure processing accuracy. In addition, data labels are determined according to the content sensitivity of fusion results, the aggregation number of data source nodes and the current processing scenario, the number of upgrade levels is nonlinearly determined according to the ratio of the aggregation number to the threshold, and the data label is dynamically downgraded when data is disaggregated, which realizes the scenario-sensitive identification of data sensitivity.

[0072] The multi-source heterogeneous data fusion processing method provided by the embodiments of the present application is applied to an energy monitoring scenario as shown in Figure 2 The scenario includes multiple data source nodes distributed in different geographical locations, including power monitoring stations, inverters of photovoltaic power stations, monitoring devices of wind power plants, and user-side power consumption terminals. These data source nodes transmit the collected energy data to an energy data fusion platform through a communication network. The energy data fusion platform acquires multiple data source nodes, calculates the impedance identifiers of fusion paths between data source nodes, and forms a data fusion topology according to the impedance identifiers. State probes are deployed at processing anchors to monitor the fusion data of the data fusion topology, and the scheduling mode of the processing anchors is determined according to the monitoring results. The fusion results are obtained based on the scheduling mode. The static data label is determined according to the content sensitivity of the fusion results, the dynamic data label is determined according to the aggregation number of data source nodes, and the scenario data label is determined based on the current processing scenario. The highest label among the static data label, the dynamic data label and the scenario data label is taken as the data label. The energy data fusion platform controls the access permission of the fusion results according to the data label, and the users access the fusion results for load prediction, fault diagnosis, operation optimization and other applications according to their access permissions.

[0073] Referring to Figure 1 is a flowchart of the multi-source heterogeneous data fusion processing method provided by the embodiments of the present application, Figure 1The execution subject of the method shown can be a software and / or hardware device. The execution subject of the present application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. Among them, the user equipment can include, but is not limited to, a computer, a smart phone, a personal digital assistant (PDA) and the above-mentioned electronic equipment, etc. The network equipment can include, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing, wherein cloud computing is a kind of distributed computing, which is a super virtual computer composed of a group of loosely coupled computers. The present embodiment does not make any limitation. Including steps 100 to 400, as follows:

[0074] Step 100: Obtain a plurality of data source nodes, calculate the impedance identifier of the fusion path between each data source node, and form a data fusion topology according to the impedance identifier.

[0075] In the energy monitoring scene, the data source nodes can be power monitoring stations, photovoltaic inverter, wind farm monitoring devices, and user-side power consumption terminals distributed in different geographical locations. The data generated by these data source nodes differ in sampling time, coverage area, and data distribution.

[0076] It should be noted that the traditional multi-source heterogeneous data fusion method adopts a full-amount fusion strategy, directly aggregates all data of the data source nodes for unified processing, and does not evaluate the compatibility between the data source nodes. When there is a time offset, spatial mismatch, or abnormal data distribution between some data source nodes, these data source nodes are still included in the fusion, resulting in an increase in the number of data source nodes and a decrease in the fusion accuracy.

[0077] In order to avoid fusing the data of incompatible data source nodes, it is necessary to calculate the impedance identifier of the fusion path between each data source node, identify mutually compatible data source nodes, and form a data fusion topology.

[0078] In some embodiments, the impedance identifier of the fusion path between each data source node in step 100 is calculated, including steps A1 to A4:

[0079] Step A1: Obtain the sampling time interval of each data source node, and calculate the ratio of the sampling time interval of any two data source nodes.

[0080] Among them, the sampling time interval represents the time period for the data source node to collect data. For example, a certain power monitoring station collects data every 1 minute, and its sampling time interval is 1 minute; a certain photovoltaic inverter collects data every 15 minutes, and its sampling time interval is 15 minutes.

[0081] It can be understood that when the sampling time interval of two data source nodes is quite different, the fusion of their data will lead to difficulty in time alignment. Therefore, the ratio of the sampling time interval of any two data source nodes is calculated to evaluate their compatibility in the time dimension.

[0082] For example, the sampling time interval of data source node A is 1 minute, and the sampling time interval of data source node B is 15 minutes. The ratio of the sampling time interval of data source node A to data source node B is 1 / 15 or 15 / 1.

[0083] Step A2: Obtain the coverage area coordinates of each data source node, and calculate the coverage area overlap area of any two data source nodes.

[0084] The coverage area represents the geographical range monitored by the data source node. Through the geographic coordinate system, the coverage area coordinates of each data source node can be represented. For example, the coverage area monitored by a certain power monitoring station is a circular area with the monitoring station as the center and a radius of 5 kilometers, and its coverage area coordinates can be represented as the center coordinates and the radius. The coverage area monitored by a certain photovoltaic power station is a rectangular plot where the power station is located, and its coverage area coordinates can be represented as the coordinates of the four vertices of the rectangle.

[0085] The coverage area overlap area of any two data source nodes is calculated to evaluate their compatibility in the spatial dimension. It is not difficult to understand that the larger the coverage area overlap area, the higher the coincidence degree of the geographical range monitored by the two data source nodes, and the data has strong spatial correlation.

[0086] For example, the coverage area of data source node A is rectangle a, and the coverage area of data source node B is rectangle b. The area of the overlapping part of rectangle a and rectangle b is calculated as the coverage area overlap area.

[0087] Step A3: Extract the data distribution features of each data source node, and calculate the data distribution feature distance of any two data source nodes.

[0088] The data distribution feature represents the distribution rule of the data generated by the data source node in the numerical value. It can be the mean, variance, kurtosis, skewness, etc. statistical features of the extracted data, or the histogram distribution, probability density function, etc. of the extracted data as the data distribution feature.

[0089] The data distribution feature distance of any two data source nodes is calculated to evaluate their compatibility in data distribution. For example, the KL divergence, JS distance, Euclidean distance, etc. method can be used to calculate the data distribution feature distance of two data source nodes. The smaller the data distribution feature distance, the more similar the data distribution generated by the two data source nodes, and the data has strong consistency.

[0090] Step A4: determining the impedance identifier of the fusion path based on the ratio of sampling time intervals, the overlapping area of coverage regions, and the distance of data distribution characteristics.

[0091] The fusion path represents a logical connection relationship between two data source nodes. The impedance identifier is used to quantify the compatibility between the two data source nodes. The smaller the impedance identifier, the better the compatibility, and the larger the impedance identifier, the worse the compatibility.

[0092] Based on the ratio of sampling time intervals, the overlapping area of coverage regions, and the distance of data distribution characteristics obtained in steps A1 to A3, the impedance identifier of the fusion path is determined.

[0093] Further, step A4 specifically includes steps A41 to A45:

[0094] Step A41: taking the logarithm of the ratio of sampling time intervals to obtain the time deviation amount.

[0095] The time deviation amount represents the difference in sampling frequency between the two data source nodes. The ratio of sampling time intervals reflects the difference in sampling frequency between the two data source nodes. In order to convert this ratio into an additive quantity, the logarithm of the ratio of sampling time intervals is taken.

[0096] For example, the sampling time interval of data source node A is 1 minute, the sampling time interval of data source node B is 15 minutes, the ratio of sampling time intervals is 15, and after taking the logarithm, the time deviation amount is log(15)≈1.18. The larger the time deviation amount, the greater the difference in sampling frequency between the two data source nodes, and the worse the temporal compatibility.

[0097] Step A42: calculating the ratio of the overlapping area of coverage regions to the union area of coverage regions of each data source node to obtain the spatial matching amount.

[0098] The spatial matching amount represents the degree of overlap of the coverage regions of the two data source nodes. The overlapping area of coverage regions represents the overlapping part of the monitoring regions of the two data source nodes, and the union area of coverage regions represents the total area of the monitoring regions of the two data source nodes.

[0099] It can be understood that the value range of the spatial matching amount is 0 to 1. The larger the spatial matching amount, the higher the degree of overlap of the coverage regions of the two data source nodes, and the better the spatial compatibility. For example, the coverage area of data source node A is 100 square kilometers, the coverage area of data source node B is 150 square kilometers, and the overlapping area of coverage regions is 50 square kilometers. Then the union area of coverage regions is 100+150-50=200 square kilometers, and the spatial matching amount is 50 / 200=0.25.

[0100] Step A43: Normalize the data distribution feature distance to obtain a distribution deviation.

[0101] The distribution deviation represents the difference between the data distributions of the two data source nodes. The data distribution feature distance can have a large range of values and no uniform dimension. In order to facilitate comprehensive evaluation with the time deviation and the spatial matching amount, the data distribution feature distance is normalized to a range of 0 to 1 to obtain the distribution deviation.

[0102] For example, the minimum-maximum normalization method can be used to map the data distribution feature distance to a range of 0 to 1. The larger the distribution deviation, the greater the difference between the data distributions of the two data source nodes, and the poorer the data consistency.

[0103] Step A44: When any one of the time deviation, the spatial matching amount, and the distribution deviation exceeds the corresponding threshold, mark the impedance identifier of the fusion path as high impedance.

[0104] It should be noted that the traditional method uses a simple weighted average method to combine evaluation indicators in multiple dimensions, which can cause a serious defect in one dimension to be covered up by good performance in other dimensions. For example, two data source nodes perform well in time and space dimensions, but have a large difference in data distribution. The traditional method may fuse them because the comprehensive score after weighted averaging is acceptable, resulting in a decrease in the quality of the fusion result.

[0105] In order to identify seriously incompatible data source node pairs, the time deviation threshold, the spatial matching amount threshold, and the distribution deviation threshold are set. When the time deviation exceeds the time deviation threshold, it indicates that the sampling frequency difference between the two data source nodes is too large to perform effective time alignment. When the spatial matching amount is less than the spatial matching amount threshold, it indicates that the overlap degree of the coverage areas of the two data source nodes is too low, and the data lacks spatial correlation. When the distribution deviation exceeds the distribution deviation threshold, it indicates that the data distribution difference between the two data source nodes is too large, and the data lacks consistency.

[0106] As long as any one of the time deviation, the spatial matching amount, and the distribution deviation exceeds the corresponding threshold, the impedance identifier of the fusion path is marked as high impedance. This reflects the principle of one vote veto, that is, serious incompatibility in any one dimension is enough to cause fusion failure.

[0107] Step A45: When the time deviation, the spatial matching amount, and the distribution deviation do not exceed the corresponding threshold, combine the time deviation, the spatial matching amount, and the distribution deviation to obtain the impedance identifier of the fusion path.

[0108] When the time deviation amount, the space matching amount and the distribution deviation amount all do not exceed the corresponding threshold, it indicates that the two data source nodes have certain compatibility in the three dimensions of time, space and data distribution. At this time, the impedance identifier of the fusion path is obtained by combining the time deviation amount, the space matching amount and the distribution deviation amount.

[0109] It can be understood that the combination method can be weighted summation. For example, the time dimension weight is set to 0.3, the space dimension weight is set to 0.3, and the data distribution dimension weight is set to 0.4. Then the impedance identifier = 0.3 x time deviation amount + 0.3 x (1-space matching amount) + 0.4 x distribution deviation amount. The space matching amount is taken inversely because the greater the space matching amount, the better the compatibility, and the smaller the impedance identifier, the better the compatibility.

[0110] Preferably, steps A41 to A45 identify the seriously incompatible data source node pair through a veto mechanism, avoiding the problem that simple weighted average may cover up the serious defects in a certain dimension with good performance in other dimensions. For the data source node pair that passes the preliminary screening, the accurate impedance identifier is calculated through weighted combination, realizing the combination of rough screening and precise evaluation, and improving the accuracy of data source node compatibility evaluation.

[0111] Preferably, steps A1 to A4 evaluate the compatibility between data source nodes through three dimensions of sampling time interval ratio, overlapping area of coverage region and data distribution feature distance. Compared with the method of considering only a single dimension, it can more comprehensively identify the conflict situations such as time asynchronization, space mismatch and data distribution inconsistency between data source nodes, and improve the accuracy of impedance identifier calculation.

[0112] In some embodiments, the data fusion topology is formed according to the impedance identifier in step 100, including steps B1 to B3:

[0113] Step B1: Mark the fusion path with an impedance identifier less than a preset impedance threshold as a low impedance fusion path.

[0114] The preset impedance threshold is used to distinguish the low impedance fusion path and the high impedance fusion path. The fusion path with an impedance identifier less than the preset impedance threshold indicates that the two connected data source nodes have good compatibility, and their data can be fused. These fusion paths are marked as low impedance fusion paths.

[0115] For example, the preset impedance threshold is set to 0.5. If the impedance identifier of the fusion path between data source node A and data source node B is 0.3, the fusion path is marked as a low impedance fusion path. If the impedance identifier of the fusion path between data source node C and data source node D is 0.8, the fusion path is a high impedance fusion path and is not marked.

[0116] Step B2: Identify the data source nodes connected by low impedance fusion paths, and group the data source nodes connected to each other into the same connected subset.

[0117] Wherein, the connected subset represents a set of data source nodes connected to each other through low impedance fusion paths. Based on the low impedance fusion paths marked in step B1, identify the data source nodes connected by these low impedance fusion paths. If multiple data source nodes are connected to each other through low impedance fusion paths, they form a connected subset.

[0118] For example, there is a low impedance fusion path between data source node A and data source node B, and a low impedance fusion path between data source node B and data source node C, then data source node A, data source node B and data source node C are connected to each other and belong to the same connected subset. If data source node D has no low impedance fusion path with other data source nodes, then data source node D forms a connected subset alone.

[0119] Step B3: Calculate the impedance identification sum of the fusion paths within each connected subset, and select the connected subset with the smallest impedance identification sum as the data fusion topology.

[0120] Wherein, the impedance identification sum represents the sum of the impedance identifications of all fusion paths within the connected subset, which is used to evaluate the overall compatibility between the data source nodes in the connected subset.

[0121] It should be noted that the traditional method tends to select the connected subset containing the most data source nodes when forming the data fusion topology, considering that the more data source nodes, the more accurate the fusion result. However, the connected subset with more data source nodes may contain data source nodes with poor compatibility, resulting in a decline in overall fusion quality.

[0122] When there are multiple connected subsets, one needs to be selected as the data fusion topology. The selection criteria is the smallest impedance identification sum, i.e. the sum of the impedance identifications of all fusion paths within the subset is the smallest. The connected subset with the smallest impedance identification sum indicates that the overall compatibility between the data source nodes in the subset is the best.

[0123] For example, connected subset 1 contains data source node A, data source node B and data source node C, and the impedance identifications of the fusion paths within the subset are 0.2 (A-B), 0.3 (B-C) and 0.25 (A-C) respectively, and the impedance identification sum is 0.75; connected subset 2 contains data source node D and data source node E, and the impedance identification of the fusion path within the subset is 0.4 (D-E), and the impedance identification sum is 0.4. Then select connected subset 2 as the data fusion topology.

[0124] This reflects the principle of less but better fusion, that is, selecting the best overall compatibility subset between data source nodes for fusion, rather than selecting the subset containing the most data source nodes.

[0125] It should be noted that after selecting the connected subset with the smallest impedance identification sum as the data fusion topology in step B3, steps B4 to B6 are also included:

[0126] Step B4: Identify isolated data source nodes outside the data fusion topology.

[0127] Among them, the isolated data source node represents the data source node not included in the data fusion topology. After forming the data fusion topology in step B3, there may be some data source nodes not included in the data fusion topology. These isolated data source nodes may be because the initial threshold is set too strictly, resulting in the impedance identification of the fusion path between the data source nodes in the data fusion topology being slightly higher than the preset impedance threshold, but actually still having certain compatibility.

[0128] Step B5: Calculate the impedance identification of the fusion path between each isolated data source node and each data source node in the data fusion topology.

[0129] For each isolated data source node, the impedance identification of the fusion path between the isolated data source node and each data source node in the data fusion topology is recalculated. The calculation method is the same as steps A1 to A4.

[0130] Step B6: When there is a fusion path with impedance identification less than the preset impedance threshold, the corresponding isolated data source node is included in the data fusion topology.

[0131] If the impedance identification of the fusion path between a certain isolated data source node and at least one data source node in the data fusion topology is less than the preset impedance threshold, the isolated data source node is included in the data fusion topology.

[0132] It can be understood that this realizes the secondary screening of isolated data source nodes, avoiding missing high-quality data source nodes due to too strict initial threshold setting. For example, the impedance identification of the fusion path between the isolated data source node F and the data source node A in the data fusion topology is 0.45, which is less than the preset impedance threshold 0.5, so the isolated data source node F is included in the data fusion topology.

[0133] Preferably, steps B1 to B3 select the connected subset with the smallest impedance identification sum as the data fusion topology, rather than the connected subset containing the most data source nodes, which reflects the fusion principle of less but better than more miscellaneous. This method avoids including data source nodes with poor compatibility in pursuit of the number of data source nodes, ensures the best overall compatibility between data source nodes in the data fusion topology, and improves the quality of subsequent fusion processing.

[0134] Preferably, step 100 realizes the selective fusion of data source nodes by calculating the impedance identifier of the fusion path, screening low-impedance fusion paths to form a connected subset, selecting the connected subset with the smallest impedance identifier sum as the data fusion topology, and performing secondary screening on isolated data source nodes. This method identifies and eliminates data source nodes that are not synchronized in time, not matched in space, or have abnormal data distribution, avoids the interference of incompatible data source nodes on the fusion result, solves the problem of decreased fusion accuracy caused by the increase of data source nodes, and improves the accuracy of data fusion.

[0135] Step 200: Processing anchor point deployment state probe monitors fusion data of the data fusion topology, and determines the scheduling mode of the processing anchor point according to the monitoring result.

[0136] Among them, the processing anchor point represents the convergence point and processing point of the data flow in the data fusion topology, which is the core position of performing fusion operation. The state probe represents the monitoring mechanism deployed in the processing anchor point, which is used to collect the data state of each data source node in the data fusion topology in real time. The scheduling mode represents the data processing strategy adopted by the processing anchor point.

[0137] It should be noted that the traditional data fusion processing method adopts a fixed stream-batch integrated architecture, regardless of the change of data state. For example, in the energy monitoring scene, when the power system is in a stable running state, the data changes slowly, at this time, the batch processing method can accumulate enough data to improve the processing accuracy; however, when a device failure or load mutation occurs, the data fluctuates violently, at this time, the batch processing method will cause response delay, and cannot discover and process abnormal situations in time. The traditional method cannot adjust the processing method according to the change of data state, and cannot balance between response speed and processing accuracy.

[0138] In order to determine the appropriate processing method according to the change of data state, the state probe is deployed in the processing anchor point to monitor the fusion data of the data fusion topology, and the scheduling mode of the processing anchor point is determined according to the monitoring result.

[0139] In some embodiments, step 200 includes steps C1 to C3:

[0140] Step C1: Deploying a state probe in the processing anchor point to collect the data fluctuation amplitude of each data source node in the data fusion topology.

[0141] Among them, the data fluctuation amplitude represents the change degree of the data generated by the data source node within a certain time. It can be understood that the data fluctuation amplitude can be represented by calculating statistical indicators such as variance, standard deviation, range, etc. of the data, or by calculating the change amount of the data values of adjacent time points.

[0142] In processing the anchor point deployment state probe, the data fluctuation amplitude of each data source node in the data fusion topology is continuously collected. For example, in an energy monitoring scenario, the current data collected by a power monitoring station in the past 10 minutes is 100A, 102A, 101A, 99A, and 100A, and the data fluctuation amplitude of the data source node can be calculated as the standard deviation of 1.58A, or the maximum value of the data change amount of adjacent time points is 3A.

[0143] Step C2: When the data fluctuation amplitude exceeds the first fluctuation threshold, the scheduling mode of the processing anchor point is determined as the stream processing mode.

[0144] The first fluctuation threshold is used to determine whether the data is in a state of violent fluctuation. The stream processing mode means that the current data of the data source node is processed in fusion one by one to achieve fast response.

[0145] When the data fluctuation amplitude exceeds the first fluctuation threshold, it indicates that the data in the data fusion topology is changing violently, which may be caused by device failure, load mutation, or other abnormal situations, and needs to be responded quickly. At this time, the scheduling mode of the processing anchor point is determined as the stream processing mode, and the current data of the data source node is processed in fusion one by one, so as to quickly obtain the fusion result and perform abnormal judgment and processing.

[0146] For example, in an energy monitoring scenario, the current data of a power monitoring station changes from 100A to 500A in 1 minute, and the data fluctuation amplitude is 400A, which exceeds the first fluctuation threshold of 200A. Therefore, the scheduling mode of the processing anchor point is determined as the stream processing mode, and the current data is immediately fused and processed to quickly identify the current abnormal situation.

[0147] Step C3: When the data fluctuation amplitude does not exceed the first fluctuation threshold, the scheduling mode of the processing anchor point is determined as the batch processing mode.

[0148] The batch processing mode means that the accumulated data of the data source node in a certain time window is processed in batch fusion to improve the processing precision.

[0149] When the data fluctuation amplitude does not exceed the first fluctuation threshold, it indicates that the data in the data fusion topology is in a relatively stable state and does not need to be responded quickly. At this time, the scheduling mode of the processing anchor point is determined as the batch processing mode, and after the data source node accumulates enough data in a certain time window, batch fusion processing is performed.

[0150] Preferably, step 200 realizes dynamic switching of processing modes according to changes in data states by collecting data fluctuation amplitudes of each data source node in the data fusion topology through processing anchor point deployment state probes, and determining the scheduling mode of the processing anchor according to whether the data fluctuation amplitude exceeds a first fluctuation threshold. When the data fluctuation amplitude exceeds the first fluctuation threshold, the stream processing mode is used to quickly respond to burst situations; when the data fluctuation amplitude does not exceed the first fluctuation threshold, the batch processing mode is used to improve processing accuracy. This method takes into account the response speed in burst scenarios and the processing accuracy in smooth scenarios, and compared with traditional methods with fixed architecture, it adapts to different needs for response speed and processing accuracy in different scenarios.

[0151] Step 300: Obtain the fusion result based on the scheduling mode.

[0152] Among them, the fusion result represents the comprehensive data obtained by fusing the data of each data source node in the data fusion topology. Based on the scheduling mode determined in step 200, the corresponding data processing mode is used to obtain the fusion result.

[0153] It should be noted that in the traditional method, whether stream processing or batch processing is used, the same fusion algorithm and parameter configuration are used when obtaining the fusion result. However, stream processing mode and batch processing mode differ in data volume, data characteristics, processing target, etc., and using the same fusion method may result in poor processing effect. For example, in the stream processing mode, the data volume is small, and a simple and efficient fusion algorithm is suitable to ensure response speed; in the batch processing mode, the data volume is large, and a complex and accurate fusion algorithm can be used to improve fusion accuracy.

[0154] In some embodiments, step 300 includes step D1 and step D2:

[0155] Step D1: When the scheduling mode is determined to be the stream processing mode, perform fusion operation on the current data of each data source node in the data fusion topology to obtain the fusion result.

[0156] Among them, the current data represents the data collected by the data source node at the current time. When the scheduling mode is the stream processing mode, it means that the data is fluctuating sharply and needs to be quickly obtained. At this time, the fusion operation is performed on the current data of each data source node in the data fusion topology to obtain the fusion result.

[0157] It can be understood that the fusion operation in the stream processing mode can adopt a simple and efficient fusion algorithm. For example, a weighted average algorithm can be adopted to perform weighted average on the current data according to the weights of the data source nodes to obtain the fusion result. The weight can be determined according to the historical accuracy, data quality score and other factors of the data source node. For example, the data fusion topology includes data source node A, data source node B and data source node C, and the current current data of them is 500A, 480A and 490A respectively, and the weight is 0.4, 0.3 and 0.3 respectively, then the fusion result is 500*0.4+480*0.3+490*0.3=491A.

[0158] It is not difficult to understand that the stream processing mode can obtain the fusion result in a short time after data acquisition by performing fusion processing on the current data piece by piece, and can realize rapid response to sudden situations. For example, in the energy monitoring scene, when an abnormal mutation of current is detected, the stream processing mode can complete the fusion processing and output the fusion result within 1 second, and timely trigger the fault warning.

[0159] Step D2: when the scheduling mode is the batch processing mode, performing fusion operation on the accumulated data of each data source node in the data fusion topology in the first time window to obtain the fusion result.

[0160] Among them, the first time window represents the time range for accumulating data, and the accumulated data represents all the data collected by the data source node in the time range. When the scheduling mode is the batch processing mode, it means that the data is in a relatively stable state, and batch processing can be performed after data accumulation. At this time, the fusion operation is performed on the accumulated data of each data source node in the data fusion topology in the first time window to obtain the fusion result.

[0161] It can be understood that the fusion operation in the batch processing mode can adopt a complex and accurate fusion algorithm. For example, the accumulated data of each data source node in the first time window can be preprocessed such as time alignment, outlier elimination, missing value filling, and then fused by using Kalman filtering, Bayesian fusion and other algorithms. For example, the data fusion topology includes data source node A, data source node B and data source node C, and the first time window is 10 minutes, and they have collected 10 pieces of current data in the time window. After time alignment, the data of the three data source nodes is fused by using Kalman filtering algorithm to obtain 10 pieces of fused current data as the fusion result.

[0162] Preferably, step 300 obtains the fusion result by adopting different data processing methods according to the scheduling mode, realizing the matching of the processing method and the data state. When the scheduling mode is the stream processing mode, a simple and efficient fusion algorithm is adopted for the current data, ensuring the response speed in the burst situation; when the scheduling mode is the batch processing mode, a complex and accurate fusion algorithm is adopted for the accumulated data, improving the fusion accuracy in the smooth scene. Compared with the traditional method adopting a fixed processing method, the method can obtain a better fusion result in different data states.

[0163] Step 400: determining a static data mark according to the content sensitivity of the fusion result, determining a dynamic data mark according to the aggregation number of the data source node, determining a scene data mark based on the current processing scene, and taking the highest mark among the static data mark, the dynamic data mark and the scene data mark as the data mark.

[0164] Among them, the data mark represents the sensitivity level of the fusion result, which is used to guide the data access permission control. The static data mark represents the sensitivity level determined based on the content attribute of the fusion result. The dynamic data mark represents the sensitivity level determined based on the aggregation number of the data source node. The scene data mark represents the sensitivity level determined based on the current processing scene.

[0165] It should be noted that the traditional data sensitivity marking method only determines the sensitivity level according to the content attribute of the data, such as marking the user electricity data as low sensitivity, marking the regional load data as medium sensitivity, and marking the power grid operation data as high sensitivity. However, this static marking method ignores the influence of data aggregation and use scene on sensitivity. For example, the electricity data of a single user has low sensitivity, but when the electricity data of 1000 users is aggregated to form regional load data, sensitive information such as the personnel distribution and commercial activities of the region can be inferred, and the data sensitivity should be improved. For another example, the same regional load data has low sensitivity when used for regular analysis in a normal operation scene, but the sensitivity should be improved when used for locating responsibility in a fault investigation scene. The traditional method cannot identify the sensitivity improvement caused by data aggregation and scene change, and there is a risk of data leakage.

[0166] In order to realize the dynamic marking of data sensitivity, the static data mark is determined according to the content sensitivity of the fusion result, the dynamic data mark is determined according to the aggregation number of the data source node, and the scene data mark is determined based on the current processing scene. The highest mark among the static data mark, the dynamic data mark and the scene data mark is taken as the data mark.

[0167] In some embodiments, the determination of the static data mark according to the content sensitivity of the fusion result in step 400 includes steps E1 to E3:

[0168] Step E1: Extract the data type identifier in the fusion result.

[0169] The data type identifier represents the data category to which the fusion result belongs. In the energy monitoring scenario, the data type identifier can be user electricity data, device operation data, regional load data, power grid topology data, etc. The data type identifier can be extracted from the metadata of the fusion result, or inferred by analyzing the data fields, data sources, etc. of the fusion result.

[0170] For example, a fusion result contains data fields such as current, voltage, power, etc., and the data source is a single user-side electricity terminal, then it can be determined that the data type identifier of the fusion result is user electricity data. For another example, a fusion result contains data fields such as bus voltage, transformer load rate, switch state, etc., and the data source is a substation monitoring device, then it can be determined that the data type identifier of the fusion result is device operation data.

[0171] Step E2: Query the sensitivity level corresponding to the data type identifier based on the sensitivity mapping table.

[0172] The sensitivity mapping table stores the mapping relationship between various data type identifiers and their corresponding sensitivity levels. The sensitivity level can be divided into L1, L2, L3, L4, etc. The larger the level value, the higher the sensitivity. The sensitivity mapping table can be formulated according to national laws and regulations, industry safety standards, enterprise data security policies, etc.

[0173] For example, in the energy monitoring scenario, the sensitivity mapping table can specify that user electricity data corresponds to sensitivity level L1, device operation data corresponds to sensitivity level L2, regional load data corresponds to sensitivity level L3, and power grid topology data corresponds to sensitivity level L4. Based on the sensitivity mapping table, the sensitivity level corresponding to the data type identifier can be determined, that is, the content sensitivity of the fusion result.

[0174] Step E3: Use the sensitivity level as a static data marker.

[0175] The sensitivity level obtained by querying step E2 is used as a static data marker. The static data marker reflects the sensitivity of the fusion result at the content level, and does not change with data aggregation and use scenario changes.

[0176] For example, the data type identifier of a certain fusion result is user electricity data, and the sensitivity level obtained based on the sensitivity mapping table is L1, then L1 is used as the static data marker of the fusion result.

[0177] Preferably, steps E1 to E3 are based on the sensitivity mapping table to query the sensitivity level corresponding to the data type identifier, and the sensitivity level is used as the static data mark to establish the mapping relationship between the content attribute of the fusion result and the sensitivity level. This method provides a basic sensitivity mark for the fusion result and provides a benchmark for determining dynamic data marks and scene data marks.

[0178] In some embodiments, the dynamic data mark is determined according to the aggregation number of the data source nodes in step 400, including steps F1 to F3:

[0179] Step F1: Count the number of data source nodes participating in the generation of the fusion result to obtain the aggregation number.

[0180] The aggregation number represents the number of data source nodes participating in the fusion operation when the fusion result is generated. It can be understood that the aggregation number reflects the degree of data aggregation of the fusion result. The more the aggregation number, the more data source nodes the fusion result aggregates, the more information it contains, and the more sensitive information it can infer.

[0181] For example, a certain fusion result is generated by the data of data source node A, data source node B and data source node C, and the aggregation number is 3. For another example, a certain fusion result is generated only by the data of data source node A without fusion operation, and the aggregation number is 1.

[0182] Step F2: When the aggregation number exceeds the first aggregation threshold, upgrade the static data mark by the first level to obtain the dynamic data mark.

[0183] The first aggregation threshold is used to determine whether the aggregation number reaches the degree of upgrading the sensitivity level. The first level represents the upgrade amplitude of the sensitivity level.

[0184] It should be noted that when the aggregation number exceeds the first aggregation threshold, it means that the fusion result aggregates data of more data source nodes, and the degree of data aggregation is higher. At this time, even if the content attribute of the fusion result corresponds to a lower static data mark, more sensitive information can be inferred due to data aggregation, and the static data mark needs to be upgraded by the first level to obtain the dynamic data mark.

[0185] For example, the static data mark of a certain fusion result is L1, the first aggregation threshold is set to 10, and the aggregation number is 15, which exceeds the first aggregation threshold. At this time, the static data mark L1 is upgraded by the first level to obtain the dynamic data mark L2. This indicates that although the user electricity data sensitivity of a single data source node is L1, the regional electricity mode and other sensitive information can be inferred after aggregating the electricity data of 15 users, and the sensitivity should be upgraded to L2.

[0186] Step F3: If the aggregation quantity does not exceed the first aggregation threshold, mark the static data label as the dynamic data label.

[0187] When the aggregation quantity does not exceed the first aggregation threshold, the data aggregation degree of the fusion result is low, which is not enough to significantly improve the sensitivity. At this time, the static data label is marked as the dynamic data label, and no upgrade is performed.

[0188] For example, the static data label of a certain fusion result is L1, the first aggregation threshold is set to 10, the aggregation quantity is 5, and does not exceed the first aggregation threshold. At this time, the static data label L1 is directly marked as the dynamic data label L1, and no upgrade is performed.

[0189] Specifically, step F2 includes steps F31 to F33:

[0190] It should be noted that the "upgrade the static data label by one level to obtain the dynamic data label" described in step F2 is a simplified description. In fact, the upgrade range of the sensitivity level should be dynamically determined according to the size of the aggregation quantity. The more the aggregation quantity, the higher the data aggregation degree, and the greater the possibility of inferring sensitive information. The upgrade range of the sensitivity level should be greater. The conventional method uses a fixed upgrade range, such as uniformly upgrading one level regardless of whether 10 or 100 data source nodes are aggregated. This method cannot accurately reflect the influence of the aggregation quantity on the sensitivity.

[0191] Step F31: Calculate the quantity ratio of the aggregation quantity to the first aggregation threshold.

[0192] The quantity ratio represents the multiple relationship of the aggregation quantity relative to the first aggregation threshold, and is used to quantify the data aggregation degree. Calculating the quantity ratio of the aggregation quantity to the first aggregation threshold can reflect the degree to which the aggregation quantity exceeds the threshold.

[0193] For example, the first aggregation threshold is set to 10, when the aggregation quantity is 15, the quantity ratio is 15 / 10=1.5; when the aggregation quantity is 50, the quantity ratio is 50 / 10=5. The greater the quantity ratio, the higher the data aggregation degree.

[0194] Step F32: Determine the upgrade level number according to the quantity ratio.

[0195] The upgrade level number represents the number of levels that the static data label should be upgraded. According to the quantity ratio, the upgrade level number is determined, which can realize the matching of the sensitivity level upgrade range and the data aggregation degree.

[0196] It can be understood that a linear or nonlinear mapping relationship can be established between the number of upgrade levels and the quantity ratio. For example, it can be stipulated that when the quantity ratio is between 1 and 2, the number of upgrade levels is 1; when the quantity ratio is between 2 and 5, the number of upgrade levels is 2; and when the quantity ratio is above 5, the number of upgrade levels is 3. This means that the more the aggregated quantity, the greater the sensitivity level upgrade range.

[0197] For example, when the quantity ratio is 1.5, the number of upgrade levels is determined to be 1; when the quantity ratio is 3, the number of upgrade levels is determined to be 2; and when the quantity ratio is 6, the number of upgrade levels is determined to be 3.

[0198] Step F33: Upgrade the static data mark to the level corresponding to the number of upgrade levels to obtain a dynamic data mark.

[0199] The static data mark is upgraded according to the number of upgrade levels to obtain a dynamic data mark. For example, the static data mark is L1, and the number of upgrade levels is 2, then the dynamic data mark is L1+2=L3. For another example, the static data mark is L2, and the number of upgrade levels is 1, then the dynamic data mark is L2+1=L3.

[0200] It is not difficult to understand that by determining the number of upgrade levels according to the quantity ratio of the aggregated quantity to the first aggregation threshold, the sensitivity level upgrade range is realized to grow nonlinearly with the aggregated quantity. When a small amount of data source nodes are aggregated, the sensitivity level upgrade range is small; when a large amount of data source nodes are aggregated, the sensitivity level upgrade range is large. This conforms to the actual law of the influence of data aggregation on sensitivity.

[0201] For example, the static data mark of a certain fusion result is L1, and the first aggregation threshold is 10. When the aggregated quantity is 15, the quantity ratio is 1.5, the number of upgrade levels is 1, and the dynamic data mark is L2; when the aggregated quantity is 100, the quantity ratio is 10, the number of upgrade levels is 3, and the dynamic data mark is L4. This means that aggregating the power consumption data of 100 users is more significant in improving sensitivity than aggregating the power consumption data of 15 users.

[0202] It should be noted that after step F33, steps F34 and F35 are also included:

[0203] It should be noted that the sensitivity improvement caused by data aggregation is temporary. When the fusion result is split or only the data of a single data source node is extracted, the degree of data aggregation decreases, and the sensitivity should decrease accordingly. The traditional method does not adjust the data mark once it is determined, resulting in that even if the data has been de-aggregated, the high sensitivity mark is still maintained, causing the data access permission control to be too strict and reducing the data utilization efficiency.

[0204] Step F34: Monitor the usage state of the fusion result.

[0205] The use state indicates the current processing situation of the fusion result, including whether the aggregation operation is being performed, whether it has been split into single-point data, etc. Monitoring the use state of the fusion result can identify changes in the degree of data aggregation.

[0206] For example, the use state of the fusion result can be monitored by recording the processing log of the fusion result, such as whether the fusion result is further aggregated, whether it is split, whether it is extracted from a single data source node, etc.

[0207] Step F35: When it is determined that the aggregation operation of the fusion result is completed and the fusion result has returned to single-point data corresponding to the single data source node, the dynamic data label is downgraded to the static data label.

[0208] The single-point data refers to unaggregated data containing only data of a single data source node. When the aggregation operation of the fusion result is completed and the fusion result has returned to single-point data, it indicates that the degree of data aggregation has been reduced to the lowest level, and there is no longer any sensitivity enhancement caused by data aggregation. At this time, the dynamic data label is downgraded to the static data label, and the sensitivity level based on the content attribute is restored.

[0209] It can be understood that this achieves dynamic adjustment of the data label and avoids overly strict access permission control. For example, a fusion result aggregates power consumption data of 100 users, and the static data label is L1, while the dynamic data label is L4. When the fusion result is split and only single-point data of user A is extracted, it is monitored that the fusion result has returned to single-point data, and the dynamic data label is downgraded from L4 to the static data label L1. This indicates that the sensitivity of the power consumption data of a single user is restored to L1, and it can be accessed by more authorized users.

[0210] Preferably, steps F1 to F3 determine the aggregation number by counting the number of data source nodes involved in generating the fusion result, determine whether to upgrade the static data label according to the relationship between the aggregation number and the first aggregation threshold, and determine the number of upgrade levels according to the number ratio, which achieves non-linear growth of the sensitivity level with the degree of data aggregation. This method identifies the sensitivity enhancement caused by data aggregation and solves the problem that the traditional static labeling method cannot reflect the aggregation effect. At the same time, by monitoring the use state of the fusion result, the dynamic data label is downgraded to the static data label when the data is de-aggregated, which achieves dynamic adjustment of the data label and avoids overly strict access permission control, thereby improving data utilization efficiency.

[0211] In some embodiments, determining the scene data label based on the current processing scenario in step 400 includes steps G1 to G3:

[0212] Step G1: Receive the scene type identifier carried in the processing request.

[0213] wherein the processing request represents a request for accessing or processing the fusion result, and the scenario type identifier represents a category of application scenario to which the processing request belongs. In the energy monitoring scenario, the scenario type identifier can be daily monitoring, statistical analysis, report generation, fault investigation, emergency response, security audit, etc. The processing request can be initiated by a user or triggered automatically by the system.

[0214] For example, an operation and maintenance personnel initiates a processing request for viewing regional load data, and the scenario type identifier is daily monitoring. For another example, when a device fault is detected, the system automatically triggers a fault positioning processing request, and the scenario type identifier is fault investigation.

[0215] Step G2: When the scenario type identifier indicates an abnormal processing scenario, upgrade the dynamic data mark to a second level to obtain a scenario data mark.

[0216] wherein the abnormal processing scenario represents an application scenario for processing special situations such as faults, abnormalities, security events, etc., including fault investigation, emergency response, security audit, etc. The second level represents the magnitude of the upgrade of the sensitivity of the scenario.

[0217] It should be noted that the same fusion result has different sensitivities in different application scenarios. In a normal processing scenario, the fusion result is used for routine monitoring, analysis, statistics, etc. The sensitivity is relatively low. However, in an abnormal processing scenario, the fusion result is used for fault positioning, responsibility tracing, security evidence, etc. It may involve sensitive matters such as responsibility identification, economic compensation, and legal proceedings. The sensitivity should be improved. The traditional method does not distinguish between application scenarios and uses the same data mark for all scenarios, which cannot reflect the influence of scenario differences on sensitivity.

[0218] When it is determined that the scenario type identifier indicates an abnormal processing scenario, the dynamic data mark is upgraded by a second level to obtain a scenario data mark. For example, the dynamic data mark of a certain fusion result is L2, and the scenario type identifier is fault investigation, which belongs to an abnormal processing scenario. At this time, the dynamic data mark L2 is upgraded by a second level to obtain a scenario data mark L4. This indicates that the sensitivity of the fusion result in the fault investigation scenario is L4, and higher access permission is required to access, preventing unauthorized personnel from obtaining fault investigation data.

[0219] Step G3: When the scenario type identifier indicates a normal processing scenario, the dynamic data mark is used as the scenario data mark.

[0220] The normal processing scene represents a regular application scene for daily operation, including daily monitoring, statistical analysis, report generation, etc. When the scene type identifier indicates the normal processing scene, it means that the fusion result is used for regular operation and does not involve sensitive matters, and does not need to upgrade the sensitivity. At this time, the dynamic data mark is marked as the scene data mark, and no upgrade is performed.

[0221] For example, the dynamic data mark of a certain fusion result is L2, the scene type identifier is daily monitoring, and it belongs to the normal processing scene. At this time, the dynamic data mark L2 is directly used as the scene data mark L2, and no upgrade is performed.

[0222] Finally, the highest mark among the static data mark, the dynamic data mark, and the scene data mark is taken as the data mark.

[0223] It can be understood that the static data mark reflects the content sensitivity of the fusion result, the dynamic data mark reflects the sensitivity improvement caused by data aggregation, and the scene data mark reflects the sensitivity improvement caused by the application scene. Taking the highest mark among the three as the final data mark can ensure that the fusion result is given a high enough sensitivity level in any case, preventing data leakage.

[0224] For example, the static data mark of a certain fusion result is L1, the dynamic data mark is L3, and the scene data mark is L4. The highest mark L4 is taken as the data mark of the fusion result. For example, the static data mark of a certain fusion result is L3, the dynamic data mark is L3, and the scene data mark is L3. The highest mark L3 is taken as the data mark of the fusion result.

[0225] It is not difficult to understand that by taking the highest mark, the principle of strict control of data security is realized, that is, in the three dimensions of content sensitivity, data aggregation degree, and application scene, as long as any one dimension needs to improve the sensitivity, the sensitivity of the fusion result will be improved to the corresponding level.

[0226] Preferably, step 400 determines the static data label according to the content sensitivity of the fusion result, determines the dynamic data label according to the aggregation number of the data source node, determines the scene data label based on the current processing scene, and takes the highest label among the three as the data label, thereby achieving multi-dimensional dynamic labeling of data sensitivity. This method not only considers the content attribute of the fusion result, but also considers the influence of data aggregation degree and application scene on sensitivity. Compared with the traditional static labeling method, this method can identify the sensitivity improvement caused by data aggregation and scene change, solve the data leakage risk caused by the unsynchronized sensitivity improvement after single-point data aggregation, and solve the security hidden danger caused by the undifferentiated sensitivity of the same data in different scenes. At the same time, by monitoring the use state of the fusion result, the data label is dynamically degraded when the data is de-aggregated, thereby avoiding excessive strict access permission control and improving data utilization efficiency. This method realizes the scene-sensitive identification of data sensitivity and improves the data security management capability.

[0227] On the basis of the above steps, the present application further includes the following embodiments:

[0228] Referring to Figure 3 , which is a structural schematic diagram of a multi-source heterogeneous data fusion processing system provided by an embodiment of the present application. The multi-source heterogeneous data fusion processing system includes:

[0229] A topology construction module is configured to acquire a plurality of data source nodes, calculate impedance identifiers of fusion paths between the data source nodes, and form a data fusion topology according to the impedance identifiers.

[0230] A state monitoring module is configured to deploy a state probe to monitor fusion data of the data fusion topology at a processing anchor point, and determine a scheduling mode of the processing anchor point according to a monitoring result.

[0231] A fusion execution module is configured to acquire a fusion result based on the scheduling mode.

[0232] A label determination module is configured to determine a static data label according to a content sensitivity of the fusion result, determine a dynamic data label according to an aggregation number of the data source node, determine a scene data label based on a current processing scene, and take a highest label among the static data label, the dynamic data label and the scene data label as a data label. Figure 3 The device of the embodiment shown can be used to execute the steps in the method embodiment shown, and the implementation principles and technical effects are similar, which will not be described here again. Figure 1

[0233] Referring to Figure 4 , which is a hardware structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 40 includes a processor 41, a memory 42 and a computer program; wherein, ​

[0234] a memory 42 for storing the computer program, which can also be a flash memory. The computer program is, for example, an application program, a function module, etc. for implementing the above method.

[0235] a processor 41 for executing the computer program stored in the memory to implement each step of the method performed by the device. Details can be referred to the relevant description in the foregoing method embodiments.

[0236] Optionally, the memory 42 can be independent or integrated with the processor 41.

[0237] When the memory 42 is a device independent of the processor 41, the device can further comprise:

[0238] a bus 43 for connecting the memory 42 and the processor 41.

[0239] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-source heterogeneous data fusion processing method, characterized in that, The method comprises the following steps: obtaining a plurality of data source nodes, wherein the data source nodes are distributed in different geographical locations, and the data source nodes transmit collected energy data to an energy data fusion platform through a communication network; calculating impedance identifiers of fusion paths between the data source nodes, comprising: obtaining sampling time intervals of the data source nodes, and calculating a ratio of the sampling time intervals of any two data source nodes; obtaining coordinates of coverage areas of the data source nodes, and calculating overlapping areas of the coverage areas of any two data source nodes; extracting data distribution characteristics of the data source nodes, and calculating data distribution characteristic distances of any two data source nodes; determining the impedance identifiers of the fusion paths based on the ratio of the sampling time intervals, the overlapping areas of the coverage areas, and the data distribution characteristic distances; forming a data fusion topology according to the impedance identifiers; deploying a state probe at a processing anchor to monitor fusion data of the data fusion topology, and determining a scheduling mode of the processing anchor according to a monitoring result, comprising: deploying the state probe at the processing anchor to collect data fluctuation amplitudes of the data source nodes in the data fusion topology; determining that the scheduling mode of the processing anchor is a stream processing mode when the data fluctuation amplitude exceeds a first fluctuation threshold; determining that the scheduling mode of the processing anchor is a batch processing mode when the data fluctuation amplitude does not exceed the first fluctuation threshold; obtaining a fusion result based on the scheduling mode; determining a static data marker according to a content sensitivity of the fusion result, determining a dynamic data marker according to an aggregated number of the data source nodes, determining a scene data marker based on a current processing scene, and taking the highest marker among the static data marker, the dynamic data marker, and the scene data marker as a data marker.

2. The method of claim 1, wherein: forming the data fusion topology according to the impedance identifiers comprises: labeling a fusion path with an impedance identifier less than a preset impedance threshold as a low-impedance fusion path; identifying data source nodes connected by the low-impedance fusion path, and grouping the data source nodes in mutual communication into a same connected subset; calculating a sum of the impedance identifiers of the fusion paths within each connected subset, and selecting a connected subset with the smallest sum of the impedance identifiers as the data fusion topology.

3. The method of claim 2, wherein: after the connected subset with the smallest sum of the impedance identifiers is selected as the data fusion topology, the method further comprises: identifying isolated data source nodes outside the data fusion topology; calculating impedance identifiers of fusion paths between each isolated data source node and each data source node in the data fusion topology; when there is a fusion path with an impedance identifier less than the preset impedance threshold, incorporating the corresponding isolated data source node into the data fusion topology.

4. The method of claim 1, wherein: determining the impedance identifiers of the fusion paths based on the ratio of the sampling time intervals, the overlapping areas of the coverage areas, and the data distribution characteristic distances comprises: taking a logarithm of the ratio of the sampling time intervals to obtain a time deviation; calculating a ratio of the overlapping areas of the coverage areas to a union area of the coverage areas of the data source nodes to obtain a spatial matching amount. Normalizing the data distribution feature distance to obtain a distribution deviation; When any one of the time deviation, the space matching amount, and the distribution deviation exceeds a corresponding threshold, marking the impedance identifier of the fusion path as high impedance; When none of the time deviation, the space matching amount, and the distribution deviation exceeds a corresponding threshold, combining the time deviation, the space matching amount, and the distribution deviation to obtain the impedance identifier of the fusion path.

5. The method of claim 1, wherein the obtaining the fusion result based on the scheduling mode comprises: when the scheduling mode is determined to be the stream processing mode, performing a fusion operation on current data of each data source node in the data fusion topology to obtain the fusion result; and when the scheduling mode is determined to be the batch processing mode, performing a fusion operation on accumulated data of each data source node in the data fusion topology within a first time window to obtain the fusion result.

6. The method of claim 1, wherein the determining the static data mark according to the content sensitivity of the fusion result comprises: extracting a data type identifier in the fusion result; querying a sensitivity level corresponding to the data type identifier based on a sensitivity mapping table; and taking the sensitivity level as the static data mark.

7. The method of claim 1, wherein the determining the dynamic data mark according to the aggregation number of the data source node comprises: counting a number of data source nodes participating in generating the fusion result to obtain an aggregation number; when the aggregation number exceeds a first aggregation threshold, upgrading the static data mark by a first level to obtain the dynamic data mark; and when the aggregation number does not exceed the first aggregation threshold, taking the static data mark as the dynamic data mark.

8. The method of claim 7, wherein the upgrading the static data mark by the first level to obtain the dynamic data mark when the aggregation number exceeds the first aggregation threshold comprises: calculating a number ratio of the aggregation number to the first aggregation threshold; determining an upgrade level number according to the number ratio; and upgrading the static data mark to a level corresponding to the upgrade level number to obtain the dynamic data mark.

9. The method of claim 8, further comprising, after the upgrading the static data mark to the level corresponding to the upgrade level number to obtain the dynamic data mark: monitoring a use state of the fusion result; and when an aggregation operation of the fusion result ends and the fusion result returns to single-point data corresponding to a single data source node, downgrading the dynamic data mark to the static data mark.

10. The method of claim 1, wherein the determining the scene data mark based on a current processing scene comprises: receiving a scene type identifier carried in a processing request; when the scene type identifier indicates an abnormal processing scene, upgrading the dynamic data mark by a second level to obtain the scene data mark; and when the scene type identifier indicates a normal processing scene, taking the dynamic data mark as the scene data mark. ​ ​ ​ ​ ​ ​ 11. A multi-source heterogeneous data fusion processing system, which adopts the multi-source heterogeneous data fusion processing method according to any one of claims 1-10, characterized in that, ​ A topology construction module is configured to acquire a plurality of data source nodes, the data source nodes are distributed in different geographical locations, and the data source nodes transmit collected energy data to an energy data fusion platform through a communication network; Impedance identification of a fusion path between each data source node is calculated, including: A sampling time interval of each data source node is acquired, and a ratio of the sampling time intervals of any two data source nodes is calculated; Coordinates of a coverage area of each data source node are acquired, and an overlapping area of the coverage areas of any two data source nodes is calculated; Data distribution features of each data source node are extracted, and a data distribution feature distance of any two data source nodes is calculated; The impedance identification of the fusion path is determined based on the ratio of the sampling time intervals, the overlapping area of the coverage areas, and the data distribution feature distance; A data fusion topology is formed according to the impedance identification; A state monitoring module is configured to monitor fusion data of the data fusion topology by deploying a state probe at a processing anchor, and determine a scheduling mode of the processing anchor according to a monitoring result, including: A data fluctuation amplitude of each data source node in the data fusion topology is collected by deploying a state probe at the processing anchor; When the data fluctuation amplitude exceeds a first fluctuation threshold, the scheduling mode of the processing anchor is determined as a stream processing mode; When the data fluctuation amplitude does not exceed the first fluctuation threshold, the scheduling mode of the processing anchor is determined as a batch processing mode; A fusion execution module is configured to acquire a fusion result based on the scheduling mode; A label determination module is configured to determine a static data label according to a content sensitivity of the fusion result, determine a dynamic data label according to an aggregated number of the data source nodes, determine a scene data label based on a current processing scene, and take a highest label among the static data label, the dynamic data label, and the scene data label as a data label.

Citation Information

Patent Citations

  • Screw air compressor fault prediction method based on multi-source data fusion

    CN118378195A

  • Power grid data management system and method based on artificial intelligence data analysis

    CN119557805A