A multi-source heterogeneous data fusion method and system based on multi-agent reinforcement learning, a terminal device, and a storage medium
Patent Information
- Application Number
- CN202611057816.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明实施例提供一种基于多代理强化学习的多源异构数据融合方法、系统、终端设备及存储介质,能有效解决现有技术无法根据电网运行工况的动态变化进行权重自适应优化,导致无法有效融合数据的问题
本发明提供一种基于多代理强化学习的多源异构数据融合方法、系统、终端设备及存储介质,其方法能够采用多代理强化学习架构,由代理单元基于数据源实时状态自主生成数据权重,不依赖人工经验;集中器通过全局融合质量评估和全局奖励信号形成闭环反馈,驱动代理单元自动优化权重生成策略,解决人工静态加权融合无法对权重进行动态优化的问题。代理单元实时根据数据源矩阵与异构数据生成局部状态向量,实时感知数据源质量、时序特征、关联关系变化;以全局奖励信号为优化目标,代理单元持续调整权重策略,权重随电网工况、数据质量实时动态适配;支持重复执行双网络迭代优化,权重生成策略网络持续收敛到当前最优,实现动态工况下的持续最优权重分配,最终根据新的数据权重向量以及数据源矩阵进行加权计算,得到多源异构数据融合结果,实现了数据的实时融合,并且能够动态调整数据权重向量以实现权重自适应优化,适配电力系统对数据融合的高时效要求。
Smart Images

Figure CN122839276A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, system, terminal device and storage medium for multi-source heterogeneous data fusion based on multi-agent reinforcement learning. Background Technology
[0002] Current power system dispatching and operation rely on multi-source heterogeneous data, including real-time grid operation data, equipment status monitoring data, environmental and meteorological data, load data, and log data. The heterogeneity of these data, due to differences in sampling frequency, format, dimension, and semantics, poses a significant challenge to real-time data processing.
[0003] Traditional data processing methods are mostly based on static weighted fusion using manual rules. They rely on the experience of dispatchers to preset fixed data weights, which cannot adaptively optimize the weights according to the dynamic changes in the power grid's operating conditions. They are difficult to effectively fuse and analyze information flows with significant differences, and cannot meet the high timeliness requirements of data fusion for dispatch and maintenance of new power systems. Summary of the Invention
[0004] This invention provides a method, system, terminal device, and storage medium for multi-source heterogeneous data fusion based on multi-agent reinforcement learning. It can effectively solve the problem that existing technologies cannot adaptively optimize weights according to the dynamic changes in power grid operating conditions, resulting in the inability to effectively fuse data.
[0005] One embodiment of the present invention provides a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning, applicable to concentrators, the method comprising: Acquire heterogeneous data from various data sources of the power grid to be processed; Heterogeneous data is matrixed and normalized to generate a data source matrix. Based on the data source matrix, a corresponding agent unit is assigned to each data source, and the data source matrix is sent to each agent unit so that each agent unit generates a local state vector based on the data source matrix and heterogeneous data. In a preset weight generation strategy network, a data weight vector is generated based on the local state vector, and the local state vector and data weight vector are transmitted to the concentrator. A global state vector is generated by concatenating the local state vector and the data weight vector. A fusion quality assessment is then performed based on the global state vector and the data source matrix to obtain a global fusion quality score. A global reward signal is generated based on the global fusion quality score and sent to each agent unit. This enables each agent unit to repeatedly perform a dual-network iterative optimization operation to optimize the weight generation strategy network, using the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. The data weight vector is then regenerated based on the optimized weight generation strategy network and sent to the concentrator. The multi-source heterogeneous data fusion result is obtained by performing weighted calculations based on the new data weight vector and the data source matrix.
[0006] Furthermore, the heterogeneous data includes real-time power grid operation data, equipment status monitoring data, environmental data, meteorological data, and log data; Heterogeneous data are matrixed and normalized separately to generate a data source matrix, including: Based on the heterogeneous data, determine the maximum and minimum values corresponding to each feature dimension of the heterogeneous data from each data source. Based on the preset interval distance and the maximum and minimum values corresponding to each feature dimension, perform equal-interval binning smoothing on the heterogeneous data of the corresponding feature dimension to obtain the first processed data. Based on the first processed data, normalization is performed according to the feature dimension to obtain the second processed data; Based on the second processed data, a two-dimensional standard data matrix is constructed using the data acquisition time sequence as the row index and the feature dimension as the column index. Missing values in the two-dimensional standard data matrix are filled with the time-series moving average of the same feature dimension to generate a data source matrix corresponding to each data source.
[0007] Further, according to the data source matrix, a corresponding proxy unit is assigned to each data source, including: Based on the data source matrix, extract the data source structure type features, time-series update features, and data quality features of heterogeneous data respectively; The corresponding agent unit is assigned based on the data source structure type characteristics, time-series update characteristics, and data quality characteristics.
[0008] Furthermore, each agent unit generates a local state vector based on the data source matrix and heterogeneous data, including: Based on the data source matrix and heterogeneous data, the number of missing values, the number of outliers, and the signal-to-noise ratio are extracted respectively. The missing rate is determined based on the ratio of the number of missing values to the total amount of heterogeneous data, and the outlier ratio is determined based on the ratio of the number of outliers to the total amount of heterogeneous data. The Pearson correlation coefficient between adjacent data collection sequences is calculated based on the row index of the data source matrix, and the Pearson correlation coefficient is used as the timestamp correlation coefficient. A quality assessment vector is constructed based on the missing rate, outlier ratio, timestamp correlation coefficient, and signal-to-noise ratio. The absolute difference, minimum difference, and maximum difference values of each heterogeneous data in the feature dimension are calculated based on the data source matrix. The gray correlation degree of the same feature dimension between different data sources is calculated based on the absolute difference, minimum difference, and maximum difference values. The associated feature vector is calculated based on the gray correlation degree. The quality assessment vector and the associated feature vector are concatenated after dimensional alignment to generate the local state vector of the current agent unit.
[0009] Furthermore, each agent unit generates a data weight vector based on the local state vector within a preset weight generation strategy network, including: In the preset weight generation strategy network, nonlinear feature extraction and high-dimensional mapping calculation are performed on the local state vector to obtain the original weight score values of each feature dimension of the data source. The original weight scores are subjected to exponential normalization to generate a data weight vector; The data weight vector satisfies the constraints that all weight elements are non-negative and the sum of all weight elements is 1.
[0010] Furthermore, based on the global state vector and the data source matrix, a fusion quality assessment is performed to obtain a global fusion quality score, including: The global state vector is decomposed according to the feature dimension, and the resulting local state vector and corresponding data weight vector are used to construct a basic dataset with the data source matrix. Using the data acquisition time sequence as the alignment benchmark, the data with the same feature dimension from each data source within the same time window is matched based on the basic dataset, and the data consistency evaluation value is calculated by combining the corresponding data weight vector. The mean of the original information entropy of heterogeneous data from each data source is calculated based on the basic dataset, and the information entropy of the fused data is calculated by weighting the current data weight vector onto the data source matrix. The difference between the mean of the original information entropy and the information entropy of the fused data is used as the information entropy gain evaluation value. Based on the absolute difference value of each heterogeneous data in the feature dimension, determine the number of data pairs that exceed the preset deviation threshold, and obtain the field conflict resolution degree based on the ratio of the number of data pairs to the total number of data pairs in the basic dataset. The global fusion quality score is obtained by weighting the data consistency assessment value, information entropy gain assessment value, and field conflict resolution degree.
[0011] Furthermore, the weight generation strategy network includes a strategy optimization network and a value evaluation network; Each agent unit optimizes the weight generation policy network by taking the global reward signal as the target, the local state vector as the input, and the data weight vector as the action, including: Using the global reward signal as the optimization objective, the local state vector as the network state input, and the data weight vector as the policy action output, a dual-network iterative optimization operation is performed until the preset loss function value converges, resulting in the optimized weight generation policy network. The dual-network iterative optimization operation includes: Based on the global reward signal, the Q-value estimate of the current state action pair is calculated in the current value evaluation network according to the current local state vector and the current data weight vector. The temporal difference error is calculated based on the Q-value estimate and the current immediate reward. Based on the aforementioned temporal difference error, the parameters of the first fully connected layer of the current value assessment network are updated using the gradient descent method. Using the temporal difference error as the strategy optimization guide, the gradient descent method is used to update the parameters of the second fully connected layer of the strategy optimization network, and the preset loss function value is calculated based on the parameters of the first and second fully connected layers. Determine whether the loss function value has converged. If it has, use the current policy optimization network and the current value evaluation network as the optimized weights to generate the policy network. Otherwise, update the current data weight vector based on the loss function value.
[0012] As an improvement to the above solution, another embodiment of the present invention provides a multi-source heterogeneous data fusion system based on multi-agent reinforcement learning, including a concentrator and several agent units; each agent unit corresponds to a data source; The concentrator includes a heterogeneous data acquisition module, a data source matrix generation module, a global fusion quality assessment module, and a multi-source data fusion module; The heterogeneous data acquisition module is used to acquire heterogeneous data from various data sources of the power grid to be processed. The data source matrix generation module is used to perform matrixing and normalization processing on heterogeneous data to generate a data source matrix, allocate corresponding proxy units to each data source according to the data source matrix, and send the data source matrix to each proxy unit. The proxy unit is used to generate a local state vector based on the data source matrix and heterogeneous data, generate a data weight vector based on the local state vector in the preset weight generation strategy network, and transmit the local state vector and data weight vector to the concentrator. The global fusion quality assessment module is used to perform a concatenation operation on the local state vector and the data weight vector to generate a global state vector, perform a fusion quality assessment on the global state vector and the data source matrix to obtain a global fusion quality score, and generate a global reward signal based on the global fusion quality score and send it to each agent unit. The agent unit is also used to repeatedly perform dual-network iterative optimization operations to optimize the weight generation strategy network with the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. Based on the optimized weight generation strategy network, the data weight vector is regenerated and the new data weight vector is sent to the concentrator. The multi-source data fusion module is used to perform weighted calculations based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result.
[0013] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning as described in the above embodiments.
[0014] Another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the multi-source heterogeneous data fusion method based on multi-agent reinforcement learning described in the above embodiment.
[0015] By implementing this invention, at least the following beneficial effects are achieved: This invention provides a method, system, terminal device, and storage medium for multi-source heterogeneous data fusion based on multi-agent reinforcement learning. The method employs a multi-agent reinforcement learning architecture, where agent units autonomously generate data weights based on the real-time state of the data source, without relying on human experience. A concentrator forms a closed-loop feedback loop through global fusion quality assessment and a global reward signal, driving the agent units to automatically optimize their weight generation strategies, thus solving the problem that manual static weighted fusion cannot dynamically optimize weights. The agent units generate local state vectors in real time based on the data source matrix and heterogeneous data, and are aware of changes in data source quality, temporal characteristics, and correlations. Using the global reward signal as the optimization target, the agent units continuously adjust their weight strategies, with weights dynamically adapting to power grid conditions and data quality in real time. It supports repeated execution of dual-network iterative optimization, with the weight generation strategy network continuously converging to the current optimal value, achieving continuously optimal weight allocation under dynamic conditions. Finally, weighted calculations are performed based on the new data weight vectors and the data source matrix to obtain the multi-source heterogeneous data fusion result. This achieves real-time data fusion and can dynamically adjust the data weight vectors to achieve adaptive weight optimization, meeting the high timeliness requirements of power systems for data fusion. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning, provided in an embodiment of the present invention. Figure 2 This is an Actor-Critic framework diagram provided in an embodiment of the present invention; Figure 3 This is the convergence curve of the global reward function provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a multi-source heterogeneous data fusion system based on multi-agent reinforcement learning, provided by an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] See Figure 1 To address the problem in existing technologies that cannot adaptively optimize weights based on dynamic changes in power grid operating conditions, thus hindering effective data fusion, this invention provides a flowchart illustrating a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning. The method is applicable to concentrators and includes: S1. Obtain heterogeneous data from various data sources of the power grid to be processed; Specifically, the concentrator refers to the centralized data processing and global coordination node deployed in the power grid dispatch center. It is the main implementer of this method, undertaking the core functions of multi-source heterogeneous data acquisition, preprocessing, agent unit allocation, global fusion quality assessment, reward signal distribution, and final data fusion. It is the global coordination hub connecting all distributed agent units. In Multi-Agent Reinforcement Learning (MARL), each agent unit completes learning decisions through interaction with the environment and collaborative information exchange among agents, ultimately achieving global goal optimization. In this invention, each agent unit corresponds to an independent intelligent agent, and distributed collaborative decision-making is carried out with the common goal of improving global fusion quality.
[0019] To illustrate, the concentrator connects to data sources such as the power grid SCADA / EMS system, equipment status monitoring system, meteorological monitoring platform, electricity consumption information collection system, and dispatch log management system through channels such as the power grid dispatch data bus, industrial Ethernet, and power wireless communication private network. It collects corresponding heterogeneous data in real time or in batches, completes the aggregation, verification and caching of raw data, and removes invalid and garbled data.
[0020] S2. Perform matrixing and normalization on the heterogeneous data to generate a data source matrix. Assign corresponding agent units to each data source according to the data source matrix and send the data source matrix to each agent unit so that each agent unit generates a local state vector according to the data source matrix and the heterogeneous data. Generate a data weight vector according to the local state vector in the preset weight generation strategy network and transmit the local state vector and the data weight vector to the concentrator. Specifically, the data source matrix refers to a two-dimensional standardized data matrix constructed after standardizing and preprocessing the original heterogeneous data. This matrix has rows representing the data acquisition time sequence and columns representing the feature dimensions. It serves as the basic data input for the agent unit's decision-making, eliminating the problems of dimensional differences, temporal misalignment, and missing values in the original data. The agent unit represents an edge intelligent decision-making unit independently allocated to each data source, possessing autonomous state awareness, weight decision-making, and policy optimization capabilities. It is a distributed intelligent agent in the multi-agent reinforcement learning framework, performing dedicated processing only on its bound data source, forming a collaborative architecture with the concentrator for distributed decision-making and centralized control. The local state vector represents a one-dimensional structured vector generated by the agent unit based on the basic data from its bound data source, used to characterize the quality of the data source itself and cross-data source relationships. It is the state input of the agent unit in the reinforcement learning framework and the sole input basis for the weight generation policy network. The weight generation policy network is a policy generation model based on a deep neural network deployed within each agent unit. It is the core decision-making module of the agent unit in the reinforcement learning framework, with the local state vector as input and the fused weight allocation result of the corresponding data source as output. The data weight vector is the fusion weight allocation result of each feature dimension of the corresponding data source, output by the weight generation strategy network. It is the action executed by the agent unit in the reinforcement learning framework, which determines the contribution ratio of the corresponding data source and each feature dimension in the final fusion result.
[0021] Indicatively, the concentrator performs standardized preprocessing on the collected raw heterogeneous data according to the data source classification. Through matrixing and normalization, the non-standardized raw data is transformed into a standardized data source matrix that can be processed by the agent unit. Subsequently, based on the structure, time series, and quality characteristics of each data source matrix, a dedicated independent agent unit is assigned to each data source, establishing a one-to-one binding relationship between the agent unit and the data source and data source matrix. Finally, the data source matrix corresponding to each data source is sent to the bound agent unit through an encrypted dedicated communication channel.
[0022] Schematic, after receiving the data source matrix from the concentrator, each agent unit combines the corresponding original heterogeneous data to extract the quality features of the data source and the cross-data source correlation features, and constructs a local state vector. The local state vector is used as the only input to the preset weight generation strategy network. After forward calculation through the network input layer, hidden layer and output layer, a data weight vector is generated. Finally, the local state vector generated by itself and the corresponding data weight vector, along with the agent unit's unique identifier, are encrypted and uploaded to the concentrator.
[0023] Preferably, the heterogeneous data includes real-time power grid operation data, equipment status monitoring data, environmental data, meteorological data, and log data; Heterogeneous data are matrixed and normalized separately to generate a data source matrix, including: Based on the heterogeneous data, determine the maximum and minimum values corresponding to each feature dimension of the heterogeneous data from each data source. Based on the preset interval distance and the maximum and minimum values corresponding to each feature dimension, perform equal-interval binning smoothing on the heterogeneous data of the corresponding feature dimension to obtain the first processed data. Based on the first processed data, normalization is performed according to the feature dimension to obtain the second processed data; Based on the second processed data, a two-dimensional standard data matrix is constructed using the data acquisition time sequence as the row index and the feature dimension as the column index. Missing values in the two-dimensional standard data matrix are filled with the time-series moving average of the same feature dimension to generate a data source matrix corresponding to each data source.
[0024] Specifically, equidistant binning smoothing refers to dividing the value range of a data feature dimension into several equally wide intervals according to a preset interval distance, and replacing all data values within the same interval with the interval mean, thereby simplifying data distribution and smoothing high-frequency fluctuations and abnormal noise. Time-series moving average filling refers to filling missing values in a certain feature dimension of the data source matrix by taking the mean of the valid data in adjacent time series within a preset window, centered on the missing location, to ensure the continuity and integrity of the time-series data and avoid interference from missing values in subsequent decisions. Real-time power grid operation data refers to structured time-series data characterizing the power grid's operating status, collected in real time by the power grid SCADA / EMS system, including core electrical quantities such as voltage, current, active power, reactive power, frequency, and phase angle. Equipment status monitoring data refers to data characterizing the operating status of equipment collected by the power grid primary equipment online monitoring system, including data such as transformer oil temperature, winding temperature, partial discharge, circuit breaker mechanical characteristics, and GIS equipment gas density. Environmental and meteorological data include ambient temperature, humidity, wind speed, wind direction, rainfall, light intensity, and icing thickness. Log data refers to semi-structured or unstructured data generated by the power grid dispatching system and automated equipment, including dispatching operation logs, equipment alarm logs, maintenance records, and abnormal event records.
[0025] Schematic illustration: For heterogeneous data from each data source, the data is split according to feature dimensions, and the maximum and minimum values of all data under each feature dimension are calculated. Based on a preset interval distance, the range of maximum and minimum values is divided into several equally wide bin intervals. All data values within each bin interval are replaced with the arithmetic mean of that interval, completing the data smoothing and denoising, resulting in the first processed data. The preset interval distance can be customized according to the data fluctuation characteristics, with 5 intervals being the preferred number.
[0026] For the first batch of processed data that has undergone smoothing, z-score standardization or min-max normalization is performed according to the feature dimensions to eliminate the differences in dimensionality and numerical scale between different feature dimensions and different data sources, mapping all data to a uniform scale interval to obtain the second batch of processed data. Among them, min-max normalization maps the data to the [0,1] interval, which is suitable for the subsequent matrix construction and weight calculation requirements.
[0027] For the second processed data that has completed normalization, the data is arranged in chronological order using the timestamp of data collection as the row index and the feature dimension of the data as the column index, and sorted by feature type. The corresponding time series and corresponding feature values are filled into the corresponding positions in the matrix to construct a two-dimensional standard data matrix, ensuring that each row of the matrix corresponds to a collection time and each column corresponds to a feature dimension.
[0028] For the constructed two-dimensional standard data matrix, all matrix elements are traversed to identify the locations of missing values. For each missing value, the effective data of adjacent time series within the feature dimension containing the missing value and a preset sliding window are taken, and their arithmetic mean is calculated. This mean is then used to fill the missing value. After filling all missing values, a standardized data source matrix corresponding to each data source is finally generated. The width of the sliding window can be customized according to the data sampling frequency, preferably 3-5 time series steps.
[0029] To illustrate, discretization and normalization can be performed concurrently. When the original data is continuous numerical and has dimensional differences, normalization is preferred. When the data fluctuates significantly, has high noise levels, and requires reducing data complexity, discretization with equal-interval binning and smoothing is preferred. Values exceeding ±3 standard deviations of the dataset mean are considered outliers.
[0030] In a preferred embodiment of the present invention, discretization can simplify data distribution and smooth data fluctuations, as shown in the following formula: Among them, heterogeneous data is , For the discretized data, The maximum value corresponding to each feature dimension. To be the minimum value, The preset interval distance is 5. Two-dimensional standard data matrix. : The first m columns represent the feature dimensions, and the last column represents the timestamp. (i=1,2,...,n), where n represents the total number of data acquisition times. (i=1,2,...,n) represents Data collected in real time Indicates the number of feature dimensions. In the matrix, M and L are simplified representations of matrices. Normalization is used to eliminate dimensional differences between the source data; the formula is as follows: ,in, For the second processing of data, and These are the mean and standard deviation of the first processed data, respectively. This is the first piece of data to be processed.
[0031] In a preferred embodiment of the present invention, based on the above-mentioned urban power grid dispatching scenario, the preprocessing process of 10kV line load data is as follows: The concentrator collects hourly load data of a certain 10kV line, extracts the maximum value of the load feature dimension as 1200kW and the minimum value as 200kW, and sets the preset interval distance as 200kW, dividing the value range into 5 equal-width sub-bin intervals; the load data in each interval is replaced with the interval mean, completing the equal-distance sub-bin smoothing process, eliminating sudden abnormal fluctuations in load data, and obtaining the first processed data. The smoothed load data is subjected to min-max normalization processing, linearly mapping the load data to the [0,1] interval, eliminating the dimensional differences with other feature dimensions, and obtaining the second processed data. Using the collection timestamp as the row index and the load, current, voltage, and other features as column indices, the normalized second processed data is filled into the corresponding positions to construct a two-dimensional standard data matrix of the line load data source. The identification matrix contains 3 missing load data points. The sliding window width is set to 3. The average effective load data of the three time series before and after each missing point is taken to fill the missing values. Finally, a standardized data source matrix of the load data source of this line is generated and sent to the bound load data agent unit.
[0032] By implementing this embodiment, high-frequency noise and abnormal fluctuations in the data are effectively eliminated through equidistant binning and smoothing; dimensional differences are eliminated through normalization; and data missing issues are resolved through time-series moving average filling, ensuring the standardization, integrity, and reliability of the data source matrix. Constructing a matrix using the data acquisition time sequence as the row index fundamentally solves the problem of time-series asynchrony in multi-source data, laying the foundation for subsequent cross-data source correlation analysis and time-series alignment and fusion. The preprocessing workflow is adaptable to different types of heterogeneous power data, including structured, semi-structured, and unstructured data, and can flexibly select processing methods according to data characteristics, covering the preprocessing needs of heterogeneous data across all scenarios in the power system.
[0033] Preferably, assigning corresponding proxy units to each data source according to the data source matrix includes: Based on the data source matrix, extract the data source structure type features, time-series update features, and data quality features of heterogeneous data respectively; The corresponding agent unit is assigned based on the data source structure type characteristics, time-series update characteristics, and data quality characteristics.
[0034] Specifically, the data source structure type feature is a characteristic used to characterize the data structure attributes of the data source. Based on the dimensionality regularity and data sparsity of the data source matrix, it is divided into three categories: structured, semi-structured, and unstructured, and is one of the core bases for agent unit allocation. The time-series update feature is a characteristic used to characterize the data update frequency and update mode of the data source. Based on the collection time interval corresponding to the row index of the data source matrix, it is divided into three categories: millisecond / second-level real-time streaming updates; minute / hour-level fixed batch updates; and non-fixed-frequency triggered updates, which determines the real-time processing requirements of the agent unit. The data quality feature is a characteristic used to characterize the data reliability and stability of the data source. Calculated based on the data source matrix, it includes four core quantitative dimensions: data missing rate, outlier ratio, temporal continuity coefficient, and signal-to-noise ratio, which determine the quality perception and weight decision logic of the agent unit.
[0035] To illustrate, a real-time data agent is assigned to the structured real-time streaming data of power grid operation monitoring; a batch data agent is assigned to the semi-structured batch update data of meteorology; and an unstructured data agent is assigned to the unstructured data of equipment status with a high missing rate. At the same time, it is specified that the agent unit is an intelligent decision-making unit with autonomous processing and collaboration capabilities, which can be deployed on edge terminals or dispatch center cloud platforms.
[0036] In a preferred embodiment of the present invention, the concentrator extracts features from the data source matrix of five types of data sources: Real-time power grid operation data: structured type, second-level real-time stream updates, data missing rate less than 0.1%, high signal-to-noise ratio; Meteorological data: semi-structured type, hourly batch updates, data missing rate 1%, some noise present; Equipment status monitoring data: unstructured type, 15-minute updates, data missing rate 3%, relatively high proportion of outliers; Log data: semi-structured type, non-fixed frequency triggered updates, high data sparsity. The concentrator assigns a high-real-time structured data agent to the real-time power grid operation data, a semi-structured batch processing agent to the meteorological data, an unstructured data noise reduction agent to the equipment status monitoring data, and a non-fixed frequency event agent to the log data.
[0037] By implementing this embodiment, adaptable proxy units are allocated based on the core characteristics of the data source, enabling the design of differentiated decision logic for data sources with different heterogeneous characteristics, significantly improving the processing capability for highly heterogeneous data; a scalable distributed proxy architecture is formed, which can automatically extract features and allocate corresponding proxy units when a new data source is added, without modifying the overall architecture, thus adapting to the ever-expanding needs of power system data sources; by classifying and allocating proxy units, parallel processing of multiple data sources is achieved, greatly improving the overall efficiency of data processing and adapting to the needs of real-time fusion.
[0038] Preferably, each agent unit generates a local state vector based on the data source matrix and heterogeneous data, including: Based on the data source matrix and heterogeneous data, the number of missing values, the number of outliers, and the signal-to-noise ratio are extracted respectively. The missing rate is determined based on the ratio of the number of missing values to the total amount of heterogeneous data, and the outlier ratio is determined based on the ratio of the number of outliers to the total amount of heterogeneous data. The Pearson correlation coefficient between adjacent data collection sequences is calculated based on the row index of the data source matrix, and the Pearson correlation coefficient is used as the timestamp correlation coefficient. A quality assessment vector is constructed based on the missing rate, outlier ratio, timestamp correlation coefficient, and signal-to-noise ratio. The absolute difference, minimum difference, and maximum difference values of each heterogeneous data in the feature dimension are calculated based on the data source matrix. The gray correlation degree of the same feature dimension between different data sources is calculated based on the absolute difference, minimum difference, and maximum difference values. The associated feature vector is calculated based on the gray correlation degree. The quality assessment vector and the associated feature vector are concatenated after dimensional alignment to generate the local state vector of the current agent unit.
[0039] Specifically, the missing value rate (RR) is an indicator of data integrity in a data source, equal to the ratio of the number of missing values to the total number of heterogeneous data samples. A higher ratio indicates poorer data integrity. The outlier ratio (ORR) is an indicator of data validity in a data source, equal to the ratio of the number of outliers exceeding a reasonable threshold to the total number of heterogeneous data samples. A higher ratio indicates lower data reliability. The timestamp correlation coefficient (SRC) is an indicator of the temporal stability of a data source, equal to the Pearson correlation coefficient between adjacent timestamp data sequences in the data source matrix. A coefficient closer to 1 indicates better temporal continuity. The signal-to-noise ratio (SNR) is an indicator of data clarity in a data source, equal to the ratio of effective signal power to noise power. A higher SNR indicates better data quality. The quality assessment vector is a one-dimensional vector composed of four dimensions: missing value rate, outlier ratio, timestamp correlation coefficient, and SNR. It quantifies the reliability and stability of the corresponding data source and is a core component of the local state vector. The absolute difference value (ADV) refers to the absolute value of the numerical difference between two different data sources on the same feature dimension. It characterizes the degree of deviation between the two data sources on that feature dimension. A smaller absolute difference value indicates a stronger correlation between the two data sources. Grey relational degree is used to calculate the degree of association between different data sources with the same feature dimension. The closer the relational degree is to 1, the stronger the potential mapping relationship between the two data sources. The association feature vector is a one-dimensional vector composed of the grey relational degrees between the current agent's bound data source and all other data sources. It is used to characterize the potential mapping relationship between the current data source and other heterogeneous data sources and is a core component of the local state vector.
[0040] Indicatively, each agent unit, based on the data source matrix bound to the data source and the original heterogeneous data, counts the number of missing values and outliers, and calculates the missing value rate and outlier ratio; based on the row index of the data source matrix, it extracts the data sequence of adjacent acquisition times and calculates the Pearson correlation coefficient as the timestamp correlation coefficient; it calculates the ratio of the effective signal power to the noise power of the data source to obtain the signal-to-noise ratio; and arranges the above four indicators in a fixed order to construct a quality assessment vector.
[0041] Schematic, each agent unit obtains the feature dimension benchmark statistics of the full data source from the concentrator, calculates the absolute difference value of the current data source and all other data sources in the same feature dimension based on its own bound data source matrix, the minimum difference value and the maximum difference value of the full data source in the same feature dimension, and then calculates the gray correlation degree of the current data source and other data sources in each feature dimension; arranges all gray correlation degrees in a fixed order to construct the correlation feature vector.
[0042] Schematic, each agent unit aligns the quality evaluation vector with the associated feature vector in terms of dimensions, ensuring that both vectors are either one-dimensional column vectors or one-dimensional row vectors. Then, the two vectors are concatenated end to end to generate a one-dimensional structured local state vector that is bound to the current agent unit, which serves as the input to the weight generation strategy network.
[0043] In a preferred embodiment of the present invention, the quality assessment vector quantifies the reliability and stability of the data source, and its formula is as follows: ,in, represents the quality assessment vector, MR represents the missing rate, OR represents the outlier ratio, TC represents the timestamp correlation coefficient, and SNR represents the signal-to-noise ratio.
[0044] The formula for the missing rate is as follows: ,in, This represents the number of missing values. Total amount of heterogeneous data; The formula for outlier ratio is as follows: ,in, The number of outliers is determined by whether a data sample value exceeds the range of ±3 times the standard deviation of the dataset mean, or exceeds the reasonable physical range of historical data. The formula for the timestamp correlation coefficient is as follows: ,in, The correlation coefficient between adjacent timestamps reflects the temporal stability of data acquisition and can be calculated using the Pearson correlation coefficient between data sequences at adjacent times.
[0045] In one embodiment, taking power grid operation monitoring data as an example, the system collects node load data at 5-minute intervals. Assuming the continuously collected data sequence is [102,105,107,110,112], then adjacent sequences [102,105,107,110] and [105,107,110,112] are constructed. The Pearson correlation coefficient between the two sequences is calculated to be 0.987, indicating that the time continuity of the data source is very high, demonstrating good temporal stability during data acquisition.
[0046] The signal-to-noise ratio formula is as follows: ,in, For signal power, This represents noise power.
[0047] Related feature vectors The underlying mapping relationships between heterogeneous data sources are revealed, and the element calculation formula is as follows: ,in, It is the absolute difference between data sources i and j on the k0th feature dimension. To minimize the difference on the k0th feature dimension, To find the maximum difference in the k0th feature dimension, The resolution coefficient is set to 0.5.
[0048] Specifically, the features of data sources i and j refer to indicators that can characterize the data content, operating status, and statistical attributes of each data source. Different types of data sources can select different features. In one embodiment, taking the power grid dispatch scenario as an example, for power grid operation monitoring data, the features include voltage, current, active power, reactive power, frequency, etc. For meteorological data, the features may include ambient temperature, humidity, wind speed, light intensity, etc. For equipment status monitoring data, the features may include equipment temperature, vibration amplitude, running time, alarm frequency, etc. In order to facilitate correlation analysis between different data sources, this invention selects features that are comparable or can reflect the operating status from each data source to form a feature vector, and calculates the correlation feature vector accordingly.
[0049] Absolute difference represents the degree of numerical deviation between two data sources on a certain feature dimension, i.e., their similarity on the corresponding feature. The smaller the absolute difference, the closer the performance of the two data sources on that feature, and the stronger their correlation. In one embodiment, taking the correlation analysis between power grid operation monitoring data and meteorological data as an example, ambient temperature is selected as the k0th feature dimension. Assuming that at a certain moment, the temperature feature value of the area associated with the power grid operation monitoring system is 26℃, and the temperature feature value corresponding to the meteorological data source is 28℃, then the absolute difference between the two data sources on the feature of ambient temperature is: The small difference indicates that the two data sources have a high degree of similarity in this feature dimension, and can be considered to have a strong potential correlation.
[0050] By implementing this embodiment, the quality attributes of the data source itself and the cross-data source correlation are covered. It takes into account both the reliability of the data source itself and the synergy between multiple data sources, providing a comprehensive decision basis for weight allocation. Through mature quantitative methods such as missing rate, outlier ratio, Pearson correlation coefficient, and grey relational degree, the data quality and correlation are accurately quantified, avoiding the subjectivity of human experience judgment.
[0051] Preferably, each agent unit generates a data weight vector based on the local state vector in a preset weight generation strategy network, including: In the preset weight generation strategy network, nonlinear feature extraction and high-dimensional mapping calculation are performed on the local state vector to obtain the original weight score values of each feature dimension of the data source. The original weight scores are subjected to exponential normalization to generate a data weight vector; The data weight vector satisfies the constraints that all weight elements are non-negative and the sum of all weight elements is 1.
[0052] Specifically, nonlinear feature extraction refers to the process of using the hidden layers of a deep neural network to perform a nonlinear transformation on the input local state vector, uncovering the potential correlations between features in different dimensions of the vector, and extracting deep features for weight decision-making. High-dimensional mapping computation refers to the process of using the fully connected layers of a neural network to map the extracted deep features to a high-dimensional space with the same number of feature dimensions as the data source, outputting the original weight score value corresponding to each feature dimension. The original weight score value is the original weight score value output by the neural network for each feature dimension of the data source, reflecting the relative importance of that feature dimension in the fusion process. Exponential normalization, also known as softmax normalization, refers to transforming the original weight score values using an exponential function, mapping all score values to the (0,1) interval, while ensuring that the sum of all weight elements is 1. This is the core processing step for generating compliant weight vectors.
[0053] Schematic, each agent unit uses the generated local state vector as input to the pre-defined weight generation strategy network, passing it into the network's input layer. The weight generation strategy network extracts non-linear features from the input local state vector through multiple fully connected hidden layers, uncovering the potential mapping relationship between data quality, correlation, and feature importance. Then, it performs high-dimensional mapping calculations through fully connected neurons in the output layer, outputting raw weight scores with the same number of feature dimensions as the data source, where each score corresponds to the relative importance of a feature dimension. The output raw weight scores are then subjected to softmax exponential normalization, mapping all scores to non-negative values through an exponential function, and normalization ensures that the sum of all weight elements is 1, ultimately generating a data weight vector. The generated weight vector strictly satisfies two constraints: first, all weight elements are greater than or equal to 0; second, the cumulative sum of all weight elements is 1.
[0054] In a preferred embodiment of the present invention, each agent unit generates a data weight vector based on a local state vector. The formula is: ,in, Let be the local state vector of the agent unit. The weight generation policy network for agent units, This is an exponentially normalized function, which ensures that the sum of the weights is 1. In one embodiment, let the local state vector of the agent unit be... The first four terms represent the missing rate, outlier ratio, timestamp correlation coefficient, and signal-to-noise ratio, respectively. The last two terms can represent the association features between this data source and other data sources. iAfter inputting the weights into the policy network, since the output layer contains three output neurons, each corresponding to the original score value of a weight component, the resulting three-dimensional original score vector is: The data weight vector is obtained after softmax normalization. .
[0055] By implementing this embodiment, the mapping relationship between local state vectors and feature importance is learned autonomously. No manual pre-setting of weight rules is required. The weight allocation can be dynamically adjusted based on changes in data source quality and correlation, exhibiting far superior adaptability compared to static weighting methods. Through exponential normalization, the weight vector strictly satisfies non-negativity and normalization constraints, allowing it to be directly used for weighted fusion calculations and avoiding distortion of fusion results caused by negative weights or abnormal weight sums. Through nonlinear feature extraction, deep feature correlations can be uncovered, accurately distinguishing the importance of different feature dimensions. Higher weights are assigned to high-quality, highly correlated features, while lower weights are assigned to low-quality, redundant features, thus enhancing key features and suppressing redundant information.
[0056] S3. Perform a concatenation operation on the local state vector and the data weight vector to generate a global state vector. Perform a fusion quality assessment on the global state vector and the data source matrix to obtain a global fusion quality score. Generate a global reward signal based on the global fusion quality score and send it to each agent unit. This enables each agent unit to repeatedly perform a dual-network iterative optimization operation to optimize the weight generation strategy network, using the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. Regenerate the data weight vector based on the optimized weight generation strategy network and send the new data weight vector to the concentrator. Specifically, the global state vector is the complete global state information obtained by concatenating the local state vectors uploaded by all agent units with the corresponding data weight vectors. It is the sole input basis for the concentrator to perform global fusion quality evaluation. The global fusion quality score is a comprehensive score obtained by the concentrator based on the global state vector, which quantitatively evaluates the fusion effect under the current weight allocation from multiple dimensions. It is the core quantitative indicator for measuring the data fusion effect. The global reward signal is a reinforcement learning feedback signal generated by the concentrator based on the global fusion quality score and synchronously fed back to all agent units. It is the sole optimization objective of the agent unit to optimize the weight generation strategy network and is the output of the global reward function. The better the fusion effect, the higher the positive incentive value of the global reward signal. The dual-network iterative optimization operation represents the operation of alternating iterative updates of the strategy optimization network and the value evaluation network contained in the weight generation strategy network based on the Actor-Critic reinforcement learning framework. It is the core link to achieve adaptive optimization of the weight generation strategy and convergence to the global optimum.
[0057] Schematic, after receiving the local state vectors and data weight vectors uploaded by all agent units, the concentrator aligns the data according to the unique identifier of the agent unit, concatenates all vectors end to end in a preset dimensional order to generate a one-dimensional global state vector; combining the global state vector with the data source matrix of each data source, it quantitatively evaluates the fusion effect under the current weight allocation from multiple dimensions and calculates the global fusion quality score; based on the preset global reward function, it maps the global fusion quality score into a global reward signal, which is synchronously distributed to all bound agent units through the broadcast channel.
[0058] Schematic, after each agent unit receives the global reward signal, it takes maximizing the global reward signal as the optimization objective, uses its own local state vector as the state input for reinforcement learning, and uses the generated data weight vector as the action to perform repeated dual-network iterative optimization operations, alternately updating the network parameters of the weight generation policy network until the network meets the preset convergence condition; after the network converges, it regenerates the optimized data weight vector based on the optimized weight generation policy network, and uploads it to the concentrator with a convergence flag.
[0059] In a preferred embodiment of the present invention, the input layer of the concentrator receives the local state vectors of all agent units. With data weight vector Concatenate them into a global state vector : ,in, This indicates a vector concatenation operation.
[0060] Preferably, a global fusion quality score is obtained by evaluating the fusion quality based on the global state vector and the data source matrix, including: The global state vector is decomposed according to the feature dimension, and the resulting local state vector and corresponding data weight vector are used to construct a basic dataset with the data source matrix. Using the data acquisition time sequence as the alignment benchmark, the data with the same feature dimension from each data source within the same time window is matched based on the basic dataset, and the data consistency evaluation value is calculated by combining the corresponding data weight vector. The mean of the original information entropy of heterogeneous data from each data source is calculated based on the basic dataset, and the information entropy of the fused data is calculated by weighting the current data weight vector onto the data source matrix. The difference between the mean of the original information entropy and the information entropy of the fused data is used as the information entropy gain evaluation value. Based on the absolute difference value of each heterogeneous data in the feature dimension, determine the number of data pairs that exceed the preset deviation threshold, and obtain the field conflict resolution degree based on the ratio of the number of data pairs to the total number of data pairs in the basic dataset. The global fusion quality score is obtained by weighting the data consistency assessment value, information entropy gain assessment value, and field conflict resolution degree.
[0061] Specifically, the base dataset is the full dataset used for fusion quality assessment, constructed by the concentrator based on the local state vectors decomposed from the global state vector, the data weight vectors, and the data source matrices of each data source. It forms the foundation for fusion quality assessment. The data consistency evaluation value represents an indicator used to quantify the degree of data consistency across different data sources in the same time series and feature dimensions. Its value ranges from [0,1]. The closer the value is to 1, the stronger the consistency of the multi-source data and the better the fusion effect. Information entropy is an indicator used to quantify data uncertainty. The higher the information entropy, the stronger the data uncertainty and the lower the effective information content; the lower the information entropy, the weaker the data uncertainty and the higher the effective information content. The information entropy gain evaluation value is an indicator used to quantify the degree of reduction in information uncertainty after data fusion. It is equal to the difference between the average original information entropy of each data source before fusion and the information entropy of the fused data. The higher the value, the higher the retention of effective information after fusion and the better the fusion effect. Field conflict resolution is an indicator used to quantify the degree of semantic conflict elimination among multi-source data. The value range is [0,1]. The closer the value is to 1, the fewer semantic conflicts there are among the multi-source data and the better the conflict resolution effect.
[0062] Schematic, the concentrator decomposes the spliced global state vector according to the unique identifier of the agent unit, extracts the local state vector and data weight vector corresponding to each agent unit, and associates and binds them with the standardized data source matrix of the corresponding data source to construct a basic dataset for fusion quality assessment.
[0063] Indicatively, using the data acquisition time sequence as the alignment benchmark, the same feature dimension data from each data source within the same time window are matched in the basic dataset. Combined with the corresponding data weight vector, the weighted variation coefficient of the multi-source data under that feature dimension is calculated. The average value of the weighted variation coefficients of all feature dimensions is taken, and normalization is performed based on the reciprocal of the average value to obtain the data consistency evaluation value, which ranges from [0,1].
[0064] Based on the basic dataset, the original information entropy of heterogeneous data from each data source is calculated, and the average of the information entropy of all data sources is taken to obtain the mean of the original information entropy. The matrix of each data source is weighted based on the current data weight vector to obtain the fused data sequence, and the information entropy of the fused data is calculated. The difference between the mean of the original information entropy and the information entropy of the fused data is normalized to obtain the information entropy gain evaluation value, which takes a value range of [0,1].
[0065] Based on the absolute difference values of each data source in the basic dataset on the same feature dimension, a reasonable deviation threshold is preset, and the number of conflicting data pairs whose absolute difference values exceed the threshold and the total number of data pairs under the feature dimension are counted. The field conflict resolution degree is obtained by subtracting the ratio of the number of conflicting data pairs to the total number of data pairs from 1 and averaging the calculation results of all feature dimensions. The value range is [0,1].
[0066] Indicatively, the global fusion quality score is obtained by weighting and summing the three dimensions of data consistency assessment value, information entropy gain assessment value, and field conflict resolution degree through a preset weighting formula. The weights corresponding to the three dimensions are adjustable hyperparameters that can be flexibly adjusted according to fusion needs, with preferred values of 0.4, 0.35, and 0.25.
[0067] In a preferred embodiment of the present invention, the output layer of the concentrator outputs a structured evaluation result, covering three dimensions: data consistency, information entropy gain, and conflict resolution. Each dimension independently quantifies a specific attribute of the data fusion result. The global reward function is used to quantify the quality of data fusion, and its calculation is based on the evaluation results of the above three dimensions, as shown in the following formula: ,in, The global reward signal generated for the global fusion quality score. These are adjustable hyperparameters, with values of: , , , This is the data consistency assessment value. This is the evaluation value of the information entropy gain after fusion. Determines the degree of field conflict resolution.
[0068] The data consistency assessment value is used to measure the degree of data consistency between different data sources in the same field or within the same time window. The formula for calculating the assessment value is as follows: Let the observed values of the k1th field in the m data sources be respectively The mean is , To prevent tiny constants with a denominator of zero, the closer the data from different data sources are to each other, the closer the consistency is to 1. The information entropy gain evaluation value is used to measure the degree of reduction in information uncertainty after data fusion. The formula for calculating the evaluation value is as follows: ,in, The average information entropy of each data source before fusion, , This represents the probability of a certain state occurring. The calculation method for the entropy of the merged data from various data sources is the same as... Similarly, when the uncertainty of the fused data decreases, Entropy_Gain increases, indicating an improvement in fusion quality; Field conflict resolution measures the degree to which field conflicts are eliminated during the fusion of multi-source data. The calculation formula for its evaluation result is as follows: ,in, This refers to the number of conflicts of the same field across multiple data sources. The number of data pairs in the base dataset is considered. The fewer the conflicts between data pairs, the smaller the Conflict_Score, indicating that the fusion result is more reliable.
[0069] By implementing this embodiment, the evaluation is carried out from three core dimensions: data consistency, information retention, and semantic conflict resolution. This comprehensively covers the core objectives of multi-source data fusion, avoids the one-sidedness of single-dimensional evaluation, and ensures that the evaluation results can truly reflect the fusion effect. The evaluation indicators of the three dimensions are all positively correlated with the fusion effect. The higher the global fusion quality score, the better the fusion effect, providing a clear and specific goal orientation for the strategy optimization of the agent unit.
[0070] Preferably, the weight generation strategy network includes a strategy optimization network and a value evaluation network; Each agent unit optimizes the weight generation policy network by taking the global reward signal as the target, the local state vector as the input, and the data weight vector as the action, including: Using the global reward signal as the optimization objective, the local state vector as the network state input, and the data weight vector as the policy action output, a dual-network iterative optimization operation is performed until the preset loss function value converges, resulting in the optimized weight generation policy network. The dual-network iterative optimization operation includes: Based on the global reward signal, the Q-value estimate of the current state action pair is calculated in the current value evaluation network according to the current local state vector and the current data weight vector. The temporal difference error is calculated based on the Q-value estimate and the current immediate reward. Based on the aforementioned temporal difference error, the parameters of the first fully connected layer of the current value assessment network are updated using the gradient descent method. Using the temporal difference error as the strategy optimization guide, the gradient descent method is used to update the parameters of the second fully connected layer of the strategy optimization network, and the preset loss function value is calculated based on the parameters of the first and second fully connected layers. Determine whether the loss function value has converged. If it has, use the current policy optimization network and the current value evaluation network as the optimized weights to generate the policy network. Otherwise, update the current data weight vector based on the loss function value.
[0071] Specifically, the policy optimization network, or Actor network, is the core execution module of the weight generation policy network. It takes a local state vector as input and outputs a data weight vector (action), responsible for generating the weight allocation policy. The value evaluation network, or Critic network, is the core evaluation module of the weight generation policy network. It is responsible for evaluating the value of the actions (data weight vectors) output by the policy optimization network, calculating the action value Q-value, and guiding the parameter updates of the policy optimization network. A state-action pair refers to the combination of the agent unit's current state (local state vector) and the action (data weight vector) performed based on that state. It is the input of the value evaluation network, used to evaluate the merits of performing the action in that state. Q-value estimation, or action value estimation, refers to the expected estimated value of the future cumulative reward for the current state-action pair output by the value evaluation network. A higher Q-value indicates a higher long-term reward and a better action in that state. Temporal difference error (TD error) is a core metric used in reinforcement learning to update network parameters. It equals the current immediate reward plus the maximum Q-value estimate for the next state, minus the Q-value estimate of the current state-action pair. It reflects the estimation bias of the current action value; a smaller bias indicates a more accurate value evaluation. The parameters of the first fully connected layer refer to the weights and biases of all fully connected layers in the value evaluation network. These are the core optimization objects of the value evaluation network and are updated based on temporal difference errors. The parameters of the second fully connected layer refer to the weights and biases of all fully connected layers in the policy optimization network. These are also the core optimization objects of the policy optimization network and are updated based on temporal difference errors. The loss function quantifies the deviation between the network output and the target value. It is the core basis for network optimization; the smaller the loss function value, the closer the network output is to the optimal target, and the more convergent the network.
[0072] Schematic, each agent unit takes the received global reward signal as the sole optimization objective, its own local state vector as the state input for reinforcement learning, and the generated data weight vector as the action to be executed. It performs a dual-network iterative optimization operation based on the Actor-Critic framework. In each iteration, the parameters of the value evaluation network are updated first, then the parameters of the policy optimization network are updated, the loss function value is calculated, and it is determined whether convergence has occurred. If convergence has not occurred, the next iteration is entered until the loss function value meets the preset convergence condition, and finally the optimized weight generation policy network is obtained.
[0073] Each agent unit inputs the state-action pair, consisting of the current local state vector and the current data weight vector, into the current value evaluation network to calculate the Q-value estimate of the current state-action pair; using the received global reward signal as the current instantaneous reward, and combining the maximum Q-value estimate of the next state with the preset discount factor, the temporal difference error is calculated.
[0074] In a schematic manner, taking the calculated temporal difference error as the core, the gradient descent method is used to calculate the gradient of the loss function with respect to the parameters of the first fully connected layer of the value assessment network. The parameters of the first fully connected layer are updated along the gradient descent direction to complete the iterative calibration of the value assessment network, thereby gradually improving the accuracy of value assessment.
[0075] Using temporal difference error as the policy optimization guide, the policy gradient algorithm is adopted to calculate the gradient of the loss function with respect to the parameters of the second fully connected layer of the policy optimization network. The parameters of the second fully connected layer are updated along the gradient ascent direction (maximizing future rewards), so that the actions output by the policy optimization network can obtain higher Q values and global rewards, thus completing the iterative update of the policy optimization network.
[0076] Based on the updated parameters of the first and second fully connected layers, a preset loss function value is calculated. The loss of the value evaluation network is the mean squared error of the temporal difference error, and the loss of the policy optimization network is the policy gradient loss. The convergence condition is then determined: the change in the loss function value over multiple consecutive iterations is less than a preset convergence threshold. If the convergence condition is met, the iteration stops, and the current policy optimization network and value evaluation network are used as the optimized weights to generate the policy network. If the convergence condition is not met, the data weight vector is regenerated based on the updated policy optimization network, and the next iteration begins.
[0077] In a preferred embodiment of the present invention, each agent unit, based on feedback, alternately optimizes the policy optimization network and the value evaluation network in the weight generation policy network using the Actor-Critic framework. Figure 2 The diagram shows the Actor-Critic framework. This framework balances exploration and exploitation by decoupling policy optimization (Actor) and value evaluation (Critic), significantly improving learning efficiency and stability. The policy optimization (Actor) network is responsible for generating action policies. It outputs the probability distribution of actions based on the current state s, and then selects the action. The formula is as follows: ,in, To optimize the parameters of the fully connected and hidden layers in the policy optimization network, their values gradually approach the optimal value as the training process progresses. State s specifically represents the state input information used by the agent to execute decisions at the current moment, which is the local state vector constructed earlier. Action... This refers to the data weight allocation result output by the agent based on the current state, which is the data weight vector of the corresponding data source in the fusion process.
[0078] The Critic network is responsible for evaluating the value of actions and guiding the action policy updates of the Actor network. The action value function formula is as follows: ,in, To evaluate the parameters of fully connected layers and hidden layers in a value assessment network, In strategy The expectation of future cumulative returns. Let be the discount factor for the reward at step t, with a value of 0.9. The instantaneous reward at time t.
[0079] Global reward signal R global The global reward signal not only serves as a feedback signal to each agent unit regarding the fusion quality, but also directly participates in the policy optimization and value evaluation processes of each agent unit. Specifically, after each agent unit completes its action output, the concentrator calculates the global reward based on the fusion result and uses this reward as the learning objective of the value evaluation network to assess the merits of the current state-action combination. Simultaneously, the policy optimization network iteratively updates the policy parameters based on the value evaluation results and the global reward feedback, enabling each agent unit to gradually learn weight allocation strategies that improve the overall fusion quality. Therefore, the global reward signal is both a feedback signal and a crucial basis for driving the joint optimization of the Actor and Critic networks.
[0080] During the alternating optimization process, the Critic network guides the Actor network's optimization by evaluating the action value function, with the core being the calculation of temporal difference error: ,in, For instant rewards, For the Critic network, the Q-value is estimated for the current state-action pair. For the Critic network, the Q-value is estimated for the next state-action pair; Based on temporal difference error, the gradient descent method is used to update the Critic network parameters. : ,in, is the learning rate for the Critic network, with a value of 0.0003. For timing difference error, For gradient operators; Similarly, based on the temporal difference error, the gradient descent method is used to update the Actor network parameters. : ,in, Here is the learning rate for the Actor network, with a value of 0.0001. For gradient operators, For the Actor policy network in state Select action The probability of.
[0081] In a preferred embodiment of the present invention, the optimization is iteratively performed until the weight generation strategy network converges, and the convergence condition is as follows: ,in, Let be the parameter vector of the agent unit's policy network parameters at the k-th iteration. Let be the parameter vector of the agent unit's policy network parameters at the (k-1)th iteration. The L2 norm is used to calculate the magnitude of change in the parameter vector before and after iteration. This means taking the maximum value among all agents to ensure that the weight generation strategy network converges for all agent units.
[0082] In a preferred embodiment of the present invention, the rationality of the optimized data weight vector is verified, and the data weight vector optimized by the proxy unit is checked. To verify whether the nonnegativity constraint (all weight values are nonnegative) and the normalization constraint (the sum of weights is 1) are satisfied, the following formula is used: ,in, This represents the j1-th element in the weight vector of agent i1. It is the minimum value in the weight vector. The sum of the elements of the weight vector. The preset threshold value is [value]. , This represents the logical AND operation, if If the weight vector obtained by the surrogate unit optimization is satisfactory, it indicates that the weight vector is reasonable and can be directly used for subsequent heterogeneous data weighted merging processing. Otherwise, local correction needs to be triggered to correct the weight vector obtained by the surrogate unit optimization. If the weight vector does not meet the reasonableness requirement, local correction is triggered: For the surrogate unit, when the reasonableness verification fails, that is, when the weight vector obtained by the surrogate unit optimization contains negative elements or the sum is not close to 1, the weight vector needs to be corrected, as shown in the following formula: ,in, This indicates the weight vector The maximum value is taken for each element to eliminate negative weights. This is the sum of the corrected non-negative weights. The corrected weight vector satisfies and .
[0083] By implementing this embodiment, a complete reinforcement learning optimization closed loop is formed, ensuring that the weight generation strategy can continuously iterate towards the global optimum. The Actor-Critic framework decouples strategy optimization and value assessment, solving the problems of large variance and slow convergence in traditional reinforcement learning algorithms, balancing strategy exploration and utilization, and significantly improving the network's learning efficiency and convergence stability. Using the global reward signal as the sole optimization objective ensures that the strategy optimization of all distributed agent units converges towards improving the global fusion effect, avoiding the deviation between the individual optimum and the global optimum of distributed agents. The direction and magnitude of strategy optimization are quantified through temporal differential error, making the logic of network parameter updates clear and traceable, adapting to the high reliability application requirements of power systems.
[0084] S4. Perform weighted calculations based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result.
[0085] Specifically, the weighted computation means that the concentrator uses the optimized data weight vector as the weighting coefficient to perform a weighted summation on the data source matrices of each data source, ultimately realizing the operation of multi-source heterogeneous data fusion.
[0086] To illustrate, after receiving the optimized data weight vectors uploaded by all agent units, the concentrator performs a validity check on the weight vectors. Using the valid weight vectors as weighting coefficients, it performs a weighted summation calculation on the standardized data source matrix corresponding to each data source, and finally generates the fusion result of multi-source heterogeneous data, which is output to the power grid intelligent dispatch and operation system as the basic data for situational awareness, fault early warning, and dispatch decision-making.
[0087] In a preferred embodiment of the present invention, a weighted merging process is performed on the data source matrix to achieve the fusion of multi-source heterogeneous data: ,in, The result of multi-source heterogeneous data fusion This is the validated data source weight vector. This is the data source matrix.
[0088] In a preferred embodiment of the present invention, taking the municipal-level dispatch center of the G province power grid as the application scenario, the specific implementation process is as follows: The concentrator connects to the SCADA system, transformer status monitoring system, regional meteorological monitoring platform, electricity consumption information collection system, and dispatch operation log system of the municipal power grid through the dispatch data network, and collects five types of heterogeneous data in real time: real-time power grid operation data, equipment status monitoring data, environmental meteorological data, load data, and dispatch log data, and completes the aggregation and verification of the original data; The concentrator performs matrixing and normalization preprocessing on the five types of heterogeneous data respectively, generating standardized data source matrices corresponding to the five data sources; based on the characteristics of each data source matrix, it assigns an independent proxy unit to each of the five data sources, deploys a total of five proxy units, establishes a one-to-one binding relationship between the proxy units and the data sources, and distributes each data source matrix to the corresponding proxy unit. After receiving the data source matrix, each agent unit combines the corresponding original heterogeneous data to generate its own local state vector. The local state vector is then input into the built-in weight generation strategy network to generate the data weight vector of the corresponding data source. Subsequently, the local state vector and the data weight vector are encrypted and uploaded to the concentrator. After receiving all the data uploaded by the five agent units, the concentrator concatenates all local state vectors and data weight vectors by dimension to generate a global state vector. Combining the global state vector with the five data source matrices, the concentrator quantitatively evaluates the current fusion effect to obtain a global fusion quality score. Based on the global reward function, it generates a global reward signal and broadcasts it to the five agent units. Each agent unit takes the received global reward signal as the optimization target, the local state vector as the input, and the data weight vector as the action, and repeatedly performs the dual-network iterative optimization operation to continuously optimize the weight generation strategy network until the network converges. After convergence, each agent unit regenerates the data weight vector based on the optimized network and uploads it to the concentrator. After verifying the five optimized data weight vectors, the concentrator uses them as weighting coefficients to perform weighted summation calculations on the five data source matrices, ultimately obtaining the fusion result of the multi-source heterogeneous data of the local power grid. This result is then output to the intelligent dispatching system of the dispatch center for power grid operation status awareness and short-term load forecasting.
[0089] In another preferred embodiment of the present invention, a provincial power grid dispatch center is taken as the research object. Power grid operation data, equipment status monitoring data, environmental data, meteorological data, log data, etc., are collected from its database to construct an experimental dataset. Based on the experimental dataset, a multi-agent distributed decision network is trained. Figure 3 The global reward function obtained during training As shown in the figure, the convergence curve rises rapidly before 10,000 rounds, with the reward function value jumping from -20 to around 48. This is because the agent strategies are highly random in the early stages of the Actor-Critic framework, resulting in a large exploration space. Furthermore, the concentrator's penalty signal for high-weight redundant data sources is strong, prompting the agent units to quickly increase the weight of high-quality data sources, leading to a sharp increase in the global reward. After 10,000 rounds, the rate of increase slows down, and the curve tends to stabilize around 20,000 rounds, indicating that redundant information has been effectively eliminated and the retention rate of key features has reached a steady state.
[0090] Based on the experimental dataset, the multi-agent reinforcement learning-based multi-source heterogeneous data fusion method of this embodiment is compared with traditional fusion methods. The comparison results are shown in Table 1. Experimental results show that, thanks to the multi-agent collaborative decision-making mechanism and the closed-loop optimization of global reward feedback, the compression ratio of this embodiment reaches 8.5, which is 3.9 times and 2.5 times higher than FL and MAP, respectively, achieving efficient removal of data redundancy. Traditional methods lack the ability to dynamically optimize weights for heterogeneous data sources, making it difficult to balance feature preservation and information fusion: FL relies on static rules, resulting in the loss of key temporal features and a fusion error of 2.1%; MAP is limited by prior distribution assumptions, resulting in insufficient semantic conflict resolution and an error as high as 3.2%. In contrast, this embodiment uses the alternating optimization mechanism of the Actor-Critic framework to enable the agent to continuously correct the weight strategy under the drive of local observation states, and introduces a consistency verification mechanism to control the fusion error to 0.8%, effectively achieving a balance between key feature preservation and fusion.
[0091] Table 1. Comparison of the fusion method in this embodiment with traditional fusion methods. This embodiment constructs a multi-agent reinforcement learning fusion architecture with distributed multi-agent decision-making and centralized global evaluation. Distributed feature perception and weight decision-making are achieved by assigning independent agent units to each heterogeneous data source, adapting to the heterogeneous characteristics of different data sources. A concentrator enables quantitative evaluation and unified feedback of the global fusion effect, ensuring that all agent decisions converge towards the global optimal goal. Adaptive dynamic adjustment of fusion weights is achieved through iterative optimization of dual networks in reinforcement learning, eliminating the need for manually preset static rules. Finally, weighted merging completes the fusion of multi-source heterogeneous data. This embodiment fully covers the entire process from data acquisition to final fusion output, solving the core problems of traditional fusion methods: reliance on static rules, inability to adapt to the strong heterogeneity of multi-source data in power systems, and difficulty in balancing redundancy removal and key feature retention.
[0092] This embodiment employs a multi-agent reinforcement learning architecture, where agent units autonomously generate data weights based on the real-time state of the data source, without relying on human experience. The concentrator forms a closed-loop feedback through global fusion quality assessment and global reward signals, driving agent units to automatically optimize weight generation strategies, thus solving the problem that manual static weighted fusion cannot dynamically optimize weights. Agent units generate local state vectors in real time based on the data source matrix and heterogeneous data, perceiving changes in data source quality, temporal characteristics, and correlations. Using the global reward signal as the optimization target, agent units continuously adjust weight strategies, with weights dynamically adapting to grid conditions and data quality in real time. It supports repeated dual-network iterative optimization, with the weight generation strategy network continuously converging to the current optimum, achieving continuously optimal weight allocation under dynamic conditions. Finally, weighted calculations are performed based on the new data weight vectors and the data source matrix to obtain the multi-source heterogeneous data fusion result, realizing real-time data fusion and enabling dynamic adjustment of data weight vectors for adaptive weight optimization, meeting the high timeliness requirements of power systems for data fusion.
[0093] See Figure 4 This is a schematic diagram of the structure of a multi-source heterogeneous data fusion system based on multi-agent reinforcement learning according to an embodiment of the present invention, including a concentrator and several agent units; each agent unit corresponds to a data source; The concentrator includes a heterogeneous data acquisition module, a data source matrix generation module, a global fusion quality assessment module, and a multi-source data fusion module; The heterogeneous data acquisition module is used to acquire heterogeneous data from various data sources of the power grid to be processed. The data source matrix generation module is used to perform matrixing and normalization processing on heterogeneous data to generate a data source matrix, allocate corresponding proxy units to each data source according to the data source matrix, and send the data source matrix to each proxy unit. The proxy unit is used to generate a local state vector based on the data source matrix and heterogeneous data, generate a data weight vector based on the local state vector in the preset weight generation strategy network, and transmit the local state vector and data weight vector to the concentrator. The global fusion quality assessment module is used to perform a concatenation operation on the local state vector and the data weight vector to generate a global state vector, perform a fusion quality assessment on the global state vector and the data source matrix to obtain a global fusion quality score, and generate a global reward signal based on the global fusion quality score and send it to each agent unit. The agent unit is also used to repeatedly perform dual-network iterative optimization operations to optimize the weight generation strategy network with the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. Based on the optimized weight generation strategy network, the data weight vector is regenerated and the new data weight vector is sent to the concentrator. The multi-source data fusion module is used to perform weighted calculations based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result.
[0094] This invention provides a multi-source heterogeneous data fusion system based on multi-agent reinforcement learning. The system acquires heterogeneous data from various power grid data sources using a concentrator's heterogeneous data acquisition module. A data source matrix generation module in the concentrator performs matrix transformation and normalization on the heterogeneous data to generate a data source matrix. Based on this matrix, corresponding agent units are assigned to each data source, and the data source matrix is sent to each agent unit. In each agent unit, a local state vector is generated based on the data source matrix and the heterogeneous data. A data weight vector is then generated based on the local state vector within a preset weight generation strategy network. The local state vector and data weight vector are transmitted to the concentrator. Finally, in the concentrator's global fusion quality evaluation module, the system evaluates the fusion quality based on the local state vectors. The vector and data weight vector are concatenated to generate a global state vector. The fusion quality is then evaluated based on the global state vector and the data source matrix to obtain a global fusion quality score. A global reward signal is generated based on this score and sent to each agent unit. In each agent unit, the weight generation strategy network is repeatedly optimized using the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. The optimized weight generation strategy network is then used to regenerate the data weight vector, which is sent to the concentrator. Finally, in the concentrator's multi-source data fusion module, a weighted calculation is performed based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result.
[0095] By employing a multi-agent reinforcement learning architecture, agent units autonomously generate data weights based on the real-time state of the data source, without relying on human experience. The concentrator forms a closed-loop feedback through global fusion quality assessment and global reward signals, driving the agent units to automatically optimize the weight generation strategy, thus solving the problem that manual static weighted fusion cannot dynamically optimize weights. Agent units generate local state vectors in real time based on the data source matrix and heterogeneous data, and are aware of changes in data source quality, temporal characteristics, and correlations. Using the global reward signal as the optimization target, the agent units continuously adjust the weight strategy, with weights dynamically adapting to grid conditions and data quality in real time. It supports repeated execution of dual-network iterative optimization, with the weight generation strategy network continuously converging to the current optimum, achieving continuously optimal weight allocation under dynamic conditions. Finally, weighted calculations are performed based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result, realizing real-time data fusion and enabling dynamic adjustment of the data weight vector for adaptive weight optimization, meeting the high timeliness requirements of power system data fusion.
[0096] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0097] Those skilled in the art will understand that, for convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning as described in the above embodiments. The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0099] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0100] The memory can be used to store the computer program. The processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device or other volatile solid-state storage device.
[0101] Another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the multi-source heterogeneous data fusion method based on multi-agent reinforcement learning described in the above embodiment.
[0102] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0103] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for fusing multi-source heterogeneous data based on multi-agent reinforcement learning, characterized in that, Applicable to concentrators; Multi-agent reinforcement learning-based methods for fusing multi-source heterogeneous data include: Acquire heterogeneous data from various data sources of the power grid to be processed; Heterogeneous data is matrixed and normalized to generate a data source matrix. Based on the data source matrix, a corresponding agent unit is assigned to each data source, and the data source matrix is sent to each agent unit so that each agent unit generates a local state vector based on the data source matrix and heterogeneous data. In a preset weight generation strategy network, a data weight vector is generated based on the local state vector, and the local state vector and data weight vector are transmitted to the concentrator. A global state vector is generated by concatenating the local state vector and the data weight vector. A fusion quality assessment is then performed based on the global state vector and the data source matrix to obtain a global fusion quality score. A global reward signal is generated based on the global fusion quality score and sent to each agent unit. This enables each agent unit to repeatedly perform a dual-network iterative optimization operation to optimize the weight generation strategy network, using the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. The data weight vector is then regenerated based on the optimized weight generation strategy network and sent to the concentrator. The multi-source heterogeneous data fusion result is obtained by performing weighted calculations based on the new data weight vector and the data source matrix.
2. The method for fusing multi-source heterogeneous data based on multi-agent reinforcement learning as described in claim 1, characterized in that, The heterogeneous data includes real-time power grid operation data, equipment status monitoring data, environmental data, meteorological data, and log data; Heterogeneous data are matrixed and normalized separately to generate a data source matrix, including: Based on the heterogeneous data, determine the maximum and minimum values corresponding to each feature dimension of the heterogeneous data from each data source. Based on the preset interval distance and the maximum and minimum values corresponding to each feature dimension, perform equal-interval binning smoothing on the heterogeneous data of the corresponding feature dimension to obtain the first processed data. Based on the first processed data, normalization is performed according to the feature dimension to obtain the second processed data; Based on the second processed data, a two-dimensional standard data matrix is constructed using the data acquisition time sequence as the row index and the feature dimension as the column index. Missing values in the two-dimensional standard data matrix are filled with the time-series moving average of the same feature dimension to generate a data source matrix corresponding to each data source.
3. The method for multi-source heterogeneous data fusion based on multi-agent reinforcement learning as described in claim 2, characterized in that, Assigning corresponding proxy units to each data source according to the data source matrix, including: Based on the data source matrix, extract the data source structure type features, time-series update features, and data quality features of heterogeneous data respectively; The corresponding agent unit is assigned based on the data source structure type characteristics, time-series update characteristics, and data quality characteristics.
4. The method for fusing multi-source heterogeneous data based on multi-agent reinforcement learning as described in claim 3, characterized in that, Each agent unit generates a local state vector based on the data source matrix and heterogeneous data, including: Based on the data source matrix and heterogeneous data, the number of missing values, the number of outliers, and the signal-to-noise ratio are extracted respectively. The missing rate is determined based on the ratio of the number of missing values to the total amount of heterogeneous data, and the outlier ratio is determined based on the ratio of the number of outliers to the total amount of heterogeneous data. The Pearson correlation coefficient between adjacent data collection sequences is calculated based on the row index of the data source matrix, and the Pearson correlation coefficient is used as the timestamp correlation coefficient. A quality assessment vector is constructed based on the missing rate, outlier ratio, timestamp correlation coefficient, and signal-to-noise ratio. The absolute difference, minimum difference, and maximum difference values of each heterogeneous data point in the feature dimension are calculated based on the data source matrix. The gray correlation degree of the same feature dimension between different data sources is calculated based on the absolute difference, minimum difference, and maximum difference values. The associated feature vector is calculated based on the gray correlation degree. The quality assessment vector and the associated feature vector are concatenated after dimensional alignment to generate the local state vector of the current agent unit.
5. The method for multi-source heterogeneous data fusion based on multi-agent reinforcement learning as described in claim 4, characterized in that, Each agent unit generates a data weight vector based on the local state vector in the preset weight generation strategy network, including: In the preset weight generation strategy network, nonlinear feature extraction and high-dimensional mapping calculation are performed on the local state vector to obtain the original weight score values of each feature dimension of the data source. The original weight scores are subjected to exponential normalization to generate a data weight vector; The data weight vector satisfies the constraints that all weight elements are non-negative and the sum of all weight elements is 1.
6. The method for multi-source heterogeneous data fusion based on multi-agent reinforcement learning as described in claim 5, characterized in that, The fusion quality is evaluated based on the global state vector and the data source matrix to obtain a global fusion quality score, including: The global state vector is decomposed according to the feature dimension, and the resulting local state vector and corresponding data weight vector are used to construct a basic dataset with the data source matrix. Using the data acquisition time sequence as the alignment benchmark, the data with the same feature dimension from each data source within the same time window is matched based on the basic dataset, and the data consistency evaluation value is calculated by combining the corresponding data weight vector. The mean of the original information entropy of heterogeneous data from each data source is calculated based on the basic dataset, and the information entropy of the fused data is calculated by weighting the current data weight vector onto the data source matrix. The difference between the mean of the original information entropy and the information entropy of the fused data is used as the information entropy gain evaluation value. Based on the absolute difference value of each heterogeneous data in the feature dimension, determine the number of data pairs that exceed the preset deviation threshold, and obtain the field conflict resolution degree based on the ratio of the number of data pairs to the total number of data pairs in the basic dataset. The global fusion quality score is obtained by weighting the data consistency assessment value, information entropy gain assessment value, and field conflict resolution degree.
7. The method for multi-source heterogeneous data fusion based on multi-agent reinforcement learning as described in claim 6, characterized in that, The weight generation strategy network includes a strategy optimization network and a value evaluation network; Each agent unit optimizes the weight generation policy network by taking the global reward signal as the target, the local state vector as the input, and the data weight vector as the action, including: Using the global reward signal as the optimization objective, the local state vector as the network state input, and the data weight vector as the policy action output, a dual-network iterative optimization operation is performed until the preset loss function value converges, resulting in the optimized weight generation policy network. The dual-network iterative optimization operation includes: Based on the global reward signal, the Q-value estimate of the current state action pair is calculated in the current value evaluation network according to the current local state vector and the current data weight vector. The temporal difference error is calculated based on the Q-value estimate and the current immediate reward. Based on the aforementioned temporal difference error, the parameters of the first fully connected layer of the current value assessment network are updated using the gradient descent method. Using the temporal difference error as the strategy optimization guide, the gradient descent method is used to update the parameters of the second fully connected layer of the strategy optimization network, and the preset loss function value is calculated based on the parameters of the first and second fully connected layers. Determine whether the loss function value has converged. If it has, use the current policy optimization network and the current value evaluation network as the optimized weights to generate the policy network. Otherwise, update the current data weight vector based on the loss function value.
8. A multi-source heterogeneous data fusion system based on multi-agent reinforcement learning, characterized in that, It includes a concentrator and several agent units; each agent unit corresponds to a data source; The concentrator includes a heterogeneous data acquisition module, a data source matrix generation module, a global fusion quality assessment module, and a multi-source data fusion module; The heterogeneous data acquisition module is used to acquire heterogeneous data from various data sources of the power grid to be processed. The data source matrix generation module is used to perform matrixing and normalization processing on heterogeneous data to generate a data source matrix, allocate corresponding proxy units to each data source according to the data source matrix, and send the data source matrix to each proxy unit. The proxy unit is used to generate a local state vector based on the data source matrix and heterogeneous data, generate a data weight vector based on the local state vector in the preset weight generation strategy network, and transmit the local state vector and data weight vector to the concentrator. The global fusion quality assessment module is used to perform a concatenation operation on the local state vector and the data weight vector to generate a global state vector, perform a fusion quality assessment on the global state vector and the data source matrix to obtain a global fusion quality score, and generate a global reward signal based on the global fusion quality score and send it to each agent unit. The agent unit is also used to repeatedly perform dual-network iterative optimization operations to optimize the weight generation strategy network with the global reward signal as the target, the local state vector as the input, and the data weight vector as the action. Based on the optimized weight generation strategy network, the data weight vector is regenerated and the new data weight vector is sent to the concentrator. The multi-source data fusion module is used to perform weighted calculations based on the new data weight vector and the data source matrix to obtain the multi-source heterogeneous data fusion result.
9. A terminal device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a multi-source heterogeneous data fusion method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.