Reinforcement learning-based data confidence dynamic weighting fusion method and system

By constructing a multi-dimensional data reliability index system and optimizing it through reinforcement learning, and dynamically adjusting the weight allocation, the problems of single reliability assessment and lack of closed-loop optimization in traditional multi-source data fusion technology are solved, thus achieving efficient energy consumption prediction and decision support.

CN121167652BActive Publication Date: 2026-03-31HUBEI UNIV OF EDUCATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional multi-source data fusion technology suffers from limitations in energy consumption forecasting, including a single reliability assessment dimension, lack of dynamic adaptability in weight allocation, and absence of a closed-loop optimization mechanism. This results in significant deviations in forecasting results, making it difficult to support the scientific validity of decisions such as grid dispatching and energy storage deployment.

Method used

We construct a multi-dimensional data reliability index system, combine reinforcement learning to achieve dynamic optimization of data weights, establish a closed-loop process of 'evaluation-fusion-prediction-error feedback', and dynamically adjust the weight allocation to minimize and eliminate prediction errors through multi-dimensional temporal feature embedding and cross-source data association evaluation.

Benefits of technology

It improves the quality of multi-source data fusion and the accuracy of consumption prediction, providing reliable data support for power grid dispatch and energy planning, and supporting the efficient operation of the energy system and the large-scale consumption of renewable energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167652B_ABST
    Figure CN121167652B_ABST
Patent Text Reader

Abstract

The application discloses a data confidence dynamic weighting fusion method and system based on reinforcement learning, and the method comprises the following steps: collecting multi-source data; firstly, the static inherent attributes of the multi-source data are quantified to obtain static reliability basic indexes; then, sub-indexes are constructed through multi-dimensional time sequence feature embedding and cross-source data association reliability evaluation; further, the calculation of data reliability scores is completed; a reinforcement learning environment is constructed with the data reliability scores as the core input; the reinforcement learning reward criterion is to minimize the prediction error, and the optimal weight distribution strategy is generated through iterative training and learning; the weighted fusion of the multi-source data is performed to obtain comprehensive data, which is input into a consumption prediction model to output a prediction result; and the parameter adjustment is completed based on the error between the prediction result and the actual consumption data. The application provides a more reliable basis for energy consumption and other related decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy forecasting technology, and more specifically, relates to a dynamic weighted fusion method and system for data confidence based on reinforcement learning. Background Technology

[0002] In critical scenarios such as energy consumption forecasting that rely on multi-source data fusion, data quality and the rationality of the fusion strategy directly determine the reliability of the forecast results. However, current technological systems have overall shortcomings in addressing practical needs and are unable to meet the requirements for efficient operation of energy systems. On the one hand, traditional data reliability assessments are singular in dimensions, mostly focusing on isolated indicators such as historical data accuracy and acquisition precision, without incorporating the correlation and stability between multi-source data into the assessment scope. However, in practical applications, the correlation between different data sources (such as the coupling relationship between meteorological data and photovoltaic output data) has a significant impact on overall data reliability. Ignoring this dimension will lead to a disconnect between reliability assessments and actual scenarios, causing subsequent fusion decisions to lack a scientific basis.

[0003] On the other hand, existing data fusion generally adopts a static weight allocation model, which pre-sets fixed weights for each data source and keeps them unchanged throughout the fusion process. However, the reliability of data sources changes dynamically over time and with environmental changes (such as fluctuations in light data due to seasonal changes and deviations in monitoring data caused by equipment aging). Static weights cannot adapt to this dynamism, often resulting in insufficient weights for high-reliability data and excessive weights for low-reliability data, directly reducing the quality of fused data and leading to significant deviations in the absorption prediction results.

[0004] More importantly, traditional technologies lack a closed-loop optimization mechanism of "evaluation-fusion-feedback." Even if some solutions can adjust weights, they cannot accurately correct reliability assessment parameters and weight allocation strategies based on the comparison between real-time prediction errors and actual consumption data. This feedback-free operating mode keeps data fusion and prediction in a "passive adaptation" state, making it difficult to continuously improve based on actual operating conditions. It cannot solve the problem of predicting "production-consumption mismatch" in energy consumption forecasting, nor can it support the scientific nature of decisions such as grid dispatching and energy storage layout. Ultimately, this affects the efficient consumption of renewable energy and restricts the process of energy system transformation towards cleaner and smarter energy. Summary of the Invention

[0005] This invention aims to address the shortcomings of traditional multi-source data fusion technologies in scenarios such as energy consumption forecasting, including a single reliability assessment dimension, lack of dynamic adaptability in weight allocation, and absence of a closed-loop optimization mechanism. By constructing a multi-dimensional data reliability index system, combining reinforcement learning to achieve dynamic optimization of data weights, and establishing a closed-loop process of "assessment-fusion-prediction-error feedback," the quality of fused data and the accuracy of consumption forecasting are improved. This provides reliable data support for grid dispatching and energy planning, contributing to the efficient operation of energy systems and the large-scale consumption of renewable energy.

[0006] To address the aforementioned deficiencies or improvement needs of existing technologies, as a first aspect of this invention, the present invention provides a dynamic weighted fusion method for data confidence based on reinforcement learning, comprising:

[0007] S1. Complete the collection of multi-source data;

[0008] S2. First, the static inherent attributes of multi-source data are summarized and quantified to obtain basic static reliability indicators; then, sub-indicators are constructed through multi-dimensional time-series feature embedding and cross-source data association reliability assessment; the three types of indicators—static, time-series, and cross-source association—are integrated to complete the calculation of data reliability score;

[0009] S3. Using the data reliability score obtained in step S2 as the core input, construct a reinforcement learning environment; set the weight allocation scheme of different data sources as the decision variable of the optimization objective, take minimizing the prediction error as the reward criterion of reinforcement learning, and learn and generate the optimal weight allocation strategy through iterative training.

[0010] S4. Based on the weights output from S3, perform weighted fusion on the multi-source data to obtain comprehensive data, input it into the absorption prediction model and output the prediction result; by comparing the prediction result with the actual absorption data, quantify the error, and dynamically adjust the calculation parameters of the data reliability index in step S2 and the core parameters of reinforcement learning in step S3 based on the error.

[0011] Furthermore, the multi-source data in S1 includes: meteorological forecast data, measured meteorological data, and energy data.

[0012] Furthermore, the method for calculating the data reliability score in S2 is as follows:

[0013] ,

[0014] in, To indicate the first Class data at time The overall reliability score; The overall weighting coefficients are optimized through reinforcement learning training. For the first Class data at time Static reliability metrics are assessments of data reliability based on the inherent attributes of the data, reflecting the fundamental reliability of the data itself. For the first Class data at time Timing reliability metrics; For the first Class data at time Cross-source correlation reliability metrics.

[0015] Furthermore, the specific method for embedding multi-dimensional temporal features in S2 is as follows:

[0016] By analyzing time series data to uncover reliability patterns over time, time series reliability metrics can be constructed.

[0017] ,

[0018] in, For trend items, For volatility, This is the balance coefficient between the trend term and volatility;

[0019] To reflect the long-term trend of data reliability, the slope of linear regression within a sliding window is used for calculation:

[0020] ,

[0021] in, Indicates the first Class data at time Trend items, The size of the sliding window; This is a time variable, representing the specific moment within the sliding window, used to iterate through each time point within the window to calculate the trend; This represents the average time within the sliding window, i.e., all times within the window. The average value; Indicates the first Class data at time Reliability score; Let be the mean of the reliability scores for the first type of data within the sliding window;

[0022] To reflect the short-term stability of data reliability, the normalized value of the standard deviation within the window is used:

[0023] ,

[0024] in, Indicates the first Class data at time Volatility indicators; The calculation starts from time 10. At the time Within this sliding window, the first Class Data Reliability Score Standard deviation; Indicates the first The maximum standard deviation obtained by calculating the class of data across all possible sliding windows in its history.

[0025] Furthermore, the specific method for assessing the reliability of cross-source data association in S2 is as follows:

[0026] By analyzing the correlations between different data sources, a correlation reliability index is constructed:

[0027] ,

[0028] in, Indicates the first Class data at time Cross-source correlation reliability metrics; Representative except the first Other data source categories besides the first type of data, used to iterate through all data related to the first type. Data sources with related categories; For data source and The association weight; Let be the dynamic correlation coefficient between the two data sources at time t; Indicates the first Class data at time The reliability score reflects the first The reliability of the data itself at that moment;

[0029] The Pearson correlation coefficient within a sliding window is used for calculation:

[0030] ,

[0031] in, To indicate at time At that time, the first Class data and the first Dynamic Pearson correlation coefficient for class data; These represent different data source categories; It is a time variable; Indicates the first Class data at time The original data values; Indicates the first Class data at time The original data values; It is the first one in the sliding window The mean of the original data values ​​for the class data; It is the first one in the sliding window The mean of the original data values ​​for the class data; It is the size of the sliding window.

[0032] Furthermore, the reward criterion for reinforcement learning in S3 is specifically as follows:

[0033] The reward criterion for reinforcement learning is centered on "minimizing prediction error," incorporating "reliability-error correlation penalty" and "cross-period reward smoothing mechanism," specifically designed as follows:

[0034]

[0035] in, For a moment Reward values ​​for reinforcement learning; For this moment The error in the absorption prediction; This is the reliability-error correlation penalty coefficient; For a moment Reliability-error correlation penalty term; This represents the change in reward at the previous moment.

[0036] Furthermore, the calculation method for the reliability-error correlation penalty term is as follows:

[0037] ,

[0038] in, Limited to reliability rating only Above the threshold The high reliability of the data source is used to calculate the penalty; For a moment For the first The actual weights assigned to each data category; The theoretically optimal weights are based on reliability. To remove the first After classifying the data, the prediction error is obtained by fusing only other data. This represents the prediction error of the original fused data.

[0039] Furthermore, the calculation parameters for the reliability index in S4 include: sliding window size. Comprehensive weighting coefficient And the balance coefficient between trend term and volatility .

[0040] As a second aspect of the present invention, the present invention provides a data confidence dynamic weighted fusion method based on reinforcement learning, comprising:

[0041] The data acquisition unit is used to collect data from multiple sources.

[0042] The reliability scoring unit is used to first summarize and quantify the static inherent attributes of multi-source data to obtain basic static reliability indicators; then, through multi-dimensional time-series feature embedding and cross-source data association reliability assessment, sub-indicators are constructed; and the three types of indicators—static, time-series, and cross-source association—are integrated to complete the calculation of data reliability scores.

[0043] The weight optimization unit is used to construct a reinforcement learning environment with the data reliability score obtained from the reliability scoring unit as the core input. The weight allocation scheme of different data sources is set as the decision variable of the optimization objective. The minimum of the elimination prediction error is used as the reward criterion for reinforcement learning. The optimal weight allocation strategy is generated through iterative training.

[0044] The parameter optimization unit is used to perform weighted fusion of multi-source data based on the weights output by the weight optimization unit to obtain comprehensive data, which is then input into the absorption prediction model to output the prediction result. By comparing the prediction result with the actual absorption data to quantify the error, the calculation parameters of the data reliability index in the reliability scoring unit and the core parameters of reinforcement learning in the weight optimization unit are dynamically adjusted based on the error.

[0045] As a third aspect of the invention, the invention provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of any step of the reinforcement learning-based data confidence dynamic weighted fusion method.

[0046] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0047] 1. The reinforcement learning-based dynamic weighted fusion method for data confidence, as described in this invention, constructs a multi-dimensional data reliability index system to quantitatively evaluate multi-source data from three dimensions: static inherent attributes, temporal dynamic changes, and cross-source correlations, resulting in a data reliability score. Specifically, the static reliability index comprehensively considers inherent attributes such as historical accuracy, quality, and source authority, providing a fundamental quantitative basis for data reliability. Multi-dimensional temporal feature embedding mines long-term trends and short-term fluctuations in data, accurately reflecting the dynamic differences in data reliability over time. Cross-source data correlation reliability assessment, based on dynamic correlation coefficients and correlation weights, effectively measures the impact of correlation stability between different data sources on reliability. This system provides a scientific and comprehensive quantitative benchmark for subsequent weighted fusion, overcoming the limitations of traditional single-dimensional data reliability assessments and ensuring that the quality of data before fusion is quantifiable and comparable.

[0048] 2. The data confidence dynamic weighted fusion method based on reinforcement learning of this invention introduces reinforcement learning into the weight allocation process, using the minimization of prediction error as the core reward criterion. It also incorporates a reliability-error correlation penalty and a cross-period reward smoothing mechanism to dynamically optimize the weight allocation of multi-source data. Reinforcement learning continuously iterates and tries different weight combinations based on data reliability scores, retaining weight allocation strategies that bring low prediction error and eliminating high-error strategies. The reliability-error correlation penalty avoids a mismatch between the weight allocation of high-reliability data and the actual error improvement, while the cross-period reward smoothing mechanism enhances the stability of the weight strategy in the time-series dimension, enabling the weight allocation to accurately match the dynamic changes in data reliability, thus improving the quality and relevance of the fused data.

[0049] 3. The reinforcement learning-based dynamic weighted fusion method for data confidence, as described in this invention, achieves continuous optimization of data fusion and prediction through a closed-loop process of multi-source data weighted fusion, energy consumption prediction, and error feedback optimization. Multi-source data is weighted and fused according to optimized dynamic weights to obtain comprehensive data, which is then input into the energy consumption prediction model to output prediction results. The prediction results are then compared with actual energy consumption data to quantify the error, and the calculation parameters of data reliability indicators and core reinforcement learning parameters are dynamically adjusted based on error feedback. This closed-loop process ensures continuous improvement based on actual performance throughout the entire process, from data reliability assessment and weight allocation to fusion prediction, effectively enhancing the accuracy and stability of energy consumption prediction and providing a more reliable basis for energy consumption and related decision-making. Attached Figure Description

[0050] Figure 1 This is a flowchart of the data confidence dynamic weighted fusion method based on reinforcement learning according to an embodiment of the present invention;

[0051] Figure 2This is a schematic diagram illustrating the reliability score calculation in an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram illustrating the adjustment of prediction feedback parameters in an embodiment of the present invention;

[0053] Figure 4 This is a system unit diagram of an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0055] Example 1

[0056] Please refer to Figure 1 This embodiment 1 provides a dynamic weighted fusion method for data confidence based on reinforcement learning, including:

[0057] S1. Complete the collection of multi-source data;

[0058] S2. First, the static inherent attributes of multi-source data are summarized and quantified to obtain basic static reliability indicators; then, sub-indicators are constructed through multi-dimensional time-series feature embedding and cross-source data association reliability assessment; the three types of indicators—static, time-series, and cross-source association—are integrated to complete the calculation of data reliability score;

[0059] S3. Using the data reliability score obtained in step S2 as the core input, construct a reinforcement learning environment; set the weight allocation scheme of different data sources as the decision variable of the optimization objective, take minimizing the prediction error as the reward criterion of reinforcement learning, and learn and generate the optimal weight allocation strategy through iterative training.

[0060] S4. Based on the weights output from S3, perform weighted fusion on the multi-source data to obtain comprehensive data, input it into the absorption prediction model and output the prediction result; by comparing the prediction result with the actual absorption data, quantify the error, and dynamically adjust the calculation parameters of the data reliability index in step S2 and the core parameters of reinforcement learning in step S3 based on the error.

[0061] This embodiment 1 further elaborates on the above steps.

[0062] (1) Data collection

[0063] In the initial implementation phase, it is essential to prioritize the systematic collection of multi-source data. Considering the practical need for data reliability to directly impact prediction accuracy in energy consumption forecasting scenarios, and the poor fusion effect caused by the single data dimension and insufficient dynamic adaptation of traditional technologies, the data collected in this step must cover key dimensions such as energy production and environmental conditions to provide comprehensive support for subsequent multi-dimensional reliability assessment and dynamic weight allocation.

[0064] The collected multi-source data includes three core categories: First, meteorological forecast data, derived from short- and medium-term forecasts issued by meteorological departments, covering parameters such as wind speed, light intensity, and temperature. This data reflects future environmental trends affecting energy production (such as wind power and photovoltaics), providing a data foundation for subsequent time-series characteristic analysis. Second, measured meteorological data, acquired in real time through monitoring equipment at and around energy sites, recording actual environmental parameters such as real-time wind speed and sunshine duration. This data can be compared with meteorological forecast data, providing a measured basis for assessing "prediction accuracy" in static reliability and cross-source correlation analysis. Third, energy data, including real-time output of new energy power generation equipment, grid transmission power, and user-side energy consumption, covering the entire energy production-transmission-consumption chain. This data is the core business data for subsequent calculation of data reliability scoring and verification of fusion effects.

[0065] During the data collection process, data needs to be continuously acquired at fixed time intervals (e.g., 5-15 minutes). At the same time, preliminary format standardization and missing value completion are completed to ensure data integrity and consistency. This process not only solves the problem of "disorganized and unusable data" in the traditional data collection stage, but also provides standardized data input for subsequent steps such as static inherent attribute quantification, temporal feature embedding, and cross-source association evaluation, ensuring the smooth connection of the entire method process.

[0066] (2) Reliability score

[0067] Please refer to Figure 2 After completing the collection of multi-source data and forming a standardized data foundation, a multi-dimensional evaluation system is constructed to calculate the data reliability score. This not only solves the problem of the single evaluation dimension of traditional technology, but also provides a quantitative basis for the subsequent dynamic weight allocation based on reinforcement learning. It is a key link connecting data collection and intelligent integration.

[0068] In a preferred embodiment, the data reliability score is calculated as follows:

[0069] ,

[0070] in, To indicate the first Class data at time The overall reliability score; The overall weighting coefficients are optimized through reinforcement learning training. For the first Class data at time Static reliability metrics are assessments of data reliability based on the inherent attributes of the data, reflecting the fundamental reliability of the data itself. For the first Class data at time Timing reliability metrics; For the first Class data at time Cross-source correlation reliability metrics.

[0071] In a preferred embodiment, the calculation of static reliability basic indicators assesses the basic reliability based on the inherent attributes of the data, focusing on three core dimensions: historical accuracy, data quality, and source authority. The resulting static indicators objectively reflect the inherent reliability of the data itself, providing a benchmark for dynamic evaluation.

[0072] Next, multi-dimensional time series features are embedded to construct a time series reliability index. This index is composed of a trend term and volatility combined with a balance coefficient. The trend term is analyzed for long-term trends through the slope of linear regression within a sliding window, while volatility is processed by the ratio of the standard deviation within the window to the historical maximum standard deviation to reflect short-term stability, thus making up for the shortcomings of traditional assessments that ignore dynamic changes in time series.

[0073] In a preferred embodiment, the specific method for embedding multi-dimensional temporal features is as follows:

[0074] By analyzing time series data to uncover reliability patterns over time, time series reliability metrics can be constructed.

[0075] ,

[0076] in, For trend items, For volatility, This is the balance coefficient between the trend term and volatility;

[0077] To reflect the long-term trend of data reliability, the slope of linear regression within a sliding window is used for calculation:

[0078] ,

[0079] in, Indicates the first Class data at time Trend items, The size of the sliding window; This is a time variable, representing the specific moment within the sliding window, used to iterate through each time point within the window to calculate the trend; This represents the average time within the sliding window, i.e., all times within the window. The average value; Indicates the first Class data at time Reliability score; Let be the mean of the reliability scores for the first type of data within the sliding window;

[0080] To reflect the short-term stability of data reliability, the normalized value of the standard deviation within the window is used:

[0081] ,

[0082] in, Indicates the first Class data at time Volatility indicators; The calculation starts from time 10. At the time Within this sliding window, the first Class Data Reliability Score Standard deviation; Indicates the first The maximum standard deviation obtained by calculating the class of data across all possible sliding windows in its history.

[0083] Finally, a reliability assessment of cross-source data association was conducted, and association indicators were constructed. First, association weights based on the degree of business closeness were set for the associated data sources. Then, the dynamic correlation coefficient of the original data of the two sources was calculated within a sliding window. The results were obtained by summing the reliability scores of the associated data sources, thus overcoming the limitation of traditional technologies that view the reliability of a single data source in isolation.

[0084] In a preferred embodiment, the specific method for assessing the reliability of cross-source data association is as follows:

[0085] By analyzing the correlations between different data sources, a correlation reliability index is constructed:

[0086] ,

[0087] in, Indicates the first Class data at time Cross-source correlation reliability metrics; Representative except the first Other data source categories besides the first type of data, used to iterate through all data related to the first type. Data sources with related categories; For data source and The association weight; Let be the dynamic correlation coefficient between the two data sources at time t; Indicates the first Class data at time The reliability score reflects the first The reliability of the data itself at that moment;

[0088] The Pearson correlation coefficient within a sliding window is used for calculation:

[0089] ,

[0090] in, To indicate at time At that time, the first Class data and the first Dynamic Pearson correlation coefficient for class data; These represent different data source categories; It is a time variable; Indicates the first Class data at time The original data values; Indicates the first Class data at time The original data values; It is the first one in the sliding window The mean of the original data values ​​for the class data; It is the first one in the sliding window The mean of the original data values ​​for the class data; It is the size of the sliding window.

[0091] After calculating the three types of indicators, the comprehensive reliability scores of various data at the corresponding time are formed by merging them according to the comprehensive weight coefficient (determined by subsequent reinforcement learning optimization). This step breaks through the limitations of traditional single-dimensional evaluation through three-dimensional collaborative analysis of static, time series, and cross-source correlation. It achieves accurate and three-dimensional quantification of data reliability, which not only fits the actual characteristics of multi-source data scenarios, but also transforms abstract reliability into specific values. This provides a core basis for the dynamic allocation of weights in subsequent reinforcement learning, avoids the subjectivity of traditional weight allocation, and ensures the quality of fused data from the source, reduces fusion bias, and lays a key foundation for the accuracy of final consumption prediction and the scientific nature of energy decision-making.

[0092] (3) Weight optimization

[0093] Please refer to Figure 3After obtaining the data reliability score, a reinforcement learning environment is built using this as the core input. The weight allocation schemes of different data sources are set as the decision variables of the optimization objective. The optimal weight allocation strategy is generated through iterative training. This process is the core link to realize dynamic weighted fusion of data. It not only builds on the reliability assessment results mentioned above, but also provides intelligent decision support for subsequent data fusion.

[0094] Specifically, reinforcement learning uses minimizing prediction error as its core reward criterion, while incorporating reliability-error correlation penalties and cross-period reward smoothing mechanisms. The former is used to avoid mismatches between the weight allocation of high-reliability data and the actual error improvement, while the latter enhances the stability of the weight strategy over time. Through continuous iteration, reinforcement learning constantly tries different weight combinations, retaining strategies that bring low prediction error and eliminating high-error strategies, ultimately forming an optimal weight allocation logic that dynamically matches the data reliability.

[0095] In a preferred embodiment, the reward criterion for reinforcement learning is specifically as follows:

[0096] The reward criterion for reinforcement learning is centered on "minimizing prediction error," incorporating "reliability-error correlation penalty" and "cross-period reward smoothing mechanism," specifically designed as follows:

[0097] ,

[0098] in, For a moment Reward values ​​for reinforcement learning; For this moment The error in the absorption prediction; This is the reliability-error correlation penalty coefficient; For a moment Reliability-error correlation penalty term; This represents the change in reward at the previous moment.

[0099] In a preferred embodiment, the reliability-error correlation penalty term is calculated as follows:

[0100] ,

[0101] in, Limited to reliability rating only Above the threshold The high reliability of the data source is used to calculate the penalty; For a moment For the first The actual weights assigned to each data category; The theoretically optimal weights are based on reliability. To remove the first After classifying the data, the prediction error is obtained by fusing only other data. This represents the prediction error of the original fused data.

[0102] The significance of this process lies in breaking through the limitations of traditional static weight allocation, which cannot adapt to dynamic changes in data reliability. Through the intelligent optimization capabilities of reinforcement learning, the weight allocation can accurately respond to real-time changes in data reliability while avoiding policy fluctuations caused by short-term random errors. This provides an adaptive decision-making basis for the efficient fusion of multi-source data and directly improves the adaptability of fused data to real-world scenarios.

[0103] (4) Parameter optimization

[0104] After the weight optimization unit outputs the optimal weight allocation strategy, this step uses it as a basis to perform multi-source data weighted fusion and error feedback optimization. This not only completes the practical application from weight decision to data fusion, but also realizes dynamic adjustment of parameters throughout the process through error feedback, forming a closed loop of "weight optimization-data fusion-prediction feedback", which solves the problem of traditional technology lacking a continuous improvement mechanism.

[0105] Specifically, firstly, multi-source data is weighted and fused based on the output dynamic weights, integrating different data sources into comprehensive data according to weight ratios matched to their reliability levels, ensuring that high-reliability data occupies a reasonable dominant position in the fusion result. Then, the comprehensive data is input into the absorption prediction model to generate absorption prediction results. Next, by comparing the prediction results with actual absorption data, the absorption prediction error is quantified and calculated, and key parameters—including the sliding window size in the reliability scoring unit—are adjusted in reverse based on this error. Comprehensive weighting coefficients of static / time-series / cross-source correlation indicators Balance coefficient between time series trend term and volatility It also covers the core parameters of reinforcement learning in the weight optimization unit.

[0106] This process, on the one hand, transforms the weight optimization results into high-quality comprehensive data through weighted fusion, providing accurate input for energy consumption forecasting and directly improving forecast accuracy; on the other hand, it dynamically calibrates parameters through error feedback, making the reliability score more consistent with the actual data characteristics and the weight optimization strategy more adaptable to forecasting needs, driving the continuous iterative optimization of the entire methodology, avoiding the decline in adaptability caused by parameter solidification, ensuring that high forecasting performance is maintained in long-term applications, and providing stable and reliable support for energy consumption decisions.

[0107] Example 2

[0108] Please refer to Figure 4 This embodiment 2 provides a dynamic weighted fusion method for data confidence based on reinforcement learning, including:

[0109] The data acquisition unit is used to collect data from multiple sources.

[0110] The reliability scoring unit is used to first summarize and quantify the static inherent attributes of multi-source data to obtain basic static reliability indicators; then, through multi-dimensional time-series feature embedding and cross-source data association reliability assessment, sub-indicators are constructed; and the three types of indicators—static, time-series, and cross-source association—are integrated to complete the calculation of data reliability scores.

[0111] The weight optimization unit is used to construct a reinforcement learning environment with the data reliability score obtained from the reliability scoring unit as the core input. The weight allocation scheme of different data sources is set as the decision variable of the optimization objective. The minimum of the elimination prediction error is used as the reward criterion for reinforcement learning. The optimal weight allocation strategy is generated through iterative training.

[0112] The parameter optimization unit is used to perform weighted fusion of multi-source data based on the weights output by the weight optimization unit to obtain comprehensive data, which is then input into the absorption prediction model to output the prediction result. By comparing the prediction result with the actual absorption data to quantify the error, the calculation parameters of the data reliability index in the reliability scoring unit and the core parameters of reinforcement learning in the weight optimization unit are dynamically adjusted based on the error.

[0113] Example 3

[0114] This embodiment 3 also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement any step of a data confidence dynamic weighted fusion method based on reinforcement learning.

[0115] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0116] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.

[0117] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A data confidence dynamic weighting fusion method based on reinforcement learning, characterized in that, The method comprises the following steps: S1. Collecting multi-source data; S2. Quantifying the static inherent properties of the multi-source data to obtain static reliability basic indicators; S3. Building a reinforcement learning environment with the data comprehensive reliability score obtained in step S2 as the core input; setting the weight allocation scheme of different data sources as the decision variable of the optimization target; taking the minimum prediction error as the reward criterion of reinforcement learning; and learning and generating the optimal weight allocation strategy through iterative training; S4. Based on the weight output by S3, performing weighted fusion on the multi-source data to obtain comprehensive data, inputting the prediction result output by the prediction model, and quantifying the error between the prediction result and the actual consumption data to dynamically adjust the calculation parameters of the data reliability indicators in step S2 and the core parameters of the reinforcement learning in step S3. The multi-source data in S1 comprises meteorological prediction data, measured meteorological data, and energy data. The reward criterion of reinforcement learning in S3 is: The reward criterion of reinforcement learning takes "minimizing prediction error" as the core, and integrates "reliability-error correlation penalty" and "cross-period reward smoothing mechanism". The specific design is: The calculation method of the reliability-error correlation penalty term is: wherein, is the time reward value of reinforcement learning; is the time forecast error of accommodation; is the reliability-error correlation penalty coefficient; is the time reliability-error correlation penalty term; is the reward change amount representing the previous time The calculation method of the data comprehensive reliability score in S2 is: wherein, defining a penalty for high reliability data sources only above a threshold value ; and for a time instant ; and a weight actually assigned to the data of class ; and a theoretical optimal weight based on reliability ; and a prediction error of the original fused data ; and a combined reliability score of the data of class at time instant ; and 2. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 1, characterized in that, The specific method of multi-dimensional time series feature embedding in S2 is: wherein, is a comprehensive reliability score of the data of the class at the time point ; is a comprehensive weight coefficient, which is trained and optimized by reinforcement learning; is a static reliability index of the data of the class at the time point , which is an evaluation of data reliability from the inherent properties of the data, reflecting the basic reliability degree of the data itself; is a time sequence reliability index of the data of the class at the time point ; is a cross-source correlation reliability index of the data of the class at the time point .

3. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 2, characterized in that, Through time series analysis, the reliability pattern of data over time is mined to construct a time series reliability indicator: The specific method of cross-source data correlation reliability evaluation in S2 is: wherein, is a trend term, is a volatility, is a balancing coefficient of the trend term and the volatility; Reflecting long-term trends in data reliability, the linear regression slope within a sliding window is calculated: wherein, represents the trend item of the data of the th class at time ; is the size of the sliding window; is the time variable, representing the specific time in the sliding window, used to traverse each time point in the window to calculate the trend; is the mean of the time in the sliding window, i.e., the average of all times in the window; represents the comprehensive reliability score of the data of the th class at time ; is the mean of the comprehensive reliability scores of the data of the th class in the sliding window; Reflecting the short-term stability of data reliability, the standard deviation normalized value within the window is adopted: wherein, represents the class data at time a volatility index; is calculated the standard deviation of the class data reliability score over a sliding window from time to time ; and represents the maximum standard deviation obtained from all possible sliding windows in the history of the class data.

4. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 2, characterized in that, By analyzing the correlation between different data sources, an association reliability indicator is constructed: The method comprises the following steps: wherein, represents the first class data at time cross-source association reliability index; represents other data source categories except the first class data, used to traverse all data sources associated with the first class data; is the association weight of the data source and ; is the dynamic correlation coefficient of the two data sources at time t; represents the reliability score of the first class data at time , reflecting the reliability degree of the first class data itself at the time. Pearson correlation coefficient within a sliding window: wherein, is the dynamic Pearson correlation coefficient between the first class data and the second class data at time ; respectively represent different data source classes; is a time variable; represents the original data value of the first class data at time ; represents the original data value of the second class data at time ; is the mean of the original data values of the first class data within the sliding window; is the mean of the original data values of the second class data within the sliding window; is the size of the sliding window.

5. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 3, characterized in that, The calculation parameters of the reliability index in S4 include: sliding window size , comprehensive weight coefficient , and balance coefficient of trend item and volatility .

6. A system for dynamic weighted fusion of data confidence based on reinforcement learning, for implementing the method for dynamic weighted fusion of data confidence based on reinforcement learning according to claim 1, characterized in that, A data collection unit is configured to collect multi-source data; A reliability scoring unit is configured to quantitatively analyze the static inherent properties of the multi-source data to obtain static reliability basic indicators; and to construct sub-indicators through multi-dimensional time series feature embedding and cross-source data correlation reliability evaluation; and to fuse static, time series, and cross-source correlation indicators to calculate the data comprehensive reliability score. A weight optimization unit is configured to build a reinforcement learning environment with the data comprehensive reliability score obtained by the reliability scoring unit as the core input; to set the weight allocation scheme of different data sources as the decision variable of the optimization target; to take the minimum prediction error as the reward criterion of reinforcement learning; and to learn and generate the optimal weight allocation strategy through iterative training. A parameter optimization unit is configured to perform weighted fusion on the multi-source data based on the weight output by the weight optimization unit to obtain comprehensive data, input the prediction result output by the prediction model, compare the prediction result with the actual consumption data to quantify the error, and dynamically adjust the calculation parameters of the data reliability indicators in the reliability scoring unit and the core parameters of the reinforcement learning in the weight optimization unit. The computer program is executed by a processor to implement the reinforcement learning-based data confidence dynamic weighted fusion method of any one of claims 1-5.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Multi-source image fusion method based on enhanced learning

    CN108447041A

  • Method and System for Association and Decision Fusion of Multimodal Inputs

    US20120290526A1