Data confidence coefficient dynamic weighted fusion method and system based on reinforcement learning
By constructing a multi-dimensional data reliability index system and a closed-loop process of reinforcement learning, the weight allocation of multi-source data is dynamically optimized, which solves the problems of single reliability assessment and lack of closed-loop optimization in traditional technologies, improves the accuracy and stability of energy consumption prediction, and supports scientific decision-making for grid dispatch and energy storage layout.
Patent Information
- Application Number
- CN202511716399.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Traditional multi-source data fusion technology suffers from limitations in energy consumption forecasting, including a single reliability assessment dimension, lack of dynamic adaptability in weight allocation, and absence of a closed-loop optimization mechanism. This results in significant deviations in forecasting results, making it difficult to support the scientific validity of decisions such as grid dispatching and energy storage deployment.
We construct a multi-dimensional data reliability index system, combine reinforcement learning to achieve dynamic optimization of data weights, establish a closed-loop process of 'evaluation-fusion-prediction-error feedback', and dynamically adjust the weight allocation to minimize and eliminate prediction errors through multi-dimensional temporal feature embedding and cross-source data association evaluation.
It improves the quality of fused data and the accuracy of absorption forecasts, providing reliable data support for power grid dispatch and energy planning, ensuring continuous optimization of data reliability assessment, weight allocation and fused forecasts, and enhancing the efficient operation of the energy system and the capacity for renewable energy absorption.
Smart Images

Figure CN121167652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of energy prediction, and more particularly relates to a data confidence dynamic weighting fusion method and system based on reinforcement learning. BACKGROUND
[0002] In key scenarios such as energy consumption prediction that rely on multi-source data fusion, the quality of data and the rationality of the fusion strategy directly determine the reliability of the prediction results. However, the current technical system has overall short boards in dealing with actual needs, and it is difficult to meet the requirements of efficient operation of the energy system. On the one hand, the traditional data reliability evaluation dimension is single, and mostly only around isolated indicators such as data historical accuracy and collection accuracy, without considering the correlation stability between multi-source data. However, in actual application, the correlation of different data sources (such as the coupling relationship between meteorological data and photovoltaic output data) has a significant impact on the overall data reliability. Ignoring this dimension will cause the reliability evaluation to be out of touch with the actual scene, making the subsequent fusion decision lose scientific basis.
[0003] On the other hand, existing data fusion generally adopts a static weight distribution mode, that is, the fixed weight of each data source is set in advance and remains unchanged throughout the fusion process. However, the reliability of data sources will change dynamically over time and with environmental changes (such as fluctuations in light data due to seasonal changes, and deviations in monitoring data caused by equipment aging). Static weights cannot adapt to this dynamic nature, often resulting in insufficient weight for high-reliability data and excessive weight for low-reliability data, directly reducing the quality of fused data and leading to large deviations in consumption prediction results.
[0004] More importantly, traditional technologies lack a closed-loop optimization mechanism of "evaluation-fusion-feedback", even if some schemes can adjust the weight, they cannot accurately correct the reliability evaluation parameters and weight distribution strategy based on the comparison results of real-time prediction errors and actual consumption data. This feedback-free operation mode makes data fusion and prediction always in a "passive adaptation" state, making it difficult to continuously improve according to actual operating conditions, and it is difficult to solve the prediction problem of "production and consumption mismatch" in energy consumption prediction, and it is also difficult to support the scientific nature of power grid dispatching and energy storage layout decisions, ultimately affecting the efficient consumption of renewable energy and restricting the process of clean and intelligent transformation of the energy system. SUMMARY
[0005] The present application aims to solve the problems of single reliability evaluation dimension, lack of dynamic adaptability of weight distribution, and lack of closed-loop optimization mechanism in traditional multi-source data fusion technology in the context of energy consumption prediction. By constructing a multi-dimensional data reliability index system, combining reinforcement learning to realize dynamic optimization of data weight, and establishing a closed-loop process of "evaluation-fusion-prediction-error feedback", the quality of fused data and the accuracy of consumption prediction are improved, providing reliable data support for power grid dispatching and energy planning, and helping efficient operation of energy system and large-scale consumption of renewable energy.
[0006] In view of the above defects or improvement needs of the prior art, as a first aspect of the present application, the present application provides a data confidence dynamic weighted fusion method based on reinforcement learning, comprising: S1. Collecting multi-source data; S2. First, the static inherent properties of multi-source data are quantitatively summarized to obtain static reliability basic indicators; then, multi-dimensional time series feature embedding and cross-source data association reliability evaluation are used to construct sub-indicators; and the static, time series and cross-source association indicators are fused to complete the calculation of data reliability score; S3. Taking the data reliability score obtained in step S2 as the core input, a reinforcement learning environment is constructed; the weight distribution scheme of different data sources is set as the decision variable of the optimization target, and the minimization of the consumption prediction error is taken as the reward criterion of reinforcement learning, and the optimal weight distribution strategy is generated through iterative training and learning; S4. Based on the weight output by S3, the multi-source data is weighted and fused to obtain comprehensive data, which is input into the consumption prediction model to output the prediction result; by comparing the quantitative error of the prediction result and the actual consumption data, the calculation parameters of the data reliability indicators in step S2 and the core parameters of the reinforcement learning in step S3 are dynamically adjusted.
[0007] Further, the multi-source data in S1 includes meteorological prediction data, measured meteorological data and energy data.
[0008] Further, the calculation method of the data reliability score in S2 is: , Wherein, is the comprehensive reliability score of the type data at time ; is the comprehensive weight coefficient, which is optimized by reinforcement learning training; is the static reliability index of the type data at time , which is an evaluation of data reliability from the inherent properties of the data, reflecting the basic reliability of the data itself; is the static reliability index of the Class data at time Timing reliability metrics; For the first Class data at time Cross-source correlation reliability metrics.
[0009] Furthermore, the specific method for embedding multi-dimensional temporal features in S2 is as follows: By analyzing time series data to uncover reliability patterns over time, time series reliability metrics can be constructed. , in, For trend items, For volatility, This is the balance coefficient between the trend term and volatility; To reflect the long-term trend of data reliability, the slope of linear regression within a sliding window is used for calculation: , in, Indicates the first Class data at time Trend items, The size of the sliding window; This is a time variable, representing the specific moment within the sliding window, used to iterate through each time point within the window to calculate the trend; This represents the average time within the sliding window, i.e., all times within the window. The average value; Indicates the first Class data at time Reliability score; Let be the mean of the reliability scores for the first type of data within the sliding window; To reflect the short-term stability of data reliability, the normalized value of the standard deviation within the window is used: , in, Indicates the first Class data at time Volatility indicators; The calculation starts from time 10. At the time Within this sliding window, the first Class Data Reliability Score Standard deviation; Indicates the first The maximum standard deviation obtained by calculating the class of data across all possible sliding windows in its history.
[0010] Further, the specific method for evaluating the cross-source data association reliability in S2 is: By analyzing the association between different data sources, the association reliability index is constructed: , Among them, represents the cross-source association reliability index of the first class data at time ; represents other data source categories except the first class data, which is used to traverse all data sources associated with the first class data; is the association weight of data source and ; is the dynamic correlation coefficient of the two data sources at time t; represents the reliability score of the first class data at time , reflecting the reliability of the first class data itself at that time; The Pearson correlation coefficient in the sliding window is calculated: , Among them, is the dynamic Pearson correlation coefficient between the first class data and the first class data at time ; respectively represent different data source categories; is a time variable; represents the original data value of the first class data at time ; represents the original data value of the first class data at time ; is the mean of the original data value of the first class data in the sliding window; is the mean of the original data value of the first class data in the sliding window; is the size of the sliding window.
[0011] Further, the reward criterion of the reinforcement learning in S3 is: The reward criterion of the reinforcement learning takes "minimizing prediction error" as the core, and integrates "reliability-error association penalty" and "cross-period reward smoothing mechanism", and is specifically designed as:
[0012] wherein, is the time reward value of reinforcement learning; is the time forecast error of consumption; is the reliability-error correlation penalty coefficient; is the reliability-error correlation penalty term of the time ; is the reward change amount representing the previous time.
[0013] Further, the calculation method of the reliability-error correlation penalty term is: , wherein, only for high reliability data sources with reliability scores higher than a threshold to calculate the penalty; is the time is the weight actually assigned to the data of the category; is the theoretically optimal weight based on reliability; is the forecast error obtained by fusing only other data after removing the data of the category; is the forecast error of the original fused data.
[0014] Further, the calculation parameters of the reliability index in S4 include: the size of the sliding window , the comprehensive weight coefficient , and the balance coefficient of the trend term and the volatility .
[0015] As a second aspect of the present application, the present application provides a data confidence dynamic weighted fusion method based on reinforcement learning, comprising: a data acquisition unit for completing the acquisition of multi-source data; a reliability scoring unit for first performing general quantitative analysis on the static inherent attributes of multi-source data to obtain static reliability basic indexes; then constructing sub-indexes through multi-dimensional time sequence feature embedding and cross-source data correlation reliability evaluation; and finally fusing the three types of indexes, i.e., static, time sequence, and cross-source correlation, to complete the calculation of data reliability scoring; a weight optimization unit for taking the data reliability score obtained by the reliability scoring unit as the core input, constructing a reinforcement learning environment, setting the weight allocation scheme of different data sources as the decision variable of the optimization target, taking the minimization of the forecast error as the reward criterion of reinforcement learning, and generating the optimal weight allocation strategy through iterative training and learning; A parameter optimization unit is configured to perform weighted fusion on the multi-source data based on the weights output by the weight optimization unit to obtain comprehensive data, and input the prediction model to output a prediction result; and by comparing the prediction result with the actual consumption data to quantify the error, the calculation parameters of the data reliability index in the reliability scoring unit and the core parameters of the reinforcement learning in the weight optimization unit are dynamically adjusted based on the error.
[0016] As a third aspect of the present application, the present application provides a computer readable storage medium having stored thereon a computer program for implementing any step of the data confidence dynamic weighted fusion method based on reinforcement learning.
[0017] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects: 1. The data confidence dynamic weighted fusion method based on reinforcement learning of the present application quantitatively evaluates multi-source data from three dimensions of static inherent attributes, time series dynamic changes and cross-source association to form a data reliability score. The static reliability index comprehensively considers data historical accuracy, quality and source authority, etc. inherent attributes to provide a basic quantitative basis for data reliability; multi-dimensional time series features are embedded to mine long-term change trends and short-term fluctuation rules of data, accurately reflecting the dynamic differences of data reliability over time; and the cross-source data association reliability evaluation is based on dynamic correlation coefficients and association weights to effectively measure the influence of the association stability between different data sources on reliability. This system provides a scientific and comprehensive quantitative scale for subsequent weighted fusion, solves the limitations of traditional single-dimensional evaluation of data reliability, and ensures the quality quantification and comparability of data before fusion.
[0018] 2. The data confidence dynamic weighted fusion method based on reinforcement learning of the present application introduces reinforcement learning into the weight distribution process, takes the minimization of prediction error as the core reward criterion, and simultaneously incorporates reliability-error correlation penalty and cross-period reward smoothing mechanism to dynamically optimize the weight distribution of multi-source data. Reinforcement learning continuously iterates different weight combinations based on the data reliability score, retains the weight distribution strategy that can bring low prediction error, and eliminates the high error strategy. The reliability-error correlation penalty can avoid the situation that the weight distribution of high reliability data does not match the actual error improvement, and the cross-period reward smoothing mechanism enhances the stability of the weight strategy in the time series dimension, so that the weight distribution can accurately match the dynamic changes of data reliability, and the quality and pertinence of the fused data are improved.
[0019] 3. The data confidence dynamic weighting fusion method based on reinforcement learning of the present application realizes continuous optimization of data fusion and prediction through a closed-loop process of multi-source data weighted fusion, consumption prediction and error feedback optimization. The multi-source data is weighted and fused according to the optimized dynamic weights, and the comprehensive data is input into the consumption prediction model to output the prediction result; then the prediction result is compared with the actual consumption data to quantify the error, and the data reliability index calculation parameters and the reinforcement learning core parameters are dynamically adjusted based on the error feedback. This closed-loop process ensures that all links from data reliability evaluation, weight distribution to fusion prediction can be continuously improved according to the actual performance, effectively improving the accuracy and stability of consumption prediction, and providing a more reliable basis for related decisions such as energy consumption. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A flowchart of the data confidence dynamic weighting fusion method based on reinforcement learning of the embodiment of the present application; Figure 2 A reliability score calculation schematic diagram of the embodiment of the present application; Figure 3 A prediction feedback parameter adjustment schematic diagram of the embodiment of the present application; Figure 4 A system unit diagram of the embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical scheme and advantages of the present application clearer and more understandable, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0022] Embodiment 1 Please refer to Figure 1 The embodiment 1 provides a data confidence dynamic weighting fusion method based on reinforcement learning, comprising: S1. Collecting multi-source data; S2. First, the static inherent properties of multi-source data are quantified to obtain static reliability basic indicators; then, sub-indicators are constructed through multi-dimensional time series feature embedding and cross-source data association reliability evaluation; and the static, time series and cross-source association indicators are fused to calculate the data reliability score; S3. Taking the data reliability score obtained in step S2 as the core input, a reinforcement learning environment is constructed; the weight distribution scheme of different data sources is set as the decision variable of the optimization target, and the minimization of consumption prediction error is set as the reward criterion of reinforcement learning, and the optimal weight distribution strategy is generated through iterative training and learning; S4. Based on the weight of the output of S3, the multi-source data is weighted and fused to obtain comprehensive data, and the prediction result is output by inputting the prediction model. By comparing the prediction result with the actual consumption data, the error is quantified, and the calculation parameters of the data reliability index in step S2 and the core parameters of the reinforcement learning in step S3 are dynamically adjusted based on the error.
[0023] This embodiment 1 further expands the above steps.
[0024] (1) Data collection In the initial stage of implementation, the systematic collection of multi-source data needs to be completed first. Combining the actual demand that data reliability directly affects the prediction accuracy in the energy consumption prediction scene, and the problem that traditional technology leads to poor fusion effect due to single data dimension and insufficient dynamic adaptation, the data collected in this step needs to cover key dimensions such as energy production and environmental conditions, providing comprehensive support for subsequent multi-dimensional reliability evaluation and dynamic weight allocation.
[0025] The specific multi-source data collected includes three types of core data: first, meteorological prediction data, which comes from short-term and medium-term forecasts published by meteorological departments, covering parameters such as wind speed, light intensity, and temperature, which can reflect the future environmental trends affecting energy (such as wind power and photovoltaic) production, providing a data basis for subsequent time series feature analysis; second, real-time meteorological data, which is obtained through monitoring equipment at energy sites and surrounding areas, recording real-time environmental parameters such as actual wind speed and light duration, which can be compared with meteorological prediction data to provide measured basis for static reliability evaluation and cross-source correlation analysis; third, energy data, including real-time output of new energy power generation equipment, power grid transmission power, and user-side energy consumption, covering the entire link of energy production-transmission-consumption, which is the core business data for subsequent calculation of data reliability score and verification of fusion effect.
[0026] During the collection process, data needs to be continuously obtained at fixed time intervals (such as 5-15 minutes), and preliminary format unification and missing value completion need to be completed to ensure data integrity and consistency - this processing not only solves the problem of "chaotic data and difficult to use directly" in traditional data collection stage, but also provides standardized data input for subsequent steps of static inherent attribute quantification, time series feature embedding, and cross-source correlation evaluation, ensuring smooth connection of the entire method process.
[0027] (2) Reliability scoring Please refer to Figure 2 After completing the collection of multi-source data and forming a standardized data basis, the data reliability score is calculated by building a multi-dimensional evaluation system, which not only solves the problem of single evaluation dimension in traditional technology, but also provides a quantitative basis for subsequent dynamic weight allocation based on reinforcement learning, and is a key link connecting data collection and intelligent fusion.
[0028] In a preferred embodiment, the data reliability score is calculated as follows: , in, To indicate the first Class data at time The overall reliability score; The overall weighting coefficients are optimized through reinforcement learning training. For the first Class data at time Static reliability metrics are assessments of data reliability based on the inherent attributes of the data, reflecting the fundamental reliability of the data itself. For the first Class data at time Timing reliability metrics; For the first Class data at time Cross-source correlation reliability metrics.
[0029] In a preferred embodiment, the calculation of static reliability basic indicators assesses the basic reliability based on the inherent attributes of the data, focusing on three core dimensions: historical accuracy, data quality, and source authority. The resulting static indicators objectively reflect the inherent reliability of the data itself, providing a benchmark for dynamic evaluation.
[0030] Next, multi-dimensional time series features are embedded to construct a time series reliability index. This index is composed of a trend term and volatility combined with a balance coefficient. The trend term is analyzed for long-term trends through the slope of linear regression within a sliding window, while volatility is processed by the ratio of the standard deviation within the window to the historical maximum standard deviation to reflect short-term stability, thus making up for the shortcomings of traditional assessments that ignore dynamic changes in time series.
[0031] In a preferred embodiment, the specific method for embedding multi-dimensional temporal features is as follows: By analyzing time series data to uncover reliability patterns over time, time series reliability metrics can be constructed. , in, For trend items, For volatility, This is the balance coefficient between the trend term and volatility; To reflect the long-term trend of data reliability, the slope of linear regression within a sliding window is used for calculation: , in, Indicates the first Class data at time Trend items, The size of the sliding window; This is a time variable, representing the specific moment within the sliding window, used to iterate through each time point within the window to calculate the trend; This represents the average time within the sliding window, i.e., all times within the window. The average value; Indicates the first Class data at time Reliability score; Let be the mean of the reliability scores for the first type of data within the sliding window; To reflect the short-term stability of data reliability, the normalized value of the standard deviation within the window is used: , in, Indicates the first Class data at time Volatility indicators; The calculation starts from time 10. At the time Within this sliding window, the first Class Data Reliability Score Standard deviation; Indicates the first The maximum standard deviation obtained by calculating the class of data across all possible sliding windows in its history.
[0032] Finally, a reliability assessment of cross-source data association was conducted, and association indicators were constructed. First, association weights based on the degree of business closeness were set for the associated data sources. Then, the dynamic correlation coefficient of the original data of the two sources was calculated within a sliding window. The results were obtained by summing the reliability scores of the associated data sources, thus overcoming the limitation of traditional technologies that view the reliability of a single data source in isolation.
[0033] In a preferred embodiment, the specific method for assessing the reliability of cross-source data association is as follows: By analyzing the correlations between different data sources, a correlation reliability index is constructed: , in, Indicates the first Class data at time Cross-source correlation reliability metrics; Representative except the first Other data source categories besides the first type of data, used to iterate through all data related to the first type. Data sources with related categories; For data source and The association weight; Dynamic correlation coefficient of two data sources at time t; The reliability score of the first class data at time t, reflecting the reliability of the first class data itself at that time; The Pearson correlation coefficient in the sliding window is calculated: , where, is the dynamic Pearson correlation coefficient of the first class data and the first class data at time t; represent different data source categories respectively; is a time variable; is the original data value of the first class data at time t; is the original data value of the first class data at time t; is the mean of the original data value of the first class data in the sliding window; is the mean of the original data value of the first class data in the sliding window; is the size of the sliding window. After the calculation of the three types of indicators, the comprehensive reliability score of each type of data at the corresponding time is formed by the comprehensive weight coefficient (determined by subsequent reinforcement learning optimization), which breaks through the limitations of traditional single-dimensional evaluation through static, time-series, and cross-source correlation three-dimensional collaborative analysis, realizes the precise stereoscopic quantification of data reliability, not only fits the actual scene characteristics of multi-source data, but also converts abstract reliability into specific numerical values, provides core basis for subsequent reinforcement learning dynamic weight distribution, avoids the subjectivity of traditional weight distribution, at the same time, guarantees the quality of fused data from the source, reduces the fusion bias, and lays a key foundation for the accuracy of the final consumption prediction and the scientific nature of energy decision-making. (3) Weight optimization
[0034] Please refer to After obtaining the data reliability score, the reinforcement learning environment is constructed with this as the core input, and the weight distribution scheme of different data sources is set as the decision variable of the optimization target, and the optimal weight distribution strategy is generated through iterative training, which is the core link of realizing dynamic weighted fusion of data, not only connecting the reliability evaluation results of the previous text, but also providing intelligent decision support for subsequent data fusion.
[0035] Figure 3
[0036] Specifically, reinforcement learning uses minimizing prediction error as its core reward criterion, while incorporating reliability-error correlation penalties and cross-period reward smoothing mechanisms. The former is used to avoid mismatches between the weight allocation of high-reliability data and the actual error improvement, while the latter enhances the stability of the weight strategy over time. Through continuous iteration, reinforcement learning constantly tries different weight combinations, retaining strategies that bring low prediction error and eliminating high-error strategies, ultimately forming an optimal weight allocation logic that dynamically matches the data reliability.
[0037] In a preferred embodiment, the reward criterion for reinforcement learning is specifically as follows: The reward criterion for reinforcement learning is centered on "minimizing prediction error," incorporating "reliability-error correlation penalty" and "cross-period reward smoothing mechanism," specifically designed as follows: , in, For a moment Reward values for reinforcement learning; For this moment The error in the absorption prediction; This is the reliability-error correlation penalty coefficient; For a moment Reliability-error correlation penalty term; This represents the change in reward at the previous moment.
[0038] In a preferred embodiment, the reliability-error correlation penalty term is calculated as follows: , in, Limited to reliability rating only Above the threshold The high reliability of the data source is used to calculate the penalty; For a moment For the first The actual weights assigned to each data category; The theoretically optimal weights are based on reliability. To remove the first After classifying the data, the prediction error is obtained by fusing only other data. This represents the prediction error of the original fused data.
[0039] The significance of this process lies in breaking through the limitations of traditional static weight allocation, which cannot adapt to dynamic changes in data reliability. Through the intelligent optimization capabilities of reinforcement learning, the weight allocation can accurately respond to real-time changes in data reliability while avoiding policy fluctuations caused by short-term random errors. This provides an adaptive decision-making basis for the efficient fusion of multi-source data and directly improves the adaptability of fused data to real-world scenarios.
[0040] (4) Parameter optimization After the weight optimization unit outputs the optimal weight distribution strategy, this step performs multi-source data weighted fusion and error feedback optimization based on it, not only completing the landing application from weight decision to data fusion, but also achieving dynamic adjustment of the whole process parameters through error feedback, forming a closed loop of "weight optimization-data fusion-prediction feedback", solving the problem of lack of continuous improvement mechanism in traditional technology.
[0041] Specifically, first, the multi-source data is weighted and fused based on the output dynamic weight, and the different data sources are integrated into comprehensive data according to the weight proportion matched with the reliability, ensuring that high reliability data occupies a reasonable dominant position in the fusion result; then the comprehensive data is input into the consumption prediction model to generate the consumption prediction result. Then by comparing the prediction result with the actual consumption data, the consumption prediction error is quantitatively calculated, and based on the error, the key parameters are adjusted reversely, including the sliding window size , the comprehensive weight coefficient of static / time series / cross-source correlation indicators , the balance coefficient of time series trend item and volatility , also covering the core parameters of reinforcement learning in the weight optimization unit.
[0042] This process, on the one hand, converts the weight optimization results into high-quality comprehensive data through weighted fusion, providing accurate input for consumption prediction and directly improving prediction accuracy; on the other hand, through error feedback dynamic calibration of parameters, the reliability score is more in line with the actual data characteristics, and the weight optimization strategy is more suitable for prediction needs, promoting the continuous iterative optimization of the whole method system, avoiding the decline of adaptability due to parameter fixation, and ensuring that high prediction performance is always maintained in long-term application, providing stable and reliable support for energy consumption decision-making.
[0043] Embodiment 2 Please refer to Figure 4 , this embodiment 2 provides a data confidence dynamic weighted fusion method based on reinforcement learning, comprising: a data acquisition unit for completing the acquisition of multi-source data; a reliability scoring unit for first performing general quantification on the static inherent attributes of multi-source data to obtain static reliability basic indicators; then constructing sub-indicators through multi-dimensional time series feature embedding and cross-source data correlation reliability evaluation; and fusing static, time series, and cross-source correlation indicators to complete the calculation of data reliability score; a weight optimization unit for taking the data reliability score obtained by the reliability scoring unit as the core input to construct a reinforcement learning environment; setting the weight distribution scheme of different data sources as the decision variable of the optimization target, taking the minimization of consumption prediction error as the reward criterion of reinforcement learning, and generating the optimal weight distribution strategy through iterative training and learning; The parameter optimization unit is configured to perform weighted fusion on the multi-source data based on the weights output by the weight optimization unit to obtain comprehensive data, and input the prediction model to output a prediction result; by comparing the prediction result with the actual consumption data to quantify the error, the calculation parameters of the data reliability index in the reliability scoring unit and the core parameters of the reinforcement learning in the weight optimization unit are dynamically adjusted based on the error.
[0044] Embodiment 3 The embodiment 3 also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement any step of the data confidence dynamic weighted fusion method based on reinforcement learning.
[0045] The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0046] For the computer readable storage medium provided in the present application, refer to the above method embodiments, which will not be repeated herein.
[0047] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data confidence dynamic weighting fusion method based on reinforcement learning, characterized in that, The method comprises the following steps: S1. Collecting multi-source data; S2. First, the static inherent properties of multi-source data are quantitatively summarized to obtain static reliability basic indicators; then, sub-indicators are constructed through multi-dimensional time sequence feature embedding and cross-source data association reliability evaluation; and finally, the calculation of data reliability score is completed by fusing three types of indicators, i.e., static, time sequence and cross-source association; S3. Taking the data reliability score obtained in step S2 as the core input, a reinforcement learning environment is constructed; the weight allocation scheme of different data sources is set as the decision variable of the optimization target; the minimum prediction error is taken as the reward criterion of reinforcement learning; and the optimal weight allocation strategy is generated through iterative training and learning; S4. Based on the weight output by S3, the multi-source data is weighted and fused to obtain comprehensive data, which is input into a consumption prediction model to output a prediction result; by comparing the prediction result with the actual consumption data to quantify the error, the calculation parameters of the data reliability indicators in step S2 and the core parameters of the reinforcement learning in step S3 are dynamically adjusted. The multi-source data in S1 comprises meteorological prediction data, actual meteorological data and energy data.
2. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 1, characterized in that, The calculation method of the data reliability score in S2 is:
3. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 1, characterized in that, The specific method of the multi-dimensional time sequence feature embedding in S2 is: , wherein, is a static reliability index of the first class data at time point, which is an evaluation of data reliability from the inherent properties of data, reflecting the basic reliability of the data itself; is a comprehensive weight coefficient, which is optimized by reinforcement learning training; is a static reliability index of the first class data at time point, which is an evaluation of data reliability from the inherent properties of data, reflecting the basic reliability of the data itself; is a time sequence reliability index of the first class data at time point; is a cross-source correlation reliability index of the first class data at time point.
4. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 3, characterized in that, The time sequence reliability indicators are constructed by mining the reliability patterns of data over time through time sequence analysis: The specific method of the cross-source data association reliability evaluation in S2 is: , wherein, is a trend term, is a volatility, is a balancing coefficient of the trend term and the volatility; Reflecting long-term trends in data reliability, the linear regression slope within a sliding window is calculated: , wherein, represents the trend item of the class data at time ; is the size of the sliding window; is the time variable, representing the specific time within the sliding window, used to traverse each time point within the window to calculate the trend; is the mean of the time within the sliding window, i.e., the average of all times within the window; represents the reliability score of the class data at time ; is the mean of the reliability scores of the class data within the sliding window; Reflecting the short-term stability of data reliability, the standard deviation normalized value within the window is adopted: , wherein, represents the class data volatility index at time ; is the standard deviation of the class data reliability score over the sliding window from time to time ; represents the maximum standard deviation of the class data over all possible sliding windows in its history.
5. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 3, characterized in that, The association reliability indicators are constructed by analyzing the association between different data sources: The reward criterion of the reinforcement learning in S3 is: , wherein, represents the first class data at time cross-source association reliability index; represents other data source categories except the first class data, used to traverse all data sources associated with the first class data; is the association weight of the data source and ; is the dynamic correlation coefficient of the two data sources at time t; represents the reliability score of the first class data at time , reflecting the reliability degree of the first class data itself at the time. Pearson correlation coefficient within a sliding window: , wherein, is the time variable; is the time variable; is the time variable; is the time variable; respectively represent different data source categories; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable; is the time variable.
6. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 1, characterized in that, The reward criterion of the reinforcement learning takes "minimization of consumption prediction error" as the core, and incorporates "reliability-error association penalty" and "cross-period reward smoothing mechanism"; and the specific design is: The calculation method of the reliability-error association penalty term is: , wherein, is the time reward value of reinforcement learning; is the time forecast error of accommodation; is the reliability-error correlation penalty coefficient; is the time reliability-error correlation penalty term; is the reward change amount indicating the previous time.
7. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 6, characterized in that, The method comprises the following steps: , wherein, defining a penalty only for high reliability data sources above a threshold ; is the time instant is the weight actually assigned to the data of class ; is the theoretically optimal weight based on reliability ; is the prediction error obtained by fusing only the other data after removing the data of class ; is the prediction error of the original fused data.
8. The data confidence dynamic weighted fusion method based on reinforcement learning according to claim 1, characterized in that, The calculation parameters of the reliability index in S4 include: sliding window size , comprehensive weight coefficient , and balance coefficient of trend item and volatility .
9. A data confidence dynamic weighting fusion method based on reinforcement learning, characterized in that, A data collection unit is configured to collect multi-source data; A reliability scoring unit is configured to first quantitatively summarize the static inherent properties of multi-source data to obtain static reliability basic indicators; then, sub-indicators are constructed through multi-dimensional time sequence feature embedding and cross-source data association reliability evaluation; and finally, the calculation of data reliability score is completed by fusing three types of indicators, i.e., static, time sequence and cross-source association; A weight optimization unit is configured to take the data reliability score obtained by the reliability scoring unit as the core input to construct a reinforcement learning environment; set the weight allocation scheme of different data sources as the decision variable of the optimization target; take the minimization of consumption prediction error as the reward criterion of reinforcement learning; and generate the optimal weight allocation strategy through iterative training and learning; A parameter optimization unit is configured to perform weighted fusion on multi-source data based on the weight output by the weight optimization unit to obtain comprehensive data, which is input into a consumption prediction model to output a prediction result; by comparing the prediction result with the actual consumption data to quantify the error, the calculation parameters of the data reliability indicators in the reliability scoring unit and the core parameters of the reinforcement learning in the weight optimization unit are dynamically adjusted. The computer program is executed by a processor to implement the reinforcement learning-based data confidence dynamic weighted fusion method of any one of claims 1-8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that,
Citation Information
Patent Citations
Multi-source image fusion method based on enhanced learning
CN108447041A
Seawater desalination load-containing energy consumption scheduling method and device
CN117200349A
New energy output evaluation system based on artificial intelligence meteorological large model
CN120181636A
New energy consumption measuring and calculating method considering hydrogen production process
CN120258482A
Multi-source heterogeneous data fusion method and system
CN120449088A