Time sequence feature enhancement system and early warning method for heterogeneous data fusion

By using high-frequency electrical waveform analysis and chaotic feature extraction, the problem of coarse information granularity in macroscopic electricity consumption data has been solved, enabling early and sensitive perception of user behavior patterns and cross-energy behavior cross-verification, thereby improving the early warning timeliness and prediction accuracy of multi-source data fusion systems.

CN121614829APending Publication Date: 2026-03-06GUANGAN POWER SUPPLY COMPANY STATE GRID SICHUANELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511807817.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

In existing technologies, the information granularity of macro-level electricity consumption data is coarse, which makes multi-source fusion systems insensitive to subtle changes in user behavior patterns, resulting in poor early warning timeliness and an inability to deeply reveal the intrinsic relationship between electricity consumption behavior and other behaviors, thus limiting the depth and effectiveness of multi-source data fusion.

Method used

By introducing high-frequency electrical waveform analysis and chaotic feature extraction, raw waveform data of voltage and current and water usage data are collected, and anomaly cleaning and spatiotemporal alignment are performed. Non-invasive load decomposition features such as load dynamic fractal dimension and transient Lyapunov exponent are constructed. Combined with the water-electricity coupling coefficient, a long short-term memory network with attention mechanism is used for time series prediction and feature optimization.

Benefits of technology

It achieves microscopic early sensitivity to user behavior patterns, reduces warning delay, improves the robustness of state recognition and prediction accuracy, and enables cross-energy behavior cross-validation, realizing three-dimensional diagnosis from micro to macro.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_28
    Figure QLYQS_28
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a time sequence feature enhancement system and early warning method for heterogeneous data fusion, and the system comprises a data perception and access module which is used for collecting heterogeneous time sequence data from different types of sensors and data platforms; the preprocessing and space-time alignment module is used for performing exception cleaning, resampling and space-time alignment on the heterogeneous time series data to generate an initial fusion data block; the multi-dimensional feature vector construction module is used for extracting and calculating a technical feature vector of a fixed dimension for each target object from the initial fusion data block; and the time sequence prediction and feature optimization module is used for carrying out time sequence prediction on the technical feature vector by utilizing a long short-term memory network based on an attention mechanism, and dynamically optimizing the weight of each technical feature in the technical feature vector according to a prediction result. According to the method, by constructing an electricity-water multi-source heterogeneous data fusion and associated feature extraction scheme, cross-energy behavior cross validation is realized, and the robustness of state recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a time-series feature enhancement system and early warning method for heterogeneous data fusion. Background Technology

[0002] With the rapid development of IoT technology and big data analytics, user behavior monitoring and analysis based on the fusion of multi-source heterogeneous data has become a key technological direction in areas such as regional governance and precision services. Among these, electricity consumption data provided by smart meters, due to its high collection density, wide coverage, and strong correlation with user activities and economic conditions, is often considered the core data source in such multi-source fusion systems. Complementing this are time-series data from smart water meters and gas meters, as well as event-based data (such as population, employment, and medical events) from government platforms.

[0003] Currently, mainstream technical solutions typically follow this process: collecting macro-level electricity consumption statistics (such as average daily electricity consumption, active power, etc.) and other auxiliary data from users through a data gateway; preprocessing these multi-source data, such as cleaning and alignment; then constructing core feature indicators (such as electricity consumption volatility, deviation rate, etc.) based on the processed macro-level electricity consumption data, and combining them with auxiliary data features; finally, inputting the resulting fused feature vector into a machine learning model for training and prediction.

[0004] However, after in-depth analysis and practice, the applicant discovered that the overall effectiveness of this type of technical solution is limited by the inherent technical bottleneck of its core data source: macro-level electricity consumption data. Due to the significant information loss inherent in macro-level electricity consumption data, even with the integration of other data, the overall system's sensing sensitivity and predictive accuracy are difficult to achieve as expected. Specific technical bottlenecks are manifested in:

[0005] First, the coarse-grained information from the core data source renders the entire fusion system insensitive to subtle changes in behavioral patterns. Existing solutions rely on macro-level data such as power and electricity consumption, which are a mixture and average of the operating states of all appliances within the household, losing detailed information about the load composition. When subtle, indicative changes occur in user behavior (e.g., ceasing the use of a small refrigerator or replacing a high-power productivity tool with a low-power device), the resulting weak signal changes are drowned out by noise from other major electrical appliances. This makes any subsequently constructed feature vectors, whether used alone or fused with other data, like drawing on a blurry film, failing to depict early, accurate behavioral pattern changes.

[0006] Second, the delayed response to core features hinders the timeliness of the entire integrated system's early warning system. Because it relies on periodic electricity consumption statistics, features based on macro data can only be identified by the system when changes in user behavior accumulate to a point where they significantly impact the total statistical data. This inherent delay of several days or even weeks means that even if government data (such as unemployment registration) has recorded an event, the system cannot obtain timely and synchronous confirmation from the electricity consumption side, resulting in a lag in the entire integrated response process.

[0007] Third, the inherent connections between electricity consumption and other behaviors are obscured, making it difficult to construct deep-level integrated features. Macro-level electricity consumption data cannot reveal why electricity consumption changes. For example, a household with stable water consumption but decreased electricity consumption may correspond to drastically different behavioral patterns than a household with simultaneous decreases in both water and electricity consumption (the former may simply be energy saving, while the latter may be due to population decline). However, because macro-level electricity consumption data cannot provide more in-depth information on appliance usage, the system struggles to construct highly discriminative technical features like the water-electricity coupling coefficient, which profoundly reflects the correlation between energy use patterns, thus limiting the depth and effectiveness of multi-source data fusion.

[0008] In summary, the fundamental dilemma of existing technologies lies in the insufficient information dimensionality and quality of macro-level electricity consumption data, which serves as the core data source for multi-source fusion systems, leading to a low performance ceiling for the entire system. Despite attempts by those skilled in the art to introduce more auxiliary data and perform complex model optimizations, they have consistently failed to address the issues of sensitivity, timeliness, and discriminative power of core features at the source of the data.

[0009] Therefore, there is an urgent need in this field for a novel technical approach: to innovate the core data source itself and extract new, high-quality feature vectors from more fundamental and abundant electrical signals. This would lay a solid technical foundation for building a truly efficient and accurate multi-source heterogeneous data fusion system. This is precisely the core technical problem that this invention aims to solve. Summary of the Invention

[0010] The purpose of this invention is to overcome the shortcomings of the prior art and provide a time-series feature enhancement system and early warning method for heterogeneous data fusion. By introducing high-frequency electrical waveform analysis and chaotic feature extraction, the quality of core data sources is fundamentally improved, thereby solving the technical problems of existing systems being insensitive to changes in user behavior patterns and having high early warning delays.

[0011] This invention is achieved through the following technical solution:

[0012] A time-series feature enhancement system for heterogeneous data fusion includes:

[0013] The data sensing and access module is used to collect heterogeneous time-series data from different types of sensors and data platforms through various communication protocols. The heterogeneous time-series data includes raw voltage and current waveform data collected from power lines through a high-frequency sampling circuit and time-series water consumption data collected from the water system through an Internet of Things protocol.

[0014] The preprocessing and spatiotemporal alignment module is communicatively connected to the data sensing and access module. It is used to perform anomaly cleaning, resampling and spatiotemporal alignment on the heterogeneous time-series data, associate power data and water data under a unified timestamp, and generate an initial fused data block.

[0015] A multi-dimensional feature vector construction module, which is communicatively connected to the preprocessing and spatiotemporal alignment module, is used to extract and calculate a fixed-dimensional technical feature vector for each target object from the initial fused data block. The technical feature vector includes a non-intrusive load decomposition feature based on voltage-current trajectory chaotic analysis and a correlation feature calculated based on multi-source data in the initial fused data block.

[0016] The temporal prediction and feature optimization module is communicatively connected to the multidimensional feature vector construction module. It is used to perform temporal prediction on the technical feature vector using a long short-term memory network based on an attention mechanism, and dynamically optimize the weights of each technical feature in the technical feature vector based on the prediction results.

[0017] As an optimization, the non-invasive load decomposition feature based on voltage-current trajectory chaotic analysis includes the load dynamic fractal dimension, which is calculated through the following steps:

[0018] A1. At a sampling rate of not less than 10kHz, synchronously acquire the instantaneous voltage value sequence U(t) and the instantaneous current value sequence I(t) of the total inlet of the target object;

[0019] A2. Using one power frequency cycle as a window, map the instantaneous voltage value sequence U(t) and the instantaneous current value sequence I(t) into a voltage-current trajectory diagram on a two-dimensional plane;

[0020] A3. The fractal dimension D of the voltage-current trajectory is calculated using the box counting method. The calculation formula is as follows: ,in, Let's say it's the side length of the box. Minimum number of boxes required to cover the voltage-current trace;

[0021] A4. The calculated fractal dimension D is used as the dynamic fractal dimension of the load.

[0022] As an optimization, the non-invasive load decomposition feature based on voltage-current trajectory chaotic analysis also includes the transient Lyapunov exponent, which is calculated through the following steps:

[0023] B1. The time series data on the voltage-current trajectory diagram Embedded into an m-dimensional phase space, constructing a phase space vector. Where τ is the time delay, , Let m represent the voltage and current values ​​corresponding to the i-th point on the voltage-current trajectory diagram, respectively, where m is an integer greater than 1.

[0024] B2. For each reference point Y(j) in the phase space, find the nearest neighbor point to the reference point among all phase points other than the reference point. ,satisfy Where d0 is the initial distance, and , The average period;

[0025] B3. Track the reference point Y(j) and its nearest neighbor. Distance after k evolutionary steps And calculate the logarithmic growth rate of the reference point Y(j), using the following formula: ;

[0026] B4. Through formula Calculate the transient Lyapunov exponent , where M is the total number of all valid index pairs (j, k).

[0027] As an optimization, the specific process of outlier cleaning in the preprocessing and spatiotemporal alignment module is as follows:

[0028] For the original voltage and current waveform data, a sliding time window of length L is set, and the mean value of the waveform effective value within the sliding time window is calculated. and standard deviation Identify and remove waveforms with valid values ​​exceeding the limit. The entire waveform segment corresponding to the data points within the range; for the data gaps after removal, linear interpolation of the effective waveform segments before and after within the sliding time window is used to fill them;

[0029] For the user's historical water consumption data, construct upper and lower limit thresholds for the user's water consumption; identify and remove data points whose instantaneous water consumption exceeds the upper and lower limit thresholds; fill the data gaps after removal by using water consumption of zero or the average of the effective water consumption data before and after.

[0030] As an optimization, the specific process of outlier cleaning in the preprocessing and spatiotemporal alignment module is as follows:

[0031] For the original voltage and current waveform data, a sliding time window of length L is set, and the mean value of the waveform effective value within the sliding time window is calculated. and standard deviation Identify and remove waveforms with valid values ​​exceeding the limit. The entire waveform segment corresponding to the data points within the range; for the data gaps after removal, linear interpolation of the effective waveform segments before and after within the sliding time window is used to fill them;

[0032] For time-series water use data, an anomaly detection algorithm based on SH-ESD is used to identify and remove global and local anomalies in the time-series water use data; the data gaps after removal are filled by a Kalman filter algorithm based on a time series prediction model.

[0033] As an optimization, the technical feature vector also includes electricity load fluctuation entropy and water-electricity coupling coefficient, wherein the electricity load fluctuation entropy is the Shannon entropy value calculated based on the daily load curve; the correlation feature is the water-electricity coupling coefficient, which is obtained by calculating the Pearson correlation coefficient within a specific time period based on the time-series electricity data and time-series water data in the initial fused data block.

[0034] As an optimization, the time-series prediction and feature optimization module further includes a dynamic weight allocator and a contribution feedback unit. The dynamic weight allocator is positioned between the multi-dimensional feature vector construction module and the attention-based long short-term memory network, and is used to assign appropriate weights to each technical feature in the technical feature vector. The contribution feedback unit receives the prediction results output by the attention-based long short-term memory network and subsequent actual data; calculates the contribution of each technical feature based on the deviation between the prediction results and the actual data; and sends the contribution to the dynamic weight allocator to update the weights of each technical feature in the next round of prediction.

[0035] As an optimization, the contribution is calculated using the Shapley value method, the specific steps of which include:

[0036] S1. Let F be the set of all N technical features of the technical feature vector, and let M be the long short-term memory network based on the attention mechanism.

[0037] S2. For the target feature i whose contribution needs to be calculated, perform the following operations:

[0038] S2.1, Traverse the feature subset S in set F that does not contain the target feature i. ;

[0039] S2.2. For each of the feature subsets S, perform the following steps:

[0040] S2.2.1 Input the technical features of the feature subset S into the prediction model M to obtain the first predicted value V(S);

[0041] S2.2.2 Input the technical feature data of the union S∪{i} of the feature subset S and the target feature i into the model M to obtain the second predicted value V(S∪{i});

[0042] S2.2.3 Calculate the marginal contribution of the target feature i to the feature subset S. : ;

[0043] S3. Based on the marginal contribution of all feature subsets S, calculate the Shapley value of target feature i using the following formula. And the Shapley value is used as the contribution of this target feature:

[0044] ;

[0045] Where Σ represents the summation of all feature subsets S that do not contain the target feature i, |S| is the size of the feature subset S, i.e. the number of features contained in the feature subset S, and |F| is the total number of features.

[0046] This invention also discloses a user status early warning method based on heterogeneous data fusion. This method uses a time-series feature enhancement system for heterogeneous data fusion as described above, and includes the following steps:

[0047] The time-series feature enhancement system is used to continuously acquire and process the technical feature vectors of multiple users in the target area;

[0048] Based on the temporal changes of the technical feature vector, at least one target user whose technical feature vector changes conform to a preset risk pattern is identified through a set of multi-level risk identification rules.

[0049] Generate and output early warning information corresponding to the target user and the risk mode.

[0050] As an optimization, the preset risk mode includes a livelihood security risk mode. When any level of the rule in the multi-level risk identification rule set is met, the user is determined to meet the livelihood security risk mode. The multi-level risk identification rule set includes a first level, a second level, and a third level. The first level is a core chaotic feature sharpening rule, which includes triggering a determination when a user's load dynamic fractal dimension continuously decreases beyond a first threshold within a preset time period, and the transient Lyapunov exponent simultaneously changes significantly beyond the corresponding threshold. The second level is a multi-dimensional energy consumption mode attenuation rule, which includes triggering a determination when both a user's load dynamic fractal dimension and electricity load fluctuation entropy continuously decrease beyond their respective thresholds within a preset time period. The third level is a cross-energy correlation rupture rule, which includes triggering a determination when a user's water-electricity coupling coefficient decreases beyond a second threshold within a preset time period, and the electricity load fluctuation entropy decreases beyond a third threshold.

[0051] The warning information includes the user's identification information, the triggered risk characteristics, and the corresponding level.

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] By introducing non-intrusive load decomposition features based on high-frequency VI trajectory chaotic analysis, early and sensitive microscopic perception of user behavior patterns is achieved, solving the technical bottlenecks of high response delay and insufficient sensitivity of traditional macroscopic power consumption data.

[0054] By constructing a scheme for the fusion and correlation feature extraction of multi-source heterogeneous data from electricity and water, cross-energy behavior cross-validation was achieved, fundamentally reducing the false alarm rate caused by analysis of a single data source and improving the robustness of state identification.

[0055] By adopting a dynamic feature weight optimization scheme based on Shapley values, the prediction model achieves self-adaptation and self-evolution, enabling the system to automatically focus on key features and continuously improve prediction accuracy and the system's intelligence level.

[0056] By designing a multi-level risk identification rule set scheme, a three-dimensional diagnosis was achieved, from microscopic physical signals to macroscopic behavioral patterns, and then to the stability of cross-energy systems, thus realizing full coverage of public welfare risks from early warning to precise assessment. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments. The illustrative embodiments and descriptions of this invention are only used to explain this invention and are not intended to limit this invention.

[0058] This embodiment 1 provides a time-series feature enhancement system for heterogeneous data fusion, including:

[0059] The data sensing and access module is used to collect heterogeneous time-series data from different types of sensors and data platforms through various communication protocols. The heterogeneous time-series data includes raw voltage and current waveform data collected from power lines through a high-frequency sampling circuit and time-series water consumption data collected from the water system through an Internet of Things protocol.

[0060] The preprocessing and spatiotemporal alignment module is communicatively connected to the data sensing and access module. It is used to perform anomaly cleaning, resampling and spatiotemporal alignment on the heterogeneous time-series data, associate power data and water data under a unified timestamp, and generate an initial fused data block.

[0061] A multi-dimensional feature vector construction module, which is communicatively connected to the preprocessing and spatiotemporal alignment module, is used to extract and calculate a fixed-dimensional technical feature vector for each target object from the initial fused data block. The technical feature vector includes a non-intrusive load decomposition feature based on voltage-current trajectory chaotic analysis.

[0062] The temporal prediction and feature optimization module is communicatively connected to the multidimensional feature vector construction module. It is used to perform temporal prediction on the technical feature vector using a long short-term memory network based on an attention mechanism, and dynamically optimize the weights of each technical feature in the technical feature vector based on the prediction results.

[0063] Next, we will introduce the functions and implementation principles of these modules in detail.

[0064] I. Data Sensing and Access Module

[0065] The data perception and access module includes the following core units:

[0066] 1. High-frequency data acquisition unit

[0067] This unit forms the basis for acquiring deep electrical characteristics. It is directly coupled to the power line via a high-precision ADC (Analog-to-Digital Converter) chip and its associated signal conditioning circuitry. The unit is configured to simultaneously acquire the instantaneous voltage sequence U(t) and instantaneous current sequence I(t) on three-phase or single-phase lines A, B, and C at a sampling rate of at least 10 kHz. This high-frequency sampling rate is sufficient to capture minute changes in the current waveform caused by events such as electrical switching and arcing, providing raw data for subsequent chaotic analysis of the VI trajectory.

[0068] 2. Communication Protocol Adaptation Unit

[0069] This unit is responsible for communicating with various IoT meters, resolving data heterogeneity issues, and is key to achieving synchronous acquisition of water and electricity data. It integrates multiple standard communication protocol parsers, enabling parallel communication. For electricity data, the unit communicates with smart meters via an RS-485 bus, following the DL / T645-2007 protocol, reading their metering data (such as hourly active energy) as supplementary and calibration data for high-frequency waveforms. For water data, the unit communicates with smart water meters via the same RS-485 physical interface or MBus interface, but switches to the CJ / T188-2004 water meter communication protocol, collecting time-series water consumption data at preset intervals (e.g., every 15 minutes).

[0070] 3. Data encryption and secure transmission unit

[0071] To protect user privacy and data security, this unit encrypts the collected raw data (especially high-frequency waveform data) before it exits the network.

[0072] In a preferred embodiment, the unit encrypts the data using the national standard SM4 algorithm and stably transmits the encrypted data stream to a remote preprocessing and spatiotemporal alignment module by establishing an IPSec VPN tunnel. This effectively prevents the data from being stolen or tampered with during transmission.

[0073] The specific workflow is as follows:

[0074] During operation, this module concurrently performs multiple tasks: the high-frequency data acquisition unit continuously buffers voltage and current waveforms; the communication protocol adaptation unit polls devices such as smart water meters at preset intervals. All data, after being accurately timestamped and tagged with its source, is sent to downstream modules by the secure transmission unit.

[0075] Through the above design, the data perception and access module successfully achieved unified, secure, and high-speed access to heterogeneous data sources, laying a reliable data foundation for the subsequent in-depth analysis of the entire system.

[0076] II. Preprocessing and Spatiotemporal Alignment Module

[0077] This module receives raw, messy, heterogeneous data and outputs a clean, well-organized initial fused data block suitable for in-depth analysis. Its core workflow is as follows:

[0078] Step 1: Outlier Cleaning

[0079] For the raw voltage and current waveform data, the module sets a sliding time window of length L (e.g., L = 100 sampling points). The mean μ and standard deviation σ of the effective values ​​of the voltage and current waveforms within the window are calculated. Data points whose effective values ​​exceed the range [μ - 3σ, μ + 3σ] are automatically identified and removed, along with their corresponding waveform segments, to eliminate transient pulse interference. Subsequently, the data gaps caused by the removal are filled using a linear interpolation algorithm between the preceding and following effective waveform segments within the window, ensuring waveform continuity.

[0080] For water consumption data, firstly, based on the user's historical water consumption data for at least the past three months, a lower limit threshold for their normal water consumption is constructed using the percentile method (e.g., taking the 5th percentile and 95th percentile). With upper limit threshold The system identifies and removes those with instantaneous water consumption values ​​exceeding [a certain threshold]. Identify obvious anomalies within the range (such as peak values ​​caused by pipe bursts or zero values ​​caused by complete sensor failure). For data gaps after elimination, fill them in using either zero water consumption (judged as no water use) or the arithmetic mean of valid water consumption data before and after, based on business logic.

[0081] Alternatively, for water consumption data, a more accurate seasonal hybrid ESD algorithm can be used. This algorithm can simultaneously consider the trend, seasonality (such as daily and weekly water consumption peaks), and residuals of water consumption data, thereby accurately identifying various outliers, including global outliers and local mutation points. After removing outliers, the module uses a Kalman filter algorithm based on a time series prediction model for data imputation. By combining historical patterns of water consumption (state prediction) with current limited observations, the Kalman filter dynamically and optimally estimates the most probable values ​​of missing points. Compared to simple mean imputation, it better preserves the time series characteristics of the data, and the imputation results are closer to the actual situation.

[0082] Step 2: Data resampling and spatiotemporal alignment:

[0083] After cleaning, the timestamp sequence of high-frequency electricity data is used as the benchmark. For lower-frequency water consumption data, a forward-filling algorithm is used for resampling, that is, each high-frequency electricity timestamp is assigned a water consumption value, which uses the most recent valid water consumption reading. Finally, the cleaned and aligned electricity data and water consumption data are associated with a joint primary key using the user's unified identifier and the precise timestamp to generate an initial fused data block.

[0084] It should be noted that the determination of terms such as "high frequency" and "low frequency" in this invention can be based on the actual situation. That is, a threshold is set according to the actual scenario. Term above the threshold is considered high frequency, and term below the threshold is considered low frequency. This will not be elaborated further here.

[0085] III. Multidimensional Feature Vector Construction Module

[0086] The multi-dimensional feature vector construction module is the core engine of this invention. It communicates with the preprocessing and spatiotemporal alignment module and is responsible for extracting and calculating a fixed-dimensional technical feature vector for each target object (such as a household user) from the regularized initial fused data block. This vector is no longer a simple electricity consumption statistic, but a multi-dimensional behavioral profile from micro to macro and from electricity to cross-energy.

[0087] More specifically, the technical feature vector includes deep electrical features, macro-level electricity consumption behavior features, and cross-energy correlation features.

[0088] 1. Deep Electrical Feature Extraction: Chaotic Analysis Based on VI Trajectory

[0089] Traditional solutions rely solely on macroscopic electricity consumption statistics, resulting in significant information loss. This invention innovatively extracts chaotic features from high-frequency power waveforms, enabling extremely sensitive capture of minute changes in the operating status of users' internal electrical appliances.

[0090] (1) Load dynamic fractal dimension

[0091] This technical feature is used to quantify the structural complexity of a user's overall electricity consumption pattern. The calculation steps are as follows:

[0092] Step A1: Based on the initial fused data block, obtain the instantaneous voltage value sequence U(t) and the instantaneous current value sequence I(t) of the total inlet of the target object, which are synchronously collected at a sampling rate of not less than 10kHz.

[0093] Step A2: Using one power frequency cycle (e.g., 20ms) as a window, map U(t) and I(t) into a voltage-current trajectory diagram on a two-dimensional plane. The shape of this trajectory diagram is determined by all the operating appliances in the household.

[0094] Step A3: Calculate the fractal dimension D of the trajectory using box counting. A trajectory full of detail and complex structure (corresponding to multiple appliances operating simultaneously) has a higher fractal dimension; while when the main appliances stop working, the trajectory becomes simpler and smoother, and the fractal dimension decreases significantly. The calculation formula is: ,in, Let's say it's the side length of the box. Minimum number of boxes required to cover the voltage-current trajectory. Step A4: Output D as the load dynamic fractal dimension.

[0095] When a household stops using core appliances such as refrigerators and air conditioners, the dynamic fractal dimension of the load can decrease by more than 20% within 24 hours, while the traditional daily average electricity consumption change is often less than 5% and is easily affected by the occasional activation of other high-power appliances. This demonstrates that this feature improves the sensitivity of detecting changes in behavioral patterns by an order of magnitude, enabling earlier insights.

[0096] (2) Transient Lyapunov index

[0097] This technical feature is used to quantify the dynamic stability and chaos of power systems, and is particularly sensitive to non-stationary states such as electrical faults.

[0098] Step B1: Transfer the time series data from the voltage-current trajectory graph. Embedded into an m-dimensional phase space, constructing a phase space vector. Where τ is the time delay, , Let m represent the voltage and current values ​​corresponding to the i-th point on the voltage-current trajectory diagram, respectively, where m is an integer greater than 1.

[0099] Step B2: For each reference point Y(j) in the phase space, find the nearest neighbor point to the reference point among all phase points other than the reference point. ,satisfy Where d0 is the initial distance, and , The average period is 1.

[0100] Step B3: Track the reference point Y(j) and its nearest neighbor. Distance after k evolutionary steps And calculate the logarithmic growth rate of the reference point Y(j), using the following formula: ;

[0101] Step B4, using the formula Calculate the transient Lyapunov exponent , where M is the total number of all valid index pairs (j,k).

[0102] A stable household electrical system, its The value fluctuates within a certain range. When an electrical appliance experiences an early fault (such as unstable operating current due to worn motor bearings), its dynamic characteristics will change, leading to... The value shows an abnormal spike or a change in trend. This provides a technical indicator for predictive equipment maintenance that traditional solutions cannot provide.

[0103] 2. Macro-level electricity consumption behavior feature extraction: Electricity load fluctuation entropy

[0104] This feature, derived from traditional electricity consumption data, is used to quantify the disorder and randomness of users' daily electricity consumption behavior.

[0105] The module extracts the daily active power sequence of the target object from the initial fused data block, forming a daily load curve. After normalizing this curve, it is treated as a probability distribution. Subsequently, the Shannon entropy of this distribution is calculated using the following formula:

[0106] ;

[0107] in, This represents the proportion of power consumption to total daily electricity consumption during the i-th time interval.

[0108] A household with a regular lifestyle and stable electricity consumption patterns exhibits a relatively regular daily load curve and a low entropy value. Conversely, a household with an irregular lifestyle and highly random electricity consumption behavior has a high entropy value. More importantly, when a household enters a standby state due to reduced economic activity, its electricity consumption pattern changes from a complex active state to a simple base load state, resulting in a significant and continuous decrease in the entropy of electricity load fluctuations. This downward trend reflects the deactivation of the household's lifestyle pattern more effectively than a simple reduction in electricity consumption, providing another stable and reliable early warning indicator.

[0109] 3. Extraction of cross-energy correlation features: water-electricity coupling coefficient

[0110] To overcome the limitations of single-energy data analysis, this invention innovatively establishes a deep correlation between electricity consumption and water usage behavior.

[0111] Based on the aligned time-series electricity and water consumption data in the initial fused data block, the Pearson correlation coefficient between the two, i.e., the water-electricity coupling coefficient, is calculated for a specific time period (e.g., one day).

[0112] This coefficient reveals the synergy of users' energy consumption patterns. In a normal household, water and electricity consumption are usually positively correlated throughout the day (for example, water and electricity consumption will peak simultaneously in the morning and evening due to activities such as cooking and washing).

[0113] When this coefficient decreases significantly, it indicates that this synergy has been broken, suggesting a decoupling of their lifestyle patterns. This decoupling is key to identifying anomalous behavior.

[0114] Scenario 1 (Non-risk): When households are unoccupied for a short period, electricity and water consumption will simultaneously and significantly decrease to baseline levels. At this time, the two still maintain a strong positive correlation, so the water-electricity coupling coefficient may remain stable or only fluctuate slightly.

[0115] Scenario 2 (Core Risk): When households voluntarily reduce non-essential spending due to a sharp decline in living standards, their electricity consumption will decrease significantly as they turn off entertainment, comfort, and even some essential appliances. However, water consumption for maintaining basic survival needs (such as drinking and sanitation) remains relatively rigid, decreasing much less than electricity consumption. This contrast between the sharp decline in one sector and the relative stability of the other directly leads to a significant reduction in the water-electricity coupling coefficient.

[0116] Therefore, by monitoring the abnormal decrease in the water-electricity coupling coefficient and analyzing it in conjunction with the absolute change in electricity consumption, the system can effectively distinguish between two fundamentally different modes: short-term unmanned operation and life-threatening situations. This cross-domain verification mechanism greatly reduces the risk of false alarms caused by fluctuations in single electricity consumption data.

[0117] IV. Time Series Prediction and Feature Optimization Module

[0118] This module receives a high-quality sequence of technical feature vectors from the multidimensional feature vector construction module and continuously improves the accuracy of its future state prediction through a dynamically optimized closed-loop process, including the following units.

[0119] 1. Attention-based Long Short-Term Memory Network (Temporal Prediction Based on LTM-ATT Network)

[0120] The temporal prediction and feature optimization module organizes each user's historical technical feature vector (such as data from the past 30 days) into an input sequence in chronological order and inputs it into a long short-term memory network based on an attention mechanism.

[0121] The network is trained to learn complex patterns in feature vectors over time and outputs predicted technical feature vectors for a specific future time window (e.g., the next 7 days). This means the system can predict the likely values ​​of key indicators such as the dynamic fractal dimension of user load and water usage fluctuation entropy over a future period.

[0122] By capturing long-term dependencies using LSTM and combining it with an attention mechanism to focus on features at key time points, this model can more accurately depict the evolution of user behavior patterns, providing direct data support for proactive early warning.

[0123] To achieve continuous optimization of the prediction model and automatically identify the most effective features, this module innovatively introduces a dynamic weight allocator and a contribution feedback unit, forming an intelligent closed loop.

[0124] 2. Dynamic weight allocator

[0125] The dynamic weight allocator is positioned before the feature vector is input into the LSTM-ATT network. It is responsible for assigning an initial weight to each technical feature (such as fractal dimension, water-electricity coupling coefficient, etc.) in the technical feature vector. The weighted feature vector is then fed into the network for training and prediction.

[0126] 3. Contribution Feedback Unit:

[0127] The workflow is as follows: The contribution feedback unit receives the predicted values ​​for the future from the LSTM-ATT network and compares them with the subsequently collected real data to calculate the prediction bias. Based on this prediction bias, the Shapley value method is used to quantitatively evaluate the contribution of each feature to the prediction accuracy. The calculated contribution values ​​of each feature are sent back to the dynamic weight allocator. The allocator dynamically adjusts the feature weights used in the next round of prediction based on the contribution values. The weights of features with high contributions are increased, while the weights of features with low contributions are decreased.

[0128] In some embodiments, the Shapley value is derived from cooperative game theory. This invention innovatively applies it to feature importance assessment to ensure the fairness and scientific rigor of the allocation. The specific calculation steps are as follows:

[0129] S1. Let F be the set of all N technical features of the technical feature vector, and let M be the long short-term memory network based on the attention mechanism.

[0130] S2. For the target feature i whose contribution needs to be calculated, perform the following operations:

[0131] S2.1, Traverse the feature subset S in set F that does not contain the target feature i. ;

[0132] S2.2. For each of the feature subsets S, perform the following:

[0133] S2.2.1 Input the technical features of the feature subset S into the prediction model M to obtain the first predicted value V(S);

[0134] S2.2.2 Input the technical feature data of the union S∪{i} of the feature subset S and the target feature i into the model M to obtain the second predicted value V(S∪{i});

[0135] S2.2.3 Calculate the marginal contribution of the target feature i to the feature subset S. : ;

[0136] S3. Based on the marginal contribution of all feature subsets S, calculate the Shapley value of target feature i using the following formula. And the Shapley value is used as the contribution of this target feature:

[0137] ;

[0138] Where Σ represents the summation over all feature subsets S that do not contain the target feature i, |S| is the size of feature subset S, i.e., the number of features contained in feature subset S, and |F| is the total number of features. Technical effect: The power of this method lies in that it does not only consider the effect of a single feature, but also evaluates the average marginal contribution of that feature when combined with all other features. This effectively identifies "catalyst" features that are weak individually but can greatly improve prediction accuracy when combined with certain features.

[0139] Through the above design, this module achieves a leap from static analysis to dynamic prediction, and then to self-optimization. The longer the system runs and the more feedback it collects, the more accurate its judgment of the importance of different features becomes. This allows it to adaptively focus on the most effective signals, ultimately achieving a continuous improvement in prediction accuracy and reliability.

[0140] Example 2 discloses a user status early warning method based on heterogeneous data fusion. The method uses a time-series feature enhancement system for heterogeneous data fusion as described in Example 1, and includes the following steps:

[0141] Step 1: Utilize the aforementioned time-series feature enhancement system to continuously acquire and process fixed-dimensional technical feature vectors for multiple users in the target area. These vectors include at least: load dynamic fractal dimension, transient Lyapunov exponent, electricity load fluctuation entropy, and water-electricity coupling coefficient.

[0142] Step 2: Based on the temporal changes of the technical feature vector, identify at least one target user whose technical feature vector changes conform to a preset risk pattern through a set of multi-level risk identification rules;

[0143] Step 3: Generate and output early warning information corresponding to the target user and the triggered risk pattern. The early warning information shall at least include the user's identification information, the specific risk characteristics triggered and their degree of change, and the risk level hit.

[0144] The essence of the multi-level risk identification rule set design lies in cross-validation from different technical dimensions to form a three-dimensional, progressively advanced risk perception network. Its principles and effects are as follows:

[0145] 1. First level: Core chaotic feature sharpening rules

[0146] This level (core chaotic feature sharpening rule) directly calls upon the deepest features parsed from high-frequency electrical signals to achieve a microscopic diagnosis of the core state of the user's electricity consumption. Its judgment logic and conclusions are as follows:

[0147] A continuous decrease in the dynamic fractal dimension of the load: This phenomenon indicates that the user's internal appliance combination is becoming simpler. This usually means that many non-essential or life-enhancing appliances (such as air conditioners, entertainment equipment, small kitchen appliances, etc.) are being continuously deactivated, which is a strong physical signal that the user is actively or passively reducing their household economic activities.

[0148] Significant changes in the transient Lyapunov exponent: This phenomenon reflects changes in the operational stability of electrical appliances and disturbances in the dynamic chaotic characteristics of the system. Possible reasons behind this include: being forced to use old, inefficient, or faulty electrical appliances to save costs, or engaging in unconventional and unstable electricity usage behaviors due to economic pressure (such as frequently manually switching on and off high-power equipment).

[0149] When both of these characteristics occur simultaneously, it strongly indicates that the user's household economic situation may be experiencing systemic difficulties. This is not simply a matter of energy-saving behavior or short-term absences from home (which typically do not cause significant changes in the transient Lyapunov index), but rather reveals a severe pattern of simultaneous simplification of electricity usage and deterioration of electricity quality. This suggests that users are not only reducing the number of appliances they use, but also that the quality of electricity they use to maintain a basic standard of living is declining.

[0150] Therefore, this level possesses extremely high sensitivity and early warning capabilities. It can detect qualitative changes in micro-level electricity consumption patterns before significant fluctuations occur in the overall macro-level electricity consumption, achieving the earliest possible risk insight. Given its indicative nature and forward-looking nature, warnings triggered at this level are defined as the highest system level (e.g., a red alert), designed to prompt managers to prioritize and precisely intervene in target households.

[0151] 2. Second level: Multi-dimensional energy consumption mode attenuation rules

[0152] This level integrates both microscopic structure and macroscopic behavior features to achieve a dual-verification diagnosis of user power consumption pattern degradation. Its judgment logic and conclusions are as follows:

[0153] The continuous decrease in the dynamic fractal dimension of the load: This phenomenon indicates that the physical structure of the user's internal electrical appliance combination is becoming simpler (same as the first level).

[0154] The continuous decline in the entropy of electricity load fluctuations indicates that users' daily electricity consumption behavior is becoming more regular, with less randomness and activity. Lifestyles are shifting from rich and varied to monotonous and fixed.

[0155] When both of the aforementioned characteristics decline simultaneously, it collectively confirms, from two independent yet complementary dimensions—physical structure and behavioral patterns—that a comprehensive deactivation of user electricity consumption behavior is occurring. This dual attenuation of structural simplification and reduced activity strongly indicates that users' socioeconomic activities are contracting, and the richness and quality of their lives are significantly decreasing. This effectively eliminates interference caused by short-term absences or single appliance malfunctions, precisely pointing to a long-term, passive downgrade of lifestyle due to limited economic resources.

[0156] This level, through cross-validation of signals across dimensions, provides a more reliable and less false-positive basis for judgment than a single feature. It constitutes the core link in the risk identification system, triggering a medium-level warning (such as a yellow warning), indicating that the target household has entered a risk state requiring attention.

[0157] 3. Third level: Cross-energy correlation fracture rules

[0158] This level, by introducing heterogeneous energy data, overcomes the limitations of single-energy analysis and achieves a systematic diagnosis of the stability of users' overall lifestyles. Its judgment logic and conclusions are as follows:

[0159] The significant decrease in the water-electricity coupling coefficient indicates that the long-term stable "water-electricity synergy" consumption pattern of users has been broken. The fact that the two energy data no longer change synchronously suggests that their lifestyles have undergone an atypical and fundamental restructuring.

[0160] The continuous decrease in the entropy of electricity load fluctuations: This phenomenon (same as the second level) further confirms the reduction in user electricity activity and the randomness of behavior.

[0161] When both of the above characteristics are present simultaneously, it reveals an abnormal state in which the user's life system is undergoing decoupling and deactivation. This phenomenon of water and electricity decoupling (i.e., water use behavior and electricity use behavior becoming detached from their historical context) is a powerful indicator for identifying systemic and disruptive changes in lifestyle caused by major upheavals. For example, it can effectively distinguish between energy-saving behavior when someone is home and a complete regression in lifestyle due to economic hardship.

[0162] This level, by introducing heterogeneous data, constructs a third chain of evidence for risk identification, greatly improving the comprehensiveness and robustness of the system's coverage and effectively reducing monitoring blind spots. The warning level it triggers can be defined as medium to high (e.g., orange alert) based on the rate of decline, indicating that the stability of the target family's life is eroding, requiring in-depth verification and intervention.

[0163] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A time series feature enhancement system for heterogeneous data fusion, characterized in that, The system comprises: a data perception and access module for collecting heterogeneous time series data from different types of sensors and data platforms through multiple communication protocols, the heterogeneous time series data including voltage and current raw waveform data collected from power lines through high-frequency sampling circuit and time series water consumption data collected from water systems through Internet of Things protocols; a preprocessing and space-time alignment module in communication connection with the data perception and access module, for performing abnormal value cleaning, resampling and space-time alignment on the heterogeneous time series data, correlating power data and water consumption data under a unified timestamp, and generating an initial fusion data block; a multi-dimensional feature vector construction module in communication connection with the preprocessing and space-time alignment module, for extracting and calculating a fixed-dimensional technical feature vector for each target object from the initial fusion data block, wherein the technical feature vector includes a non-intrusive load decomposition feature based on voltage-current trajectory chaos analysis and a correlation feature calculated based on multi-source data in the initial fusion data block; a time series prediction and feature optimization module in communication connection with the multi-dimensional feature vector construction module, for performing time series prediction on the technical feature vector by using a long short-term memory network based on an attention mechanism, and dynamically optimizing the weights of each technical feature in the technical feature vector according to the prediction result.

2. The system according to claim 1, wherein, The non-intrusive load decomposition feature based on voltage-current trajectory chaos analysis includes a load dynamic fractal dimension, which is calculated by the following steps: A1, synchronously collecting voltage instantaneous value sequence U(t) and current instantaneous value sequence I(t) of the total inlet of the target object at a sampling rate not lower than 10 kHz; A2, mapping the voltage instantaneous value sequence U(t) and the current instantaneous value sequence I(t) into a voltage-current trajectory graph on a two-dimensional plane with one power frequency cycle as a window; A3, the fractal dimension D of the voltage-current trajectory graph is calculated by using the box-counting method, and the calculation formula is: wherein, is the side length of the box, is the minimum number of boxes required to cover the voltage-current trajectory; A4, taking the calculated fractal dimension D as the load dynamic fractal dimension.

3. The system of claim 2, wherein, The non-intrusive load decomposition feature based on voltage-current trajectory chaos analysis also includes a transient Lyapunov exponent, which is calculated by the following steps: B1, time series data on the voltage-current trajectory graph embedded into an m-dimensional phase space, constructing a phase space vector where τ is a time delay, , respectively represent the voltage value and the current value corresponding to the i-th point on the voltage-current trajectory graph, and m is an integer greater than 1; B2. For each reference point Y(j) in the phase space, find the nearest neighbor to the reference point among the phase points other than the reference point , satisfies where is the initial distance, and , is the average period; B3, tracking the reference point Y(j) and the nearest neighbor point Distance after k evolution steps and calculating the logarithmic growth rate of the reference point Y(j) as follows: ; B4, by the formula computing the transient Lyapunov exponent where M is the total number of all valid index pairs (j, k).

4. The system of claim 1, wherein, The specific process of abnormal value cleaning in the preprocessing and space-time alignment module is as follows: For the voltage and current raw waveform data, a sliding time window with length L is set, and the mean value of the waveform effective value in the sliding time window is calculated and the standard deviation ; identify and reject entire waveform segments corresponding to data points having values outside the range of valid values for the waveform effective values For the data gaps after removal, linear interpolation of the effective waveform segments before and after in the sliding time window is used for filling; For the user's historical water consumption data, upper and lower threshold values of the user's water consumption are constructed; data points with instantaneous water consumption values exceeding the upper and lower threshold values are identified and removed; for the data gaps after removal, the mean of the effective water consumption data before and after or zero water consumption is used for filling.

5. The system of claim 4, wherein, The specific process of abnormal value cleaning in the preprocessing and space-time alignment module is as follows: For the voltage and current raw waveform data, a sliding time window with length L is set, and the mean value of the waveform effective value in the sliding time window is calculated and the standard deviation ; identify and reject entire waveform segments corresponding to data points having values outside the range of valid values for the waveform effective values For the data gaps after removal, linear interpolation of the effective waveform segments before and after in the sliding time window is used for filling; For time series water consumption data, an abnormality detection algorithm based on S-H-ESD is used to identify and remove global and local abnormal points in the time series water consumption data; for the data gaps after removal, a Kalman filter algorithm based on a time series prediction model is used for filling.

6. The system of claim 1, wherein, The technical feature vector further comprises a power load fluctuation entropy and a water-power coupling coefficient, wherein the power load fluctuation entropy is a Shannon entropy value calculated based on a daily load curve; the correlation feature is the water-power coupling coefficient, which is a Pearson correlation coefficient calculated based on time-series power data and time-series water data in the initial fusion data block within a specific time period.

7. The system of claim 1, wherein, The time-series prediction and feature optimization module further comprises a dynamic weight allocator and a contribution feedback unit, wherein the dynamic weight allocator is arranged between the multi-dimensional feature vector construction module and the attention mechanism-based long short-term memory network, and is configured to assign appropriate weights to each technical feature in the technical feature vector; the contribution feedback unit is configured to receive a prediction result output by the attention mechanism-based long short-term memory network and subsequent actual data; based on the deviation between the prediction result and the actual data, the contribution of each technical feature is calculated; and the contribution is sent to the dynamic weight allocator for updating the weight of each technical feature in the next round of prediction.

8. The system of claim 7, wherein, The contribution is calculated by the Shapley value method, and the specific steps include: S1, a set of all N technical features of the technical feature vector is denoted as F, and the attention mechanism-based long short-term memory network is denoted as a prediction model M; S2, for a target feature i whose contribution is to be calculated, the following operations are performed: S2.1, traverse the subset S of features in the set F that does not contain the target feature i, ; S2.2, for each feature subset S, the following steps are performed: S2.2.1, input the technical features of the feature subset S into the prediction model M to obtain a first prediction value V(S); S2.2.2, input the technical features of the union S∪{i} of the feature subset S and the target feature i into the model M to obtain a second prediction value V(S∪{i}); S2.2.3, compute the marginal contribution of the target feature i to the feature subset S : ; S3, based on the marginal contribution of all feature subsets S, calculate the Shapley value of the target feature i through the following formula and take the Shapley value as the contribution degree of the target feature ; Wherein, Σ represents the summation of all feature subsets S not containing the target feature i, |S| is the size of the feature subset S, i.e. the number of features contained in the feature subset S, and |F| is the total number of features.

9. A user state early warning method based on heterogeneous data fusion, characterized in that, The method uses a time-series feature enhancement system for heterogeneous data fusion according to any one of claims 1-8, comprising the following steps: Using the time-series feature enhancement system, continuously acquiring and processing technical feature vectors of multiple users in a target area; Based on the time-series changes of the technical feature vectors, at least one target user whose technical feature vector changes conform to a preset risk pattern is identified through a set of multi-level risk identification rules; Generating and outputting early warning information corresponding to the target user and the risk pattern.

10. The user state early warning method based on heterogeneous data fusion according to claim 9, characterized in that, The preset risk mode includes a livelihood security risk mode. When any one of the multi-level risk identification rules is satisfied, it is determined that the user meets the livelihood security risk mode. The multi-level risk identification rule set includes a first level, a second level, and a third level. The first level is a core chaotic feature sharpening rule. When the load dynamic fractal dimension of a certain user continuously decreases by more than a first threshold value within a preset time period, and the transient Lyapunov index synchronously changes by more than a corresponding threshold value, the rule is triggered. The second level is a multi-dimensional energy consumption mode attenuation rule. When the load dynamic fractal dimension and the power load fluctuation entropy of a certain user both continuously decrease by more than respective threshold values within a preset time period, the rule is triggered. The third level is a cross-energy correlation fracture rule. When the water-electricity coupling coefficient of a certain user decreases by more than a second threshold value within a preset time period, and at the same time, the power load fluctuation entropy decreases by more than a third threshold value, the rule is triggered. The early warning information includes the identification information of the user and the triggered risk feature and the corresponding level.