Efficient compression, storage method and device of energy storage monitoring data
By performing semantic classification and frequency grading on energy storage monitoring data, and by dynamically filtering and extracting features before writing the data, the problems of redundancy and low query efficiency in energy storage system monitoring data storage are solved, achieving efficient data management and storage optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-29
- Publication Date
- 2026-06-09
AI Technical Summary
Existing energy storage system monitoring and data storage solutions fail to manage data in a refined manner based on its semantics and value density, resulting in redundant data occupying storage space, high hardware costs, low query efficiency, and fixed strategies being unable to cope with the problem of sudden changes in sampling frequency.
Semantic classification is performed on energy storage monitoring data, dividing it into setting objects, status objects, and time series objects, and then classifying them according to sampling frequency, configuring independent storage strategies and lifecycles; dynamic filtering is performed before data is written, including source filtering and value density filtering; for time series object data whose lifecycle has expired, feature extraction and re-filtering based on time windows are performed, retaining only the data with incremental information.
It achieves "hot-warm-cold" data separation, reduces storage hardware and bandwidth costs, improves query efficiency, solves the problem that fixed strategies cannot cope with sudden changes in sampling frequency, and ensures that critical data is not lost.
Smart Images

Figure CN122178921A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of energy storage system monitoring technology, specifically to a method and apparatus for efficient compression and storage of energy storage monitoring data. Background Technology
[0002] With the widespread use and large-scale deployment of energy storage systems, the flow of monitoring data has surged, and this monitoring data is characterized by strong periodicity, correlation, and dynamic changes in sampling frequency.
[0003] In related technologies, traditional monitoring data storage solutions typically write all data indiscriminately into a single storage device or perform simple hierarchical classification, failing to manage data granularly based on its semantics and value density. This results in a large amount of redundant data occupying storage space, high hardware costs, or mixed storage of high-frequency details and low-frequency aggregated data, leading to low query efficiency. Therefore, there is an urgent need for a method that can achieve intelligent data filtering, hierarchical compression, and efficient storage. Summary of the Invention
[0004] In view of this, this disclosure provides a method and apparatus for efficient compression and storage of energy storage monitoring data, in order to solve problems such as how to achieve intelligent data filtering, hierarchical compression and efficient storage, thereby reducing the hardware cost of energy storage monitoring and improving the efficiency of monitoring data query.
[0005] This disclosure provides a method for efficient compression and storage of energy storage monitoring data, the method including: The collected energy storage monitoring data is semantically classified into setting objects, status objects, and time-series objects. For time-series objects, they are divided into at least two frequency levels according to their sampling frequency, and independent storage strategies and lifecycles are configured for time-series objects of different frequency levels. Before data is written to storage, time-series objects and state objects are dynamically filtered. Dynamic filtering includes source filtering based on the validity status of the data collection source and value density filtering based on preset data tolerance judgment rules. For time-series object data that has reached the end of its life cycle, it is compressed and archived according to its frequency level. Among them, for low-frequency time-series object data, feature extraction based on time windows is performed, and the extraction results are further filtered according to data tolerance judgment rules, retaining only the data with information increment.
[0006] This disclosure also provides a high-efficiency compression and storage device for energy storage monitoring data, the device comprising: The semantic classification module is used to perform semantic classification on the collected energy storage monitoring data, dividing it into setting objects, status objects, and time-series objects; The frequency partitioning module is used to divide time-series objects into at least two frequency levels according to their sampling frequency, and to configure independent storage strategies and lifecycles for time-series objects of different frequency levels. The dynamic filtering module is used to dynamically filter time-series objects and state objects before data is written to storage. Specifically, the dynamic filtering module is used to filter sources based on the validity status of the data acquisition source and to filter value density based on preset data tolerance judgment rules. The compression and archiving module is used to compress and archive time-series object data that has reached the end of its life cycle, based on its frequency level. Specifically, for low-frequency time-series object data, feature extraction based on time windows is performed, and the extraction results are further filtered according to data tolerance judgment rules, retaining only the data with information increment.
[0007] This disclosure also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described methods for efficiently compressing and storing energy storage monitoring data.
[0008] This disclosure also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described methods for efficiently compressing and storing energy storage monitoring data.
[0009] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for efficiently compressing and storing energy storage monitoring data.
[0010] In the efficient compression and storage method and apparatus for energy storage monitoring data of the above embodiments of this disclosure, data is classified and distributed according to setting-state-time semantics, and time-series data is divided into multiple levels and stored independently according to sampling frequency. This enables "hot-warm-cold" separation of data at the physical level, thereby overcoming storage redundancy and query path confusion caused by indiscriminate or simple mixed storage, reducing query latency, and thus improving query efficiency. By sequentially performing source filtering based on device online status and value density filtering based on preset tolerance before data writing, and combined with secondary compression at the end of the life cycle, active redundancy discarding is achieved throughout the entire process from data entry to long-term archiving, significantly reducing storage hardware and bandwidth costs.
[0011] Furthermore, by applying the variation tolerance judgment rules to real-time filtering and subsequent secondary compression, and establishing an adaptive lifecycle management strategy linked to frequency levels, the system can dynamically adjust the storage granularity according to the data value density and business status. This solves the problem that fixed strategies cannot cope with dynamic operating conditions such as sudden changes in sampling frequency, and continuously optimizes storage efficiency while ensuring that critical data is not lost. Attached Figure Description
[0012] To more clearly illustrate the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating an efficient compression and storage method for energy storage monitoring data provided in this embodiment of the disclosure; Figure 2 A schematic diagram illustrating the specific process of an efficient compression and storage method for energy storage monitoring data provided in this embodiment of the disclosure; Figure 3 A schematic diagram of a high-efficiency compression and storage device for energy storage monitoring data provided in this embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0014] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this disclosure.
[0015] As energy storage systems develop towards large-scale and high-density deployments, the volume of monitoring data is growing exponentially. It is predicted that the scale of energy storage systems will continue to expand in the coming years, with the daily data volume per station reaching GB levels or even higher. Energy storage system data has the following characteristics: 1. Periodic redundancy: The battery charge and discharge curves exhibit 24-hour periodic fluctuations, with a large amount of data overlap between adjacent periods.
[0016] 2. Spatiotemporal correlation: The data similarity of the same cluster of battery packs is >85%, indicating a strong correlation between the data.
[0017] 3. Dynamic abrupt changes: When the sampling frequency suddenly increases from 1Hz to 10Hz during power grid frequency regulation, the frequency and amplitude of data changes abruptly.
[0018] 4. Value decay: The value of data decreases exponentially over time (half-life ≈ 7 days), with recent data having high value and long-term data having low value.
[0019] The above data characteristics lead to three major challenges for storage solutions: 1. Storage Costs: Full storage leads to a surge in hardware investment. As data volume continues to grow, the cost of storage hardware will become a major burden for enterprises.
[0020] 2. Query efficiency: Mixed storage creates I / O bottlenecks. Different types of data are stored together, and queries require scanning a large amount of irrelevant data, resulting in high query latency.
[0021] 3. Dynamic response lag: Fixed compression parameters cannot cope with sudden changes in sampling frequency. When the sampling frequency changes suddenly, the fixed compression parameters cannot be adjusted in time, resulting in poor data compression or loss of important information.
[0022] In related technologies, the following solutions are often adopted: 1. Traditional integrated architecture: Treats all monitoring data as records of the same level, writes and indexes them uniformly using a single storage medium. All data, regardless of its semantic type or sampling frequency, is stored and indexed according to the same rules.
[0023] 2. Cache-and-Swap Architecture: High-frequency writes are first absorbed by a high-speed cache, and then periodically batched and dumped into persistent storage. When new data is written, it is first written to the high-speed cache. When the cache reaches a certain threshold or is triggered periodically, the data in the cache is batched and dumped into persistent storage.
[0024] 3. Cloud-Edge Hierarchical Architecture: The edge is responsible for data collection and initial processing, while the cloud is responsible for long-term archiving and big data analysis. Edge devices perform initial screening and processing on the collected data, and then upload the processed data to the cloud for long-term storage and further analysis.
[0025] However, the aforementioned technologies often have the following problems: 1. Indiscriminate writing: Traditional monolithic architectures, cache-storage architectures, and cloud-edge hierarchical architectures all fail to "identify value first, then decide where to write to disk," resulting in a high proportion of redundant data. For example, in traditional monolithic architectures, all data, regardless of its value, is stored uniformly, wasting a significant amount of storage space.
[0026] 2. Single-level lifecycle: Existing technologies use the same retention strategy to manage high-frequency and low-frequency data, resulting in two types of waste: "low-frequency data being deleted prematurely" or "high-frequency data being deleted too late." For example, in a cache-storage architecture, no different lifecycle strategies are formulated based on the frequency and value of the data, leading to some low-frequency data being deleted prematurely, while some high-frequency data occupies storage space for a long time.
[0027] 3. Static filtering rules: The threshold, sampling frequency, and device status are disconnected, making it impossible to adaptively adjust to changing operating conditions. In a cloud-edge hierarchical architecture, the filtering rules for edge devices are fixed and cannot be dynamically adjusted based on actual operating conditions, resulting in poor filtering performance.
[0028] 4. Query Path Differentiation: Raw data and downsampled data are physically isolated, requiring cross-interface assembly for fault backtracking, increasing latency and consistency risks. In traditional integrated architectures, different types of data are stored together, requiring the scanning of a large amount of irrelevant data during queries, resulting in high query latency. In cloud-edge hierarchical architectures, raw data and downsampled data are stored at the edge and in the cloud respectively, requiring cross-interface assembly during queries, increasing query complexity and latency.
[0029] To address the aforementioned issues, various embodiments of this disclosure provide an efficient method for compressing and storing energy storage monitoring data. The method includes: semantically classifying the collected energy storage monitoring data into setting objects, status objects, and time-series objects; dividing time-series objects into at least two frequency levels based on their sampling frequency, and configuring independent storage strategies and lifecycles for time-series objects at different frequency levels; dynamically filtering time-series objects and status objects before writing the data to storage; wherein, dynamic filtering includes: source filtering based on the validity status of the data acquisition source, and value density filtering based on preset data tolerance judgment rules; compressing and archiving time-series object data whose lifecycle has expired based on its frequency level; wherein, for low-frequency level time-series object data, performing feature extraction based on a time window, and further filtering the extraction results based on data tolerance judgment rules, retaining only data with incremental information.
[0030] It should be noted that, in the description of this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this disclosure are used to distinguish similar objects and are not used to describe a particular order or sequence.
[0031] To enable those skilled in the art to better understand the present disclosure, the present disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an efficient compression and storage method for energy storage monitoring data provided in an embodiment of the present disclosure. The method may include the following steps: Step S101: Semantic classification is performed on the collected energy storage monitoring data, dividing it into setting objects, status objects, and time-series objects.
[0033] In this embodiment, energy storage monitoring data (hereinafter referred to as data) can come from various sensors and control devices in the energy storage system, such as battery management system and power conversion system. The data can have different business meanings and change characteristics when it is generated.
[0034] Specifically, in order to manage data efficiently, it can first be classified according to the semantic information expressed by the data.
[0035] The setting object can refer to data describing the static configuration parameters of the system. For example, the setting object can include, but is not limited to, the rated capacity of the battery and the system operating threshold. Its characteristics are that the values are stable and the frequency of change is extremely low.
[0036] A status object can refer to a discrete status flag that describes the operating condition of a device, such as, but not limited to, the device start / stop status and fault alarm flags. Its characteristic is that it only changes when the status changes.
[0037] Time-series objects can refer to physical quantity measurements that are continuously collected in chronological order, such as continuously changing quantities such as voltage, current, and temperature. Their characteristics are that they change continuously over time and the amount of data is huge.
[0038] Understandably, through the above semantic classification, the originally mixed data stream can be split into three independent data channels, which facilitates differentiated management based on their respective characteristics.
[0039] Step S102: For time-series objects, divide them into at least two frequency levels according to their sampling frequency, and configure independent storage strategies and lifecycles for time-series objects of different frequency levels.
[0040] In this embodiment, the time-series object can contain data with different sampling frequencies ranging from milliseconds to minutes. To achieve fine-grained management of the time-series object, it can be further divided into multiple frequency levels based on a preset frequency threshold.
[0041] For example, millisecond-level data can be classified as high-frequency, second-level data as medium-frequency, and minute-level data as low-frequency.
[0042] Each frequency level corresponds to different data "popularity" and business value. For example, high-frequency data can reflect instantaneous details, and its value decays quickly; while low-frequency data reflects trend changes and needs to be retained for a long time.
[0043] Furthermore, for different frequency levels, independent physical or logical storage spaces can be established to achieve physical data isolation. Simultaneously, dedicated storage and lifecycle management strategies can be configured for each frequency level. For example, high-frequency levels have shorter data retention periods, so high-throughput, low-latency storage media and indexing methods can be used; medium- and low-frequency levels have progressively longer data retention periods and can employ storage solutions with higher compression ratios and lower costs.
[0044] Understandably, the above-mentioned hierarchical mechanism can naturally reflect the "hot-warm-cold" attributes of data at the storage level, avoiding the query efficiency bottleneck caused by mixed storage of data with different frequencies.
[0045] Step S103: Before writing the data to storage, perform dynamic filtering on the time-series objects and state objects; wherein, dynamic filtering includes: source filtering based on the validity status of the data collection source, and value density filtering based on preset data tolerance judgment rules.
[0046] In this embodiment, dynamic filtering can be a key preprocessing step before data enters the storage system. Its purpose is to proactively identify and discard invalid or redundant data, thereby reducing the storage load from the source.
[0047] Specifically, this step can include two independent but coordinated filtering mechanisms.
[0048] Source validity filtering can be based on the status of data acquisition devices or communication links. The system continuously monitors the validity signals of the acquisition source, such as heartbeat signals or link quality indicators sent periodically by the device. If the acquisition source is determined to be offline, faulty, or experiencing a communication interruption, all data generated during this period is deemed to have lost business significance and can therefore be discarded entirely, thus preventing invalid data from consuming storage space and processing resources.
[0049] Value density filtering, on the other hand, allows for more precise judgments based on changes in the business value of the data itself.
[0050] The system pre-sets corresponding data tolerance judgment rules for each type of data. These rules define the minimum amount of change that must be met to determine whether data has storage value.
[0051] Understandably, through the two layers of filtering mentioned above, the system can significantly filter out meaningless noise data and duplicate business data before the data is written to disk, thereby achieving the initial purification of data value.
[0052] Step S104: For time series object data whose life cycle has expired, compression and archiving are performed according to their frequency level; among them, for low frequency level time series object data, feature extraction based on time window is performed, and the extraction results are further filtered according to data tolerance judgment rules, retaining only the data with information increment.
[0053] In this embodiment, this step can be backend optimization and in-depth management of the stored data, aiming to further reduce long-term storage costs while meeting regulatory requirements for the retention of critical data.
[0054] Different lifecycles can be preset for time-series data at different frequency levels. When a data fragment reaches the end of its preset lifecycle, the system can automatically trigger a targeted compression and archiving process.
[0055] Specifically, for high-frequency or mid-frequency data, the archiving strategy can be to transfer the entire data to a low-cost storage medium; while for low-frequency data whose value has significantly deteriorated but which needs to be stored in compliance with regulations for a long time, a more proactive "secondary compression" mechanism can be adopted.
[0056] The "secondary compression" mechanism first divides the data from the same measurement point within the same lifecycle into continuous time windows of a preset length (e.g., 1 hour, 1 day). Then, it extracts features from the original sampling points within each window, such as calculating the maximum, minimum, and average values within the window, or directly selecting the sampling values at the beginning and end of the window as representatives.
[0057] Furthermore, after feature extraction, instead of simply retaining the aggregated results of all windows, a second round of filtering can be performed, including: The extraction results of each window (such as the difference between the first and last values) are substituted into the preset data tolerance judgment rules of the measurement point for evaluation. If the evaluation shows that the window has not undergone significant changes beyond the tolerance range throughout the entire time span, it is determined that there is no new information increment in the data within the window, and only a representative data point (such as the starting value of the window) is retained. If the change is significant, the aggregated feature value of the window is retained.
[0058] This mechanism allows for the elimination of static redundancy in the time dimension to the greatest extent possible, while preserving long-term data trends and key turning points, thus achieving deep optimization of compliance volume.
[0059] In the efficient compression and storage method and apparatus for energy storage monitoring data in the above embodiments of this disclosure, data is classified and distributed according to setting-state-time semantics, and time-series data is divided into multiple levels and stored independently according to sampling frequency. This enables "hot-warm-cold" separation of data at the physical level, thereby overcoming storage redundancy and query path confusion caused by indiscriminate or simple mixed storage, reducing query latency, and thus improving query efficiency. By sequentially performing source filtering based on device online status and value density filtering based on preset tolerance before data writing, and combined with secondary compression at the end of the life cycle, active redundancy discarding is achieved throughout the entire process from data entry to long-term archiving, significantly reducing storage hardware and bandwidth costs. By applying variable tolerance judgment rules to real-time filtering and subsequent secondary compression, and establishing an adaptive life cycle management strategy linked to frequency level, the system can dynamically adjust the storage granularity according to data value density and business status, solving the problem that fixed strategies cannot cope with dynamic operating conditions such as sudden changes in sampling frequency, and continuously optimizing storage efficiency while ensuring that critical data is not lost.
[0060] In one possible implementation of step S101 above, the collected energy storage monitoring data is semantically classified into setting objects, status objects, and time-series objects, including: The energy storage monitoring data describing the static configuration parameters of the system is divided into setting objects, which are represented in the form of key-value pairs. Energy storage monitoring data, which describes discrete state indicators of equipment operation, is divided into state objects, which are represented by an enumeration type. Energy storage monitoring data, which consists of continuously collected physical quantity measurements in chronological order, is divided into time-series objects, which are represented by timestamp-value sequences.
[0061] In this embodiment, semantic classification can be the foundation for building an efficient data management system. It can clearly and operably summarize massive and heterogeneous raw monitoring stream data into three core abstract types based on their inherent business logic and change patterns.
[0062] Specifically, the settings can correspond to the inherent attributes and static parameters of the system or device, such as the rated capacity of the battery pack, the power limit of the inverter, and the alarm threshold of the protection function.
[0063] This data can be organized and managed in the form of "key-value pairs," and its core characteristic is an extremely low change frequency, changing only when the system configuration is updated.
[0064] By categorizing settings objects independently, it means that a "change only to drive storage" strategy can be designed for them, thereby avoiding periodic full duplicate recording.
[0065] Furthermore, state objects can be used to characterize discrete operating conditions of equipment or subsystems during operation, such as "run / stop" switch status, "normal / warning / fault" health status, and on / off network status.
[0066] This type of data is naturally suitable for representation using enumeration types, and its value lies in state transition events, while continuous sampling that maintains the same state does not have incremental information.
[0067] After being independently categorized, "change-triggered" write rules can be applied to them, fundamentally eliminating redundant storage of state information.
[0068] Furthermore, time-series objects can carry dynamic physical quantities of a continuously operating system, such as voltage, current, temperature, and power, which take the form of a sequence of timestamps and sampled values as their basic form.
[0069] This type of data has the largest volume and changes continuously, making it a primary target for compression and storage optimization.
[0070] By clearly classifying them, it is easier to design refined management, including frequency classification and dynamic filtering, specifically targeting their continuity, timeliness and value decay characteristics.
[0071] In the efficient compression and storage method and apparatus for energy storage monitoring data in the above embodiments of this disclosure, logical diversion can be achieved at the data entry point through the above-mentioned refined semantic classification; it assigns data with different change modes, query requirements and processing logic to different management channels, thereby overcoming the inherent contradictions caused by the traditional solution of indiscriminately mixing and processing various types of data.
[0072] In one possible implementation of step S102 above, the time-series object is divided into at least two frequency levels according to its sampling frequency, and independent storage strategies and lifecycles are configured for time-series objects of different frequency levels, including: Based on a preset sampling frequency threshold, the time series object is divided into high-frequency, mid-frequency and low-frequency levels; among them, the sampling frequency of the high-frequency level is higher than that of the mid-frequency level, and the sampling frequency of the mid-frequency level is higher than that of the low-frequency level. Configure independent and mutually exclusive physical storage spaces for timing objects at high frequency, mid frequency, and low frequency levels respectively; Configure a first lifecycle and a first storage strategy for high-frequency time-series objects, configure a second lifecycle and a second storage strategy that are longer than the first lifecycle for mid-frequency time-series objects, and configure a third lifecycle and a third storage strategy that are longer than the second lifecycle for low-frequency time-series objects.
[0073] In this embodiment, the frequency classification mechanism can be the core operation for fine-grained and differentiated management of time-series objects.
[0074] Among them, the frequency grading mechanism can break away from the traditional thinking of treating time-series data as a single whole, and instead perform vertical grading based on its time granularity and business value, thereby achieving the separation of data "hotness" at the storage architecture level.
[0075] Specifically, a set of sampling frequency thresholds is first preset according to business rules. For example, data with a sampling interval of less than 1 second (such as milliseconds) is classified as high frequency, data with a sampling interval between 1 second and 1 minute is classified as medium frequency, and data with a sampling interval greater than 1 minute is classified as low frequency.
[0076] Here, this division directly corresponds to the information density and level of detail of the data. For example, high-frequency data captures instantaneous details and rapid dynamics, medium-frequency data reflects normal operational fluctuations, and low-frequency data depicts long-term trends.
[0077] Furthermore, to achieve true isolation, independent and mutually exclusive physical storage spaces can be established for each of the three frequency levels. This means that data of different levels is stored in different physical files, database tables, or storage volumes, thus avoiding data mixing at the source.
[0078] In addition, storage strategies and lifecycles can be customized for each frequency level.
[0079] Preferably, a first lifecycle (e.g., 7 days) and a first storage strategy can be configured for high-frequency levels. Typically, low-latency storage media with high input / output operations per second (IOPS) are used, and the index granularity is fine to meet the query requirements for real-time monitoring and second-level fault traceability.
[0080] A second lifecycle (e.g., 30 days) can be configured for mid-frequency tiers, which is longer than the first lifecycle. This second storage strategy may strike a balance between storage cost and query performance, such as using more cost-effective solid-state drives (SSDs) or high-speed hard drives.
[0081] A third lifecycle (e.g., 180 days or longer) can be configured for low-frequency tiers, which is longer than the second lifecycle. This third storage strategy focuses on high compression ratios and low costs, such as using high-density hard drives or object storage, and may apply coarser-grained indexes.
[0082] In the efficient compression and storage method and apparatus for energy storage monitoring data in the above embodiments of this disclosure, the isolation of physical storage space through the aforementioned frequency hierarchy and independent strategy configuration completely eliminates I / O interference between data of different "hotness" levels. This allows queries targeting a specific frequency range to avoid scanning irrelevant data, thereby significantly reducing query latency from the second level under traditional hybrid storage to the sub-second level. Differentiated lifecycles and storage strategies achieve a precise match between cost and value. "Hot data" occupies expensive but high-performance resources for a short period, while "cold data" is stored long-term in cost-effective media, optimizing the overall system storage cost. Furthermore, this hierarchical architecture provides a clear path and operational unit for the natural flow of data from "hot" to "cold" and subsequent compression operations, making the entire data management system more flexible and efficient.
[0083] In one possible implementation of step S103 above, source filtering is performed based on the validity status of the data acquisition source, including: Monitor the heartbeat signal of the data acquisition device corresponding to the energy storage monitoring data. If the difference between the timestamp of the heartbeat signal and the current time exceeds the preset validity judgment threshold, the data acquisition device is determined to be offline, and the time sequence objects and status objects generated by the data acquisition device during the offline period are discarded. Value density filtering is performed based on preset data tolerance judgment rules, including: For a setting object, calculate the hash value of the current setting object data packet and perform a first comparison with the hash value of the previously stored setting object data packet; if the result of the first comparison is consistent, it is determined that the current setting object data packet has not been changed and storage is skipped. For a state object, the current state bit value is compared with the previously stored state bit value. The current state object is written to the storage only if the result of the second comparison is inconsistent. For time series objects, obtain the sampled value of the time series object at the current moment and the previously stored sampled value at the same measurement point. If the absolute value of the difference between the sampled value at the current moment and the previously stored sampled value does not exceed the preset variation tolerance bandwidth of the measurement point, the time series object at the current moment is determined to be redundant and storage is skipped.
[0084] In this embodiment, dynamic filtering can be a core element in improving data quality and efficiency. Before data is written to storage, it can actively screen and purify the data from two dimensions: effectiveness and value density, transforming the traditional passive full storage mode into an active selection storage mode.
[0085] Among them, source validity filtering constitutes the first gate for data access, and its core logic can be to determine that data from invalid collection sources has no business significance.
[0086] Specifically, this is achieved by continuously monitoring the heartbeat signals periodically sent by the data acquisition device. The system can preset a validity judgment threshold (e.g., 300 seconds). When the system receives new data, it will check the timestamp of the latest heartbeat of the corresponding device. If the difference between the current time and the timestamp exceeds the preset threshold, the device is determined to be offline.
[0087] Once the device is determined to be offline, all time-series data and status data generated during the offline period will be discarded.
[0088] This mechanism can effectively prevent invalid data from polluting storage space due to network outages, equipment failures, etc.
[0089] Furthermore, value density filtering can, based on the validity of the data, further determine whether it contains new information increments.
[0090] Specifically, value density filtering employs different judgment rules based on data type: For configuration objects, a hash comparison trigger mechanism is used. A unique hash value is calculated for the currently received complete configuration parameter data packet (such as system rated parameters, protection settings, etc.). This hash value is compared with the hash value of the configuration object data packet that was successfully stored last time. If the two match, it means that the static configuration of the system has not been modified or updated since the last storage. The current data packet is considered completely redundant and the storage operation is skipped directly.
[0091] Only when the hash values are inconsistent does it mean that the configuration has been effectively changed, and only then will the storage of the new version configuration data be triggered.
[0092] This mechanism ensures that configuration data is updated only when necessary, avoiding periodic full duplication of records.
[0093] For status objects, a bit-flip trigger mechanism can be used. By comparing the value of the current status bit with the value that was successfully stored in the last time, a meaningful change in operating condition (such as changing from "running" to "fault") is only indicated when the two are inconsistent. Only then will the current status object be written. Consecutive identical statuses are only recorded the first time, fundamentally eliminating redundant recording of status information.
[0094] For time-series objects, a variable tolerance bandwidth mechanism is adopted. The system presets a variable tolerance bandwidth (e.g., ±0.5V) for each measurement point (such as "battery pack A voltage") to represent its normal fluctuation range or measurement error. When a new sampling point arrives, the system calculates the absolute value of the difference between its value and the previously stored sampled value at the same measurement point. If the difference does not exceed the preset bandwidth, the sampled value is considered to be natural fluctuation or noise, does not provide new information, is judged as redundant data, and is skipped from storage.
[0095] In the efficient compression and storage method and apparatus for energy storage monitoring data of the above embodiments of this disclosure, significant data simplification can be achieved before data is written into the database through the aforementioned two-level dynamic filtering. Source filtering eliminates the consumption of storage resources by invalid data, while value density filtering precisely removes redundant data that is irrelevant to the business. The combination of these two methods ensures that the data ultimately written to storage is all "effective and necessary" information increments. Verification based on actual business scenarios shows that this mechanism can reduce the effective data write volume of the system by more than 70%, directly alleviating the pressure on network transmission bandwidth, storage I / O, and backend processing links, thereby reducing storage costs and improving storage efficiency.
[0096] In one possible implementation of step S103 above, feature extraction based on a time window is performed on low-frequency time-series object data, including: Low-frequency time-series object data within the same lifecycle and at the same measurement point are divided into multiple data windows according to a preset time window length. Aggregate calculations are performed on multiple sampled values within each data window to generate aggregated feature values representing that data window; wherein, the aggregation calculations include at least one of the following: calculating the maximum value, calculating the minimum value, calculating the average value, taking the first point value within the window, and taking the last point value within the window.
[0097] In this embodiment, secondary compression of low-frequency time-series object data can be a core optimization step in long-term data archiving management.
[0098] When low-frequency data reaches the end of its lifecycle, the system does not directly delete it as a whole or transfer it as is. Instead, it can initiate a deep feature extraction and re-filtering process to preserve the long-term value of the data while eliminating static redundancy in the time dimension to the greatest extent.
[0099] Specifically, the processing scope can be defined first, and low-frequency time-series data from the same measurement point (such as the voltage of battery cluster 1) within the same life cycle can be used as processing units. Based on the preset time window length, the continuous time-series data within this unit is divided into a series of continuous data windows, where each data window can cover a fixed time period and can contain multiple raw sampled values arranged in chronological order.
[0100] Furthermore, multiple original sampled values within each data window are aggregated and calculated to generate an aggregated feature value that can represent the overall characteristics of the data window.
[0101] Preferably, the aggregation calculation method may include, but is not limited to: calculating the maximum or minimum value of all sampled values within the window to capture the peak or trough value within that period; calculating the average value to reflect the average level of that period; or directly taking the first and last values of the window to characterize the start and end states of that period. Through aggregation, the dozens or hundreds of original data points within each data window can be reduced to one or a few highly generalized feature values, achieving preliminary dimensionality reduction of the data.
[0102] In the efficient compression and storage method and apparatus for energy storage monitoring data in the above embodiments of this disclosure, by converting the low-frequency raw data that needs to be stored for a long time from a continuous sampling point sequence into a feature value sequence organized by a time window, this conversion process can directly remove the redundancy in the time details of the raw data while preserving the long-term trend and key extreme value information of the data. This results in a significant reduction in the volume of data entering the long-term archiving stage for the first time, thereby reducing the amount of writing, reducing the burden on each hardware component, and eliminating the need for additional computing power for the edge gateway.
[0103] In one possible implementation of step S104 above, the extraction results are further filtered according to data tolerance judgment rules, retaining only data with information increments, including: For each data window, calculate the difference between the first and last sampled values of that data window; Determine whether the difference exceeds the preset variation tolerance bandwidth for the measurement point; If the limit is not exceeded, it is determined that there is no information increment in the data within the data window, and only the first sample value of the data window is retained; If the value exceeds the limit, the aggregated feature value of that data window will be retained.
[0104] In this embodiment, this step can be a mitigation of secondary compression at the end of the data lifecycle, and can be used to perform final value screening on low-frequency time-series data after aggregation through a time window, thereby optimizing the volume of long-term archived data.
[0105] After completing the time window segmentation and aggregate feature extraction, a general description of each data window can be obtained.
[0106] Not all windows contain meaningful trend changes; within certain window periods, the measured physical quantity may fluctuate steadily within a very small range. Therefore, to further eliminate storage redundancy in such steady-state data, this step can introduce a re-filtering mechanism based on the original tolerance rules.
[0107] Specifically, the re-filtering mechanism can perform the following operations on a data window basis: The first and last sampled values of the data window are taken, and the difference between them is calculated to determine the characteristic change within the window. This difference reflects the net change of the physical quantity from the start to the end within the time window. The absolute value of this difference is compared and judged with the variation tolerance bandwidth preset for the measurement point.
[0108] If the difference does not exceed the variation tolerance bandwidth, it is determined that the data has not undergone significant changes beyond normal fluctuations or measurement error ranges within the entire window time span. This means that the data within this window period does not contain new information increments and belongs to a static or stable period.
[0109] At this point, only the first sampled value of the data window is retained as the sole representative of the entire time period, thereby compressing a window that may contain dozens of aggregated features or original points into a single data point.
[0110] If the difference exceeds the variation tolerance bandwidth, it indicates that the physical quantity has undergone a business-significant trend change within that window period.
[0111] At this point, the data in this window is considered to have preservation value, and the aggregated feature values of the window (such as the previously calculated maximum, minimum, or average values) can be retained to characterize the key statistical features of this period of change.
[0112] In the efficient compression and storage method and apparatus for energy storage monitoring data in the above embodiments of this disclosure, the re-filtering mechanism accurately identifies and eliminates a large number of stable periods in long-term historical data through window-level trend judgment. On the basis of already reducing the amount of data through aggregation, it can further reduce the volume of long-term archived data exponentially, thereby further improving storage query efficiency and reducing storage volume.
[0113] In one embodiment, please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating a specific process for an efficient compression and storage method for energy storage monitoring data provided in an embodiment of this disclosure, as shown below. Figure 2 As shown: The data acquisition module obtains raw sampled data from protocols such as Modbus, CAN, and IEC104, and sends it to a unified node for processing. First, the data is semantically classified into setting objects, status objects, and timing objects.
[0114] For a setting object, its hash value is calculated and compared with the previously stored hash value. If there is no change, it is discarded, and only setting objects that have changed are stored.
[0115] For state objects, a bit comparison method is used to determine whether their state value has changed. Storage is triggered only when the previous stored value is inconsistent with the current value; otherwise, the value is discarded.
[0116] For time-series data, it is further divided into high-frequency, mid-frequency, and low-frequency data based on the sampling frequency. Before storage, an online status check is performed on the peripheral device. The data acquisition device is checked for online status based on the heartbeat signal. If it is offline, the data for that time period is discarded. If it is online, a bandwidth change check is performed. That is, it is checked whether the difference between the current sampled value and the previously stored value exceeds the preset bandwidth for that measurement point. If it does not exceed the bandwidth, it is considered redundant data and discarded. Only the data exceeding the bandwidth is stored.
[0117] In addition, the system also includes a lifecycle management and secondary compression scheduling mechanism: when the storage period of time-series data shards expires, aggregation processing is triggered to extract features from the data within the window and perform a change bandwidth judgment again, retaining only the aggregation results with information increments for archiving and storage.
[0118] In one embodiment, a high-efficiency compression and storage device 300 for energy storage monitoring data is provided, which corresponds one-to-one with the high-efficiency compression and storage method for energy storage monitoring data in the above embodiments. For example... Figure 3 As shown, the device includes: The semantic classification module 301 is used to perform semantic classification on the collected energy storage monitoring data, dividing it into setting objects, status objects and time series objects; The frequency partitioning module 302 is used to divide time-series objects into at least two frequency levels according to their sampling frequency, and to configure independent storage strategies and lifecycles for time-series objects of different frequency levels. The dynamic filtering module 303 is used to dynamically filter time-series objects and state objects before data is written to storage; specifically, the dynamic filtering module is used to: filter sources based on the validity status of the data acquisition source, and filter value density based on preset data tolerance judgment rules. The compression and archiving module 304 is used to compress and archive time-series object data that has reached the end of its life cycle, according to its frequency level. Specifically, for low-frequency time-series object data, feature extraction based on time windows is performed, and the extraction results are further filtered according to data tolerance judgment rules, retaining only the data with information increment.
[0119] In one embodiment, the semantic classification module 301 is specifically used to divide the energy storage monitoring data describing the static configuration parameters of the system into setting objects, and the setting objects are represented in the form of key-value pairs; Energy storage monitoring data, which describes discrete state indicators of equipment operation, is divided into state objects, which are represented by an enumeration type. Energy storage monitoring data, which consists of continuously collected physical quantity measurements in chronological order, is divided into time-series objects, which are represented by timestamp-value sequences.
[0120] In one embodiment, the frequency division module 302 is specifically used to divide the timing object into high-frequency level, mid-frequency level and low-frequency level based on a preset sampling frequency threshold; wherein the sampling frequency of the high-frequency level is higher than that of the mid-frequency level, and the sampling frequency of the mid-frequency level is higher than that of the low-frequency level. Configure independent and mutually exclusive physical storage spaces for timing objects at high frequency, mid frequency, and low frequency levels respectively; Configure a first lifecycle and a first storage strategy for high-frequency time-series objects, configure a second lifecycle and a second storage strategy that are longer than the first lifecycle for mid-frequency time-series objects, and configure a third lifecycle and a third storage strategy that are longer than the second lifecycle for low-frequency time-series objects.
[0121] In one embodiment, the dynamic filtering module 303 is specifically used to monitor the heartbeat signal of the data acquisition device corresponding to the energy storage monitoring data. If the difference between the timestamp of the heartbeat signal and the current time exceeds a preset validity judgment threshold, the data acquisition device is determined to be offline, and the time sequence objects and status objects generated by the data acquisition device during the offline period are discarded. The dynamic filtering module 303 is also used to calculate the hash value of the current setting object data packet for the setting object, and perform a first comparison between the hash value and the hash value of the previously stored setting object data packet; if the result of the first comparison is consistent, it is determined that the current setting object data packet has not changed and storage is skipped. For a state object, the current state bit value is compared with the previously stored state bit value. The current state object is written to the storage only if the result of the second comparison is inconsistent. For time series objects, obtain the sampled value of the time series object at the current moment and the previously stored sampled value at the same measurement point. If the absolute value of the difference between the sampled value at the current moment and the previously stored sampled value does not exceed the preset variation tolerance bandwidth of the measurement point, the time series object at the current moment is determined to be redundant and storage is skipped.
[0122] In one embodiment, the compression archiving module 304 is specifically used to divide low-frequency time-series object data within the same life cycle and at the same measurement point into multiple data windows according to a preset time window length; Aggregate calculations are performed on multiple sampled values within each data window to generate aggregated feature values representing that data window; wherein, the aggregation calculations include at least one of the following: calculating the maximum value, calculating the minimum value, calculating the average value, taking the first point value within the window, and taking the last point value within the window.
[0123] In one embodiment, the compression archiving module 304 is specifically used to calculate the difference between the first sample value and the last sample value of each data window. Determine whether the difference exceeds the preset variation tolerance bandwidth for the measurement point; If the limit is not exceeded, it is determined that there is no information increment in the data within the data window, and only the first sample value of the data window is retained; If the value exceeds the limit, the aggregated feature value of that data window will be retained.
[0124] It should be noted that the high-efficiency compression and storage device for energy storage monitoring data provided in the above embodiments is only illustrated by the division of the above-described program modules when implementing the corresponding high-efficiency compression and storage method for energy storage monitoring data. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the above device can be divided into different program modules to complete all or part of the processing described above. In addition, the device provided in the above embodiments and the corresponding Figure 1 The embodiments of the methods shown belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0125] This disclosure also provides an electronic device having the above-described features. Figure 3 The device shown is an efficient compression and storage device for energy storage monitoring data.
[0126] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure.
[0127] The following is a detailed reference. Figure 4 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 401, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 402 or a program loaded from memory 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of the electronic device. The processor 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0128] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a memory 408, or installed from a ROM 402. When the computer program is executed by the processor 401, it performs the functions defined in the efficient compression and storage method for energy storage monitoring data according to embodiments of this disclosure.
[0130] Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0131] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the efficient compression and storage method for energy storage monitoring data shown in the above embodiments is implemented.
[0132] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0133] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for efficiently compressing and storing energy storage monitoring data, characterized in that, The method includes: The collected energy storage monitoring data is semantically classified into setting objects, status objects, and time-series objects. The time-series objects are divided into at least two frequency levels according to their sampling frequency, and independent storage strategies and lifecycles are configured for time-series objects of different frequency levels. Before the data is written to storage, the time-series object and the state object are dynamically filtered; wherein, the dynamic filtering includes: source filtering based on the validity status of the data acquisition source, and value density filtering based on preset data tolerance judgment rules; For time-series object data that has reached the end of its life cycle, it is compressed and archived according to its frequency level. Among them, for low-frequency time-series object data, feature extraction based on time window is performed, and the extraction results are further filtered according to the data tolerance judgment rules, retaining only the data with information increment.
2. The method according to claim 1, characterized in that, The collected energy storage monitoring data is semantically classified into setting objects, status objects, and time-series objects, including: The energy storage monitoring data describing the static configuration parameters of the system is divided into setting objects, which are represented in the form of key-value pairs; Energy storage monitoring data that describes discrete state indicators of equipment operating conditions are divided into state objects, which are represented by an enumeration type. Energy storage monitoring data, which are continuously collected in chronological order, are divided into time-series objects, which are characterized in the form of timestamp-value sequences.
3. The method according to claim 1 or 2, characterized in that, The process of dividing the time-series objects into at least two frequency levels according to their sampling frequency, and configuring independent storage strategies and lifecycles for time-series objects at different frequency levels, includes: Based on a preset sampling frequency threshold, the time series object is divided into high-frequency level, mid-frequency level and low-frequency level; wherein the sampling frequency of the high-frequency level is higher than that of the mid-frequency level, and the sampling frequency of the mid-frequency level is higher than that of the low-frequency level. Configure independent and mutually exclusive physical storage spaces for the timing objects of the high-frequency level, the mid-frequency level, and the low-frequency level respectively; Configure a first lifecycle and a first storage strategy for the high-frequency time-series objects, configure a second lifecycle and a second storage strategy that are longer than the first lifecycle for the mid-frequency time-series objects, and configure a third lifecycle and a third storage strategy that are longer than the second lifecycle for the low-frequency time-series objects.
4. The method according to claim 3, characterized in that, The source filtering based on the validity status of the data collection source includes: The heartbeat signal of the data acquisition device corresponding to the energy storage monitoring data is monitored. If the difference between the timestamp of the heartbeat signal and the current time exceeds the preset validity judgment threshold, the data acquisition device is determined to be offline, and the time sequence objects and status objects generated by the data acquisition device during the offline period are discarded. The value density filtering based on preset data tolerance judgment rules includes: For a setting object, calculate the hash value of the current setting object data packet and perform a first comparison with the hash value of the previously stored setting object data packet; if the result of the first comparison is consistent, it is determined that the current setting object data packet has not been changed and storage is skipped. For a state object, the current state bit value is compared with the previously stored state bit value. The current state object is written to the storage only if the result of the second comparison is inconsistent. For a time series object, obtain the sampled value of the time series object at the current moment and the previously stored sampled value at the same measurement point. If the absolute value of the difference between the sampled value at the current moment and the previously stored sampled value does not exceed the preset variation tolerance bandwidth of the measurement point, then the time series object at the current moment is determined to be redundant and storage is skipped.
5. The method according to claim 3, characterized in that, The process of performing time-window-based feature extraction on low-frequency time-series object data includes: Low-frequency time-series object data within the same lifecycle and at the same measurement point are divided into multiple data windows according to a preset time window length. Aggregate calculations are performed on multiple sampled values within each data window to generate aggregated feature values representing that data window; wherein, the aggregation calculations include at least one of the following: calculating the maximum value, calculating the minimum value, calculating the average value, taking the first point value within the window, and taking the last point value within the window.
6. The method according to claim 5, characterized in that, The process of further filtering the extraction results according to the data tolerance judgment rules, retaining only data with incremental information, includes: For each data window, calculate the difference between the first and last sampled values of that data window; Determine whether the difference exceeds the preset variation tolerance bandwidth for the measurement point; If the limit is not exceeded, it is determined that there is no information increment in the data within the data window, and only the first sample value of the data window is retained; If the value exceeds the limit, the aggregated feature value of that data window will be retained.
7. A high-efficiency compression and storage device for energy storage monitoring data, characterized in that, The device includes: The semantic classification module is used to perform semantic classification on the collected energy storage monitoring data, dividing it into setting objects, status objects, and time-series objects; The frequency partitioning module is used to divide the time-series object into at least two frequency levels according to its sampling frequency, and to configure independent storage strategies and lifecycles for time-series objects of different frequency levels. A dynamic filtering module is used to dynamically filter the time-series object and the state object before the data is written to storage; specifically, the dynamic filtering module is used to: perform source filtering based on the validity status of the data acquisition source, and perform value density filtering based on preset data tolerance judgment rules. The compression and archiving module is used to compress and archive time-series object data that has reached the end of its life cycle, according to its frequency level. Specifically, for low-frequency time-series object data, feature extraction based on time windows is performed, and the extraction results are further filtered according to the data tolerance judgment rules to retain only the data with information increment.
8. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the computer instructions to perform the efficient compression and storage method for energy storage monitoring data as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which are used to cause the computer to execute the efficient compression and storage method for energy storage monitoring data according to any one of claims 1-6.
10. A computer program product, characterized in that, It includes computer instructions, which are used to cause a computer to execute the efficient compression and storage method for energy storage monitoring data according to any one of claims 1-6.