Data storage method, device and equipment, storage medium and computer program product
By initializing and dynamically adjusting the survival time and data level of the power plant data storage nodes, the problem that storage nodes in the existing technology cannot be dynamically adjusted is solved, and efficient data management and resource utilization are achieved.
Patent Information
- Application Number
- CN202510464264.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Existing storage nodes cannot be dynamically adjusted with the importance of data or fluctuations in the system load, resulting in early elimination of key data or long-term retention of redundant data, affecting the storage and overall storage performance of new data.
By obtaining the data generated in the current time period of the power plant, initialize the storage node and survival time based on the data compression results of the previous time period, determine the data level, and adjust the storage set based on the monitored changes, including data migration and node recycling, and dynamically optimize the storage structure.
Dynamic optimization is achieved based on changes in data timeliness and importance, ensuring that important data is in suitable storage nodes, avoiding resource waste, and improving storage system performance and management rationality.
Smart Images

Figure CN120491888A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to a data storage method, apparatus, device, storage medium, and computer program product. Background Art
[0002] In the field of power plant data management, with the rapid development of smart grids and the Industrial Internet of Things, the amount of data generated by power plant systems is growing exponentially. Existing storage nodes cannot dynamically adjust to changes in data importance or fluctuations in system load, resulting in the premature elimination of critical data or the long-term retention of redundant data, affecting the storage of new data and overall storage performance. This system cannot meet the data management and application needs of different sources, multiple types (including relational data, time series data, and unstructured data), and different tenants. Summary of the Invention
[0003] The main purpose of this application is to provide a data storage method, apparatus, equipment, storage medium and computer program product, aiming to solve the technical problem that existing storage nodes cannot be dynamically adjusted with changes in data importance or fluctuations in system load, resulting in premature elimination of key data or long-term retention of redundant data, affecting the storage of new data and overall storage performance.
[0004] To achieve the above objectives, the present application proposes a data storage method, which includes:
[0005] Obtaining power plant data generated in the current time period of the power plant, and initializing storage nodes and the corresponding survival times of the storage nodes based on the data compression results of the previous time period;
[0006] determining a data level of each data object in the power plant data;
[0007] storing the data object in the storage node according to the data level and activating the survival time to obtain an initial storage set;
[0008] When the survival time starts to shorten, the initial storage set is adjusted based on the monitored change of the survival time to obtain a target storage set.
[0009] Optionally, the initial storage set includes core nodes, buffer nodes and eliminated nodes;
[0010] The step of adjusting the initial storage set based on the monitored change of the survival time to obtain the target storage set when the survival time begins to shorten includes:
[0011] When the survival time begins to shorten, monitoring the change of the survival time in real time;
[0012] If the change is that the first survival time of the eliminated node decreases to a first safety threshold, migrating the first data object in the eliminated node to the buffer node, and recycling the eliminated node to obtain a first adjustment result;
[0013] If the change is that the second survival time of the core node decreases to a second safety threshold, migrating the second data object of the core node to the eliminated node to obtain a second adjustment result;
[0014] The initial storage set is adjusted based on the first adjustment result and / or the second adjustment result to obtain a target storage set.
[0015] Optionally, if the change is that the first survival time of the eliminated node decreases to a first safety threshold, migrating the first data object in the eliminated node to the buffer node and reclaiming the eliminated node, after obtaining the first adjustment result, the step further includes:
[0016] Determine whether the fluctuation of the current survival time of the buffer node exceeds a preset safety range;
[0017] When the fluctuation of the lifetime of the buffer node exceeds a preset safety range, calculating a Shannon entropy value based on the distribution of the third data object in the buffer node;
[0018] Detecting whether the buffer node meets a preset reallocation condition based on the Shannon entropy value, and obtaining an entropy value detection result;
[0019] If the entropy value detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is greater than the high entropy alarm threshold, the anti-entropy storage optimization algorithm is started to reallocate the buffer node;
[0020] If the entropy detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is less than the low entropy optimization threshold, the buffer node is reallocated according to the data concentration of the buffer node.
[0021] Optionally, if the change is that the second survival time of the core node decreases to a second safety threshold, migrating the second data object of the core node to the eliminated node to obtain a second adjustment result includes:
[0022] If the change is that the second survival time of the core node decreases to a second safety threshold, generating a migration queue based on the data level of the second data object stored in the core node;
[0023] Setting a third survival time for the migration queue according to the second survival time;
[0024] The second data object is migrated to the eliminated node based on the set migration queue, and the first survival time of the eliminated node after the migration is completed is adjusted according to the third survival time to obtain a second adjustment result.
[0025] Optionally, the step of obtaining power plant data generated in a current time period of the power plant and initializing the storage node and the lifetime corresponding to the storage node based on the data compression result of the previous time period includes:
[0026] Obtain power plant data within the current time period from the business system through the data acquisition interface;
[0027] Analyze the data compression results of the previous time period to determine the remaining available resources and status of the storage node;
[0028] Determine the initial lifetime and storage type of the storage node based on the remaining available resources and the status;
[0029] Initialize the storage node according to the storage type.
[0030] Optionally, the step of determining the data level of each data object in the power plant data includes:
[0031] Obtaining data sensitivity of each data object in the power plant data;
[0032] Determining a first degree of association between the data object and a production business and a second degree of association between the data object and a business system;
[0033] Determine the impact scope of the data object based on a preset assessment process;
[0034] Each of the data objects is graded according to the data sensitivity, the first relevance, the second relevance, and the impact range to obtain a data level of each of the data objects.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a data storage device, which includes:
[0036] A node initialization module is used to obtain power plant data generated in the current time period of the power plant and initialize the storage nodes and the corresponding lifespan of the storage nodes based on the data compression results of the previous time period;
[0037] a level determination module, configured to determine the data level of each data object in the power plant data;
[0038] a data storage module, configured to store the data object in the storage node according to the data level, and activate the survival time to obtain an initial storage set;
[0039] The storage adjustment module is configured to adjust the initial storage set based on the monitored change of the survival time when the survival time starts to shorten, so as to obtain a target storage set.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a data storage device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data storage method described above.
[0041] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the data storage method described above are implemented.
[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the data storage method described above.
[0043] This application discloses obtaining power plant data generated in the current time period of the power plant, and initializing the storage node and the corresponding life span of the storage node based on the data compression result of the previous time period; determining the data level of each data object in the power plant data; storing the data object in the storage node according to the data level, and activating the life span to obtain an initial storage set; when the life span begins to shorten, adjusting the initial storage set based on the monitored change of the life span to obtain a target storage set. By adjusting the initial storage set according to the change of the life span, the storage structure is dynamically optimized in real time according to the timeliness and importance of the data, ensuring that important data is always in the appropriate storage node, avoiding the waste of storage resources, and improving the overall performance of the storage system and the rationality of data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1This is a flow chart of the first embodiment of the data storage method of the present application;
[0047] Figure 2 This is a flow chart of the second embodiment of the data storage method of the present application;
[0048] Figure 3 This is a flowchart of the third embodiment of the data storage method of the present application;
[0049] Figure 4 This is a schematic diagram of the module structure of the data storage device according to an embodiment of the present application;
[0050] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the data storage method in the embodiment of the present application.
[0051] The purpose, features and advantages of this application will be further explained with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solution of the embodiment of the present application is: obtaining the power plant data generated in the current time period of the power plant, and initializing the storage node and the survival time corresponding to the storage node based on the data compression result of the previous time period; determining the data level of each data object in the power plant data; storing the data object in the storage node according to the data level, and activating the survival time to obtain an initial storage set; when the survival time begins to shorten, adjusting the initial storage set based on the monitored changes in the survival time to obtain a target storage set.
[0055] In the field of power plant data management, with the rapid development of smart grids and the Industrial Internet of Things, the volume of data generated by power plant systems is growing exponentially, encompassing high-value information such as real-time monitoring data, equipment logs, and user electricity usage records. Traditional data storage solutions typically employ static tiered storage architectures (e.g., hot / cold data separation) and fixed Time To Live (TTL) management. However, these solutions suffer from significant deficiencies in terms of dynamism, resource efficiency, and security. Specifically, they: 1. Rigid storage tiers rely on predefined storage tiers (e.g., core storage and archive storage), with fixed data classification rules that cannot be dynamically adjusted based on data value. For example, the co-storage of highly sensitive real-time monitoring data and low-priority log data can lead to excessive storage pressure on core nodes and underutilization of buffer nodes. 2. Fixed Time To Live management: Traditional solutions set a uniform TTL for storage nodes, failing to account for data level differences. Critical data may be accidentally deleted due to TTL expiration, while low-value data may remain and occupy resources for a long time.
[0056] Therefore, this application provides a data storage method that coordinates dynamic hierarchical storage and lifetime optimization. Through dynamic data level classification, lifetime self-initialization and real-time feedback adjustment mechanism, it realizes flexible allocation of storage resources, precise management and control of data life cycle and self-optimization operation of the system, thereby meeting the high concurrency, high security and low latency storage requirements of power plant data.
[0057] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, data migration, and program execution functions, such as a production operation system, or an electronic device capable of implementing the above functions. This embodiment and the following embodiments will be described below using a power plant production operation management system as an example.
[0058] Based on this, the embodiment of the present application provides a data storage method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the data storage method of the present application.
[0059] In this embodiment, the data storage method includes:
[0060] Step S10: acquiring power plant data generated in the current time period of the power plant, and initializing storage nodes and the corresponding survival times of the storage nodes based on the data compression results of the previous time period.
[0061] It should be noted that power plant data refers to various types of data generated during power plant operation, including but not limited to real-time monitoring data (such as equipment temperature and voltage), historical operation logs, operation records, fault alarms, etc. Data compression results are storage efficiency indicators obtained by compressing data from the previous time period, such as compression ratio, compression time, and remaining space utilization. Storage nodes are the logical storage units underlying the storage system and can be divided into core nodes (high-value data), buffer nodes (medium- and short-term data), and retired nodes (low-value data) based on data value or access frequency. The lifetime can be the maximum time that data is retained in a storage node or the time that a storage node can operate normally under the current load. It is related to factors such as the storage node's performance and remaining capacity. Based on the changes in the lifetime of each storage node, the value of the data currently stored in the storage node can be dynamically matched and storage node capacity allocation can be adjusted to avoid resource waste.
[0062] It is understandable that when the original data within the current time period can be obtained in real time from the power plant business system through a standardized interface, it is necessary to analyze the data compression results of the storage nodes in the previous time period, extract key indicators, and predict the storage space requirements of the current data volume based on the key indicators, so as to initialize the storage nodes and the corresponding survival time of the storage nodes.
[0063] Step S20: determining the data level of each data object in the power plant data.
[0064] It should be noted that data level is a quantitative indicator of the importance of a data object.
[0065] In actual applications, the data level of each data object can be determined based on the degree of impact on the power plant's operations, safety, or compliance; it can also be determined based on the data and the power plant's core business processes. Of course, in order to reflect the scope of business links that may be affected by data anomalies or loss, the data level can also be determined by the frequency and dependencies of cross-module calls in the power plant information system.
[0066] Step S30: storing the data object in the storage node according to the data level, and activating the lifetime to obtain an initial storage set.
[0067] It should be understood that when storing the data object in the storage node, node type matching can be used, and the target storage node can be selected according to the data level. Important data can be stored in the core node for high-performance storage, supporting real-time reading and writing; medium data can be stored in the buffer node, supporting short-term caching while balancing performance and cost; and edge data (lower data level) needs to be stored in the eliminated node for low-cost storage, which can support fast erasure.
[0068] It's understandable that different types of storage nodes have different lifetimes, and even the same type of storage node can have different lifetimes depending on the data level of the stored data objects. Each storage node has a corresponding lifetime, and when a data object is stored on the corresponding storage node, the lifetime of that storage node is activated.
[0069] After activating the storage node's corresponding lifetime, the TTL will dynamically decay over time as the data passes. This decay mechanism ensures the rapid elimination of low-value data, reduces the attack surface, and improves data storage security. The system can also take timely measures to migrate data or adjust storage policies, improving the efficiency and flexibility of data management.
[0070] In one example, when real-time monitoring data of a thermal power plant is stored based on a preset detection cycle, boiler temperature data and turbine pressure data are set as Class A data, and fan operation logs are set as Class C data. The target storage nodes are selected according to the data level (A / B / C level), such as Class A: core nodes (high-performance storage, supporting real-time reading and writing); Class B: buffer nodes (balancing performance and cost, supporting short-term caching); Class C: obsolete nodes (low-cost storage, supporting fast erasure).
[0071] After the real-time monitoring data is stored in the corresponding storage node, the corresponding TTL needs to be set for each storage node using the dynamic TTL formula. The formula is as follows:
[0072] T=T base ×α×β
[0073] Where T represents the time to live (TTL), T base is the basic survival time, which can be preset according to the storage node type. α is the coefficient of the data level, and β is the node load coefficient, which is negatively correlated with the current load rate (L) of the node.
[0074] After data is written to a storage node, a countdown timer starts, decaying the TTL every second. The remaining node space is updated in real time, and the TTL decay rate is automatically adjusted when a threshold is triggered. If the storage node failure rate exceeds 5%, the data is migrated to a backup node and the TTL is reset.
[0075] In high-load scenarios, such as when the load on core node 1 rises to 85%, the TTL decay rate is accelerated, triggering the migration queue to relegate Class A data (such as boiler temperature data) with a TTL of less than 12 hours to the buffer node. By migrating the core node load down to 60%, the buffer node utilization rate increases from 40% to 65%.
[0076] Step S40 : When the survival time starts to shorten, the initial storage set is adjusted based on the monitored change of the survival time to obtain a target storage set.
[0077] It is understood that when data objects and storage nodes in the initial storage set are reallocated and adjusted based on changes in lifetime, adjustments may include data migration, node recycling, and node activation to ensure that data is properly stored in appropriate storage nodes. When the lifetime of a storage node shortens, timely migration of data to more reliable nodes can avoid data loss or corruption due to storage node failures and improve data reliability and availability.
[0078] In this embodiment, the power plant data generated in the current time period of the power plant is obtained, and the storage node and the corresponding life span of the storage node are initialized based on the data compression result of the previous time period; the data level of each data object in the power plant data is determined; the data object is stored in the storage node according to the data level, and the life span is activated to obtain an initial storage set; when the life span begins to shorten, the initial storage set is adjusted based on the monitored changes in the life span to obtain a target storage set. By adjusting the initial storage set according to the changes in the life span, the storage structure can be dynamically optimized in real time according to the timeliness and importance of the data, ensuring that important data is always in the appropriate storage node, while timely cleaning or migrating unimportant data to avoid wasting storage resources, further improving the overall performance of the storage system and the rationality of data management.
[0079] Reference Figure 2 , Figure 2 This is a flow chart of the second embodiment of the data storage method of the present application. Based on the above-mentioned first embodiment, the second embodiment of the data storage method of the present application is proposed.
[0080] In the second embodiment, step S40 includes:
[0081] Step S401 : When the survival time starts to shorten, monitor the change of the survival time in real time.
[0082] Step S402: If the change is that the first survival time of the eliminated node drops to a first safety threshold, the first data object in the eliminated node is migrated to the buffer node, and the eliminated node is recycled to obtain a first adjustment result.
[0083] It should be noted that the first survival time is the initial survival time of the data in the retired node. The first safety threshold is the critical value that triggers data migration and resource recovery at the retired node, which can be a percentage of the first survival time. The first data object is the data unit to be processed in the retired node that meets the migration conditions. It is usually sorted by priority, such as generation time and access frequency.
[0084] It is understood that when the first data object from the eliminated node is migrated to the buffer node, all data within the node is deleted after the data migration is complete, freeing up physical resources and marking the node status as "idle" to facilitate storage allocation in the next cycle. The buffer node receiving the migrated data can appropriately extend its TTL for troubleshooting and short-term backtracking.
[0085] Furthermore, in order to improve the stability and storage efficiency of the buffer nodes and ensure the performance of the entire storage system through the adaptive optimization strategy, after step S402, the following steps are further included:
[0086] Determine whether the fluctuation of the current survival time of the buffer node exceeds a preset safety range;
[0087] When the fluctuation of the lifetime of the buffer node exceeds a preset safety range, calculating a Shannon entropy value based on the distribution of the third data object in the buffer node;
[0088] Detecting whether the buffer node meets a preset reallocation condition based on the Shannon entropy value, and obtaining an entropy value detection result;
[0089] If the entropy value detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is greater than the high entropy alarm threshold, the anti-entropy storage optimization algorithm is started to reallocate the buffer node;
[0090] If the entropy detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is less than the low entropy optimization threshold, the buffer node is reallocated according to the data concentration of the buffer node.
[0091] It should be noted that fluctuation indicates the magnitude of the change in survival time over time, reflecting the stability of the storage load. The preset safety range can be derived based on historical data statistics (such as the mean ± 2 times the standard deviation). The third data object is a specific category of data stored in the buffer node, usually Class B or temporary migration data. The Shannon entropy value is used to quantify the degree of chaos in the data distribution. The larger the value, the more dispersed the data, and the smaller the value, the more concentrated it is. The high entropy alarm threshold and the low entropy optimization threshold are used to determine whether the data in the buffer node is chaotically distributed or over-aggregated, respectively.
[0092] In one example, when counting TTL fluctuations in a time window (5 minutes), if the fluctuation is greater than a preset safety range, the entropy value H(X) calculation is triggered. The calculation formula is as follows:
[0093]
[0094] Among them, p(x i ) represents the proportion of the i-th type of data in the buffer node.
[0095] When H(X) is greater than the high entropy alarm threshold, it indicates that the data distribution is chaotic. It is necessary to start the anti-entropy storage optimization algorithm, split the same type of data into multiple nodes, use the consistent hashing algorithm to allocate storage locations, and update the TTL to 80% of the original value to balance the load. The entropy value is reduced through distributed data storage.
[0096] When H(X) is less than the low entropy optimization threshold, it indicates excessive data aggregation and requires redistribution based on the degree of aggregation. This can include merging consecutive storage blocks and adjusting the TTL to 120% of the original value to extend the retention period. Data aggregation measures the degree of data concentration in the storage space and can be determined by the ratio of the number of consecutive storage nodes to the total number of data nodes.
[0097] Step S403: If the change is that the second survival time of the core node decreases to a second safety threshold, the second data object of the core node is migrated to the eliminated node to obtain a second adjustment result.
[0098] It should be noted that the second survival time is the dynamic lifetime of data in the core node, representing the remaining retention time from the time the data is written until it needs to be downgraded or cleared. The second safety threshold is the critical value that triggers the core node to perform data downgrade migration and can be a percentage of the second survival time.
[0099] Specifically, the system monitors the second time to live (TTL) of core nodes in real time. When it detects that its value drops to a preset second safety threshold (e.g., 10% of the remaining time), the system triggers the data migration process. Data to be downgraded is screened based on the data object's level. For example, data with a lower level or long-term inaccessibility is prioritized for migration. Alternatively, an orderly migration queue is generated based on data attributes such as last access time, size, and associated business importance, ensuring that critical data is migrated last to minimize impact.
[0100] Furthermore, in order to achieve precise control over the data migration process and target node time management and ensure the security and timeliness of important data during the migration process, the step S403 may include:
[0101] If the change is that the second survival time of the core node decreases to a second safety threshold, generating a migration queue based on the data level of the second data object stored in the core node;
[0102] Setting a third survival time for the migration queue according to the second survival time;
[0103] The second data object is migrated to the eliminated node based on the set migration queue, and the first survival time of the eliminated node after the migration is completed is adjusted according to the third survival time to obtain a second adjustment result.
[0104] It should be noted that the migration queue is a list of data to be migrated, sorted by priority, such as data level and access frequency, and is used to control the order of demotion. The third survival time is the new survival time of the data after migration to the decommissioned node. It is usually shorter than the original survival time but longer than the default value of the decommissioned node.
[0105] It is understood that core nodes typically store high-priority, highly sensitive data, supporting real-time access and strong encryption. When the second lifetime of a core node falls below the second security threshold, the downgrade process is initiated, screening data that meets the downgrade criteria, such as data objects that were previously classified as low-level but were assigned to core nodes. Simultaneously, a migration queue is generated in reverse order of access time, and a third lifetime is generated based on the original lifetime and the decay coefficient. During data migration, the data in the migration queue is first copied to the corresponding decommissioned node, followed by a data integrity check and deletion of the original data from the core node.
[0106] It should be understood that the above-mentioned hierarchical degradation mechanism can ensure the long-term retention of high-value data and the release of low-value data on demand. At the same time, setting a third survival time can also prevent the migrated data from being cleared prematurely, achieving resource balance and data value continuation between core nodes and eliminated nodes, and significantly improving the efficiency and compliance of the power plant storage system.
[0107] In one example, the third lifetime is dynamically set based on the decay rate of the original remaining TTL of the core node and the current load of the eliminated node:
[0108] T new =T remaining ×γ
[0109] Where γ is the attenuation coefficient, and γ<1, T remaining Indicates the original remaining TTL of the core node. If the core node stores boiler pressure data (level A), the initial TTL is 24 hours. When the remaining TTL is monitored to be 2.4 hours (i.e., it drops to the 10% threshold):
[0110] 1. Filter out Class B historical data that has not been accessed in the past week and generate a migration queue.
[0111] 2. Set the third survival time to 1.2 hours (γ=0.5) and migrate to the eliminated node.
[0112] 3. Eliminate nodes and adjust the default cleanup cycle to 1 hour based on the TTL of new data to accelerate the elimination of low-value data.
[0113] 4. The core node releases 50GB of space to store newly generated real-time monitoring data.
[0114] Data migration can be performed using block-by-block transfer, copying data in the migration queue to the decommissioned node in blocks. Incremental synchronization is used to reduce bandwidth usage. A consistency check is performed using hash validation to ensure data integrity. Failures are retried or rolled back. After the migration is complete, core node space is released and marked as available. Failed data blocks are logged and retried after the system recovers. If the migration takes too long or the failure rate exceeds a certain limit, an alarm is triggered and the administrator is notified.
[0115] Step S404: Adjust the initial storage set based on the first adjustment result and / or the second adjustment result to obtain a target storage set.
[0116] It is understood that during the adjustment process, when the first and second adjustment results are received to form an adjustment instruction set, if both are triggered simultaneously, the core node downgrade migration, i.e., the second adjustment result, will be prioritized to avoid the risk of high-value data being retained. If only a single adjustment result is used, it will be directly applied to the storage set.
[0117] Specifically, after entering the adjustment results, the global storage topology needs to be updated with data status synchronization, recording the latest data distribution, remaining capacity, and TTL of each node. When adjusting core nodes, the migrated data is removed and the remaining capacity and load rate of the core nodes are recalculated. If the core node load rate is less than 50%, the baseline TTL for newly written data is extended (for example, from 8 hours to 12 hours). Conversely, the TTL is shortened to accelerate data flow. During buffer node adjustments, if the buffer node load exceeds 80% due to receiving retired data, horizontal expansion is triggered, and a new buffer node is automatically created. A consistent hashing algorithm is used to migrate some data to the new node. The routing table is simultaneously updated to ensure that access requests are distributed to the new node, and the buffer node TTL is dynamically adjusted based on data access frequency. During retired node adjustments, the TTL (third time to live) of data migrated from the core node is compared with the default TTL of the retired node. If the median TTL of the new data is greater than the original TTL of the retired node, the overall TTL of the retired node is increased to the new median. For example, if the original TTL of the retired node is 1 hour and the median TTL of the new data is 1.5 hours, the TTL is updated to 1.5 hours. A background process is then launched to scan the retired node for data with a TTL less than 0.5 hours, immediately deleting it and marking the storage block as "overwritable." Finally, storage set optimization and data consistency verification are required. If the buffer node's load is less than 30% for a long period (e.g., three consecutive cycles), it will be downgraded to a decommissioned node. If the decommissioned node's load is continuously greater than 90%, it will be upgraded to a buffer node. Of course, cold data can also be automatically divided into cold data (migrated to decommissioned nodes) and hot data (retained in core / buffer nodes) based on data access patterns. After data migration is complete, the data hash values of the source and target nodes are compared to ensure integrity. All adjustment operations are logged and rollback operations are supported. For example, if a node crashes during migration, the system can roll back to the last stable state based on the log.
[0118] In this embodiment, when the survival time begins to shorten, the change of the survival time is monitored in real time; if the change is that the first survival time of the eliminated node drops to the first safety threshold, the first data object in the eliminated node is migrated to the buffer node, and the eliminated node is recycled to obtain a first adjustment result; if the change is that the second survival time of the core node drops to the second safety threshold, the second data object of the core node is migrated to the eliminated node to obtain a second adjustment result; based on the first adjustment result and / or the second adjustment result, the initial storage set is adjusted to obtain a target storage set. In response to different situations where the survival time of the eliminated node and the core node drops to the safety threshold, the operations of migrating the eliminated node data to the buffer node and recycling the eliminated node, and migrating the core node data to the eliminated node are respectively adopted to effectively balance the load of various types of storage nodes, ensure the stability of data storage, and realize dynamic optimization of storage resources.
[0119] Reference Figure 3 , Figure 3 This is a flow chart of the third embodiment of the data storage method of the present application. Based on the above second embodiment, the third embodiment of the data storage method of the present application is proposed.
[0120] In the third embodiment, step S10 includes:
[0121] Step S101: acquiring power plant data within the current time period from a business system through a data acquisition interface.
[0122] It should be noted that business systems refer to the various information systems used in a power plant for daily production, operations, and management. These systems cover all aspects of the power plant's operations, such as the power generation management system, equipment monitoring system, fuel management system, and financial management system. When acquiring power plant data, the time period can be determined based on specific business needs and data collection strategies.
[0123] Understandably, the data collection interface requires multi-protocol interface adaptation, automatically matching the interface protocol based on the business system type and establishing a protocol-system mapping table. Alternatively, a factory model can be used to dynamically load protocol adapters. The data collection window is configured to a preset collection period (e.g., 5-minute micro-batches, 1-hour batches). The clocks of each business system are synchronized using the NTP service, triggering an alarm if the time deviation exceeds 500ms. Of course, power plant data can also be read through transaction logs.
[0124] Step S102: Analyze the data compression result of the last time period to determine the remaining available resources and status of the storage node.
[0125] It's understood that remaining available resources refer to the amount of resources a storage node can use to store new data in its current state. This primarily includes the remaining storage space, but may also include the node's computing resources. Status includes the storage node's current operational state or health index, such as performance status and data storage status.
[0126] In one example, analyzing the data compression results can involve parsing the compression logs generated by the storage system, extracting key fields, and aggregating the compression indicators of each node using a time window. Calculating the deviation between the node compression rate and the historical baseline (the average of the same time period over the past 7 days) is used to evaluate the compression rate performance, and predicting future storage requirements based on the compression results. remaining ,like:
[0127]
[0128] Among them, R i (t) is the expected compression rate of node i at time t, which can be set based on historical compression. compressed,i is the expected compression demand of node i, S total For the total storage space.
[0129] Step S103: determining the initial lifetime and storage type of the storage node based on the remaining available resources and the status.
[0130] It should be understood that the more remaining available resources there are, the longer the storage node can continue to work normally. When setting the basic life span of the storage node, it is necessary to make a determination based on the remaining available resources of the current storage node.
[0131] It's understandable that different types of storage nodes are suitable for different storage types. If a storage node's remaining available resources show a large amount of storage space but relatively slow read and write speeds, it's suitable for large-capacity data storage with low read and write speed requirements, such as tape storage or standard mechanical hard disk storage, and is suitable for scenarios like data backup. If the remaining resources show a storage node with high read and write speeds but relatively small storage space, it's more suitable for flash storage, used for critical data storage with high read and write speed requirements, such as database cache data storage.
[0132] In one example, the remaining available resources can be calculated by the future storage demand S remaining The status is indicated by a health score, which is determined by failure rate and response time. Different storage types have different classification logic. High-performance storage, such as SSDs, offers low latency and high IOPS, and is suitable for hot data. Balanced storage balances performance and cost. Archive storage, such as object storage, offers high capacity and low cost, and is suitable for infrequently accessed data. The formula for calculating the initial time to live (Base_TTL) is as follows:
[0133]
[0134] Among them, P health Score health, T max is the preset maximum lifetime. space and W health They are the preset weights of the remaining available resources and health score, which are initially 0.6 and 0.4, respectively, and can be adjusted according to the adaptation of the last initial survival time.
[0135] In special circumstances, the initial lifetime also needs to be dynamically adjusted. For example, when the CPU or memory load is greater than 70%, the Base_TTL is shortened by 20%; when the bandwidth is less than 50%, the Base_TTL is extended by 15%.
[0136] Step S104: Initialize the storage node according to the storage type.
[0137] It is understandable that the storage node is initialized based on the determined storage type. Storage types can be diverse, such as local disk storage, distributed file system storage, cloud storage, etc., and different storage types have different characteristics and requirements. In this step, the system will perform a series of initialization tasks on the storage node based on the specified storage type. This may include creating the necessary storage directories, configuring storage parameters, establishing connections to storage services, and other operations. In this way, the storage node can be correctly configured and prepared according to the established storage type, laying the foundation for subsequent data storage and management.
[0138] In the third embodiment, step S20 includes:
[0139] Step S201: Acquire the data sensitivity of each data object in the power plant data.
[0140] It should be understood that data sensitivity refers to the sensitivity of the information contained in each data object in the power plant data. It is determined by the type of each data object. The sensitivity of each data is preset by the management personnel. For example, Top Secret 5: core control system parameters, which directly affect the security of the power grid; Confidential 4: user electricity bills, contract electricity prices, involving user privacy or commercial secrets; Internal 3: business operation data such as equipment maintenance records and fuel purchase volume; Restricted Public 2: non-sensitive data that requires authorized access, such as power generation statistical reports; Public 1: environmental monitoring summaries and other data that can be released to the public.
[0141] Step S202 : determining a first correlation between the data object and the production business and a second correlation between the data object and the business system.
[0142] It should be noted that the first degree of association represents the closeness of a data object to the power plant's production operations. For example, the real-time operating parameters of a generator set can be determined by assessing its role and impact on the production process. The second degree of association refers to the degree of association between the data object and the power plant's business systems, including power generation management systems, equipment monitoring systems, and financial management systems. This data is frequently used and interacts with multiple business systems. The second degree of association of the data object can be determined by analyzing its usage in various business systems, including data input, output, interaction, and sharing.
[0143] Step S203: determining the impact scope of the data object based on a preset evaluation process.
[0144] It should be noted that the preset assessment process refers to a series of standardized and normalized operating steps and methodologies that are pre-established before the impact scope assessment of the power plant's data objects is conducted. It can be designed based on the power plant's business characteristics, management needs, and industry-related standards and best practices.
[0145] It is understandable that, for a data object of a power plant, the scope of impact refers to the extent of the impact on various aspects of the power plant when a problem occurs with the data object (such as data loss, error, leakage, etc.).
[0146] Step S204 : grading each of the data objects according to the data sensitivity, the first relevance, the second relevance, and the impact range to obtain a data level of each of the data objects.
[0147] It is understandable that when rating each of the data objects, different weights can be set for data sensitivity, first correlation, second correlation and scope of influence according to the rating rules determined by historical data, and each factor can be scored according to the specific situation, and then the total score is calculated, and the data level is determined based on the total score.
[0148] In one example, the data sensitivity is F. The first relevance is the production relevance F1, which is related to the production business and is determined by the production number I and the production time α of the data object. T Determine, that is, F1 = I × log 10 (α T +1). The first correlation is the business correlation F2, which can be calculated by the number of system calls U and the number of cross-system references C in a unit time (hour). ref Calculation, that is, F2=U+C ref Impact range F dIt is expressed as the product of data sensitivity and the number of system calls F U. Finally, based on the average of the data sensitivity, the first correlation degree, the second correlation degree, and the impact range, and rounded up, the rated data is obtained, and the data level corresponding to the rated data is selected in the statistical level interval.
[0149] In this embodiment, the remaining available resources and status of storage nodes are determined by analyzing the data compression results of the previous period, and the storage nodes, their lifetimes, and storage types are then initialized. This approach, based on actual resource and status information, makes storage node initialization more scientific and reasonable, reducing performance issues or resource waste caused by improper initialization. Data objects are also graded based on their sensitivity, relevance to production operations and business systems, and scope of impact. This allows for a comprehensive and accurate assessment of data importance, enabling differentiated data storage management and improving the storage system's ability to protect data of varying importance.
[0150] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the data storage method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0151] This application also provides a data storage device, please refer to Figure 4 , the data storage device comprises:
[0152] The node initialization module 10 is used to obtain the power plant data generated in the current time period of the power plant and initialize the storage nodes and the corresponding life time of the storage nodes based on the data compression result of the previous time period;
[0153] a level determination module 20, configured to determine the data level of each data object in the power plant data;
[0154] a data storage module 30, configured to store the data object in the storage node according to the data level, and activate the survival time to obtain an initial storage set;
[0155] The storage adjustment module 40 is configured to adjust the initial storage set based on the monitored change of the survival time when the survival time starts to shorten, so as to obtain a target storage set.
[0156] The data storage device provided by this application, which adopts the data storage method of the above-mentioned embodiment, can solve the technical problem that existing storage nodes cannot dynamically adjust with changes in data importance or fluctuations in system load, resulting in the premature elimination of key data or the long-term retention of redundant data, affecting the storage of new data and overall storage performance. Compared with the existing technology, the beneficial effects of the data storage device provided by this application are the same as the beneficial effects of the data storage method provided by the above-mentioned embodiment, and the other technical features of the data storage device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0157] The present application provides a data storage device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data storage method in the above-mentioned embodiment one.
[0158] Reference below Figure 5 , which shows a schematic diagram of the structure of a data storage device suitable for implementing the embodiments of the present application. The data storage device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The data storage device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0159] like Figure 5As shown, the data storage device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. Random access memory 1004 also stores various programs and data required for the operation of the data storage device. Processing device 1001, read-only memory 1002, and random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the data storage device to communicate with other devices wirelessly or by wire to exchange data. Although the figures show a data storage device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have instead.
[0160] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0161] The data storage device provided in this application, utilizing the data storage method of the aforementioned embodiment, can resolve the technical problem that existing storage nodes cannot dynamically adjust with changes in data importance or fluctuations in system load, resulting in the premature elimination of critical data or the long-term retention of redundant data, affecting the storage of new data and overall storage performance. Compared to the prior art, the beneficial effects of the data storage device provided in this application are the same as those of the data storage method provided in the aforementioned embodiment, and the other technical features of the data storage device are the same as those disclosed in the method of the previous embodiment, and are not further described here.
[0162] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0163] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0164] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, and the computer-readable program instructions are used to execute the data storage method in the above embodiment.
[0165] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0166] The computer-readable storage medium may be included in a data storage device, or may exist independently without being assembled into a data storage device.
[0167] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the data storage device, the data storage device executes the data storage method described above.
[0168] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0170] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0171] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned data storage method. It can solve the technical problem that existing storage nodes cannot dynamically adjust with changes in data importance or system load fluctuations, resulting in the premature elimination of key data or the long-term retention of redundant data, affecting the storage of new data and overall storage performance. Compared with the existing technology, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the data storage method provided in the above-mentioned embodiment, and will not be repeated here.
[0172] The present application also provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned data storage method when executed by a processor.
[0173] The computer program product provided in this application can address the technical problem that existing storage nodes cannot dynamically adjust with changes in data importance or fluctuations in system load, resulting in the premature elimination of critical data or the long-term retention of redundant data, which affects the storage of new data and overall storage performance. Compared with the existing technology, the beneficial effects of the computer program product provided in this application are the same as those of the data storage method provided in the above-mentioned embodiments, and will not be elaborated here.
[0174] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A data storage method, characterized in that: The data storage method includes: Obtaining power plant data generated in the current time period of the power plant, and initializing storage nodes and the corresponding survival times of the storage nodes based on the data compression results of the previous time period; determining a data level of each data object in the power plant data; storing the data object in the storage node according to the data level and activating the survival time to obtain an initial storage set; When the survival time starts to shorten, the initial storage set is adjusted based on the monitored change of the survival time to obtain a target storage set.
2. The data storage method according to claim 1, wherein: The initial storage set includes core nodes, buffer nodes and eliminated nodes; The step of adjusting the initial storage set based on the monitored change of the survival time to obtain the target storage set when the survival time begins to shorten includes: When the survival time begins to shorten, monitoring the change of the survival time in real time; If the change is that the first survival time of the eliminated node decreases to a first safety threshold, migrating the first data object in the eliminated node to the buffer node, and recycling the eliminated node to obtain a first adjustment result; If the change is that the second survival time of the core node decreases to a second safety threshold, migrating the second data object of the core node to the eliminated node to obtain a second adjustment result; The initial storage set is adjusted based on the first adjustment result and / or the second adjustment result to obtain a target storage set.
3. The data storage method according to claim 2, wherein: If the change is that the first survival time of the eliminated node decreases to a first safety threshold, migrating the first data object in the eliminated node to the buffer node and recycling the eliminated node, after obtaining the first adjustment result, the step further includes: Determine whether the fluctuation of the current survival time of the buffer node exceeds a preset safety range; When the fluctuation of the lifetime of the buffer node exceeds a preset safety range, calculating a Shannon entropy value based on the distribution of the third data object in the buffer node; Detecting whether the buffer node meets a preset reallocation condition based on the Shannon entropy value, and obtaining an entropy value detection result; If the entropy value detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is greater than the high entropy alarm threshold, the anti-entropy storage optimization algorithm is started to reallocate the buffer node; If the entropy detection result shows that the buffer node meets the preset reallocation condition and the Shannon entropy value is less than the low entropy optimization threshold, the buffer node is reallocated according to the data concentration of the buffer node.
4. The data storage method according to claim 2, wherein: If the change is that the second survival time of the core node decreases to a second safety threshold, the step of migrating the second data object of the core node to the eliminated node to obtain a second adjustment result includes: If the change is that the second survival time of the core node decreases to a second safety threshold, generating a migration queue based on the data level of the second data object stored in the core node; Setting a third survival time for the migration queue according to the second survival time; The second data object is migrated to the eliminated node based on the set migration queue, and the first survival time of the eliminated node after the migration is completed is adjusted according to the third survival time to obtain a second adjustment result.
5. The data storage method according to any one of claims 1 to 4, wherein: The step of obtaining power plant data generated in the current time period of the power plant and initializing the storage node and the survival time corresponding to the storage node based on the data compression result of the previous time period includes: Obtain power plant data within the current time period from the business system through the data acquisition interface; Analyze the data compression results of the previous time period to determine the remaining available resources and status of the storage node; Determine the initial lifetime and storage type of the storage node based on the remaining available resources and the status; Initialize the storage node according to the storage type.
6. The data storage method according to any one of claims 1 to 4, characterized in that: The step of determining the data level of each data object in the power plant data includes: Obtaining data sensitivity of each data object in the power plant data; Determining a first degree of association between the data object and a production business and a second degree of association between the data object and a business system; Determine the impact scope of the data object based on a preset assessment process; Each of the data objects is graded according to the data sensitivity, the first relevance, the second relevance, and the impact range to obtain a data level of each of the data objects.
7. A data storage device, characterized in that The device comprises: A node initialization module is used to obtain power plant data generated in the current time period of the power plant and initialize the storage nodes and the corresponding lifespan of the storage nodes based on the data compression results of the previous time period; a level determination module, configured to determine the data level of each data object in the power plant data; a data storage module, configured to store the data object in the storage node according to the data level, and activate the survival time to obtain an initial storage set; The storage adjustment module is configured to adjust the initial storage set based on the monitored change of the survival time when the survival time starts to shorten, so as to obtain a target storage set.
8. A data storage device, characterized in that The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data storage method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data storage method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the data storage method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Cluster head selection method, cluster head selection system and multi-hop network
CN105744584A
Multi-copy data object management method, medium and distributed storage system
CN115562581A
Data storage method, computing device and data storage system
CN115686925A
Storage strategy optimization method based on data life cycle
CN118466858A
Data storage method and device for power internet of things, and storage medium
CN119415035A
Cited By
MES-based production data management system and method
CN121303569A