Data processing method, server, medium and program product for power consumption data
By using a distributed file system and data sharding technology, the performance bottleneck of centralized servers when processing large-scale electricity consumption data is solved, achieving efficient storage and real-time querying, reducing storage costs, and optimizing data lifecycle management.
Patent Information
- Application Number
- CN202411006262.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-07-25
AI Technical Summary
In existing technologies, centralized servers experience a decrease in data processing and response speed when handling large-scale electricity consumption data, affecting the real-time performance of data query and analysis. Furthermore, the lack of effective data lifecycle management leads to high storage costs and insufficient performance.
By employing a distributed file system and data sharding technology, electricity consumption data is sharded and stored across multiple data nodes. Through metadata management, combined with methods for handling missing values and identifying cold data, efficient data storage and retrieval are achieved.
It improves data processing efficiency and system response speed, ensures real-time data query and analysis, reduces storage costs, and optimizes data lifecycle management.
Smart Images

Figure CN118689860B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of electric digital data processing, and in particular to a data processing method for electric data, a server, a medium and a program product. BACKGROUND
[0002] With the development of data technology, the application of big data in multiple industries has become more and more widespread, especially in energy management and power monitoring systems. Using big data technology to analyze and process electric data can not only optimize the allocation of electric power resources, but also improve the efficiency of energy use, thereby helping to reduce energy waste and improve the economy and reliability of the power system. The processing of electric data usually involves collecting and storing a large amount of electric data in a centralized server or database system.
[0003] With the increase of electric information collection frequency and the large-scale deployment of smart meters, the amount of data is growing explosively. On the one hand, due to the demand of electric planning, the electric data of a single user accumulates over time, and there is a large amount of historical electric data in the time dimension; on the other hand, with the popularization of smart grid, the number of users connected to the server is also increasing, and the data base is increasing in the space dimension.
[0004] In related technologies, the server usually processes and analyzes the collected data uniformly to identify the electric mode, predict future power demand, etc. However, when the amount of electric data is extremely large, the processing and response speed of a single server or centralized database system in related technologies will be greatly reduced, affecting the real-time performance of data query and analysis. SUMMARY
[0005] The present application provides a data processing method for electric data, a server, a medium and a program product, which are used to store and query large-scale electric data, and can improve the efficiency of data processing and the response speed of the system, and ensure the real-time performance of data query and analysis.
[0006] In a first aspect, the present application provides a data processing method for power consumption data, applied to a server, the method comprising: obtaining power consumption data of a plurality of users; dividing the power consumption data based on a preset data sharding rule to obtain a plurality of data shards, such that each data shard contains a preset number of power consumption data and the data volume is less than a preset threshold; storing the plurality of data shards on a plurality of data nodes in a distributed file system respectively, and recording the correspondence between each data shard and the data node; in a management node of the distributed file system, generating metadata based on the correspondence; the metadata is used to represent the data node where each data shard is located; when receiving a power consumption data query request for a target time period sent by a query client, obtaining a plurality of target data shards corresponding to the target time period; according to the metadata, sending a data reading request to the data node where the target data shard is located to obtain target power consumption data of the target data shard; according to the target power consumption data, generating a power consumption data analysis result corresponding to the target time period, and returning the power consumption data analysis result to the query client.
[0007] In the above embodiment, the server effectively shards and distributes large-scale power consumption data, thereby improving data processing efficiency and scalability of the server. By storing data shards on multiple data nodes and generating metadata in the management node that describes the location of each data shard, the data retrieval process can be optimized. The ability of the server to process large data is enhanced, thereby ensuring the real-time performance of data query and analysis.
[0008] In combination with some embodiments of the first aspect, in some embodiments, the plurality of data shards are stored on the plurality of data nodes in the distributed file system respectively, and the correspondence between each data shard and the data node is recorded, specifically comprising: for each data shard, obtaining a list of currently available data nodes in the distributed file system; based on a consistent hashing algorithm, selecting a preset number of candidate data nodes from the list of available data nodes; obtaining historical load data, network delay data, and geographic location distance data in the candidate data nodes; according to a preset node selection strategy, combining the historical load data, network delay data, and geographic location distance data, selecting a target data node from the candidate data nodes; storing the data shard to the target data node, and storing the correspondence between the data shard and the target data node in a preset metadata management system.
[0009] In the above embodiment, for each data shard to be stored, the server first obtains a list of currently available data nodes, and then selects a plurality of candidate nodes, which can avoid node overload and ensure the balance of data distribution. In the candidate nodes, further obtain multi-dimensional attribute data such as historical load, network delay, and geographic location distance, and comprehensively consider these factors to determine the final target storage node, which optimizes the data shard storage from multiple perspectives.
[0010] In some embodiments of the first aspect, according to the target electricity consumption data, the electricity consumption data analysis result corresponding to the target time period is generated and returned to the query client, specifically including: calculating the target electricity consumption data according to the user dimension to obtain the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption and electricity consumption curve of each user in the target time period; calculating the target electricity consumption data according to the time dimension to obtain the total electricity consumption and electricity peak period of each day in the target time period; calculating the target electricity consumption data according to the region dimension to obtain the electricity consumption distribution of users in different regions; generating a multi-dimensional electricity consumption data analysis report according to the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption, electricity consumption curve, total electricity consumption, electricity peak period and electricity consumption distribution; the electricity consumption data analysis report includes an electricity consumption statistics table, an electricity load curve diagram and an electricity region distribution diagram; generating a report webpage link according to the electricity consumption data analysis report and sending the report webpage link to the query client; generating an electricity consumption data analysis result according to the electricity consumption data analysis report and returning the electricity consumption data analysis result to the query client.
[0011] In the above embodiments, the server performs comprehensive summary analysis on the data from different dimensions such as users, time and regions, obtains various electricity consumption indicators and macro electricity consumption patterns of users, generates diversified data analysis reports including intuitive charts such as electricity consumption statistics, load curve and region distribution, and facilitates users to overview the characteristics and laws of their electricity consumption data, so that the originally complex mass data becomes clear and easy to read.
[0012] In some embodiments of the first aspect, after the step of obtaining the electricity consumption data of the plurality of users, the method further includes: detecting missing values in the electricity consumption data, and determining the electricity consumption data value corresponding to the missing values based on other electricity consumption data of the same user in the time period in which the missing values are located; detecting duplicate records in the electricity consumption data, and retaining one of the completely duplicated records in the duplicate records and merging the records with partially duplicated fields; detecting illegal values in the electricity consumption data, and replacing the illegal values with a threshold value or performing format conversion; the illegal values include values exceeding a preset value range and non-numeric values.
[0013] In the above embodiments, after obtaining the original electricity consumption data, the server improves the data quality through a data cleaning process. First, missing values are detected and processed, the missing values are estimated by analyzing other complete data of the target user in the same time period, and the continuity of the data is ensured. Then, duplicate data is identified, and for completely duplicated records, only one is retained, and for partially duplicated records, field merging is performed to eliminate data redundancy. Then, data that exceeds the value range or has illegal format is corrected, the value is limited within a reasonable range and the format is standardized, and the standardization and credibility of data analysis are enhanced.
[0014] In some embodiments of the first aspect, in some embodiments, missing values in the electricity consumption data are detected, and based on other electricity consumption data of the same user in the time period where the missing values are located, the electricity consumption data values corresponding to the missing values are determined, specifically comprising: obtaining the original electricity consumption data of the target user in the target time period, and sorting the original electricity consumption data according to the collection time to obtain a sequence set of electricity consumption data; traversing the sequence set of electricity consumption data, and after determining that the interval period of the collection times of two adjacent original electricity consumption data is greater than a preset time interval, taking a number of electricity consumption data corresponding to the interval period as missing values; obtaining the historical electricity consumption data of the same user in the time period corresponding to the interval period, and calculating the average value of the historical electricity consumption data as an initial estimated value corresponding to the missing values; obtaining the electricity consumption data values of a number of adjacent non-missing sampling points before and after the interval period, and correcting the initial estimated value of the missing values by interpolation fitting to obtain a corrected estimated value; filling the corrected estimated value to the corresponding missing value position of the interval period to determine the electricity consumption data values corresponding to the missing values.
[0015] In the above embodiments, when the server detects that there are missing values in the electricity consumption data, it first obtains the original data of the target user and sorts it according to the collection time to form a time series. Then, the server traverses the time series, identifies data segments with a collection time interval exceeding a predetermined threshold, and marks them as missing values. For these missing values, the server first obtains the historical electricity consumption data of the target user in the time range corresponding to the missing values from the history library, calculates the average value of these data as the initial estimate of the missing values, and then corrects the initial estimate value by interpolation fitting to obtain a more accurate missing value.
[0016] In some embodiments of the first aspect, in some embodiments, after the step of sending a data reading request to the data node where the target data shard is located according to the metadata to obtain the target electricity consumption data of the target data shard, the method further comprises: recording the creation time, update time and access time of each data shard; marking a data shard as cold data when the access frequency of the data shard is lower than a preset threshold and the time since the last access exceeds a preset time length; migrating the data shard marked as cold data from the distributed file system to the object storage system to save storage costs; archiving or deleting the cold data whose storage time exceeds a preset deadline after the preset deadline.
[0017] In the above embodiment, the server can master the life cycle information of the data by recording the creation time, update time and last access time of the data shards. If the access frequency of a certain shard is lower than a set threshold and the time interval from the last access exceeds a specified period, the shard is marked as cold data. For the cold data, the server migrates the cold data from the distributed file system to the object storage system with lower cost, to realize the tiered storage of hot and cold data, and significantly save the storage cost while ensuring the data security.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of archiving or deleting the cold data whose storage time exceeds the preset period, the method further comprises: storing the data feature information of the deleted cold data into a preset data feature management server; when receiving a user query request and the requested data is the deleted cold data, obtaining the data feature similar to the requested data from the data feature management server and returning the data feature to the user as the substitute information of the requested data; when receiving a user query request and the requested data only partially belongs to the deleted cold data, obtaining the summary statistical information of the deleted part from the data feature management server, and returning the original data of the non-deleted part and the summary statistical information of the deleted part to the user; when receiving a user query request and the requested time range exceeds a preset threshold, dividing the data in the requested range according to a preset time interval, returning the original data of the non-deleted time period, and obtaining the summary statistical information of the deleted time period from the data feature management server and returning the summary statistical information.
[0019] In the above embodiment, the deletion of the cold data by the server can cause data loss and query failure. When deleting the cold data, the server extracts the key feature information of the cold data and stores the key feature information into an independent metadata management system. When a user queries the deleted data, the server compares the query condition and finds the data most similar to the feature information, and returns the data to the user as the substitute. If the query condition spans multiple data shards, the original detailed data is returned for the non-deleted shards, and the summary statistical information is returned for the deleted shards, to meet the query requirement as much as possible. When the time span of the query is too long, the server divides the data into multiple sub-intervals according to the data feature, and returns the data in different granularities for the deleted and non-deleted intervals, to maximize the queryability of the data.
[0020] In the second aspect, the embodiments of the present application provide a server, which comprises one or more processors and a memory. The memory is coupled to the one or more processors, and is used to store computer program codes, the computer program codes comprising computer instructions. The one or more processors invoke the computer instructions to enable the server to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0021] In a third aspect, the embodiments of the present application provide a computer program product comprising instructions which, when executed on a server, cause the server to carry out the method according to the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, the embodiments of the present application provide a computer-readable storage medium comprising instructions which, when executed on a server, cause the server to carry out the method according to the first aspect and any possible implementation of the first aspect.
[0023] It can be understood that the server provided by the second aspect, the computer program product provided by the third aspect, and the computer storage medium provided by the fourth aspect are all used to execute the method provided by the embodiments of the present application. Therefore, the beneficial effects that can be achieved thereby can refer to the beneficial effects in the corresponding method, which will not be described here again.
[0024] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0025] 1. Since the distributed file system and the data sharding technology are adopted, the method can store the power consumption data in a distributed manner on different data nodes, effectively solving the data processing bottleneck and single point failure problem in the prior art, and thereby realizing high scalability and high availability of the server. In a traditional centralized data storage system, all data is processed on a single or a few servers, which not only easily leads to server overload, but also affects the availability of the entire server when the server fails. The method greatly reduces the burden on a single node and improves the processing speed and efficiency by uniformly distributing data to multiple nodes and processing corresponding data shards by each node.
[0026] 2. Since the data repair method based on historical data of the same user in the time period where the missing value is located and interpolation correction is adopted, after detecting the missing value in the power consumption data whose collection time interval exceeds the threshold, the average value of the historical data of the same user in the same time period is first calculated as the initial estimation of the missing value, and then the complete sampling points before and after the missing value are selected to correct the initial estimation by interpolation fitting to obtain a more accurate estimation of the missing value, and finally the estimated value is backfilled to the missing point, so that the overall characteristics of the user's own power consumption mode can be reflected, and the local features of the context of the missing value can be calibrated, effectively solving the problem that the actual power consumption of the user cannot be reflected in the prior art, resulting in a large estimation deviation and affecting the accuracy of data analysis.
[0027] 3、Due to the cold data identification method based on access frequency and time interval, and the cold data-oriented hierarchical storage and life management strategy, by tracking the creation, update and access time of each data shard, when it is found that the access frequency and access interval of a certain shard exceed the set threshold, it can be determined as cold data, and it is migrated from the distributed file system to the low-cost object storage, and the cold data exceeding the archival period is periodically deleted, so that the hot and cold data can be dynamically divided according to the actual access mode of the data, and matched to the appropriate storage medium, thereby reducing the storage cost while improving the access efficiency of hot data, and effectively solving the problem that in the prior art, for massive multi-source heterogeneous power consumption data, there is a lack of data life cycle management means, the data retention period is difficult to control, and the overall performance is affected. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a flowchart of a data processing method of power consumption data in an embodiment of the present application;
[0029] Figure 2 is another flowchart of a data processing method of power consumption data in an embodiment of the present application;
[0030] Figure 3 is an entity device structure diagram of a server in an embodiment of the present application. DETAILED DESCRIPTION
[0031] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be limiting to the present application. As used in the specification of the present application, the singular expression "one", "a", "the", "said" and "this" are intended to include the plural expression, unless there is clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application means any or all possible combinations of one or more listed items.
[0032] Hereinafter, the terms "first", "second" are only for the purpose of description, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more.
[0033] In order to facilitate understanding, the application scenarios of the embodiments of the present application are introduced as follows.
[0034] A large power company needs to collect, store and analyze the meter data of millions of users in its jurisdiction in order to realize the functions of power consumption monitoring, load forecasting, electricity fee calculation and other business functions in the whole region. However, with the large-scale deployment of smart meters and the increase of power consumption information collection frequency, the power consumption data generated every day is growing explosively, with a data volume reaching TB level; at the same time, with the expansion of the user scale and the refinement of the power consumption information collection granularity, the amount of data added every day is as high as tens of billions, and the historical data reaches hundreds of billions, which brings great challenges to data management and calculation analysis. In the face of such a large amount of data, the traditional centralized storage architecture and relational database system has been difficult to meet the performance requirements of real-time writing and fast querying.
[0035] In the related art, centralized management and query analysis of power consumption data can be realized by using centralized relational database storage. The following introduces the scenario of using the power consumption data processing method in the related art.
[0036] A power company uses a centralized storage scheme based on a relational database to process power consumption data. At the beginning, the system runs smoothly because the data volume is small. However, with the increase in the number of users and the upgrade of collection equipment, the data volume and concurrency are constantly increasing, and the original system begins to show bottlenecks: on the one hand, the storage capacity and processing capacity of a single database server are close to saturation, and cannot cope with the growing data size; on the other hand, the write and query performance of the database gradually decreases, especially in high concurrency scenarios, the response time is greatly extended, and even the system crashes. The operation and maintenance personnel try to alleviate the pressure by vertically expanding the database server configuration and introducing data partitioning and read-write separation mechanisms. In addition, due to the multi-dimensional nature of power consumption data, many business-oriented queries and analysis requirements, such as cross-table aggregation, ad hoc queries, etc., are difficult to efficiently implement in a relational database.
[0037] However, by using the power consumption data processing method in the embodiments of the present application, the massive power consumption data is divided into several small shards, and the shards are vertically divided into independent column files according to fields such as degrees, voltage, etc., and then uniformly distributed and stored in a large-scale cluster composed of ordinary PCs, which can make full use of the high compression ratio and flexible coding method of columnar storage, and better support complex query analysis of multi-dimensional attributes.
[0038] The following introduces the scenario of using the power consumption data processing method in the present application.
[0039] In the face of the above challenges, the power company decided to introduce the distributed storage scheme based on data sharding and columnar storage of the present application. First, the server recommends and determines reasonable data sharding rules according to the time and space attributes of the data, for example, sharding by day and region, so that the data volume of a single shard is controllable, while the local spatio-temporal characteristics are retained. Then, the server organizes each shard in a columnar storage format, and the data is vertically divided into independent column files according to fields such as collection time, user ID, and degree, and for each column, a suitable encoding and compression algorithm is used to reduce storage space while optimizing query performance.
[0040] On the storage medium, the server uses HDFS (Hadoop Distributed File System) to uniformly store the data shards into a large-scale cluster composed of hundreds of ordinary PCs, and through the multi-replica fault-tolerant mechanism, the high reliability of the data is guaranteed. The metadata service is responsible for maintaining the mapping relationship between the shards and the data nodes, and the client can quickly locate and retrieve data according to the shard key. Thanks to the distributed storage architecture and columnar storage model, the storage capacity and write throughput of power consumption data have been greatly improved, and the server can handle tens of thousands of high-concurrency writes per second. At the same time, the query engine can read multiple shard data in parallel according to the metadata, and quickly generate results through distributed aggregation operations, even for complex multi-table association and statistical analysis, it can return within milliseconds.
[0041] As can be seen, the data processing method for power consumption data in the embodiments of the present application can effectively solve the problems of capacity limitation, expansion difficulty and performance bottleneck of traditional single-machine databases in the scenario of massive data, while realizing high-concurrency writing and real-time querying of power consumption data, thereby realizing performance improvement of power consumption data storage and management.
[0042] For ease of understanding, the method provided by the present embodiment will be described in the following flow in combination with the above scenario. Please refer to Figure 1 , a flowchart of the data processing method for power consumption data in the embodiments of the present application.
[0043] S101, acquire power consumption data of a plurality of users.
[0044] Wherein, the user represents an electric energy consumer in the power server, and is usually identified by a unique user identifier such as a house number, a meter number, etc. The power consumption data refers to the electric energy consumption record of the user within a period of time, and usually includes multiple indexes such as power consumption, voltage, current, power factor, etc. and corresponding timestamp information. Specifically, during the operation of the power server, a large number of intelligent meter devices distributed in various regions continuously monitor and record the power consumption information of the user. The server will maintain a connection with these devices, and periodically obtain incremental data according to a preset collection frequency. The obtained original power consumption data is usually stored in a source file or a message queue in a text, binary or other format, and needs to be subjected to a series of preprocessing operations such as parsing, conversion and cleaning, so as to form a standardized data format before being imported into a subsequent storage system.
[0045] S102, divide the power consumption data based on a preset data slicing rule to obtain a plurality of data slices, so that each data slice contains a preset number of power consumption data and the data volume is less than a preset threshold.
[0046] Wherein, the data slicing refers to the process of dividing a large-scale data set into a plurality of small pieces of data that are easy to manage and store according to a certain rule. The preset data slicing rule refers to a condition or strategy defined in advance for slicing data according to the characteristic attributes of the data, such as slicing according to time range, spatial region, numerical interval or Hash value, etc. The preset number and the preset threshold refer to the number of data rows and the upper limit of the data volume that each slice should contain, which are set by the developer based on experience and are usually matched with the processing capacity and management granularity of the storage system.
[0047] Specifically, this step is performed after a large amount of power consumption data is obtained, and the purpose is to divide the data into small pieces for parallel storage and processing. For the power consumption scenario, since the number of user groups and power consumption behaviors in different regions are significantly different, the data distribution is usually uneven. If simply sliced according to a fixed time or space span, it may cause some slice data to be too large to affect storage performance, and some slice data to be too small to cause storage space waste. In order to fully parallelize while considering load balancing, the present application adopts an adaptive data slicing method: first, the server divides the data into a plurality of large partitions according to the space-time attributes of the data, such as by day or by county-level administrative division, and then dynamically slices the data in each partition based on the quantity and volume thresholds. In this way, the local correlation of slice data in the same partition is ensured, and the uneven distribution of slice data is avoided. For example, the power consumption data within the city range every day is first divided by county, and then the data of each county is sliced until the number of users contained in each slice is less than 100,000 and the number of data rows is not more than 1 million. The slicing process can be realized by means of a parallel computing framework such as MapReduce.
[0048] In some embodiments, the adaptive sharding of power consumption data can be achieved in various ways: optionally, a multi-level Hash sharding method is adopted. For each piece of power consumption data, the user ID is extracted first and mapped to a first-level shard through a Hash function, and then the date part of the data timestamp is extracted and mapped to a second-level shard corresponding to the first-level shard through Hash. The server records the data size of each second-level shard, and triggers the splitting of the second-level shard when it exceeds the preset threshold. When splitting, a new shard key such as the hour part of the data timestamp is introduced on the basis of the original second-level shard, the second-level shard data is re-mapped through Hash, and the metadata record is updated.
[0049] S103, store the plurality of data shards to the plurality of data nodes in the distributed file system respectively, and record the correspondence between each data shard and the data node.
[0050] The distributed file system includes file system managed physical storage resources, which are not necessarily directly connected to the local node, but are distributed on multiple nodes connected through a network, on which data is stored in blocks to multiple data nodes, and a unified file system view and access interface is provided, with good scalability and fault tolerance. The data node represents a storage node in the distributed file system, and each node stores a part of the file data block. The correspondence indicates the mapping of which data nodes a data shard is actually stored in, and this mapping relationship is usually saved in a special metadata management node.
[0051] Specifically, after data sharding, in order to further improve the storage and query performance of data, the server uses columnar storage instead of traditional row storage, and divides the data of each shard vertically according to fields such as collection time, user ID, power consumption and voltage, organizes them into independent column files respectively, and stores them on the distributed file system. Columnar storage is a data storage method optimized for analytical workloads, as opposed to traditional row storage. Columnar storage divides a data set into multiple independent column files according to columns (attributes), and each column only contains data of one field. The column name and offset can be used to quickly locate and extract the specified column value of any row, which is very suitable for multi-dimensional aggregation computing scenarios. Compared with row storage, column storage significantly compresses the data volume and speeds up query response.
[0052] At the same time, with the help of the massive storage capacity of the distributed file system, data shards can be evenly stored on multiple inexpensive general-purpose hardware storage nodes, achieving linear expansion of storage capacity. The distributed file system ensures high reliability of data storage through shard replica management, fault recovery and other mechanisms, so that even if individual storage nodes fail, data loss will not occur.
[0053] S104, in the management node of the distributed file system, generating metadata based on the correspondence; the metadata is used to represent each data node where the data shard is located.
[0054] Wherein, the management node is the core node responsible for metadata management and access control in the distributed file system, usually holding the directory tree of the entire file system, the mapping table of data blocks and storage nodes, and other key information. Metadata is data used to describe data, in this step, it specifically refers to a data structure used to record the storage location of each data shard, usually in the form of key-value pairs, tables, etc. The correspondence between data shards and data nodes has been obtained in the previous step, which is the basis for generating metadata. For example, in HDFS, NameNode as the management node maintains the mapping metadata of file data blocks and DataNode, and can directly obtain the DataNode information where the shard is located when querying the data shard.
[0055] Specifically, this step is executed after the data shard is stored, and the purpose is to provide data routing basis for subsequent distributed query. In order to realize the association of cross-node shard data, it is necessary to record and query the mapping relationship between each shard and storage node in a unified metadata management component. The server uses the built-in metadata management function of the distributed file system to extract shard ID and storage node IP information based on the shard storage log in the management node, build an easily searchable metadata structure (such as HashMap), and persist it to the meta database. In this way, when an application queries the data of a certain shard, it can find the storage node where the target shard is located through the metadata service, thereby avoiding one-time full table scan and improving data retrieval efficiency.
[0056] In some embodiments, the generation and management of shard metadata can be implemented in various ways: optionally, for the scenario of storing data shards based on HDFS and other general distributed file systems, the metadata management function of HDFS NameNode is used. When a data node uploads a shard data block, it will register with the NameNode and submit metadata such as data block ID, storage node, and checksum. The NameNode is responsible for updating the metadata mapping table in memory and synchronizing it to the FsImage and EditLog files on disk. When querying, the client requests the metadata corresponding to a certain shard ID from the NameNode, and the NameNode reads and returns all data blocks and their storage locations of the shard from the in-memory table. Since metadata access is coordinated by the NameNode, distributed consistency problems are avoided, but the NameNode may also become a single point of bottleneck for the server.
[0057] It can be understood that according to different storage modes of data shards, the specific form and management method of metadata are also different. In addition to the above two common implementations, a Redis or other in-memory database can be used to centrally store a mapping table, or shard metadata can be directly embedded into a data file. As long as the mapping between the data shard ID and the storage location can be established and efficient metadata retrieval is supported, it is a feasible metadata management scheme, and this step is not limited.
[0058] S105, when receiving the power consumption data query request of the target time period sent by the query client, obtaining a plurality of target data shards corresponding to the target time period.
[0059] Among them, the query client generally refers to an application program that needs to query and analyze power consumption data, and is usually deployed in a data center node or a user terminal. The query request refers to a query call sent to the server according to an agreed communication protocol, which contains the limiting conditions of the target data, such as power consumption data in a certain time period. The target data shard refers to one or more data shards containing the data of the target time period found by matching the time range metadata of the target time period and the data shard, which is the basic unit for cross-shard association.
[0060] Specifically, this step is triggered after receiving a data query request initiated by a user. In order to improve the real-time performance of the query, the server needs to quickly locate the relevant data shards according to the query conditions. The present application uses a time period matching method to find the target shard: first, the query time period parameter input by the user is parsed into a unified timestamp representation, then the time range metadata of the data shard is obtained, whether it overlaps with the target time period is judged by the timestamp interval, and finally all shard IDs containing the data of the target time period are determined. This method takes advantage of the continuity and orderliness of time series data shards, avoiding full table scanning, but may need to introduce auxiliary indexes (such as interval trees) to speed up the filtering operation of a large number of shards. After obtaining the target data shard ID, the storage location of these shards can be queried through the metadata service, and a data reading request is sent to the corresponding data node. For example, to query the power consumption data in January 2024, first find all data shards from January 1, 2024 to January 31, 2024 according to the shard metadata, then obtain the node addresses where these shards are located, and pull the shard data from each node in parallel.
[0061] In some embodiments, the retrieval of data shards by time period can be implemented in various ways: optionally, a SQL on Hadoop server such as Hive can be used to implement shard retrieval. All data shard information is organized into a Hive wide table, which contains fields such as shard ID, start time, end time, HDFS path, etc. When querying, the Where statement in Hive SQL is used to filter the wide table by time period. Hive converts the SQL into a MapReduce task, which scans the wide table in parallel in the Map phase, selects the target shard according to the time period condition, and aggregates the results in the Reduce phase and returns them to the client. Optionally, a HBase + Phoenix secondary index can be used to implement shard retrieval. A composite RowKey is established for the data shard table, in the format of "reverse timestamp shard ID", and auxiliary indexes are established for the start time and end time. Each row saves the metadata of a shard.
[0062] S106, according to the metadata, a data reading request is sent to the data node where the target data shard is located, and target power consumption data of the target data shard is obtained.
[0063] Among them, the data reading request refers to the query client calling the data reading interface provided by the distributed storage system according to the specific data access protocol, inputting the target shard ID and other parameters, and requiring to return part or all of the data of the shard.
[0064] Specifically, after obtaining the storage location information of the data shards corresponding to the target time period, the server queries the network address recorded in the metadata, establishes a connection with each data node where the shard is located, and sends a data reading request in parallel, inputting the shard ID, user ID range, sampling timestamp and other query conditions. After receiving the request, the distributed storage node finds the corresponding data file according to the shard ID, filters the file data using the query conditions, reads only the data in the target time period, and returns the structured data object after deserialization. After obtaining the shard data, the client needs to calculate the data returned by multiple nodes, which is usually to sort the data of each shard according to time, and assemble the complete target time period power consumption detail data set. If the shard data uses columnar storage, the data reading can be done more granularly, that is, only the field column appearing in the query statement is scanned, further reducing the network transmission amount. For example, to query the power consumption load of a certain enterprise in January 2024, the storage nodes corresponding to the enterprise's January power consumption records need to be found through the metadata, and then the timestamp and power consumption columns in the data of each node are requested, and finally the data returned by multiple nodes is sorted according to the timestamp to obtain the hourly power consumption load data of the enterprise in January.
[0065] S107, generating, according to the target power consumption data, power consumption data analysis results corresponding to the target time period, and returning the power consumption data analysis results to the query client.
[0066] The power consumption data analysis results are usually presented in the form of reports, charts, etc. visualized results of the aggregated calculation, so as to facilitate the user to overview the whole situation, find rules, and analyze abnormalities, and are a key means of data mining and decision support.
[0067] Specifically, this step is executed after the distributed reading of the target shard data is completed. Since the shard data read from different nodes in the previous step is independent and unordered, the data of multiple shards needs to be first sorted according to the timestamp and other fields, so as to obtain complete and continuous target time period raw data. Then, according to the aggregation analysis conditions in the query request, the grouping field (such as user type, belonging region, power supply unit, etc.), the measurement field (such as total power, average load, peak-to-valley ratio, etc.), and the aggregation function (such as sum, avg, max, etc.) of the aggregation operation are determined, and the sorted data is executed by using a streaming calculation or a batch processing engine to perform the aggregation operation, so as to form a multi-dimensional cross aggregation result data set. Finally, the aggregation result data is converted into a general format such as JSON, CSV, etc. suitable for front-end display, and is returned to the query client through a Web service interface. The aggregation calculation generally refers to the operation of grouping, filtering, and summarizing a group of data records according to certain dimensions and indicators, and in the field of power consumption data analysis, is commonly used for the statistics of power, electricity charges, load, etc. with multiple users, multiple indicators, and multiple time granularities. The client uses report engine, BI tool, and other visualization components to render the aggregation result into intuitive data reports and statistical charts for business personnel to browse and analyze. For example, the monthly power consumption aggregation of residential users is to use the user ID in the original data as the grouping field, use the collection time as the filtering condition, use the power as the measurement value, use the sum function to accumulate by household, and finally generate a data table of “total power consumption of each user in January”.
[0068] In some embodiments, the aggregation calculation of the target power consumption data can be implemented in various ways: optionally, the DataSet / DataFrame API of Spark SQL is used to perform power consumption data aggregation in a distributed cluster. First, SparkSQL is used to read the target shard data from columnar storage such as ORC, so as to obtain a DataSet containing various fields. Then, the groupBy( ), agg( ), sum( ), etc. conversion operators on the DataSet are called to group and aggregate the data according to the specified fields, such as:
[0069] ds.groupBy(“user_id”).agg(sum(“kwh”).as(“total_kwh”)).
[0070] Customized UDAF aggregation functions such as average voltage, mode power factor, etc. can also be registered. The aggregation results are returned in the form of a new DataFrame. Finally, the toJSON( ) or collect( ) operator of the DataFrame is called to collect the final results and return them to the client. This way can take advantage of the memory computing and lazy execution optimization of Spark SQL to realize distributed parallelization and data localization of aggregation operations, but requires joining the aggregation results of multiple shards on the client side. Alternatively, the distributed table and materialized view mechanism of ClickHouse is used to directly complete multi-dimensional aggregation on the database side. First, a logical distributed table is created using the distributed DDL statement of ClickHouse, which maps the local shard tables on multiple nodes together. Then, using the CREATE MATERIALIZED VIEW statement of ClickHouse, multiple query-oriented materialized aggregation views are defined in advance for the distributed table, such as device daily / monthly / yearly power consumption, regional 15 / 30 / 60 minute average load, etc. The sum, avg, etc. aggregation functions are used in the materialized view to precompute the aggregation metrics.
[0071] In practical applications, specific business characteristics and needs can also be customized and optimized in terms of data sharding, data compression, query optimization, etc. to further tap the application value of electricity big data. The scenario of the present embodiment is supplemented as follows.
[0072] After successfully applying the present solution for a period of time, the power company further optimized the server to cope with more demanding business requirements. On the one hand, the size and distribution strategy of the data shards were dynamically adjusted, and mechanisms such as data node load balancing, access hotspot identification, network topology awareness, etc. were introduced to improve the resource utilization of the cluster and the adaptive ability of the server. For example, when the user group in a certain region rapidly expands, the server can migrate the shard data of that region to a node with lower load and trigger a shard split operation, thereby coping with the sudden increase in data volume without interrupting service. On the other hand, the server made further optimizations in data organization and compression based on the periodic characteristics of the electricity consumption curve.
[0073] In combination with the above scenario, the method provided by the present embodiment is further described in more detail. Please refer to Figure 2 , another flowchart of the data processing method of electricity data in the present embodiment.
[0074] S201, obtaining electricity data of a plurality of users.
[0075] Referring to step S101, the server obtains electricity data.
[0076] S202, detect the missing value in the electricity consumption data, and determine the electricity consumption data value corresponding to the missing value based on other electricity consumption data of the same user in the time period where the missing value is located.
[0077] The missing value refers to invalid data such as null value and NULL value caused by missing filling and loss of some fields in the data set. The electricity consumption data usually includes multiple fields such as user number, collection time, meter reading, voltage, current, power factor, etc. Some fields may be missing due to transmission failure, equipment abnormality, etc. The time period specifically refers to the time interval of electricity consumption data collection, such as 15 minutes, 1 hour, etc.
[0078] Specifically, in order to ensure the accuracy of subsequent electricity quantity statistics and load analysis, the server must process the missing values. The server first uses NULL, empty string, special placeholder, etc. to identify various missing values in the electricity consumption data; then for each missing record, the user identifier and missing time period are extracted; then the other normal records of the user within a certain range before and after the missing time period are retrieved, and these normal values are used to estimate the missing values. Common single-variable time series interpolation methods include using the average, previous value, and next value of the missing time point to fill, using the average of the missing time period to fill, using the average of the same time of the user for multiple days to fill, using the average of the same period of similar users to fill, etc.; after the estimation is completed, it is backfilled to the empty field of the original missing record. For example, for the missing current value of an industrial user at 8:00 on May 10, 2023, the normal current values at 8:00 on May 9 and 8:00 on May 11 can be used to fill in the average. For example, for the missing electricity quantity value of a residential user on March 10, 2023, the average electricity quantity values of the same period in January and February can be used to fill in.
[0079] In some embodiments, the detection and repair of missing power consumption data can be achieved in various ways: optionally, the missing values of power consumption data are processed in a distributed cluster using the DataSet API of Spark. First, the missing records of each field with NULL, blank, and other special values are found using the filter() operator of DataSet to form a missing dataset; then the schema of the missing dataset is extracted, an empty DataFrame is created, only the necessary columns such as user ID, timestamp, and missing field name are retained, and the records of the missing dataset are aggregated into a sparse empty DataFrame according to user ID and timestamp; the window functions and LAG() and LEAD() functions of Spark SQL are used to find the previous and next n normal records of the user within a certain time range for each missing record of the user and time, and the normal records are aggregated and filled into the corresponding positions of the empty DataFrame; finally, the filled DataFrame and the original complete dataset are fully outer joined according to user ID and timestamp, and the non-empty values are selected in turn by coalesce() to obtain a new dataset with repaired missing values.
[0080] In some embodiments, the server will: obtain the original power consumption data of the target user in the target time period, and sort the original power consumption data according to the collection time to obtain a power consumption data sequence set; traverse the power consumption data sequence set, and determine that the interval period of the collection time of the adjacent two original power consumption data is greater than the preset time interval, and then use the number of power consumption data corresponding to the interval period as the missing value; obtain the historical power consumption data of the same user in the time period corresponding to the interval period, and calculate the average value of the historical power consumption data as the initial estimated value of the missing value; obtain the power consumption data values of the adjacent number of non-missing sampling points before and after the interval period, and correct the initial estimated value of the missing value by interpolation fitting to obtain a corrected estimated value; and fill the corrected estimated value to the corresponding missing value position of the interval period to determine the power consumption data value corresponding to the missing value.
[0081] Specifically, the server first calculates the average value of the historical power consumption data of the user in the missing time period as the initial estimate of the missing value. This is based on the periodicity of the user's power consumption behavior. Then, in order to further improve the accuracy of the estimate, the real sampling data of several time points before and after the missing time period are selected, and the initial estimate is corrected by interpolation fitting. Common interpolation methods include linear interpolation, spline interpolation, etc. Finally, the corrected estimated value is filled into the missing time points to obtain a relatively complete and accurate power consumption data sequence. This missing value processing method makes full use of the historical data characteristics of the user himself, rather than simply filling with fixed values.
[0082] S203, detecting repeated records in the electricity consumption data, and retaining one of the completely repeated records in the repeated records, and merging the records with partially repeated fields in the repeated records.
[0083] wherein the complete repetition refers to all fields of the two records being the same. The partial field repetition refers to only part of the primary keys or the measurement value fields of the two records being the same, that is, there is a repeated primary key, but the measurement value associated therewith is inconsistent.
[0084] Specifically, due to data retransmission, network jitter and other abnormalities in the electricity information collection process, or data introduced by customer information changes, a part of repeated records flow into the data set. The repeated records will make the electricity statistics high, and must be identified and cleaned in advance. The server will use the multi-field grouping comparison deduplication method, that is, the server will first define the key field combination of the electricity data, usually including user ID, collection time, electric energy meter ID, etc., as the joint primary key for repeated judgment; then grouping the data set according to the joint primary key, obtaining a plurality of repeated record sets; then comparing the non-primary key measurement value fields in each set, deleting the records that are completely repeated (i.e. the measurement values are also the same), and retaining only one; for the records that are partially repeated (i.e. the measurement values are inconsistent), taking the average of each measurement field to merge and update to one record and delete the rest. After deduplication and merging, the electricity data set without repeated data is obtained. For example, for two electricity details of an enterprise on June 1, 2023, 8:00, if the user number, meter number, and electricity quantity and electricity fee fields of the two records are the same, it means complete repetition, and the latter one is directly deleted; for two electricity records of the same residential user, if the user's name and address information are inconsistent, the new and old addresses can be spliced, and the record with the latest time is retained.
[0085] S204, detecting illegal values in the electricity consumption data, and replacing the illegal values with a threshold value or performing format conversion.
[0086] wherein the illegal value refers to data that exists in the field value but does not meet the required specifications, such as negative electricity quantity, collection time not within a reasonable range, missing mandatory information, decimal precision overflow, type mismatch, etc. The threshold value refers to the upper and lower boundaries set for quantitative indicators, such as voltage should be between 0 and 500 volts, and exceeding this range is abnormal. Format conversion refers to changing the data type of the original value to make it conform to the type definition of the target field, such as parsing the string type date and time into a timestamp, parsing the numerical value containing units and thousand separators into a floating point number, etc.
[0087] Specifically, the server adopts an illegal value identification and processing strategy based on data specifications. First, for each electricity consumption data indicator, the server will determine a complete set of data quality rules based on the design of the developer, including value range, accuracy requirement, format template and constraint condition, etc., to form a rule check table. Then a series of quality check functions are determined, and the server will use regular expressions, value range judgment, type conversion and assertion comparison, etc. to perform legality verification on each field of each record one by one, and mark the verification results in the newly added label column. Next, the server will filter out all the data marked as illegal to form a set of data to be cleaned, while the legal data directly enters the subsequent link. The server will then develop appropriate processing strategies according to the illegal types of the data to be cleaned, and use truncation, mapping, parsing, correction and other means to convert illegal values to legal values. Finally, the server will splice the cleaned data set with the legal data set filtered out previously to obtain a set of electricity consumption data that is regular and acceptable to the downstream link. For example, for abnormal voltage values exceeding 0-500V, if the order of magnitude difference is not large, it can be truncated to the nearest boundary value; if the difference is large, it may be a decimal point position offset, which can be divided by 10 or 100 for correction. For example, for electricity data with high precision, the first three decimal places can be truncated; for electricity data containing unit characters, the numerical part can be extracted and the unit can be unified.
[0088] Optionally, for structured and semi-structured data sources, the server can directly perform data quality management in the ETL tool through a visual method. Common open source ETL tools such as Kettle, DataX, etc., and commercial ETL tools such as Informatica, DataStage, etc., all have built-in graphical data quality checking and processing components. These tools use a drag-and-drop method to connect each node into a DAG job pipeline according to the data flow. For structured data sources such as relational tables, a data quality analysis node can be connected after the source table input node to configure the quality rules of each field through the interface, and to mark the records that violate the rules during job execution. Then a legal record filtering node is connected to import the illegal records into the repair node for value replacement, data standardization, etc. Finally, a merging node is used to converge the two data streams and output to the target table node. For JSON, XML and other semi-structured data, replace the data quality analysis node with a data parsing and verification node to complete illegal value checking and repair during the parsing process, and the subsequent process is similar to structured data. ETL jobs integrated with data quality capabilities can achieve end-to-end data cleaning with low code.
[0089] S205, divide the electricity consumption data based on a preset data slicing rule to obtain a plurality of data slices, so that each data slice contains a preset number of electricity consumption data and the data volume is less than a preset threshold.
[0090] Referring to step S102, the server divides the power consumption data to obtain multiple data shards.
[0091] S206, respectively store the multiple data shards to multiple data nodes in the distributed file system, and record the correspondence between each data shard and the data node.
[0092] Referring to step S103, the server records the correspondence between each data shard and the data node.
[0093] In some embodiments, the server will: correspond to each data shard, obtain a list of currently available data nodes in the distributed file system; based on a consistent hashing algorithm, select a preset number of candidate data nodes from the list of available data nodes; obtain historical load data, network delay data and geographic location distance data in the candidate data nodes; according to a preset node selection strategy, combine the historical load data, network delay data and geographic location distance data to select a target data node from the candidate data nodes; store the data shard to the target data node, and store the correspondence between the data shard and the target data node in the preset metadata management system.
[0094] Specifically, for each data shard to be stored, the server obtains a list of currently available data nodes in the file system. Then, a preset number of candidate nodes are selected from the node list using a consistent hashing algorithm. Consistent hashing can ensure uniform distribution of shards among nodes, and only affect limited shard migration when nodes are added or deleted. In order to further optimize the best target storage node from the candidate nodes, the historical load of the node, the network delay with the client and the geographic distance between the node and the client, etc. need to be considered. Among them, the load data reflects the busy degree of the node, and the delay data and distance data help to optimize data transmission performance and improve the locality of access. According to the preset node selection strategy, the index data is weighted and calculated to obtain the comprehensive score of each candidate node. The node with the highest score will be selected as the target storage node of the data shard. Common strategies such as load balancing first, distance closest first, etc. can also set multiple priority rules. Finally, the data shard is uploaded and stored to the selected target node, and the correspondence between the shard and the target node is recorded as metadata in a dedicated metadata management system, providing a basis for subsequent data retrieval and node fault recovery.
[0095] Node ID CPU load Memory load IO load Latency Distance N1 85% 65% 56% 45 ms 500 km N3 38% 73% 27% 89 ms 800 km N5 61% 82% 43% 63 ms 300 km
[0096] According to the preset strategy:
[0097] Priority 1: load balancing first, that is, the one with the lowest average load is given priority;
[0098] Priority 2: Close to distance priority, i.e. the closest one is preferred.
[0099] For N1, the average load is (85%+65%+56%) / 3=68.7%,
[0100] For N3, the average load is (38%+73%+27%) / 3=46%,
[0101] For N5, the average load is (61%+82%+43%) / 3=62%;
[0102] Therefore, according to the load balancing, N3>N5>N1;
[0103] According to the distance proximity, N5(300km)>N1(500km)>N3(800km);
[0104] Since the load balancing priority is higher, N3 is finally selected as the target storage node of shard 1. Shard 1 is uploaded to N3, and the mapping relationship (shard 1 to N3) is recorded in the metadata system. The other 9 shards also perform node selection and storage according to the process, so as to realize the distributed storage of the entire power consumption data file in the cluster. As can be seen, the data shard storage optimization method takes into account network delay and geographical distance and other factors while ensuring load balancing, and can efficiently and reliably distribute massive data to multiple storage nodes.
[0105] S207, in the management node of the distributed file system, generating metadata based on the corresponding relationship; the metadata is used to represent the data node where each data shard is located.
[0106] Referring to step S104, the server generates metadata.
[0107] S208, when receiving the power consumption data query request of the target time period sent by the query client, obtaining a plurality of target data shards corresponding to the target time period.
[0108] Referring to step S105, the server obtains a plurality of target data shards.
[0109] S209, according to the metadata, sending a data reading request to the data node where the target data shard is located, and obtaining the target power consumption data of the target data shard.
[0110] Referring to step S106, the server determines the target power consumption data.
[0111] S210, recording the creation time, update time and access time of each data shard.
[0112] The creation time represents the time point at which the data shard is initially generated and persistently stored in the server. The update time refers to the time point at which the content in the data shard is changed and re-persistently stored. The access time is used to represent the time point at which the data shard is last read or written by the application.
[0113] Specifically, in the distributed big data system of the server, the original data is cut, copied into multiple shards, and stored on different nodes. The application accesses the shard data through a shard routing strategy. However, the access frequency of each shard is often uneven, and there is a hot and cold data differentiation phenomenon. To reduce storage costs, it is necessary to identify cold data shards and migrate them to a low-cost object storage layer. The time attributes of data shards are an important basis for identifying cold data. The server collects key time points such as creation, update, and access of shards: when a new shard is first generated and persistently stored in a distributed file system such as HDFS, the server records the current timestamp as the creation time; when the shard is rewritten and re-persistently stored due to data update, merging, compression, etc., the server refreshes the current timestamp as the update time; when the shard is loaded into memory by a computing framework such as MapReduce or Spark for calculation, the server refreshes the current timestamp as the access time; when the shard is persistently stored or the application ends, the server persistently stores the corresponding timestamp in the shard metadata management system such as HDFSNameNode, HBase HMaster, etc.
[0114] S211、In the access frequency of the data shard is lower than the preset threshold, and the time from the last access exceeds the preset time length, the data shard is marked as cold data.
[0115] The access frequency refers to the number of times a data shard is read or written by an application per unit time. The preset threshold represents a human-set upper limit on access frequency. If the actual access frequency of a shard is lower than this threshold, it is considered that the shard is not active enough and has a tendency to become cold data. The preset time length is used to represent a period of idleness set by development and operation personnel. If the idleness time of a shard exceeds this period, it is considered that the shard has been idle for a long time and has become cold data. Cold data refers to a data set that is rarely accessed within a certain period of time, and hot data corresponds to frequently accessed data. For example, if a shard has an average daily access frequency of less than 2 times in the past month, and the last access occurred 20 days ago, and the threshold is set to 3 accesses per day and the time length is set to 14 days, then the access frequency and idleness time of the shard both meet the cold data criteria, and the shard is marked as cold data.
[0116] Specifically, common hot data includes real-time business data and interactive query data, which are usually frequently read, calculated and updated. Cold data is usually historical data such as log archiving and business snapshots, which are rarely accessed after being generated for a period of time, but cannot be completely deleted and need to be separated from hot data. The administrator of the server can manually set two cold data determination indicators, access frequency threshold and idle duration, according to the business scenario to form a set of cold data identification rules. Then, the server will periodically scan and load the time attribute data of each shard from the shard metadata management system, including creation time, update time and access time. Next, the server will calculate the access frequency of each shard in the scanning window period, and calculate the difference between the current time and the latest access time to obtain the idle duration. Then, the frequency and duration of each shard are compared with the preset threshold. For shards with access frequency lower than the threshold and idle duration longer than the limit, they are determined as cold data and marked with a "cold" label. Finally, the shard ID list with the "cold" label is persisted to the cold data marking table in the metadata management system, which is used to guide subsequent storage layering, data life cycle management and other operations.
[0117] It can be understood that, in addition to access frequency and idle duration, other time characteristics can also be used as discriminators for cold data identification, such as setting the earliest and latest update time range, marking shards that have only been modified within the specified time window as cold data; setting the minimum and maximum life cycle, marking shards with too short or too long survival time as cold data; setting the creation time limit, marking shards created earlier than the specified date as cold data, etc. Time attributes reflect the freshness of data, but are not always positively correlated with access heat. Sometimes, historical data is frequently queried, such as annual purchase report of e-commerce. Therefore, in addition to the time dimension, cold data identification also needs to consider multiple dimensions such as data content, user portrait and business rules for comprehensive judgment.
[0118] S212, migrating the data shard marked as cold data from the distributed file system to the object storage system to save storage costs.
[0119] Among them, the distributed file system is a file system that allows files to be stored on multiple servers and allows multiple users to share access. Common ones are HDFS, Ceph, etc. The object storage system is a server that stores and manages data in units of objects. Each data object has a unique identifier and carries rich metadata attributes. The storage interface is usually REST API. Common ones are Amazon S3, Aliyun OSS, etc. Migration refers to the process of copying or moving data from one storage system to another. Storage cost is used to represent the cost of occupying storage resources in the storage system. It is usually related to data volume, storage type, access frequency, and redundancy strategy. For example, migrating a 100GB cold data file from HDFS to S3 can reduce the local disk capacity of the HDFS cluster, while S3 uses online disks at a much lower price than HDFS, thereby saving storage costs.
[0120] Specifically, this step starts after the cold and hot data identification is completed and the cold data shard list is generated. The purpose is to unload the cold data that is no longer frequently accessed from the high-speed storage layer to the low-speed storage layer, reducing the waste of high-speed storage capacity. The distributed file system provides high-throughput data access capability through network-connected local disks, but its expansion capability is limited, and the disk price is expensive. The object storage system uses a global namespace and flat storage, which can easily break through the capacity bottleneck through horizontal expansion, and the medium cost is relatively low. Hierarchical storage of cold and hot data can provide massive data access while controlling the overall TCO of storage. This step uses an incremental migration strategy to periodically migrate newly identified cold data shards from the distributed file system to the object storage: first, load the cold data shard ID list to be migrated from the metadata management system; then, on the NameNode or similar master node of the distributed file system, use hdfs dfs -get or similar data pulling commands to read the cold data shard corresponding to the ID from the DataNode or similar storage node of the distributed file system, and temporarily store it in the local buffer of the master node; then, on the master node, use aws s3 cp or similar data pushing commands to upload the cold data shard in the buffer to the specified bucket or prefix of the object storage system; again on the master node, use hdfs dfs -rm or similar data deletion commands to delete the migrated cold data shard from the distributed file system, freeing up storage space; finally, update the data location index table in the metadata management system to change the location of the migrated cold data shard from the path of the distributed file system to the URI (Uniform Resource Identifier, a string pointing to a resource, usually used in URL to specify the specific location of a resource file on the Web) of the object storage system. This step transfers the cold data shard from the distributed file system to the object storage system, realizing the cold and hot layering of data.
[0121] S213, archiving or deleting the cold data whose storage time exceeds the preset time limit after the preset time limit.
[0122] The preset time limit refers to a cold data retention time limit set by a person, which is divided into different levels such as short-term, medium-term, long-term, and the like according to the decay rate of data value. Archiving refers to the process of transferring data that is rarely accessed to a medium or area with longer access latency but lower storage cost, such as a tape library, cold storage, and the like. Deleting refers to the process of completely removing data that is no longer needed from the storage system.
[0123] For example, a certain e-commerce website sets the following storage time limit for transaction snapshot data: T+1 for 3 months, T+3 for 6 months, and T+6 for 1 year. For cold snapshots that have been migrated to OSS for 3 months, they can be archived to the deep archive storage class of OSS, which has a storage unit price of only 1 / 5 of the standard storage; for 6 months, they can be archived to the archive storage, which has a retrieval latency of minutes but a further reduced unit price of 1 / 10 of the standard storage; and for 1 year, they are directly deleted from the OSS.
[0124] Specifically, although cold data is no longer frequently accessed, its business value will gradually decay over time, and continued storage will result in a large cumulative storage cost. On the other hand, cold data is a historical snapshot of the business, and centralized deletion may cause data loss and compliance risks. Therefore, the server balances access demand and storage cost and performs phased storage life cycle management on cold data. First, the server administrator sets a set of hierarchical storage time limits based on the business attributes of the data, such as 3-6-9-12 representing short-term 3 months, medium-term 6 months, long-term 9 months, and permanent 12 months. Then, the metadata management table of the object storage system is periodically scanned to identify the storage time stamp of each cold data object, and the object storage time is obtained by subtracting the storage time stamp from the current time. Then, the stored time of each object is compared with the preset hierarchical storage time limit, and the object is labeled with a time label according to the most matched time limit level, such as a cold object that has been stored for 83 days, which is labeled as “medium-term”. The corresponding life cycle management actions are performed on objects with different time labels, such as deleting short-term objects, transferring medium-term objects to archive storage, migrating long-term objects to deep archive, and keeping permanent objects in place. Finally, the object name, original path, target path, and processing method of the objects processed in this round are recorded in the life cycle log of the metadata management system, which facilitates the administrator to check.
[0125] It can be understood that the object storage life cycle management rule provides a declarative cold data storage SLA configuration scheme, but for the scene of cross-storage system migration or heterogeneous data source aggregation, a customized cold data life cycle management component also needs to be developed. For example, a Hadoop ecological component can be developed to periodically scan the cold data of storage systems such as HDFS and HBase as MapReduce jobs, selectively push them to object storage, and set the corresponding storage life cycle label. For another example, an external service program can be developed to receive cold data consumption of different data pipelines such as Flume and Logstash, dump them to object storage, and determine their storage level and validity period according to the metadata label.
[0126] In some embodiments, the server will: store the data feature information of the deleted cold data to a preset data feature management server; when receiving a user query request, and the requested data is the deleted cold data, obtain the data features similar to the requested data from the data feature management server, and return the data features to the user as the substitute information of the requested data; when receiving a user query request, and the requested data only partially belongs to the deleted cold data, obtain the summary statistical information of the deleted part from the data feature management server, and return the original data of the non-deleted part and the summary statistical information of the deleted part to the user; when receiving a user query request, and the requested time range span is greater than a preset threshold, divide the data in the request range according to a preset time interval, return the original data for the non-deleted time period, and obtain the summary statistical information of the deleted time period from the data feature management server and return it.
[0127] Specifically, the server extracts the data feature information of the deleted cold data and stores it in a dedicated feature management server when deleting the cold data. The data features usually include data summary statistics (such as record number, mean, variance, etc.), data distribution, key value range, and other meta information. When a user queries the deleted cold data, the server first determines whether the requested data is all or part deleted. If the requested data is all deleted, the server obtains similar feature information from the feature management server as a substitute result and returns it to the user, so that the user can still get a general data profile even if the original detailed data cannot be provided. If only part of the requested data is deleted, the request is divided into deleted and non-deleted parts. For the non-deleted part, the server obtains and returns the data from the original database. For the deleted part, the server obtains its summary statistics information from the feature management server. Finally, the server combines the results of the two parts and returns them to the user. In this way, the accuracy of non-deleted data is guaranteed, and statistical level information is provided for deleted data. If the user's query time span is very large and exceeds the preset range threshold, in order to avoid the impact of large-scale scanning on server performance, the method divides the requested time range into segments according to the preset time interval (such as each month), and judges whether the data has been deleted for each segment. For the non-deleted time period, the server returns the original detailed data; for the deleted time period, the server returns its statistical feature information. Finally, the server splices the results of each time period and returns them to the user.
[0128] Suppose a user submits a query request to obtain the electricity consumption data of the user with ID 001 from January 2023 to June 2024. The server sets one month of data as a shard, and cold data over 12 months will be deleted and archived.
[0129] Shard ID Record count Total power consumption Max power consumption Min power consumption 202301 2976 281 kWh 2.24 kWh 0.32 kWh 202302 2688 250 kWh 2.31 kWh 0.35 kWh ... ... ... ... ... 202306 2880 359 kWh 2.96 kWh 0.41 kWh
[0130] The data from July 2023 to June 2024 has not yet reached 12 months, and the original electricity consumption record details are still retained.
[0131] When the server receives the user's query request, it finds that the requested time span is 18 months, which exceeds the preset threshold of 12 months, so it splits the request according to the monthly shard.
[0132] For the request from January to June 2023, the corresponding data shard is identified as deleted, so the server obtains the summary statistics information of the 6 shards from the feature management server and returns it:
[0133] {"202301": {"recordCount": 2976, "totalUsage": 281, "maxUsage": 2.24, "minUsage": 0.32}, "202302": {"recordCount": 2688, "totalUsage": 250, "maxUsage": 2.31, "minUsage": 0.35}, "202303": {"recordCount": 2700, "totalUsage": 250, "maxUsage": 2.31, "minUsage": 0.35}, "202304": {"recordCount": 2700, "totalUsage": 250, "maxUsage": 2.31, "minUsage": 0.35}, "202305": {"recordCount": 2700, "totalUsage": 250, "maxUsage": 2.31, "minUsage": 0.35}, "202306": {"recordCount": 2880, "totalUsage": 359, "maxUsage": 2.96, "minUsage": 0.41}}
[0134] For requests from July 2023 to June 2024, since the original data still exists, the server directly retrieves the power consumption detail records of these 12 shards from the database and returns them. Finally, the server combines the feature statistics of the deleted cold data with the original details of the non-deleted data and integrates them into a complete JSON result to return to the user, realizing a power consumption query spanning 18 months. As can be seen, by deleting cold data and extracting features, we can save storage costs while flexibly returning statistical data or raw details according to different query requests, better balancing server performance and data availability.
[0135] In some embodiments, the server will also dynamically adjust the storage strategy of the relevant power consumption data in combination with the user's historical data retrieval and query records.
[0136] Among them, the user's historical data retrieval and query records refer to the log information of the user's past access and query of power consumption data saved by the server, including the time range of the query, the frequency of the query, and the data shards involved in the query, etc. The storage strategy represents the storage management scheme adopted for different data shards, including storage medium selection, cold and hot layering rules, and life cycle management, etc. Dynamic adjustment means automatically optimizing the storage strategy matched by the data according to its actual access pattern, so that the configuration of storage resources is adapted to the characteristics of data access. For example, data shards that are frequently queried by users are migrated back from the cold data area to the hot data area, and their life cycle is adjusted to extend the retention period; for data that is rarely queried by users, its storage level is reduced and the retention period is shortened to save storage space.
[0137] Specifically, different users have different degrees of dependence on historical data, showing personalized data access patterns, and it is difficult to fully meet the use of static storage strategies. The application automatically adjusts the most matching storage parameters by continuously recording and analyzing the access of each user to the data shards under his name. First, the server designs a user access log table in the metadata management system to record the time of each user-initiated data retrieval and query request, the data shards involved, and the time span of the query, etc. Then, the server will aggregate and analyze the log table by user and data shard to obtain statistical indicators such as query frequency and average query span of each shard. According to the preset threshold rule, the matching degree of the actual access mode of each shard and the current storage strategy is judged, and for the shards with large deviations, storage strategy adjustment suggestions are generated, such as adjusting the frequently accessed cold data shard to hot data, adjusting the long-term non-accessed hot data to cold data, etc. Finally, the server submits the adjustment suggestions to the distributed storage management system to trigger the corresponding data migration and parameter update tasks. In this way, the storage strategy of the data shard can be dynamically optimized according to the actual access needs of the user, improving the utilization of storage resources while taking into account data availability.
[0138] In some embodiments, the access log-based personalized storage strategy adjustment can be implemented in various ways: optionally, using the YARN scheduler and HDFS storage management functions in the Hadoop ecosystem. First, the data input path information of each task is extracted from the container monitoring log of YARN, the involved HDFS directory or file is associated with the user who submitted the task, and a user data access log is formed; then the log is counted using Spark, Hive and other analysis tools, and the access frequency and access user distribution of each directory or file are calculated; when the indicators of a certain shard file exceed the threshold of the HDFS hot and cold data identification rule, a storage strategy change request is initiated to the HDFS NameNode, and the FileSystem API supports dynamic modification of the storage strategy of the file, such as Archive, Lazy_Persist and All_SSD; HDFS will automatically schedule data migration and replica layout update in the storage cluster according to the new strategy until the data layout matches the strategy. Optionally, using the Lifecycle rules and S3 API interfaces provided by object storage services such as CephRGW or OpenstackSwift to implement. First, parse the object path, access time, access user and other information of each GET / PUT object request in the access log of the object storage server to generate a fine-grained data access log stream; then use Storm, Flink and other stream computing frameworks to do real-time statistics on the access log, and write the aggregated results to the metadata database of the object storage such as MongoDB, ElasticSearch; then, deploy a custom lifecycle rule engine to match the access statistics data with the user's preset tiered storage rules, automatically generate LifecycleRules and push them to the object storage service through the S3 API interface, and the rules define the time nodes for converting objects with different access frequencies between the Standard, Warm and Cold storage categories; the object storage service will automatically execute the lifecycle rules, periodically scan the metadata of the objects in the bucket, and trigger storage type changes or deletion for objects that meet the rule conditions.
[0139] S214、According to the target electricity data, the electricity data analysis result corresponding to the target time period is generated, and the electricity data analysis result is returned to the query client.
[0140] Referring to step S107, the server returns the electricity data analysis result to the query client.
[0141] In some embodiments, the server will: calculate the target electricity consumption data according to the user dimension to obtain the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption, and electricity consumption curve for each user within the target time period; calculate the target electricity consumption data according to the time dimension to obtain the total daily electricity consumption and peak electricity consumption periods within the target time period; calculate the target electricity consumption data according to the region dimension to obtain the electricity consumption distribution of users in different regions; generate a multi-dimensional electricity consumption data analysis report based on the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption, electricity consumption curve, total electricity consumption, peak electricity consumption periods, and electricity consumption distribution; the electricity consumption data analysis report includes an electricity consumption statistics table, an electricity load curve, and an electricity consumption regional distribution map; generate a report webpage link based on the electricity consumption data analysis report and send the report webpage link to the query client; generate electricity consumption data analysis results based on the electricity consumption data analysis report and return the electricity consumption data analysis results to the query client.
[0142] Specifically, the server comprehensively summarizes electricity consumption data for the target time period according to different dimensions. At the user level, it calculates the total electricity consumption, average electricity consumption, maximum electricity consumption, and minimum electricity consumption for each user, and plots a 24-hour electricity consumption curve. At the time level, it calculates the total daily electricity consumption and peak consumption periods. At the regional level, it statistically analyzes the total electricity consumption and its distribution across different regions. These indicators characterize the statistical properties of electricity consumption data from different perspectives. Next, based on these statistical indicators, the server designs various intuitive visualization charts and generates multi-dimensional electricity consumption analysis reports. Common chart types include electricity consumption statistical summary tables, electricity load curves reflecting 24-hour power changes, and pie charts showing regional differences in electricity consumption distribution. Rich data visualization methods can transform obscure numbers into easily understandable and analyzable graphs, facilitating the discovery of correlations and patterns in the data. Considering the potentially large size of the data reports, which is inconvenient for direct transmission, the server deploys the complete analysis reports as web pages, providing a unique access link, and sending this link address to the query client. Users can access the online dynamic report page at any time through the link. At the same time, the server will also extract the key statistical results of the report and return them directly to the client in JSON and other formats, making it convenient for other programs to carry out secondary development and utilization.
[0143] Suppose a provincial branch of the State Grid receives a data analysis request from its superior department, requiring the calculation and analysis of electricity consumption data across the province in the first quarter of 2024 and the submission of a visualization report. The server first extracts the raw electricity consumption data from January 1st to March 31st, 2024, and aggregates and calculates it according to three dimensions: user, time, and region. At the user level, it finds that the province has 5.32 million residential users and 87,000 industrial and commercial users. It calculates the total electricity consumption, average daily electricity consumption, and maximum electricity consumption for each user in the first quarter, and plots electricity consumption curves for typical users.
[0144] In terms of time, statistics show that the total electricity consumption in the first quarter was 13.2 billion kWh, with January accounting for 27%, February for 30%, and March for 43%. The analysis also indicates that the average daily peak electricity consumption occurred between 11:00 and 21:00.
[0145] Based on regional latitude, a pie chart showing the electricity consumption share of each city in the province was drawn. The provincial capital city had the highest share, reaching 35% of the province's total electricity consumption, followed by the industrial city of City A, accounting for 25%, while the remaining cities all had a share of less than 10%.
[0146] Based on the above analysis, the server generates an electricity consumption data analysis report, which includes three key charts: an electricity consumption statistics table, which lists key indicators such as the number of users, total electricity consumption, average electricity consumption, and the largest electricity user; a daily 24-hour electricity load curve for the first quarter, which reflects the fluctuation pattern of electricity consumption in the province over time; and a pie chart showing the proportion of electricity consumption in various cities, which intuitively shows the current uneven distribution of electricity consumption in different regions.
[0147] In this embodiment, by adopting a data cold and hot separation storage and automated storage management strategy, different storage media are matched for data with different access frequencies. Cold data is migrated to low-cost object storage, while hot data is retained in the distributed file system. Cold data that has not been accessed for a long time is deleted. Therefore, while meeting the needs of online data query, storage costs are effectively reduced and storage resource utilization is improved. At the same time, automated storage management of historical data is realized, which effectively solves the problems of storage resource waste and difficulty in controlling data retention period in the scenario of massive electricity consumption data. In this way, the elastic and flexible expansion and contraction and intelligent operation and maintenance of the storage system are realized.
[0148] The server in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of a server in an embodiment of this application.
[0149] It should be noted that, Figure 3 The server structure shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0150] like Figure 3 As shown, the server includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 302 or a program loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 303. The CPU 301, ROM 302, and RAM 303 are interconnected via bus 304. An Input / Output (I / O) interface 305 is also connected to bus 304.
[0151] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including hard disks, etc.; and communication section 309 including network interface cards such as LAN (Local Area Network) cards, modems, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0152] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.
[0153] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0155] Specifically, the server in this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the data processing method for electricity consumption data provided in the above embodiment.
[0156] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above embodiments; or it may exist independently and not assembled into the server. The storage medium carries one or more computer programs that, when executed by a processor of the server, cause the server to implement the data processing method for electricity consumption data provided in the above embodiments.
[0157] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0158] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0159] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for processing electricity consumption data, characterized in that, Applied to a server, the method includes: Obtain electricity consumption data from multiple users; Detecting missing values in the electricity consumption data and determining the corresponding electricity consumption data value based on other electricity consumption data of the same user within the time period of the missing value; specifically, this includes: acquiring the original electricity consumption data of the target user within a target time period, and sorting the original electricity consumption data according to the collection time to obtain an electricity consumption data sequence set; traversing the electricity consumption data sequence set, and after determining that the interval between the collection times of two adjacent original electricity consumption data is greater than a preset time interval, taking several electricity consumption data corresponding to the interval as missing values; acquiring historical electricity consumption data of the same user within the time period corresponding to the interval, and calculating the average value of the historical electricity consumption data as the initial estimate corresponding to the missing value; acquiring the electricity consumption data values of several non-missing sampling points adjacent before and after the interval, and correcting the initial estimate of the missing value by interpolation fitting to obtain a corrected estimate; filling the corrected estimate into the corresponding missing value position of the interval to determine the electricity consumption data value corresponding to the missing value. Detect duplicate records in the electricity consumption data, retain one completely duplicate record among the duplicate records, and merge records with duplicate fields among the duplicate records; Detect illegal values in the electricity consumption data and replace them with a threshold or perform format conversion; the illegal values include values that exceed a preset range and non-numeric values. The electricity consumption data is divided based on a preset data segmentation rule to obtain multiple data segments, such that each data segment contains a preset number of electricity consumption data and the data volume is less than a preset threshold. The multiple data shards are stored on multiple data nodes in a distributed file system, and the correspondence between each data shard and data node is recorded. In the management node of the distributed file system, metadata is generated based on the correspondence; the metadata is used to characterize the data node where each data shard is located. When a query request for electricity consumption data for a target time period is received from a query client, several target data fragments corresponding to the target time period are obtained; Based on the metadata, a data read request is sent to the data node where the target data segment is located to obtain the target electricity consumption data of the target data segment. Based on the target electricity consumption data, generate electricity consumption data analysis results corresponding to the target time period, and return the electricity consumption data analysis results to the query client; specifically, generating electricity consumption data analysis results corresponding to the target time period based on the target electricity consumption data and returning the electricity consumption data analysis results to the query client includes: calculating the target electricity consumption data according to the user dimension to obtain the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption, and electricity consumption curve for each user within the target time period; and calculating the target electricity consumption data according to the time dimension to obtain the total daily electricity consumption and peak electricity consumption during the target time period. The system calculates the electricity consumption data of the target user based on the regional dimension to obtain the electricity consumption distribution of users in different regions. Based on the stage electricity consumption, average electricity consumption, peak electricity consumption, valley electricity consumption, electricity consumption curve, total electricity consumption, peak electricity consumption periods, and electricity consumption distribution, a multi-dimensional electricity consumption data analysis report is generated. The electricity consumption data analysis report includes an electricity consumption statistics table, an electricity load curve, and an electricity consumption regional distribution map. A report webpage link is generated based on the electricity consumption data analysis report, and the report webpage link is sent to the query client. Electricity consumption data analysis results are generated based on the electricity consumption data analysis report, and the electricity consumption data analysis results are returned to the query client.
2. The method according to claim 1, characterized in that, The step of storing the multiple data shards on multiple data nodes in a distributed file system and recording the correspondence between each data shard and data node specifically includes: For each data shard, obtain the list of currently available data nodes in the distributed file system; Based on the consistent hashing algorithm, a preset number of candidate data nodes are selected from the list of available data nodes; Obtain historical load data, network latency data, and geographic distance data from the candidate data nodes; Based on a preset node selection strategy, and in combination with the historical load data, the network latency data, and the geographical location distance data, a target data node is selected from the candidate data nodes. The data shards are stored in the target data nodes, and the correspondence between the data shards and the target data nodes is stored in a preset metadata management system.
3. The method according to claim 1, characterized in that, After the step of sending a data read request to the data node where the target data shard is located based on the metadata to obtain the target electricity consumption data of the target data shard, the method further includes: Record the creation time, update time, and access time of each data shard; If the access frequency of a data shard is lower than a preset threshold and the time since the last access exceeds a preset duration, the data shard will be marked as cold data. Migrate data shards marked as cold data from the distributed file system to the object storage system to save storage costs; Cold data that has been stored for longer than the preset period will be archived or deleted.
4. The method according to claim 3, characterized in that, After the step of archiving or deleting cold data whose storage time has exceeded the preset period after the preset period, the method further includes: The data feature information of the deleted cold data is stored in a preset data feature management server; When a user query request is received, and the requested data is deleted cold data, data features similar to the requested data are obtained from the data feature management server and returned to the user as alternative information for the requested data. When a user query request is received, and the requested data only partially belongs to deleted cold data, the summary statistics of the deleted part are obtained from the data feature management server, and the original data of the undeleted part and the summary statistics of the deleted part are returned to the user. When a user query request is received and the time range of the request exceeds a preset threshold, the data within the request range is divided according to a preset time interval. For the time period that has not been deleted, the original data is returned, and for the time period that has been deleted, the summary statistics are obtained from the data feature management server and returned.
5. A server, characterized in that, The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the server to perform the method as described in any one of claims 1-4.
6. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the server, it causes the server to perform the method as described in any one of claims 1-4.
7. A computer program product, characterized in that, When the computer program product is run on the server, the server performs the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Data query method and device, server and medium
CN117785952A
Method and system for managing distributed content and related metadata
US20020133491A1