Power consumer activity data processing method and device based on storage integrated system

By combining Redis and RocksDB in an integrated storage system, Redis's memory limitation and persistence performance in data management is solved, and high-performance and high-reliability power user activity data processing is achieved, improving the overall performance and stability of the system.

CN120215834APending Publication Date: 2025-06-27GUANGDONG ELECTRIC POWER COMM CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510353414.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In traditional technology, Redis has problems such as memory limitations, insufficient persistence performance, long data recovery time, single point of failure and complex cluster expansion in data management, resulting in poor data management of database systems.

Method used

The power user activity data processing method based on an integrated storage system is adopted, and the hotspot service separation and multi-level cache processing of data are realized through the combination of the Redis engine and the RocksDB engine. The write-pre-logging mechanism is used for failure recovery and data persistence, reducing the dependence on memory and improving the reliability and scalability of the system.

Benefits of technology

By combining Redis's memory speed and RocksDB persistence capabilities, high-speed data reading and writing and reliable data persistence are achieved, which effectively improves the performance and stability of the system and reduces hardware costs and operation and maintenance complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215834A_ABST
    Figure CN120215834A_ABST
Patent Text Reader

Abstract

The invention relates to a power consumer activity data processing method and device based on a storage integrated system. The method comprises the following steps: carrying out hotspot service separation on an obtained real-time data stream of a power consumer activity through a storage integrated system to obtain service hot data and service cold data; storing the business hot data to a first storage medium by utilizing a Redis engine, and storing the business cold data to a second storage medium by utilizing a RocksDB engine; determining a storage level of business hot data and a storage level of business cold data based on the data hierarchical storage information of the storage integrated system; the data hierarchical storage information is used for indicating different storage media corresponding to each hierarchy; and when a data access event of the power client is detected, performing storage medium conversion processing on the business hot data and the business cold data according to the cold and hot data exchange condition, and feeding back a data reading result for the data access event. By adopting the method, high-speed data reading and writing and data persistence can be realized, and the system performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a method, device, computer device, computer-readable storage medium, and computer program product for processing power user activity data based on a storage integrated system. Background Art

[0002] With the increasing number of power grid users year by year, the collection, storage, and analysis of relevant power user activity data have become increasingly important, posing many challenges to the processing of database systems facing high-concurrency scenarios and big data scenarios.

[0003] In traditional technologies, Redis, as a high-performance in-memory database, has been widely used in scenarios such as caching and session storage due to its low latency and high throughput characteristics. However, Redis has the following deficiencies in data management: Redis mainly uses memory to store data. When facing large-scale data sets with massive information, a large amount of RAM resources are required, which will increase the hardware cost and may also lead to memory overflow problems; Redis has weak data persistence capabilities, especially in the event of system crashes or power outages, which may result in data loss.

[0004] Therefore, there are problems with poor data management effects in database systems in related technologies. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for processing power user activity data based on a storage integrated system that can improve the data management effect of the system.

[0006] In a first aspect, the present application provides a method for processing power user activity data based on a storage integrated system, the method including:

[0007] Through the storage integrated system, perform hot business separation on the acquired real-time data stream of power user activities to obtain business hot data and business cold data; the storage integrated system includes a Redis engine and a RocksDB engine;

[0008] Use the Redis engine to store the business hot data into a first storage medium, and use the RocksDB engine to store the business cold data into a second storage medium; the first storage medium is memory, and the second storage medium is a disk;

[0009] Based on the data hierarchical storage information of the storage integrated system, determine the storage levels of the business hot data and the business cold data; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0010] When a data access event of the power client is detected, perform storage medium conversion processing on the service hot data and the service cold data according to the hot and cold data exchange conditions, and feedback the data reading result for the data access event.

[0011] In one embodiment, the method further includes:

[0012] Execute distributed multi-level caching processing on the real-time data stream through the storage integrated system;

[0013] Among them, the distributed multi-level caching processing includes collaborative processing among data hierarchical storage, data access optimization, distributed storage and access, data synchronization mechanism, and cache policy setting.

[0014] In one embodiment, the method further includes:

[0015] During the process of writing data through the storage integrated system, write the power user activity data to be processed into the data connection pool for storage;

[0016] Determine the target storage system of the power user activity data according to the service identification information of the power user activity data;

[0017] Asynchronously write the stored power user activity data into the target storage system through the data connection pool.

[0018] In one embodiment, the storage integrated system is connected to multiple power service systems. Through the storage integrated system, separating the hot spot services from the obtained real-time data stream of power user activities includes:

[0019] Identify the hot spot services in each of the power service systems; the hot spot services include high-traffic services and high-load services;

[0020] Separate the hot spot services according to the service attribute information of each power user activity data in the real-time data stream; the service attribute information is used to determine whether it belongs to a hot spot service.

[0021] In one embodiment, the method further includes:

[0022] During the process of processing the real-time data stream through the storage integrated system, utilize the write-ahead log mechanism to perform fault recovery processing on the monitored system fault events and execute immediate analysis operations on the real-time data stream.

[0023] In one embodiment, the method further includes:

[0024] Based on the log-structured merge tree in the storage-in-one system, process the power user activity data stream in the high-concurrency data scenario; the log-structured merge tree is used to optimize the write performance of the storage-in-one system.

[0025] In a second aspect, the present application also provides a power user activity data processing device based on a storage-in-one system, and the device includes:

[0026] A data separation module, configured to separate hot services from the obtained real-time data stream of power user activities through the storage-in-one system to obtain service hot data and service cold data; the storage-in-one system includes a Redis engine and a RocksDB engine;

[0027] A data storage module, configured to store the service hot data in a first storage medium by using the Redis engine and store the service cold data in a second storage medium by using the RocksDB engine; the first storage medium is a memory, and the second storage medium is a disk;

[0028] A storage level determination module, configured to determine the storage level of the service hot data and the storage level of the service cold data based on the data hierarchical storage information of the storage-in-one system; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0029] A data exchange module, configured to perform storage medium conversion processing on the service hot data and the service cold data according to the hot and cold data exchange conditions when detecting a data access event of a power client, and feedback a data reading result for the data access event.

[0030] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0031] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0032] In a fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0033] The above-mentioned method, device, computer equipment, computer-readable storage medium, and computer program product for processing power user activity data based on a storage-integrated system separate hot services from the real-time data stream of power user activities obtained through the storage-integrated system to obtain service hot data and service cold data. The storage-integrated system includes a Redis engine and a RocksDB engine. Then, the Redis engine is used to store the service hot data in a first storage medium, and the RocksDB engine is used to store the service cold data in a second storage medium. The first storage medium is memory, and the second storage medium is a disk. Based on the data hierarchical storage information of the storage-integrated system, the storage levels of the service hot data and the service cold data are determined. The data hierarchical storage information is used to indicate different storage media corresponding to each level. Furthermore, when a data access event of the power client is detected, the storage media conversion process is performed on the service hot data and the service cold data according to the hot and cold data exchange conditions, and the data reading result for the data access event is fed back, realizing the optimization of the processing of power user activity data based on the storage-integrated system. By combining the memory speed of Redis and the persistence ability of RocksDB, high-speed data reading and writing and reliable data persistence can be achieved, effectively improving the performance and stability of the system. Brief Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic flowchart of a method for processing power user activity data based on a storage-integrated system in an embodiment;

[0036] Figure 2 It is a schematic diagram of the hot and cold data exchange process in an embodiment;

[0037] Figure 3 It is a schematic diagram of the read / write and compression operation process in an embodiment;

[0038] Figure 4 It is a schematic flowchart of a method for processing power user activity data based on a storage-integrated system in another embodiment;

[0039] Figure 5 It is a structural block diagram of a device for processing power user activity data based on a storage-integrated system in an embodiment;

[0040] Figure 6Internal structure diagram of a computer device in an embodiment. Detailed implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting this application.

[0042] Redis has the following deficiencies in data management:

[0043] 1. Memory limitation: Redis mainly uses memory to store data. When facing large-scale data sets with massive information, a large amount of RAM resources are required, which not only increases the hardware cost but also may lead to memory overflow problems.

[0044] 2. Persistence performance: The persistence mechanisms of Redis (such as the RDB method and the AOF method) may become performance bottlenecks in high-concurrency scenarios. Especially when facing large-scale data sets with massive information, the persistence operations will affect the system's response speed.

[0045] 3. Data recovery time: In case of a failure, the RDB persistence method of Redis takes a long time to recover data. Although the AOF method has a fast recovery speed, the file size is large and the recovery time is still long.

[0046] 4. Single point of failure: Although Redis supports master-slave replication and sentinel mode, in high-concurrency scenarios, the load on the master node is high and it is prone to becoming a single point of failure.

[0047] 5. Cluster expansion: As the business grows, the expansion of the Redis cluster becomes increasingly complex.

[0048] This application provides a method for processing power user activity data based on a storage integrated system. By proposing an optimization scheme of Redis on RocksDB (ROR), this scheme combines the memory speed of Redis and the persistence ability of RocksDB, enabling high-speed data reading and writing and reliable data persistence, effectively improving the performance and stability of the system.

[0049] In an exemplary embodiment, as Figure 1 shown, a method for processing power user activity data based on a storage integrated system is provided. In this embodiment, this method is exemplified by being applied to a terminal. It can be understood that this method can also be applied to a server and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, this method includes the following steps 101 to 104. Among them:

[0050] Step 101: Through the storage integrated system, perform hot service separation on the real-time data stream of power user activities to obtain service hot data and service cold data.

[0051] Among them, the storage integrated system may include a Redis engine and a RocksDB engine.

[0052] In practical applications, based on high-speed solid-state disk technology, by combining the technologies of memory and high-speed solid-state disk, the redis kernel can be enhanced to support data cold and hot exchange, and the solid-state disk can be used to expand the cache capacity, thereby saving a large amount of costs, ensuring data durability and making the performance comparable to that of a full-memory database.

[0053] Exemplarily, by proposing an optimization scheme of Redis on RocksDB (ROR), it can combine the memory speed of Redis and the persistence ability of RocksDB. Through cold and hot data separation technology, ROR can store frequently accessed hot data in memory and infrequently accessed cold data on disk. Storing hot data in memory allows for quick access, improving the system's response speed, while storing cold data on disk reduces the dependence on expensive memory resources and lowers the storage cost. Thus, by reasonably allocating memory and disk resources, the overall resource utilization rate of the system can be improved.

[0054] Step 102: Use the Redis engine to store the service hot data in the first storage medium, and use the RocksDB engine to store the service cold data in the second storage medium.

[0055] As an example, the first storage medium can be memory, and the second storage medium can be a disk.

[0056] In specific implementation, data can be divided into cold data and hot data for cold and hot multi-level storage. For example, hot data can be stored in memory based on the redis engine, and cold data can be stored using a disk-based storage engine, such as the RocksDB engine optimized for SSD disk storage.

[0057] Step 103: Based on the data hierarchical storage information of the storage integrated system, determine the storage levels of the service hot data and the service cold data.

[0058] Among them, the data hierarchical storage information can be used to indicate different storage media corresponding to each level.

[0059] In one example, ROR supports hierarchical data storage. Data can be divided into multiple levels, with each level corresponding to a different storage medium (such as SSD, HDD). By storing frequently accessed data (i.e., business hot data) on high-performance SSDs and infrequently accessed data (i.e., business cold data) on low-cost HDDs, read and write performance can be optimized. Additionally, by reasonably allocating the use of different storage media, storage costs can be reduced while ensuring performance, and flexible storage expansion is also supported. The storage medium configuration can be dynamically adjusted according to business requirements.

[0060] Step 104: When a data access event of the power client is detected, perform storage medium conversion processing on the business hot data and the business cold data according to the hot and cold data exchange conditions, and feedback the data reading result for the data access event.

[0061] Specifically, as Figure 2 shown in the ROR hot and cold data exchange process, when the client accesses cold data (i.e., a data access event of the power client is detected), the data in RocksDB can be swapped into Redis. ROR can swap the command-dependent data into Redis so that subsequent command execution is the same as that of native Redis. After the memory usage exceeds maxmemory (a configuration parameter used to set the maximum amount of memory that a Redis instance can use), the business hot data can be swapped out to RocksDB. The ROR hot and cold exchange algorithm can swap the data originally eliminated by Redis into memory by adopting the native LFU algorithm of Redis.

[0062] Compared with traditional methods, ROR has achieved storage engine optimization. By using RocksDB as the storage engine and combining the high performance of Redis and the persistent storage ability of RocksDB, it can solve the memory limitation problem of Redis. Moreover, through the LSM (Log-Structured Merge-Tree) tree structure, ROR can efficiently store a large amount of data on disk, reducing the dependence on memory and improving the storage capacity of the system at the same time; ROR has achieved optimization of the persistence mechanism and data recovery optimization. By adopting the WAL (Write-Ahead Logging) mechanism, it can ensure data persistence and consistency, reduce the latency of persistence operations, and improve the persistence performance through the LSM tree structure, providing a better persistence experience in high-concurrency scenarios; ROR has achieved high-availability optimization. It supports distributed deployment. Through a multi-node architecture, it can effectively avoid single-point failures and improve the high availability of the system. Moreover, through master-slave replication and sentinel mode, it can achieve automatic data synchronization and failover, ensuring the stable operation of the system; ROR has achieved performance optimization innovation. By combining the LSM tree structure and the WAL mechanism, it has optimized the write and read performance of data. In high-concurrency scenarios, its write and read performance is significantly better than that of pure-memory Redis, especially when dealing with large amounts of data, it performs even more outstandingly; ROR has achieved cost-effective innovation. By using RocksDB as the storage engine, it can reduce the dependence on memory and hardware costs. Moreover, through an efficient persistence and data recovery mechanism, it has improved the overall performance and reliability of the system and reduced the operation and maintenance costs.

[0063] In an optional embodiment, as Figure 3 shown, ROR supports data compression and caching mechanisms. By compressing data, it can reduce storage space occupancy and storage costs. Based on the caching mechanism, it can improve data access speed, achieve fast access to frequently used data, and improve the response speed of the system. Thus, by reasonably using caching and compression technologies, it can improve the resource utilization rate and overall performance of the system; ROR supports dynamic data migration. It can automatically migrate cold data to disk and hot data to memory according to the access frequency of data. The system can automatically adjust the storage location of data without manual intervention, improving the intelligence level of the system, reducing the workload of manually adjusting the data storage location, effectively reducing the operation and maintenance costs. By dynamically adjusting the data storage location, it can better handle sudden access peaks and improve the stability and reliability of the system.

[0064] In the above method for processing power user activity data based on a storage integrated system, through the storage integrated system, the real-time data stream of power user activities is separated into hot business data and cold business data. Then, the Redis engine is used to store the hot business data in the first storage medium, and the RocksDB engine is used to store the cold business data in the second storage medium. Based on the data hierarchical storage information of the storage integrated system, the storage levels of the hot business data and the cold business data are determined. Furthermore, when a data access event of the power client is detected, the storage medium conversion process is performed on the hot business data and the cold business data according to the hot and cold data exchange conditions, and the data reading result for the data access event is fed back, realizing the optimization of the power user activity data processing based on the storage integrated system. By combining the memory speed of Redis and the persistence ability of RocksDB, high-speed data reading and writing and reliable data persistence can be achieved, effectively improving the performance and stability of the system.

[0065] In an exemplary embodiment, the following steps may further be included:

[0066] Execute distributed multi-level caching processing on the real-time data stream through the storage integrated system; wherein, the distributed multi-level caching processing includes collaborative processing among data hierarchical storage, data access optimization, distributed storage and access, data synchronization mechanism, and cache policy setting.

[0067] In practical applications, through distributed caching, data can be cached on multiple nodes to improve data access speed. For example, the cached data is scattered to multiple nodes, each node stores a part of the cached data, and the consistent hashing algorithm is used to achieve distributed storage and access of the data. When a node needs to access data, it can check whether the required data exists in the cache of this node. If not, it can send requests to other nodes to obtain the data. Thereby, the access speed and reliability of the cache can be effectively improved, and at the same time, the load pressure on a single node can be reduced, improving the performance and scalability of the entire system.

[0068] Specifically, through data hierarchical storage processing, data can be stored at different levels, and each level may include memory cache (such as JVM cache), local cache (such as cache on the application server), remote cache (such as Redis cluster), etc. Each level of cache has its specific capacity, access speed, and usage, so as to achieve optimization for different access requirements and data characteristics.

[0069] For access optimization processing, when the client initiates a data access request, the system can search according to a preset caching policy. For example, the system can check whether the required data exists in the local or memory cache. If it exists, it can be directly returned, thus avoiding the overhead of remote access or database queries. If the local cache misses, the system can sequentially check other levels of caches until the required data is found or the data source is finally accessed.

[0070] For distributed storage and access processing, in a distributed environment, data is scattered and stored on multiple nodes. Each node is responsible for storing part of the cached data and uses a consistent hashing algorithm or other distributed algorithms to achieve fast data location. When the client needs to access data, the system can request routing to the node storing the data according to the routing policy to achieve distributed access.

[0071] For the data synchronization mechanism, since data may exist in multiple levels of caches, it is necessary to maintain data consistency. When the data in the data source (such as a database) changes, the system needs to ensure that the data in each level of cache is also updated in a timely manner. For example, it can be achieved through mechanisms such as cache invalidation, cache update, or message notification. Moreover, to ensure data consistency, complex issues such as concurrent control and transaction management also need to be considered.

[0072] For cache policy setting processing, a distributed multi-level cache system can adopt a series of caching policies and algorithms to optimize performance. For example, the cache eviction policy can be set according to the access frequency and importance of the data, such as LRU (Least Recently Used), LFU (Least Frequently Used), etc.; use the cache preheating mechanism to load popular data during system startup or off-peak periods; enable a degradation policy to ensure system availability when the cache system fails, etc.

[0073] In one example, based on the distributed multi-level cache, through the collaboration of multiple aspects such as data hierarchical storage, access optimization, distributed storage and access, data synchronization and consistency, and cache policies and algorithms, the system performance can be improved and the user experience can be optimized; and the cache policy has significant advantages in dealing with high-concurrency access, reducing latency, and increasing data access speed.

[0074] In an exemplary embodiment, the following steps may further be included:

[0075] During the process of writing data through the storage integrated system, the to-be-processed power user activity data is written into the data connection pool for storage; according to the service identification information of the power user activity data, the target storage system of the power user activity data is determined; through the data connection pool, the stored power user activity data is asynchronously written into the target storage system.

[0076] In specific implementation, based on the distributed multi-level cache, it is necessary to consider how to ensure data consistency during asynchronous writing. Through asynchronous writing consistency processing, the data can be first written into an intermediate layer or a connection pool, and then the data is asynchronously written into the target storage system (such as a database). Based on the asynchronous feature, the writing operation of the data is not immediately synchronized to the target system, but is performed later in the background.

[0077] Exemplarily, when an application needs to write data, the data can be sent to a connection pool or an intermediate layer, and the connection pool or the intermediate layer can temporarily store the data. At the same time, the application can continue to execute other tasks without waiting for the completion step of writing the data into the target system. Then the connection pool or the intermediate layer can asynchronously write the data into the target storage system. For example, the corresponding storage system can be determined according to the service identification, and this writing process may occur at a certain time point after the data is sent to the connection pool, specifically depending on the system scheduling and load conditions. Since this process is asynchronous, it does not block the task execution of the application and can ensure the consistency and integrity of the data when it is finally written into the target system. It can be achieved through a series of mechanisms, such as using transactions to ensure the atomicity and consistency of the data, and using checksums or other data integrity check mechanisms to ensure the accuracy of the data.

[0078] In an optional embodiment, since asynchronous writing may introduce certain data latency and potential data inconsistency risks while improving the throughput and response speed of the system. When designing and using an asynchronous writing system, it is necessary to carefully weigh various factors and select a suitable solution according to the specific application scenario and requirements. Thus, by first writing the data into the intermediate layer and then asynchronously writing it into the target system, the consistency and integrity of the data can be ensured simultaneously, realizing an efficient and reliable data writing operation.

[0079] In an exemplary embodiment, the storage integrated system is connected to multiple power service systems. The hot service separation of the real-time data stream of the power user activities obtained through the storage integrated system may include the following steps:

[0080] Identify the hot-spot services in each of the power business systems; the hot-spot services include high-traffic services and high-load services; separate the hot-spot services according to the service attribute information of each power user activity data in the real-time data stream; the service attribute information is used to determine whether it belongs to a hot-spot service.

[0081] In one example, an optimization strategy for hot-spot service separation can be adopted to improve the performance and stability of the entire system by identifying and independently processing the high-traffic or high-load parts (i.e., hot-spot services) in the system. Optionally, for resource isolation, hot-spot resources can be isolated from other resources by physical or logical means to avoid the impact of hot-spot requests on other services. For example, it can be achieved by setting up different server clusters, database instances, or service instances; for load balancing, load balancing technology can be used to evenly distribute hot-spot traffic to multiple processing units to avoid single-point overload. Based on the dynamic load balancing algorithm, it can be automatically adjusted according to the real-time traffic situation, so as to ensure that hot-spot requests are effectively dispersed.

[0082] In another example, through the caching strategy, a caching mechanism can be implemented for hot-spot data or business logic to reduce direct access to the backend system (such as the database), thereby improving the response speed and reducing system pressure. For example, the caching strategy can include local caching, distributed caching, etc.; for cold-hot separation processing, at the database level, frequently accessed hot data and infrequently accessed cold data can be separated and stored. For example, in Elasticsearch, by setting up hot and cold nodes, active data is stored in the hot node and archived data is stored in the cold node, and different hardware configurations and indexing strategies can be adopted for each; for asynchronous processing, for hot-spot tasks with non-real-time requirements, a message queue or event-driven architecture can be used for asynchronous processing to avoid blocking the main thread and improve system throughput; for automatic scaling, the resource scale can be automatically adjusted according to system monitoring data. When it is detected that the resource pressure caused by hot-spots increases, processing resources are automatically added, and resources can be released after the pressure is relieved to maintain the efficient use of resources.

[0083] In an exemplary embodiment, the following steps may further be included:

[0084] In the process of processing the real-time data stream through the storage integrated system, use the write-ahead log mechanism to perform fault recovery processing on the monitored system fault events and execute the immediate analysis operation on the real-time data stream.

[0085] In practical applications, through real-time data stream processing technology, it is possible to perform immediate analysis on continuous and uninterrupted data streams without waiting for the data to accumulate to a certain amount before performing batch processing. It can support the efficient collection, transmission, and processing of large-scale real-time data streams and ensure low latency and high throughput of data processing.

[0086] Specifically, based on real-time data stream processing technology, it can process data and generate results at a speed ranging from milliseconds to seconds, ensuring near-instantaneous response and decision-making, achieving the effect of low latency; by adopting an event-driven approach, the system can trigger processing logic based on events in the data stream rather than according to a preset schedule; through continuous processing, data can be processed as it is generated without waiting for all the data to be collected, making it suitable for processing infinite data streams; based on the window mechanism, concepts such as time windows or count windows can be used to divide continuous data streams into manageable small chunks for processing, facilitating analysis and aggregation; it can dynamically adjust the processing capacity according to the data traffic to ensure the stable operation of the system during high concurrency or data surges, achieving the effect of dynamic expansion.

[0087] In one example, in terms of fault tolerance processing, it has the ability to recover from failures. Even in the case of partial system component failures (i.e., system fault events), it can ensure the continuity and integrity of data processing. For example, the WAL mechanism of RocksDB can be utilized to record logs before writing data to the main storage to ensure data consistency and persistence, enabling rapid data recovery, reducing the recovery time, and improving the system's recovery speed and availability; the real-time stream processing framework and platform can include Kafka message queue systems, Flink real-time computing engines, Spark Streaming stream computing, Dataflow data stream services, Kinesis real-time data stream services, etc., which can provide distributed processing capabilities and support high throughput and complex data processing logics.

[0088] In an exemplary embodiment, the following steps may further be included:

[0089] Based on the log-structured merge tree in the storage-integrated system, process the power user activity data stream in the high-concurrency data scenario; the log-structured merge tree is used to optimize the write performance of the storage-integrated system.

[0090] In specific implementation, an LSM tree structure can be adopted. It is suitable for write-intensive applications. By writing data into the MemTable in memory and then periodically merging the data in the MemTable into the SSTable on disk, the write performance can be optimized. ROR can efficiently store a large amount of data on disk, reducing the dependence on memory, improving the system's storage capacity while reducing hardware costs, and performing even more outstandingly especially in high-concurrency scenarios.

[0091] In an exemplary embodiment, as Figure 4 shown, a flowchart of another method for processing power user activity data based on a storage-integrated system is provided. In this embodiment, the method includes the following steps:

[0092] In step 401, through the storage integrated system, the real-time data stream of power user activities obtained is separated into hot business data and cold business data. In step 402, during the process of processing the real-time data stream through the storage integrated system, the pre-write log mechanism is used to perform fault recovery processing on the monitored system fault events and execute the immediate analysis operation on the real-time data stream. In step 403, based on the log-structured merge tree in the storage integrated system, the data stream of power user activities in the high-concurrency data scenario is processed; the log-structured merge tree is used to optimize the write performance of the storage integrated system. In step 404, the Redis engine is used to store the hot business data in the first storage medium, and the RocksDB engine is used to store the cold business data in the second storage medium. In step 405, based on the data hierarchical storage information of the storage integrated system, the storage levels of the hot business data and the cold business data are determined. In step 406, when a data access event of the power client is detected, the storage medium conversion processing is performed on the hot business data and the cold business data according to the hot and cold data exchange conditions, and the data reading result for the data access event is fed back.

[0093] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a method for processing power user activity data based on a storage integrated system described above, and will not be elaborated here.

[0094] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0095] Based on the same inventive concept, an embodiment of the present application also provides a device for processing power user activity data based on a storage integrated system for implementing the method for processing power user activity data based on a storage integrated system described above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for processing power user activity data based on a storage integrated system provided below can refer to the limitations of the method for processing power user activity data based on a storage integrated system described above, and will not be elaborated here.

[0096] In an exemplary embodiment, as Figure 5 shown, a power user activity data processing device based on a storage integrated system is provided, including:

[0097] A data separation module 501, configured to perform hot service separation on the acquired real-time data stream of power user activities through the storage integrated system to obtain service hot data and service cold data; the storage integrated system includes a Redis engine and a RocksDB engine;

[0098] A data storage module 502, configured to store the service hot data into a first storage medium by using the Redis engine and store the service cold data into a second storage medium by using the RocksDB engine; the first storage medium is a memory, and the second storage medium is a disk;

[0099] A storage level determination module 503, configured to determine the storage level of the service hot data and the storage level of the service cold data based on the data hierarchical storage information of the storage integrated system; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0100] A data exchange module 504, configured to perform storage medium conversion processing on the service hot data and the service cold data according to the hot and cold data exchange conditions when detecting a data access event of a power client, and feed back a data reading result for the data access event.

[0101] In one of the embodiments, the device further includes:

[0102] A distributed multi-level cache module, configured to perform distributed multi-level cache processing on the real-time data stream through the storage integrated system;

[0103] Wherein, the distributed multi-level cache processing includes collaborative processing among data hierarchical storage, data access optimization, distributed storage and access, data synchronization mechanism, and cache policy setting.

[0104] In one of the embodiments, the device further includes:

[0105] A connection pool writing module, configured to write the to-be-processed power user activity data into a data connection pool for storage during the process of data writing through the storage integrated system;

[0106] A target storage system determination module, configured to determine the target storage system of the power user activity data according to the service identification information of the power user activity data;

[0107] An asynchronous writing module for asynchronously writing the stored power user activity data into the target storage system through the data connection pool.

[0108] In one embodiment, the integrated storage system is connected to multiple power business systems. The data separation module 501 is specifically configured to identify the hot business in each of the power business systems; the hot business includes high-traffic business and high-load business; separate the hot business according to the service attribute information of each power user activity data in the real-time data stream; the service attribute information is used to determine whether it belongs to the hot business.

[0109] In one embodiment, the device further includes:

[0110] A fault recovery module for performing a fault recovery process on the monitored system fault event by using the write-ahead log mechanism during the process of processing the real-time data stream through the integrated storage system, and performing an immediate analysis operation on the real-time data stream.

[0111] In one embodiment, the device further includes:

[0112] A high-concurrency processing module for processing the power user activity data stream in a data high-concurrency scenario based on the log-structured merge tree in the integrated storage system; the log-structured merge tree is used to optimize the writing performance of the integrated storage system.

[0113] Each module in the above power user activity data processing device based on the integrated storage system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0114] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for processing power user activity data based on a storage integrated system.

[0115] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0116] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0117] Through the storage integrated system, perform hot service separation on the acquired real-time data stream of power user activities to obtain service hot data and service cold data; the storage integrated system includes a Redis engine and a RocksDB engine;

[0118] Use the Redis engine to store the service hot data in the first storage medium, and use the RocksDB engine to store the service cold data in the second storage medium; the first storage medium is memory, and the second storage medium is a disk;

[0119] Based on the data hierarchical storage information of the storage integrated system, determine the storage levels of the service hot data and the service cold data; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0120] When a data access event of a power client is detected, perform storage medium conversion processing on the service hot data and the service cold data according to the hot and cold data exchange conditions, and feedback a data reading result for the data access event.

[0121] In one embodiment, when the processor executes the computer program, it also implements the steps of the power user activity data processing method based on the storage integrated system in the above-mentioned other embodiments.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0123] Through the storage integrated system, perform hot spot service separation on the acquired real-time data stream of power user activities to obtain service hot data and service cold data; the storage integrated system includes a Redis engine and a RocksDB engine;

[0124] Use the Redis engine to store the service hot data in a first storage medium, and use the RocksDB engine to store the service cold data in a second storage medium; the first storage medium is memory, and the second storage medium is a disk;

[0125] Based on the data hierarchical storage information of the storage integrated system, determine the storage levels of the service hot data and the service cold data; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0126] When a data access event of a power client is detected, perform storage medium conversion processing on the service hot data and the service cold data according to the hot and cold data exchange conditions, and feedback a data reading result for the data access event.

[0127] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the power user activity data processing method based on the storage integrated system in the above-mentioned other embodiments.

[0128] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0129] Through the storage integrated system, perform hot spot service separation on the acquired real-time data stream of power user activities to obtain service hot data and service cold data; the storage integrated system includes a Redis engine and a RocksDB engine;

[0130] Use the Redis engine to store the business hot data in the first storage medium, and use the RocksDB engine to store the business cold data in the second storage medium; the first storage medium is memory, and the second storage medium is disk;

[0131] Based on the data hierarchical storage information of the storage integrated system, determine the storage levels of the business hot data and the business cold data; the data hierarchical storage information is used to indicate different storage media corresponding to each level;

[0132] When a data access event of the power client is detected, perform storage medium conversion processing on the business hot data and the business cold data according to the hot and cold data exchange conditions, and feedback the data reading result for the data access event.

[0133] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the power user activity data processing method based on the storage integrated system in the above other embodiments.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0137] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for processing power user activity data based on a storage integrated system, characterized in that: The method comprises: Through the integrated storage system, the real-time data stream of the power user activities obtained is separated into hot business, and hot business data and cold business data are obtained; the integrated storage system includes a Redis engine and a RocksDB engine; The business hot data is stored in a first storage medium using the Redis engine, and the business cold data is stored in a second storage medium using the RocksDB engine; the first storage medium is a memory, and the second storage medium is a disk; Based on the data hierarchical storage information of the storage-in-one system, determining the storage level of the business hot data and the storage level of the business cold data; the data hierarchical storage information is used to indicate different storage media corresponding to each level; When a data access event of the power client is detected, storage medium conversion processing is performed on the business hot data and the business cold data according to the hot and cold data exchange condition, and a data reading result for the data access event is fed back.

2. The method according to claim 1, characterized in that The method further comprises: Performing distributed multi-level cache processing on the real-time data stream through the storage-in-one system; The distributed multi-level cache processing includes collaborative processing among data hierarchical storage, data access optimization, distributed storage and access, data synchronization mechanism and cache strategy setting.

3. The method according to claim 1, characterized in that The method further comprises: In the process of writing data through the integrated storage system, the power user activity data to be processed is written into the data connection pool for storage; Determining a target storage system for the power user activity data according to the service identification information of the power user activity data; The stored power user activity data is asynchronously written to the target storage system via the data connection pool.

4. The method according to claim 1, characterized in that The integrated storage system is connected to a plurality of power business systems. The integrated storage system is used to separate hotspot business from the acquired real-time data stream of power user activities, including: Identifying hotspot services in each of the power service systems; the hotspot services include high-visit services and high-load services; Hotspot services are separated according to the service attribute information of each power user activity data in the real-time data stream; the service attribute information is used to determine whether it belongs to a hotspot service.

5. The method according to claim 1, characterized in that The method further comprises: In the process of processing the real-time data stream by the integrated storage system, a write-ahead log mechanism is used to perform fault recovery processing on the monitored system fault events and execute an instant analysis operation on the real-time data stream.

6. The method according to claim 1, characterized in that The method further comprises: Based on the log structure merge tree in the integrated storage system, the power user activity data stream in a high data concurrency scenario is processed; the log structure merge tree is used to optimize the write performance of the integrated storage system.

7. A device for processing power user activity data based on a storage integrated system, characterized in that: The device comprises: A data separation module is used to separate hot business from real-time data streams of power user activities acquired through the integrated storage system to obtain hot business data and cold business data; the integrated storage system includes a Redis engine and a RocksDB engine; A data storage module, used to store the hot business data in a first storage medium using the Redis engine, and to store the cold business data in a second storage medium using the RocksDB engine; the first storage medium is a memory, and the second storage medium is a disk; A storage level determination module, used to determine the storage level of the business hot data and the storage level of the business cold data based on the data hierarchical storage information of the storage integrated system; the data hierarchical storage information is used to indicate different storage media corresponding to each level; The data exchange module is used to perform storage medium conversion processing on the business hot data and the business cold data according to the hot and cold data exchange conditions when a data access event of the power client is detected, and to feed back the data reading result for the data access event.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Intelligent power transaction data processing method, system, medium and equipment

    CN121166711A