Data management solution
Through the method of data integration, synchronization and automatic update of cache version numbers, the problem of inconsistent data storage and cache expiration time in traditional data storage and cache management methods is solved, and unified storage and cache management of business data is realized, and system performance and user experience are improved.
Patent Information
- Application Number
- CN202510053076.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
Smart Images

Figure CN119961293A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management, and specifically is a data management solution. Background Art
[0002] In a complex business data environment, data statistics of different date dimensions (such as weekly, monthly, quarterly, and annual) are stored in different data tables. Since the project involves collaborative development by multiple people, the same data table may need to support query requirements of multiple different businesses. In order to improve query performance, some queries are added with a cache mechanism and different cache expiration times are set. These time settings vary depending on business differences and developers.
[0003] When dealing with complex data query requirements, the traditional data storage and cache management methods store data of different time dimensions of the same business in different tables, and the cache expiration time of different business queries in the same table is inconsistent. However, this method has problems such as decentralized data storage, inconsistent cache expiration time, and inconvenient data management. These problems not only increase the complexity of development and maintenance, but may also lead to inconsistent data and user confusion, thereby increasing the interpretation cost of the platform. Therefore, it needs to be improved. Summary of the invention
[0004] The purpose of the present invention is to provide a data governance solution to solve the problems raised in the above background technology.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a data governance solution, comprising the following steps:
[0006] S1. Data integration and synchronization: Collect and preprocess data, create Hive tables, load preprocessed data into Hive tables, and use data synchronization tools to synchronize data in Hive tables to TiDB.
[0007] S2. Update cache version number: Design an interface and configure synchronization tasks. After data synchronization is completed, automatically call the interface to update the cache version number.
[0008] S3. Business query: Design a cache mechanism in the business system. When a user initiates a query request, check whether there is a query result in the cache that matches the cache version number specified in the request, and store the query result and the latest cache version number in the cache.
[0009] Preferably, when synchronizing multi-time dimension data, first create the corresponding table structure in Hive, write a HiveQL script, clean and convert the data, and use the data synchronization tool to synchronize the data in Hive to the TiDB table.
[0010] Preferably, when the version number needs to be updated, first configure a scheduled task in the scheduling system, write an interface to receive the call of the Job, and modify the cached version number of the corresponding table in Apollo, configure the Job to call the above interface after data synchronization is completed, and update the cached version number in Apollo.
[0011] Preferably, when a user requests data, the system first generates a cache key based on the user-defined key and the current table cache version number in Apollo, and uses the cache key to query data in the COIDS cache system.
[0012] Preferably, when a user queries data based on a new cache key, if there is data matching the cache key in the cache system, the data is directly returned to the user. If there is no data matching the cache key in the cache system, the data is queried from TiDB based on the conditions requested by the user.
[0013] Preferably, after querying the corresponding data from the TiDB table, the queried data is stored in the COIDS cache system using the previously generated cache key, the data queried from the TiDB table is returned to the user, and the cache is updated at the same time.
[0014] Preferably, in S2:
[0015] The interface should receive the table name and the new cache version number as parameters. The interface should be implemented on the backend to update the cache version number of the specified table by calling the API of the Apollo configuration center.
[0016] Preferably, in S3:
[0017] Design cache update and invalidation strategies based on business needs and data change frequency, monitor cache usage and performance, and optimize and adjust as needed.
[0018] The beneficial effects of the present invention are as follows:
[0019] The present invention achieves unified business display by uniformly storing multi-time dimension business data, simplifying program logic, and dynamically changing cache version numbers, thereby improving system performance and stability, reducing development and maintenance costs, and improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a functional architecture diagram of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] like Figure 1 As shown, an embodiment of the present invention provides a data governance solution, including the following steps:
[0023] S1. Data integration and synchronization: Collect and preprocess data, create Hive tables, load preprocessed data into Hive tables, and use data synchronization tools to synchronize data in Hive tables to TiDB.
[0024] S2. Update cache version number: Design an interface and configure synchronization tasks. After data synchronization is completed, automatically call the interface to update the cache version number.
[0025] S3. Business query: Design a cache mechanism in the business system. When a user initiates a query request, check whether there is a query result in the cache that matches the cache version number specified in the request, and store the query result and the latest cache version number in the cache.
[0026] By uniformly storing multi-time dimension business data, simplifying program logic, and dynamically changing cache version numbers, the unified business display is achieved, the system performance and stability are improved, the development and maintenance costs are reduced, and the user experience is improved.
[0027] When synchronizing multi-time dimension data, first create the corresponding table structure in Hive, write a HiveQL script, clean and convert the data, and use the data synchronization tool to synchronize the data in Hive to the TiDB table.
[0028] Creating a table structure in Hive can ensure that data is stored in a unified format. By writing HiveQL scripts, you can clean and transform the original data, remove invalid, duplicate, or erroneous data, and convert the data into a form suitable for analysis, which helps improve work efficiency, ensure data accuracy and consistency, and meet the needs of future business growth.
[0029] When the version number needs to be updated, first configure a scheduled task in the scheduling system, write an interface to receive the call of the Job, and modify the cached version number of the corresponding table in Apollo. Configure the Job to call the above interface after data synchronization is completed to update the cached version number in Apollo.
[0030] Realize automatic update of version numbers, reduce manual operations, and improve work efficiency; ensure that version numbers are updated in a timely manner after data synchronization is completed to avoid data inconsistency issues; enhance system reliability and stability through scheduled tasks and Job calls, and ensure the accuracy and timeliness of version number updates.
[0031] Among them, when a user requests data, the system first generates a cache key based on the user-defined key and the current table cache version number in Apollo, and uses the cache key to query data in the COIDS cache system.
[0032] By quickly locating data through cache keys, the number of database queries is reduced, and the system response speed is improved. Combined with the cache version number, it ensures that users obtain the latest and consistent data. The use of the cache system enables the system to more flexibly respond to the growth of data volume and user requests.
[0033] When a user queries data based on a new cache key, if there is data matching the cache key in the cache system, the data will be directly returned to the user. If there is no data matching the cache key in the cache system, the data will be queried from TiDB based on the conditions requested by the user.
[0034] When the cache hits, data is returned quickly, significantly reducing query latency and improving user experience. When the cache misses, although the database needs to be queried, overall direct access to the database is reduced, which helps protect database performance, rationally utilizes cache resources, reduces unnecessary calculations and operations, and improves the overall efficiency of the system.
[0035] Among them, after querying the corresponding data from the TiDB table, the queried data is stored in the COIDS cache system using the previously generated cache key, and the data queried from the TiDB table is returned to the user, and the cache is updated at the same time.
[0036] Storing query results in cache can speed up subsequent processing of the same request and improve system response speed; reduce the number of direct queries to TiDB, help protect database resources and extend its service life; quickly respond to user requests, provide a smooth user experience, and enhance user satisfaction.
[0037] Among them, in S2:
[0038] The interface should receive the table name and the new cache version number as parameters. Implement the interface on the backend and update the cache version number of the specified table by calling the API of the Apollo configuration center.
[0039] Through interface implementation, you can easily expand the update operations on different tables and cache version numbers, avoid errors that may be caused by manual updates, and improve the accuracy and efficiency of update operations.
[0040] Among them, in S3:
[0041] Design cache update and invalidation strategies based on business needs and data change frequency, monitor cache usage and performance, and optimize and adjust as needed.
[0042] Ensure that the data in the cache is always up to date, optimize cache usage and improve system response speed; avoid cache resource waste and achieve efficient resource utilization through monitoring and optimization.
[0043] Apollo Configuration Center is an open source configuration management center developed by Ctrip's framework department. It can centrally manage the configuration of different application environments and clusters. After the configuration is modified, it can be pushed to the application end in real time, and it has standardized permissions, process governance and other features.
[0044] Hive is a data warehouse tool based on Hadoop, which is used for data extraction, transformation and loading. The hive data warehouse tool can map structured data files into a database table and provide SQL query functions.
[0045] Tidb is an open source distributed relational database that supports both online transaction processing and online analytical processing.
[0046] Codis is a distributed Redis solution. For upper-level applications, there is no obvious difference between connecting to Codis Proxy and connecting to the native Redis Server. Upper-level applications can use it just like a stand-alone Redis. The underlying Codis will handle request forwarding, data migration without downtime, and other tasks. All the subsequent things are transparent to the front-end client, and you can simply think of it as a Redis service with infinite memory connected to it.
[0047] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0048] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data governance solution, characterized by: The following steps are involved: S1. Data integration and synchronization: Collect and preprocess data, create Hive tables, load preprocessed data into Hive tables, and use data synchronization tools to synchronize data in Hive tables to TiDB. S2. Update cache version number: Design an interface and configure synchronization tasks. After data synchronization is completed, automatically call the interface to update the cache version number. S3. Business query: Design a cache mechanism in the business system. When a user initiates a query request, check whether there is a query result in the cache that matches the cache version number specified in the request, and store the query result and the latest cache version number in the cache.
2. A data governance solution according to claim 1, characterized in that: When synchronizing multi-time dimension data, first create the corresponding table structure in Hive, write a HiveQL script, clean and convert the data, and use the data synchronization tool to synchronize the data in Hive to the TiDB table.
3. A data governance solution according to claim 1, characterized in that: When the version number needs to be updated, first configure a scheduled task in the scheduling system, write an interface to receive the call of the Job, and modify the cached version number of the corresponding table in Apollo. Configure the Job to call the above interface after data synchronization is completed to update the cached version number in Apollo.
4. A data governance solution according to claim 1, characterized in that: When a user requests data, the system first generates a cache key based on the user-defined key and the current table cache version number in Apollo, and uses the cache key to query data in the COIDS cache system.
5. A data governance solution according to claim 1, characterized in that: When a user queries data based on a new cache key, if there is data matching the cache key in the cache system, the data will be returned directly to the user. If there is no data matching the cache key in the cache system, the data will be queried from TiDB based on the conditions requested by the user.
6. A data governance solution according to claim 1, characterized in that: After querying the corresponding data from the TiDB table, the queried data is stored in the COIDS cache system using the previously generated cache key, and the data queried from the TiDB table is returned to the user, and the cache is updated at the same time.
7. A data governance solution according to claim 1, characterized in that: In S2: the interface receives the table name and the new cache version number as parameters, implements the interface on the backend, and updates the cache version number of the specified table by calling the API of the Apollo configuration center.
8. A data governance solution according to claim 1, characterized in that: In the S3: Design cache update and invalidation strategies based on business needs and data change frequency, monitor cache usage and performance, and optimize and adjust as needed.