A method for managing data product versions based on a data grid
Through distributed data grid architecture and dual-temporal data modeling, the efficiency and reliability of data product version management are solved, precise control and historical traceability of data products are achieved, and the uniqueness and security of the version are ensured.
Patent Information
- Application Number
- CN202311144559.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-06
AI Technical Summary
The prior art is difficult to efficiently manage versions of data products in data grids, especially when data changes frequently, errors are prone to and lack of structured management, resulting in data consistency and integrity problems.
A distributed data grid architecture is adopted, and the actual and processing time of data products is recorded using dual-temporal data modeling, immutable versions are generated, and the uniqueness and security of the version are ensured through global unique identifiers and blockchain technology, and consensus authentication is carried out between nodes.
It realizes efficient management of data product versions, ensures accurate control of data modification, deletion and addition, supports historical traceability and restoration, and improves the efficiency and reliability of version management.
Smart Images

Figure CN117633735B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and relates to a method for managing data product versions based on a data grid. Background Art
[0002] With the continuous expansion of data scale and the increasing complexity of organizational scale, the existing monolithic, centralized data ownership, and technology-oriented data management solutions can no longer meet the strategic development of the company, and the investment in data far fails to reach the expected return; the data management architecture of the new generation of data grid realizes true data distribution, and proposes to regard data as a product. The data grid is a way of decentralized, autonomous, and scalable data management by decentralizing data ownership and responsibility to domain teams, and adopting principles such as product thinking, platform-based data infrastructure, and standardized governance. In the data grid, a data product is an architectural quantum. It is the smallest architectural unit that can be independently deployed and managed. It has high functional cohesion, that is, it performs specific analysis and transformation and securely shares the results as domain-oriented analysis data. It has all the structural components required to execute its functions: transformation code, data, metadata, policies for managing data, and its dependencies on the infrastructure. During the initialization process of a new data product, it is automatically registered in the grid. It is assigned a unique global identifier and address, and makes the data product visible to the grid and governance processes. Therefore, the management of data versions is an important link, which involves the management and association of distributed data products, and the trustworthy traceability of data. The traditional implementation method is to use a centralized registry or directory to list available data sets and provide some additional information about each data set, owner, location, sample data, etc. Usually, this information is retrospectively planned by a centralized data team or governance team. There are some problems to be solved in such a data version management method. First of all, the current data version management method is usually based on the file system, that is, different versions of data are stored as different files. However, this method has some difficulties. For example, when data products change frequently, managing multiple file versions becomes cumbersome and error-prone. In addition, tracking and restoring specific versions of data also becomes complex, especially when multiple users modify the data simultaneously. Secondly, the current version management method lacks structured management of data products. Data products usually consist of multiple data sets and objects, and the current method fails to provide an effective mechanism to manage the versions of each data set and object. Therefore, when it is necessary to modify, delete, or add a specific data set or object, version control becomes difficult, easily leading to data consistency and integrity problems. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for managing data product versions based on a data grid, which solves the technical problem of improving the efficiency and reliability of version management by using a data grid.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A data product version management method based on a data grid, comprising the following steps:
[0006] Step 1: Construct a distributed data grid architecture management system, including a distributed network composed of several nodes. Each node obtains service data through a local device, and the local device is connected to the data grid through the distributed network;
[0007] Step 2: Each node stores and manages the service data in its respective domain, realizes the distributed management of data through the data grid. When a node obtains service data, it records it at the actual time to generate a data product;
[0008] Step 3: The node uses two-temporal data modeling to record the change time of the data product, so that each data product records two timestamps and creates a new version of the data product;
[0009] The two timestamps include the actual time and the processing time;
[0010] Step 4: When processing a data product, the node selects service data for processing within a fixed interval of the actual time to obtain the processing time, and generates a new version of the data product based on the processing time. This version of the product is finalized and immutable;
[0011] Step 5: Subsequent data products are subjected to secondary analysis and processing based on the previous data products with fixed processing times to generate new data products and their corresponding versions of processing times;
[0012] Step 6: The node uses two-temporal data modeling to track the actual time and processing time of each version of the data product;
[0013] Step 7: A globally unique identifier is assigned to each version of the data, and consensus authentication is performed between different nodes through distributed data management.
[0014] Preferably, the nodes communicate with each other through a preset protocol.
[0015] Preferably, the actual time is the time when the event related to the data product actually occurs or the actual time of the state of the service data, and the processing time is the time when the data product is processed, specifically including the time for observing, processing, recording, and providing knowledge or understanding of its state or event within a specific time.
[0016] Preferably, the processing time monotonically increases whenever a new data entity is processed, and the processing time is the only trusted monotonically advancing time.
[0017] Preferably, the data attribute convention of the data product includes:
[0018] Convention 1: The data attributes of each version are agreed to be read-only through the system protocol;
[0019] Convention 2: Once the data occurs, the record cannot be changed, unless the system as a whole determines that the data is deleted or recycled;
[0020] Convention 3: The unique version of the data is ensured through bi-temporal data modeling, Convention 1, and Convention 2.
[0021] Preferably, only data products with the same processing time are jointly used to generate a new data product.
[0022] A method for managing data product versions based on a data grid according to the present invention solves the technical problem of improving version management efficiency and reliability by using a data grid, can effectively manage multiple data product versions, precisely control data modification, deletion, and addition, realize historical backtracking and restoration of data products, and the global unique identifier and blockchain technology ensure the uniqueness and security of version identification, improving version management efficiency and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is the main flowchart of the present invention;
[0024] Figure 2 is the schematic diagram of version management of the data product of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] As Figure 1 - Figure 2 a method for managing data product versions based on a data grid as described above, includes the following steps:
[0026] Step 1: Construct a distributed data grid architecture management system, including a distributed network composed of several nodes. Each node obtains service data through a local device, and the local device is connected to the data grid through the distributed network; the nodes communicate with each other through a preset protocol.
[0027] In this embodiment, the local device can be a data storage server for storing service data. For example, the health data of users, including various versions of steps, heart rate, and sleep conditions. Each version will be stored in the database and associated with the corresponding user account.
[0028] The data grid is the core of the distributed architecture, composed of multiple nodes, forming a distributed network. Each node stores and manages the service data in its respective domain to achieve distributed management of data. The nodes communicate with each other through protocols to jointly maintain data consistency and integrity.
[0029] Step 2: Each node stores and manages business data in its respective domain, and realizes distributed management of data through a data grid. When a node obtains business data, it records it at the actual time to generate a data product.
[0030] Step 3: The node uses bi-temporal data modeling to record the change time of the data product, so that each data product records two timestamps and creates a new version of the data product.
[0031] The two timestamps include the actual time and the processing time.
[0032] In this embodiment, the actual times of the same data product can be the same, and the processing times can be different. Different data products can have the same actual time and processing time.
[0033] The actual time is the time when the event related to the data product actually occurs or the actual time of the state of the business data. The processing time is the time when the data product is processed, specifically including the time to observe, process, record, and provide knowledge or understanding of its state or event within a specific time.
[0034] The processing time monotonically increases whenever a new data entity is processed, and the processing time is the only trusted monotonically advancing time.
[0035] Only data products with the same processing time are jointly used to generate a new data product.
[0036] Step 4: When processing a data product, the node selects business data for processing within a fixed interval of the actual time to obtain the processing time, and generates a new version of the data product based on the processing time. This version of the product is finalized and immutable.
[0037] Step 5: Subsequent data products perform secondary analysis and processing based on the previous data products with fixed processing times to generate new data products and their corresponding versions of processing times.
[0038] Step 6: The node tracks the actual time and processing time of each version of the data product by using bi-temporal data modeling.
[0039] Step 7: A globally unique identifier is assigned to each version of the data, and nodes authenticate each other through distributed data management for consensus.
[0040] In this embodiment, blockchain technology can be used to record the change history of data to ensure the immutability of data.
[0041] The data attribute agreements of the data product include:
[0042] Agreement 1: It is agreed through the system protocol that the data attributes of each version are read-only;
[0043] Agreement 2: Once the data occurs, the record cannot be changed, unless the system as a whole determines that the data is deleted or recycled;
[0044] Agreement 3: The unique version of the data is ensured through bi-temporal data modeling, Agreement 1 and Agreement 2.
[0045] As Figure 2 shown, in this embodiment, the business data flows into the system metadata pool through the terminal, and the time of its occurrence is recorded, that is, the actual generation time of each piece of data. Metadata is the necessary input to form a certain basic data product. For example, Data Product 1 will regularly capture business data from the metadata pool of Business Data 1 to form a data product. The range of metadata captured by each version of the data product, the strategy, and the processing method of the metadata can all be different. Here, the metadata, strategy, transformation code, and dependencies can all change according to the version.
[0046] Based on Business Data 1, Data Product 1 is generated, and based on Business Data 2, Data Product 2 is generated. Data Product 1 and Data Product 2 are two independent data products, reflecting different functions. On this basis, the output data of Data Product 1 and Data Product 2 can also be regarded as the business data adopted by Data Product 3. According to the principle of data products, data products with the same processing time can be combined to generate a new data product, that is, Data Product 1 with a processing time of 1 and Data Product 2 with a processing time of 1 can be combined to generate Data Product 3.
[0047] In an application scenario of this embodiment, the same data product can have different versions, and the judgment of different versions increases according to different processing times. For example, for Data Product 1 with a version of January 1, 2023 as the processing time, all the data with an actual time from June 2022 to December 2022 in the business data is adopted. This actual time can change in each version of the data product. For example, Figure 2 in addition to the processing time increasing from top to bottom representing the unique version identifier, the remaining attributes in the data product are variable.
[0048] Taking an actual application scenario of a health data management data product as an example in this embodiment, the health data management data product is used to track the health data of users, including the number of steps, heart rate, and sleep conditions. The specific process of managing the data product version is as follows:
[0049] Step S1: Data inflow and recording, including the user using the mobile phone (local device) application to record the daily number of steps, heart rate, and sleep conditions. These data are recorded at the actual time, that is, the time when the user completes the record.
[0050] Step S2: Version management based on the time dimension, including that whenever a user records data, the data flows into the system metadata pool as business data at the actual time of occurrence; the data product uses two-temporal data modeling to regularly build a new version of the data product. For example, the user automatically records the cumulative number of steps of 8,000 steps from 0:00 today to 8:00 am at the current time, and the heart rate value of 75 bpm at the current moment (the local device automatically records the above information every minute). The system defines that a new version of the data product is established at each whole hour, and the data used at this time is all the recorded data in the actual time interval of [0:00, 8:00]. The processing time is the time when this version is created, such as '8:00 on August 1, 2023'. The data product analyzes through the data used at this time and returns to the user the calorie consumption during exercise and the fluctuation analysis of the heart rate.
[0051] Step S3: Version identification and association, each data product version will be assigned a globally unique identifier (GUID) to ensure the uniqueness and stability of the version. For example, the version of the health data management data product at 7:00 on August 1, 2023 is identified as GUID-X1, and the version of the health data management data product at 8:00 on August 1, 2023 is represented as GUID-X2, and so on.
[0052] Step S4: Historical traceability of two-temporal data, by using two-temporal data modeling, the system can accurately track the actual time and processing time of each data version. If the user wants to view the change in the number of steps in the past week, the system can trace back to the past week according to the actual time and provide the data of each version during that period.
[0053] Step S5: Distributed data verification, the system uses a globally unique identifier (GUID) to identify each data version to ensure the uniqueness of the version. In addition, the system can adopt a distributed algorithm to enable different nodes to cooperate in processing data and ensure data consistency and integrity.
[0054] A method for managing data product versions based on a data grid according to the present invention solves the technical problem of improving the efficiency and reliability of version management by using a data grid, can effectively manage multiple data product versions, precisely control the modification, deletion, and addition of data, realize the historical traceability and restoration of data products, and the globally unique identifier and blockchain technology ensure the uniqueness and security of version identification, improving the efficiency and reliability of version management.
Claims
1. A method for managing data product versions based on a data grid, characterized in that: It includes the following steps: Step 1: Construct a distributed data grid architecture management system, including a distributed network composed of several nodes. Each node obtains service data through a local device, and the local device is connected to the data grid through the distributed network; Step 2: Each node stores and manages the service data in its respective domain, realizes the distributed management of data through the data grid. When a node obtains service data, it records it at the actual time to generate a data product; The data attribute conventions of the said data product include: Convention 1: The data attributes of each version are agreed to be read-only through the system protocol; Convention 2: Once the data occurs, the record is immutable, unless the system as a whole determines that the data is deleted or recycled; Convention 3: Ensure the unique version of the data through bi-temporal data modeling, Convention 1 and Convention 2; Step 3: The node uses bi-temporal data modeling to record the change time of the data product, so that each data product records two timestamps and creates a new version of the data product; The two timestamps include the actual time and the processing time; Step 4: When the node processes the data product, it selects service data for processing within a fixed interval of the actual time to obtain the processing time, and generates a new version of the data product based on the processing time. The version of this product is set to be immutable; Step 5: Subsequent data products perform secondary analysis and processing based on the previous data products with fixed processing times to generate new data products and their corresponding versions of processing times; Step 6: The node uses bi-temporal data modeling to track the actual time and processing time of each version of the data product; Step 7: A globally unique identifier is assigned to each version of the data. The nodes authenticate each other's identities through distributed data management in a consensus manner.
2. The data product version management method based on a data grid according to claim 1, wherein: The nodes communicate with each other through a preset protocol.
3. The data product version management method based on a data grid according to claim 1, characterized in that: The actual time is the time when the event related to the data product actually occurs or the actual time of the state of the service data. The processing time is the time when the data product is processed, specifically including the time to observe, process, record and provide its knowledge or understanding of the state or event within a specific time.
4. The data product version management method based on a data grid according to claim 1, characterized in that: The processing time monotonically increases whenever a new data entity is processed, and the processing time is the only trusted monotonically advancing time.
5. The data product version management method based on a data grid according to claim 1, characterized in that: Only data products with the same processing time are combined to generate new data products.
Citation Information
Patent Citations
Handling temporal data in append-only databases
US20180300377A1