A data management method and system based on data blood relationship tracing

CN117667912BActive Publication Date: 2026-09-15ZHENGZHOU DIWEILEPU TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311821657.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2026-09-15
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

[0002]在不同的业务场景中,企业内部的不同数据往往存在一定的关联关系,例如在电信企业中,用户的消费数据、流量使用数据、基站的网络流量数据存在一定的关联关系,但是由于不同类型的数据之间往往存放于不同的数据库,因此若不能基于不同的数据之间的数据血缘关系进行关联数据的确定,则无法实现基于不同类型的关联数据对问题数据进行识别以及校正处理,导致企业内部的数据质量难以满足要求

Benefits of technology

1、在本发明中通过使用情况和更新情况进行数据中的根节点数据的确定,综合考虑到由于使用频繁程度以及更新频繁程度的差异导致的不同的数据的可靠性的差异,实现了对可靠性较低以及使用可靠性要求较高的数据的筛选,也为进一步提升数据处理的可靠性奠定了基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117667912B_ABST
    Figure CN117667912B_ABST
Patent Text Reader

Abstract

The application provides a data management method and system based on data blood relationship tracing, and belongs to the technical field of data processing, and specifically comprises the following steps: determining root node data in data by use and update, taking the number of sector layers as a constraint condition, taking the root node data as a basis, determining associated data in the associated sector of the root node data and the data label of the associated data through the data blood relationship between different data, determining the weight value of different associated data through the number of sector layers where different associated data are located, and determining the data verification period of different root node data and associated data in the associated sector of the root node data in combination with the change of the associated data, so that the accuracy of data is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, and in particular relates to a data management method and system based on data lineage tracing. Background Technology

[0002] In different business scenarios, different data within an enterprise often have certain relationships. For example, in a telecommunications company, user consumption data, data usage data, and base station network traffic data are related. However, since different types of data are often stored in different databases, if the relationship between different data cannot be determined, it is impossible to identify and correct problematic data based on the relationship between different types of related data, resulting in the data quality within the enterprise failing to meet requirements.

[0003] To address the aforementioned technical problems, this invention provides a data management method and system based on data lineage tracing. Summary of the Invention

[0004] To achieve the objectives of this invention, the following technical solution is adopted: According to one aspect of the present invention, a data management method based on data lineage tracing is provided.

[0005] A data management method based on tracing data lineage, characterized in that it specifically includes: S1 determines the usage of different data by extracting and using different data records within the enterprise, and determines the update status of different data based on the update records of different data. The root node data in the data is determined by the usage and update status. S2 determines the data management type of different root node data based on the data anomaly situation in different update processes of different root node data, and uses the data management type to determine the number of traceability sector layers of different root node data; S3 uses the number of traceable sector layers as a constraint and the root node data as a basis to determine the associated data within the associated sector of the root node data and the data tags of the associated data through the data lineage relationship between different data. S4 determines the weight values ​​of different associated data based on the sector level of different root node data, and determines the data verification period of different root node data and associated data in the associated sectors of the root node data based on the changes in the associated data.

[0006] The beneficial effects of this invention are as follows: 1. In this invention, the root node data in the data is determined by the usage and update conditions. Taking into account the differences in the reliability of different data due to the differences in the frequency of use and the frequency of update, the data with low reliability and the data with high reliability requirements are filtered out, which also lays the foundation for further improving the reliability of data processing.

[0007] 2. In this invention, the number of traceability sector layers for different root node data is determined by using data management types. This enables the filtering of the number of associated data that need to be traced and processed for root node data in case of anomalies during the data update process. This ensures the reliability of data update management for data with more serious anomalies, while also avoiding the technical problem of slow processing efficiency caused by a large amount of associated data.

[0008] 3. In this invention, the data verification period of different root node data and related data in the related sectors of the root node data is determined by utilizing the changes in related data. This realizes the determination of differentiated data verification period from the perspective of data update. At the same time, by determining the weight value, the difference in the degree of impact of the update of related data on the data due to the difference in the degree of association between different related data is also fully considered.

[0009] A further technical solution is that the extraction and usage records are determined based on the analysis results of the database logs, specifically by determining the extraction and usage methods of different data in the database through the analysis results of the database logs.

[0010] A further technical solution is that the usage includes the number of times the data is used, the number of business scenarios in which it is used, and the frequency of use of different business scenarios.

[0011] A further technical solution is that the update status includes the number of times the data is updated and the time interval between different update counts.

[0012] A further technical solution involves determining the root node data in the data based on the comprehensive application frequency, specifically including: The frequency threshold is determined based on the amount of data within the enterprise and the frequency of data usage. The root node data in the data is determined based on the frequency threshold and the comprehensive application frequency.

[0013] A further technical solution is that the data anomaly includes the number of data anomalies and the anomaly types of different numbers of data anomalies.

[0014] A further technical solution is that the method for determining the data management type of the root node data is as follows: The number of data anomalies in the root node data and the anomaly type of each anomaly number are determined by the data anomalies in the root node data during different update processes. The degree of anomaly of each anomaly type is determined by the proportion of the number of anomalies of each anomaly type in different update processes. The overall anomaly level of the root node data is determined based on the anomaly level of different anomaly types, and the data management type of the root node data is determined based on the overall anomaly level.

[0015] A further technical solution involves determining the data management type of the root node data based on the comprehensive anomaly level, specifically including: The anomaly level of the root node data is determined by the comprehensive anomaly level, and the data management type of the root node data is determined based on the anomaly level range.

[0016] A further technical solution involves using the data management type to determine the number of traceability sector layers for different root node data, specifically including: The range of traceability sector layers for the root node data is determined by the data management type, and the number of traceability sector layers for the root node data is determined based on the overall anomaly level of the root node data.

[0017] On the other hand, the present invention provides a computer system comprising: a memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, characterized in that: when the processor runs the computer program, it executes the above-described data management method based on data lineage tracing.

[0018] Other features and advantages will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0020] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0021] Figure 1 A flowchart of a data management method based on tracing data lineage; Figure 2This is a flowchart illustrating the method for determining the root node data in the data. Figure 3 This is a flowchart illustrating the method for determining the data verification cycle. Detailed Implementation

[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that the invention will be thorough and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and therefore their detailed description will be omitted.

[0023] The terms “a,” “one,” “the,” and “the” are used to indicate the existence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended meaning of inclusion and that other elements / components / etc. may exist in addition to the listed elements / components / etc.

[0024] Example 1 To solve the above problems, according to one aspect of the present invention, such as Figure 1 As shown, according to one aspect of the present invention, a data management method based on data lineage tracing is provided, characterized in that it specifically includes: S1 determines the usage of different data by extracting and using different data records within the enterprise, and determines the update status of different data based on the update records of different data. The root node data in the data is determined by the usage and update status. Specifically, the extraction and usage records are determined based on the analysis results of the database logs. Specifically, the extraction and usage of different data in the database are determined based on the analysis results of the database logs.

[0025] Furthermore, the usage details include the number of times the data is used, the number of business scenarios in which it is used, and the frequency of use of different business scenarios.

[0026] It is understood that the update status includes the number of times the data is updated and the time interval between different update counts.

[0027] Specific examples, such as Figure 2 As shown, the method for determining the root node data in the data in step S1 is as follows: The frequency of data usage is determined by the usage data, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. The frequency of data usage is determined by the number of times the data is used, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. The update frequency of the data is determined by the update status and the average time interval between different update counts. The overall application frequency of the data is determined based on the frequency of updates and usage of the data, and the root node data in the data is determined based on the overall application frequency.

[0028] Furthermore, determining the root node data in the data based on the comprehensive application frequency specifically includes: The frequency threshold is determined based on the amount of data within the enterprise and the frequency of data usage. The root node data in the data is determined based on the frequency threshold and the comprehensive application frequency.

[0029] In another possible embodiment, the method for determining the root node data in the data in step S1 is as follows: The number of times the data is used is determined based on the usage data, and it is determined whether the data does not belong to the root node data based on the number of times it is used. If it does, the data does not belong to the root node data; otherwise, proceed to the next step. The update count of the data is determined based on the update status, and it is determined whether the data does not belong to the root node data based on the update count. If so, the data does not belong to the root node data; otherwise, proceed to the next step. Based on the number of business scenarios using the data and the frequency of use of different business scenarios, the frequency of use of the data is determined by the number of times the data is used, the number of business scenarios using the data, and the frequency of use of different business scenarios. Based on the frequency of use of the data, it is determined whether the data does not belong to the root node data. If so, the data is determined not to belong to the root node data. If not, proceed to the next step. The update frequency of the data is determined based on the number of times the data is updated and the average time interval between different update frequencies. Based on the update frequency of the data, it is determined whether the data does not belong to the root node data. If so, the data does not belong to the root node data. If not, proceed to the next step. The overall application frequency of the data is determined based on the frequency of updates and usage of the data, and the root node data in the data is determined based on the overall application frequency.

[0030] In another possible embodiment, the method for determining the root node data in the data in step S1 is as follows: The number of times the data is used is determined based on the usage data, and it is determined whether the data does not belong to the root node data based on the number of times it is used. If it does, the data does not belong to the root node data; otherwise, proceed to the next step. The update count of the data is determined based on the update status, and it is determined whether the data does not belong to the root node data based on the update count. If so, the data does not belong to the root node data; otherwise, proceed to the next step. Obtain the number of business scenarios in which the data is used and the frequency of use of different business scenarios. Determine the frequency of use of the data by the number of times the data is used, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. Based on the frequency of use of the data, determine whether the data does not belong to the root node data. If yes, then determine that the data does not belong to the root node data. If no, proceed to the next step. The update frequency of the data is determined based on the number of times the data is updated and the average time interval between different update frequencies. Based on the update frequency of the data, it is determined whether the data does not belong to the root node data. If so, the data does not belong to the root node data. If not, proceed to the next step. The overall application frequency of the data is determined based on the frequency of updates and usage of the data, and the root node data in the data is determined based on the overall application frequency.

[0031] In another possible embodiment, the method for determining the root node data in the data in step S1 is as follows: S11 determines the number of times the data is used based on the usage situation, and judges whether the number of times the data is used is greater than the preset number of times. If yes, proceed to step S13; otherwise, proceed to the next step. S12 obtains the number of business scenarios in which the data is used and the frequency of use of different business scenarios. The frequency of use of the data is determined by the number of times the data is used, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. It is then determined whether the frequency of use of the data is greater than the preset usage busyness. If yes, proceed to the next step; otherwise, proceed to step S15. S13 determines the number of times the data is updated based on the update status, and determines whether the number of times the data is updated is greater than the preset number of updates. If so, it is determined that the data belongs to the root node data; otherwise, proceed to the next step. S14 determines the update frequency of the data based on the number of times the data is updated and the average time interval between different update times. Based on the update frequency of the data, it determines whether the data belongs to the root node data. If yes, the data belongs to the root node data. If no, proceed to the next step. S15 determines the comprehensive application frequency of the data based on the update frequency and usage frequency of the data, and determines the root node data in the data based on the comprehensive application frequency.

[0032] S2 determines the data management type of different root node data based on the data anomaly situation in different update processes of different root node data, and uses the data management type to determine the number of traceability sector layers of different root node data; Furthermore, the data anomalies include the number of data anomalies and the anomaly types for different numbers of data anomalies.

[0033] In one possible embodiment, the method for determining the data management type of the root node data in step S2 is as follows: The number of data anomalies in the root node data and the anomaly type of each anomaly number are determined by the data anomalies in the root node data during different update processes. The degree of anomaly of each anomaly type is determined by the proportion of the number of anomalies of each anomaly type in different update processes. The overall anomaly level of the root node data is determined based on the anomaly level of different anomaly types, and the data management type of the root node data is determined based on the overall anomaly level.

[0034] Furthermore, determining the data management type of the root node data based on the comprehensive anomaly level specifically includes: The anomaly level of the root node data is determined by the comprehensive anomaly level, and the data management type of the root node data is determined based on the anomaly level range.

[0035] In another possible embodiment, the method for determining the data management type of the root node data in step S2 is as follows: The number of data anomalies in the root node data is determined by analyzing the data anomalies in different update processes. It is then determined whether the number of data anomalies in the root node data is greater than the preset number of anomalies. If so, the data management type of the root node data is determined to be the set data type. If not, proceed to the next step. Based on the number of anomalies of different anomaly types in the root node data and the proportion of the number of anomalies in the number of updates, the degree of anomaly of different anomaly types is determined, and it is determined whether the number of anomaly types whose degree of anomaly does not meet the requirements is met. If yes, the data management type of the root node data is determined to be the set data type; otherwise, proceed to the next step. The anomaly types that do not meet the requirements are designated as anomaly types of concern. The degree of concern for the root node data is determined based on the number of anomaly types of concern and the degree of anomaly of different anomaly types of concern. It is then determined whether the degree of concern for the root node data meets the requirements. If yes, the data management type of the root node data is determined to be the set data type. If no, proceed to the next step. The number of anomalies in the root node data and the percentage of the number of anomalies in the number of updates are obtained. The overall anomaly level of the root node data is determined by combining the anomaly level of different anomaly types and the anomaly level of the root node data. The data management type of the root node data is determined by the overall anomaly level.

[0036] It should be noted that determining the number of traceability sector layers for different root node data using the aforementioned data management type specifically includes: The range of traceability sector layers for the root node data is determined by the data management type, and the number of traceability sector layers for the root node data is determined based on the overall anomaly level of the root node data.

[0037] S3 uses the number of traceable sector layers as a constraint and the root node data as a basis to determine the associated data within the associated sector of the root node data and the data tags of the associated data through the data lineage relationship between different data. S4 determines the weight values ​​of different associated data based on the sector level of different root node data, and determines the data verification period of different root node data and associated data in the associated sectors of the root node data based on the changes in the associated data.

[0038] In one possible embodiment, such as Figure 3 As shown, the method for determining the data verification period in step S4 is as follows: The frequency of updates for different related data is determined by the number of changes in different related data and the number of updates within a preset time period. The frequency of updates for different related data is determined by the number of updates for different related data and the number of updates within a preset time period. The frequency of updates for root node data is determined by the number of updates for root node data and the number of updates within a preset time period. The data verification cycle is determined by the data update frequency of different related data, the weight values ​​of different related data, and the data update frequency of the root node data.

[0039] In another possible embodiment, the method for determining the data verification period in step S4 is as follows: S41 determines the data update frequency of the root node data by the number of times the root node data is updated and the number of times it is updated within a preset time. It then determines whether the data update frequency of the root node data is greater than a preset frequency threshold. If so, the data verification period is determined by the data update frequency of the root node data. If not, the process proceeds to the next step. S42 determines the number of updates of different related data and the number of updates within a preset time period based on the different changes of related data, and determines the data update frequency of different related data based on the number of updates of different related data and the number of updates within a preset time period. It then determines whether there is related data whose data update frequency is greater than the preset frequency threshold. If so, proceed to the next step; otherwise, proceed to step S44. S43 determines the screening frequency assessment quantity by the proportion of related data whose data update frequency is greater than the preset frequency threshold and the weight value, and the data update frequency of related data whose data update frequency is greater than the preset frequency threshold. It then determines whether the screening frequency assessment quantity meets the requirements. If yes, proceed to the next step; otherwise, determine the data verification cycle by using the screening frequency assessment quantity. S44 determines the data verification cycle by considering the data update frequency of different related data, the weight values ​​of different related data, the data update frequency of the root node data, and the evaluation of the screening frequency.

[0040] Example 2 On the other hand, the present invention provides a computer system comprising: a memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, characterized in that: when the processor runs the computer program, it executes the above-described data management method based on data lineage tracing.

[0041] Through the above embodiments, this application achieves the following technical effects: The beneficial effects of this invention are as follows: 1. In this invention, the root node data in the data is determined by the usage and update conditions. Taking into account the differences in the reliability of different data due to the differences in the frequency of use and the frequency of update, the data with low reliability and the data with high reliability requirements are filtered out, which also lays the foundation for further improving the reliability of data processing.

[0042] 2. In this invention, the number of traceability sector layers for different root node data is determined by using data management types. This enables the filtering of the number of associated data that need to be traced and processed for root node data in case of anomalies during the data update process. This ensures the reliability of data update management for data with more serious anomalies, while also avoiding the technical problem of slow processing efficiency caused by a large amount of associated data.

[0043] 3. In this invention, the data verification period of different root node data and related data in the related sectors of the root node data is determined by utilizing the changes in related data. This realizes the determination of differentiated data verification period from the perspective of data update. At the same time, by determining the weight value, the difference in the degree of impact of the update of related data on the data due to the difference in the degree of association between different related data is also fully considered.

[0044] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0045] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0046] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A data management method based on tracing data lineage, characterized in that, Specifically, it includes: By examining the extraction and usage records of different data within the enterprise, the usage of different data is determined, and the update status of different data is determined based on the update records of different data. The root node data in the data is then determined based on the usage and update status. The different data refers to data in the company's database; Based on the data anomalies of different root node data in different update processes, the data management type of different root node data is determined, and the number of traceability sector layers of different root node data is determined using the data management type. Using the number of traceable sector layers as a constraint and the root node data as a basis, the associated data within the associated sector of the root node data and the data tags of the associated data are determined through the data lineage relationship between different data. The weight values ​​of different associated data are determined by the sector level of different root node data, and the data verification period of different root node data and associated data in the associated sector of the root node data is determined by combining the changes of the associated data. The method for determining the data management type of the root node data is as follows: The number of data anomalies in the root node data and the anomaly type of each anomaly number are determined by the data anomalies in the root node data during different update processes. The degree of anomaly of each anomaly type is determined by the proportion of the number of anomalies of each anomaly type in different update processes. The overall anomaly level of the root node data is determined based on the anomaly level of different anomaly types, and the data management type of the root node data is determined based on the overall anomaly level. The determination of the number of traceability sector layers for different root node data using the aforementioned data management type specifically includes: The range of traceability sector layers for the root node data is determined by the data management type, and the number of traceability sector layers for the root node data is determined based on the overall anomaly level of the root node data. The extraction and usage records are determined based on the analysis results of the database logs. Specifically, the extraction and usage records of different data in the database are determined based on the analysis results of the database logs. The usage data includes the number of times the data is used, the number of business scenarios in which it is used, and the frequency of use of different business scenarios. The data anomalies include the number of times the root node data is abnormal and the types of anomalies for different numbers of anomalies; The method for determining the data verification period is as follows: The frequency of updates for different related data is determined by the number of changes in different related data and the number of updates within a preset time period. The frequency of updates for different related data is determined by the number of updates for different related data and the number of updates within a preset time period. The frequency of updates for root node data is determined by the number of updates for root node data and the number of updates within a preset time period. The data verification cycle is determined by the data update frequency of different related data, the weight values ​​of different related data, and the data update frequency of the root node data.

2. The data management method based on data lineage tracing as described in claim 1, characterized in that, The update information includes the number of times the data has been updated and the time interval between different update counts.

3. The data management method based on data lineage tracing as described in claim 1, characterized in that, The method for determining the root node data in the data is as follows: The frequency of data usage is determined by the usage data, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. The frequency of data usage is determined by the number of times the data is used, the number of business scenarios in which the data is used, and the frequency of use of different business scenarios. The update frequency of the data is determined by the update status and the average time interval between different update counts. The overall application frequency of the data is determined based on the frequency of updates and usage of the data, and the root node data in the data is determined based on the overall application frequency.

4. The data management method based on data lineage tracing as described in claim 3, characterized in that, Determining the root node data in the data based on the comprehensive application frequency specifically includes: The frequency threshold is determined based on the amount of data within the enterprise and the frequency of data usage. The root node data in the data is determined based on the frequency threshold and the comprehensive application frequency.

5. A computer system, comprising: A memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, characterized in that: when the processor runs the computer program, it executes a data management method based on data lineage tracing as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Data management method and device, equipment and medium

    CN111008192A