Data collaboration management method and device, terminal equipment and storage medium

By identifying and processing changed data, generating and sending data change information, the consistency of data in the flow chain is ensured, the problem of data inconsistency is solved, and production accidents are reduced.

CN117194455BActive Publication Date: 2026-01-13CHINA MERCHANTS BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311208874.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2026-01-13
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

In big data organizations, when data flows between different systems, it is difficult to accurately identify and process data changes, leading to data inconsistency, which can cause business logic errors and production accidents.

Method used

By identifying changed data, generating data change information, and sending it to data users based on data lineage for change response processing, and after accepting the change response information, collaborating to go online with the changed data, the consistency between data producers and users is ensured.

Benefits of technology

It improves data consistency along the flow path, reduces business logic errors and system failures, and lowers the incidence of production accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117194455B_ABST
    Figure CN117194455B_ABST
Patent Text Reader

Abstract

The application discloses a data collaborative management method and device, a terminal equipment and a storage medium, and the method comprises the following steps: in response to identifying first change data, processing the first change data to obtain first data change information; based on the data blood relationship corresponding to the first change data, the first data change information is sent to the corresponding data user, so that the data user generates change response information according to the first data change information and sends the change response information to the data producer; the change response information is accepted, and the first change data is put online in cooperation with the data user based on the accepted change response information, so as to ensure the consistency of data change of the data producer and the data user, improve the consistency of data on the flow link, reduce the business logic error and system failure caused by data inconsistency, and reduce the occurrence of production accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data governance, and in particular to a data collaborative governance method, apparatus, terminal equipment, and storage medium. Background Technology

[0002] Currently, organizations with big data often employ a multi-layered data system architecture, including business systems that produce data, data platform systems that collect data, thematic data systems that analyze data, and application systems that utilize data. Large amounts of data flow between these different systems. When data changes, the change information is difficult to accurately identify and process across the various systems, leading to inconsistencies in the data flow. These inconsistencies can cause business logic errors and system failures, potentially triggering production accidents.

[0003] Therefore, it is necessary to propose a solution to improve the consistency of data in the flow link.

[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this invention is to provide a data collaborative governance method, apparatus, terminal equipment, and storage medium, which aims to improve the consistency of data in the flow link and reduce the occurrence of production accidents.

[0006] To achieve the above objectives, the present invention provides a data collaborative governance method, the data collaborative governance method comprising:

[0007] In response to the identification of the first changed data, the first changed data is processed to obtain the first data change information;

[0008] Based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer;

[0009] Accept the change response information, and based on the accepted change response information, collaborate with the data user to put the first changed data online.

[0010] Optionally, before the step of processing the first changed data in response to identifying the first changed data to obtain the first data change information, the method further includes:

[0011] Obtain system data;

[0012] Identify key data within the system data that is associated with the data user;

[0013] Key data in the system data are tagged to obtain labeled data;

[0014] Identify whether the first change data exists in the identification data.

[0015] Optionally, the step of processing the first changed data to obtain the first data change information includes:

[0016] Based on the first changed data, generate a data change pending task;

[0017] Based on the aforementioned data change pending tasks, design a data change framework;

[0018] The first changed data is compared with a version to obtain the difference data;

[0019] The data change framework is populated based on the difference data to obtain the first data change information, wherein the first data change information includes the data change task and the deadline for the data user to complete the data change task.

[0020] Optionally, the step of accepting the change response information includes:

[0021] Determine whether the change response information is the first change response information corresponding to the first data change information:

[0022] If the change response information is the first change response information corresponding to the first data change information, then the acceptance is deemed successful, and the first change response information is used to coordinate with the data user to launch the first changed data online;

[0023] If the change response information is not the first change response information corresponding to the first data change information, the acceptance is deemed unsuccessful, and the corresponding incomplete change information is recorded in the current business system.

[0024] This invention also provides a data collaborative governance method, which is applied to data users and includes the following steps:

[0025] Receive the first data change information sent by the data producer;

[0026] Based on the first data change information, perform change response processing to generate change response information;

[0027] The change response information is sent to the data producer so that the data producer can accept the change response information and, based on the accepted change response information, collaborate with the data user to launch the first changed data online.

[0028] Optionally, the first data change information includes a data change task and a deadline for the data user to complete the data change task; the change response information includes a first change response information or a second change response information; and the step of processing the change response based on the first data change information to generate change response information includes:

[0029] Identify the scope of impact of the data change task in the first data change information;

[0030] Determine whether the scope of the impact includes the current business system;

[0031] If the scope of impact includes the current business system, then the data change task is executed within the deadline to obtain the second changed data and generate the first change response information;

[0032] If the scope of impact does not include the current business system, and / or the data change task is not executed within the deadline, then the second change response information is generated.

[0033] Optionally, after the step of performing the data change task and obtaining the second changed data, the method further includes:

[0034] Based on the second changed data, generate second data change information;

[0035] Based on the data lineage corresponding to the second changed data, the second data change information is sent to the downstream data user so that the downstream data user can perform change response processing according to the second data change information, generate and return the corresponding change response information.

[0036] Furthermore, to achieve the above objectives, the present invention also provides a data collaborative governance device, the data collaborative governance device comprising:

[0037] The response module is used to respond to the detection of first changed data, process the first changed data, and obtain first data change information;

[0038] The sending module is used to send the first data change information to the corresponding data user based on the data lineage relationship corresponding to the first changed data, so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer;

[0039] The acceptance module is used to accept the change response information and, based on the accepted change response information, to coordinate with the data user to put the first changed data online.

[0040] In addition, to achieve the above objectives, the present invention also provides a terminal device, the terminal device including a memory, a processor, and a data collaborative governance program stored in the memory and executable on the processor, wherein the data collaborative governance program, when executed by the processor, implements the steps of the data collaborative governance method as described above.

[0041] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a data collaborative governance program, which, when executed by a processor, implements the steps of the data collaborative governance method as described above.

[0042] This invention proposes a data collaborative governance method, apparatus, terminal device, and storage medium. The data collaborative governance method is applied to a data producer. Upon identifying first changed data, the producer processes the first changed data to obtain first data change information. Based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user, allowing the data user to perform change response processing, generate change response information, and send it back to the data producer. The change response information is then accepted, and based on the accepted change response information, the producer collaborates with the data user to deploy the first changed data online. After identifying the first changed data, the data producer processes it to obtain first data change information and sends it to the data user, enabling the data user to receive the data change information, perform change response processing, and then return change response information to the data producer. This allows the data producer to collaborate with the data user to deploy the first changed data online based on the accepted change response information. This ensures the consistency of data changes between the data producer and the data user, improving data consistency along the data flow path, reducing business logic errors and system failures caused by data inconsistency, and thus reducing the occurrence of production accidents. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the functional modules of the terminal device to which the data collaborative governance device of this invention belongs;

[0044] Figure 2 This is a flowchart illustrating an exemplary embodiment of the data collaborative governance method of the present invention;

[0045] Figure 3 This is a flowchart illustrating another exemplary embodiment of the data collaborative governance method of the present invention;

[0046] Figure 4 for Figure 2 A schematic diagram illustrating the specific process of processing the first changed data to obtain the first data change information in this embodiment;

[0047] Figure 5 for Figure 2 A schematic diagram illustrating the specific process for accepting the change response information in this embodiment;

[0048] Figure 6 This is a flowchart illustrating another exemplary embodiment of the data collaborative governance method of the present invention;

[0049] Figure 7 for Figure 6 A detailed flowchart of step A20 in the embodiment;

[0050] Figure 8 This is a schematic diagram of the overall process in an embodiment of the present invention.

[0051] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0052] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0053] The main solution of this invention is as follows: In response to the identification of first changed data, the first changed data is processed to obtain first data change information; based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user, so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer; the change response information is accepted, and based on the accepted change response information, the data user is coordinated to go online with the first changed data; after receiving and processing the data change information, the data user sends the data change response information to the data producer.

[0054] Organizations with large amounts of data in their current technology often employ a multi-layered data system architecture, including business systems for producing data, data platform systems for collecting data, thematic data systems for analyzing data, and application systems for utilizing data. Large amounts of data flow between these different systems. When data changes, this change information must be accurately identified and processed along the data flow path; otherwise, inconsistencies and production accidents may occur. Currently, data changes in the data flow path are mainly handled manually, which is difficult, prone to omissions, costly, and prone to accidents, reducing data accuracy and hindering data sharing and value realization. Therefore, finding effective governance methods for data flow change management is crucial.

[0055] This invention provides a solution to improve data consistency across the data flow path, thereby reducing the occurrence of production accidents.

[0056] Specifically, refer to Figure 1 , Figure 1 This is a schematic diagram of the functional modules of the terminal device to which the data collaborative governance device of the present invention belongs. The data collaborative governance device can be an independent device capable of data collaborative governance, and it can be implemented on the terminal device in hardware or software form. The terminal device can be a smart mobile terminal with data processing capabilities, such as a mobile phone or tablet computer, or it can be a fixed terminal device or server with data processing capabilities.

[0057] In this embodiment, the terminal device to which the data collaborative governance device belongs includes at least an output module 110, a processor 120, a memory 130, and a communication module 140.

[0058] The memory 130 stores the operating system and data collaborative management program. The data collaborative management device can store the first changed data, the first data change information, the data lineage corresponding to the first changed data and change response information, the second changed data, the second data change information, the data lineage corresponding to the second changed data and change response information, and the data change report in the memory 130. The output module 110 can be a display screen, etc. The communication module 140 can include a WIFI module, a mobile communication module, and a Bluetooth module, etc., and communicates with external devices or servers through the communication module 140.

[0059] When the data coordination and governance program in memory 130 is executed by the processor, it performs the following steps:

[0060] In response to the identification of the first changed data, the first changed data is processed to obtain the first data change information;

[0061] Based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer;

[0062] Accept the change response information, and based on the accepted change response information, collaborate with the data user to put the first changed data online.

[0063] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0064] Obtain system data;

[0065] Identify key data within the system data that is associated with the data user;

[0066] The key data is tagged to obtain labeled data;

[0067] Identify whether the first change data exists in the identification data.

[0068] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0069] Based on the first changed data, generate a data change pending task;

[0070] Based on the aforementioned data change pending tasks, design a data change framework;

[0071] The first changed data is compared with a version to obtain the difference data;

[0072] The data change framework is populated based on the difference data to obtain the first data change information, wherein the first data change information includes the data change task and the deadline for the data user to complete the data change task.

[0073] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0074] Determine whether the change response information is the first change response information corresponding to the first data change information:

[0075] If the change response information is the first change response information corresponding to the first data change information, then the acceptance is deemed successful, and the first change response information is used to coordinate with the data user to launch the first changed data online;

[0076] If the change response information is not the first change response information corresponding to the first data change information, the acceptance is deemed unsuccessful, and the corresponding incomplete change information is recorded in the current business system.

[0077] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0078] Receive the first data change information sent by the data producer;

[0079] Based on the first data change information, perform change response processing to generate change response information;

[0080] The change response information is sent to the data producer so that the data producer can accept the change response information and, based on the accepted change response information, collaborate with the data user to launch the first changed data online.

[0081] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0082] Identify the scope of impact of the data change task in the first data change information;

[0083] Determine whether the scope of the impact includes the current business system;

[0084] If the scope of impact includes the current business system, then the data change task is executed within the deadline to obtain the second changed data and generate the first change response information;

[0085] If the scope of impact does not include the current business system, and / or the data change task is not executed within the deadline, then the second change response information is generated.

[0086] Furthermore, when the data coordination management program in memory 130 is executed by the processor, it also performs the following steps:

[0087] Based on the second changed data, generate second data change information;

[0088] Based on the data lineage corresponding to the second changed data, the second data change information is sent to the downstream data user so that the downstream data user can perform change response processing according to the second data change information, generate and return the corresponding change response information.

[0089] This embodiment, through the above-described scheme, specifically responds to the identification of first changed data, processes the first changed data to obtain first data change information; based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user, so that the data user can perform change response processing according to the first data change information, generate change response information, and send it to the data producer; the change response information is accepted, and based on the accepted change response information, the data producer collaborates with the data user to go online with the first changed data. After the data producer identifies the first changed data, it processes the first changed data to obtain first data change information and sends it to the data user, so that the data user receives the data change information, performs change response processing, and then returns change response information to the data producer, so that the data producer collaborates with the data user to go online with the first changed data based on the accepted change response information. This ensures the consistency of data changes between the data producer and the data user, that is, it improves the consistency of data in the flow link, reduces business logic errors and system failures caused by data inconsistency, and thus reduces the occurrence of production accidents.

[0090] Based on, but not limited to, the terminal device architecture described above, embodiments of the method of the present invention are proposed.

[0091] The subject executing the method in this embodiment can be a data collaborative governance device or a terminal device, etc. This embodiment takes a data collaborative governance device as an example.

[0092] Reference Figure 2 , Figure 2 This is a flowchart illustrating an exemplary embodiment of the data collaborative governance method of the present invention, as shown below. Figure 2 As shown, the data collaborative governance method includes:

[0093] Step S10: In response to the identification of the first changed data, the first changed data is processed to obtain the first data change information;

[0094] Specifically, organizations with big data often employ a multi-layered data system architecture, including business systems that produce data, data platform systems that collect data, thematic data systems that analyze data, and application systems that utilize data. Large amounts of data flow between these different systems. When data changes, this change information must be accurately identified and processed along the data flow to improve data consistency throughout the process.

[0095] This invention proposes a data collaborative governance method, which first acquires system data from various systems, then tags the acquired system data to identify the first changed data, processes the first changed data, and obtains the corresponding first data change information.

[0096] Optionally, the data producer can generate a data change task based on the first change data, design the data change, and perform a version comparison based on the first change data to obtain the difference data. Then, the data change framework can be filled in based on the difference data to obtain the first data change information.

[0097] Optionally, the first data change information can comprehensively reflect the data changes that have occurred by the data producer, so that the data user can evaluate and respond accordingly.

[0098] Step S20: Based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer.

[0099] Furthermore, after identifying the first changed data and processing it to obtain the first data change information, the first data change information can be sent to the corresponding data user based on the data lineage corresponding to the first changed data.

[0100] Optionally, data lineage refers to a relationship similar to human social kinship formed between data during the process of data generation, processing, flow, and extinction. Based on the data lineage corresponding to the first changed data, the downstream data related to the first changed data can be identified, and then the first changed data can be sent to the data users corresponding to the relevant downstream data so that the data users can perform change response processing based on the first data change information, generate change response information, and send it to the data producer.

[0101] Optionally, after receiving data change information, the data user assesses the scope of the data change based on the upstream change information and the actual data usage, confirms whether to accept the change, and provides change response information to the data producer.

[0102] Optionally, data users can also automatically identify and transmit change information to downstream data users based on the design changes and data lineage.

[0103] Step S30: Accept the change response information, and based on the accepted change response information, collaborate with the data user to bring the first changed data online;

[0104] Furthermore, once the data producer receives the change response information from the data user, it can accept the change response information and, based on the accepted change response information, collaborate with the data user to launch the first changed data online.

[0105] Optionally, during the data acceptance testing phase, the data producer completes the acceptance of data changes based on the information collected and identified by the system. Once it is confirmed that the data changes have been received and processed by the data link, the data changes can be accepted and then collaboratively deployed online. The collaborative deployment of upstream, midstream, and downstream data change parties is managed based on the deployment date to ensure the consistency and accuracy of data changes.

[0106] In this embodiment, in response to the identification of first changed data, the first changed data is processed to obtain first data change information. Based on the data lineage corresponding to the first changed data, the first data change information is sent to the corresponding data user, so that the data user can perform change response processing according to the first data change information, generate change response information, and send it to the data producer. The change response information is accepted, and based on the accepted change response information, the data producer collaborates with the data user to launch the first changed data. After the data producer identifies the first changed data, it processes the first changed data to obtain first data change information and sends it to the data user, so that the data user receives the data change information and performs change response processing, and then returns change response information to the data producer. This allows the data producer to collaborate with the data user to launch the first changed data based on the accepted change response information, ensuring the consistency of data changes between the data producer and the data user, that is, improving the consistency of data in the flow link, reducing business logic errors and system failures caused by data inconsistency, thereby reducing the occurrence of production accidents.

[0107] Reference Figure 3 , Figure 3 This is a flowchart illustrating another exemplary embodiment of the data collaborative governance method of the present invention. Based on the above... Figure 2 In the embodiment shown, prior to step S10, the data collaborative governance method further includes:

[0108] Step S001: Obtain system data;

[0109] Specifically, system data can be obtained from system-level data in a multi-layered data system architecture, including business systems that produce data, data platform systems that collect data, thematic data systems that analyze data, and application systems that utilize data.

[0110] Optionally, methods for obtaining system data include:

[0111] (1) Database Export: If the system uses a database to store data, you can use a database export tool or SQL query to export the data. For example, you can use the mysqldump command in MySQL to export the entire database or a specific table;

[0112] (2) System Log: The system usually generates log files to record the system's operation and data operations. You can view the system log files to obtain relevant data;

[0113] (3) Network packet capture: If the system transmits data over a network, network packet capture tools can be used to capture data packets and parse their contents. For example, tools such as Wireshark can be used for network packet capture;

[0114] (4) System Interface: If the system provides a system interface (API), the corresponding interface can be used to obtain system data. For example, data can be obtained via HTTP requests using a REST API;

[0115] (5) Data export tools: Some systems may provide data export tools, which can easily export data from the system. For example, some ERP systems or CRM systems provide data export functions.

[0116] Step S002: Identify key data in the system data that is associated with the data user;

[0117] Furthermore, after acquiring system data from multiple layers of data system architecture, including business systems that acquire production data, data platform systems that collect data, thematic data systems that analyze data, and application systems that utilize data, key data can be identified from the acquired system data.

[0118] Optionally, key data refers to data that has a wide impact or high importance in various business systems. Among them, key data associated with data users refers to data in the current system that has a wide impact or high importance on downstream business systems.

[0119] Step S003: Tag key data in the system data to obtain identification data;

[0120] Furthermore, after identifying data in the system that has a wide impact or high importance on downstream business systems, key data used downstream can be dynamically tagged to obtain labeled data, which will then enter the data link change process.

[0121] Specifically, data changes and flows through multiple systems, but not all data changes need to be transmitted to other business systems. Therefore, it is very important for data producers to clearly understand which changes will affect data users. Thus, in this invention, the key data obtained is further tagged to obtain the identified data that needs to be monitored and enter the data link change process.

[0122] Step S004: Identify whether the first change data exists in the identification data;

[0123] Furthermore, key data in the system data is tagged to obtain labeled data. This labeled data is then monitored closely to identify any changes. The identified changed data is then used as the first changed data to generate the first change information, so that downstream data users are aware of the data changes and can respond accordingly.

[0124] This embodiment, through the above-described scheme, specifically involves acquiring system data; identifying key data in the system data that is associated with the data user; tagging the key data in the system data to obtain labeled data; identifying whether the labeled data contains the first changed data. By identifying data in the business system data that has a wide impact or high importance on downstream business systems as key data, and then tagging the key data, if the key data changes, it is confirmed that the current system data contains the first changed data, enabling the data producer to clearly understand the changed data and thus accurately publish data change information.

[0125] Reference Figure 4 , Figure 4 for Figure 2 This embodiment illustrates the specific process of processing the first changed data to obtain the first data change information. This embodiment is based on the above... Figure 2 In the embodiment shown, the step of processing the first changed data to obtain the first data change information includes:

[0126] Step S101: Generate a data change pending task based on the first changed data;

[0127] Specifically, the steps to generate to-do tasks include: clarifying the goals and requirements of data changes, understanding the existing data structure and usage, as well as the scope and content of the changes, and generating to-do tasks.

[0128] Step S102: Based on the data change pending tasks, design a data change framework;

[0129] Furthermore, the data manager of the data producer needs to design and develop systems or tools that meet the data change requirements based on the results of the requirements analysis. In this process, the structure, format and business rules of the data need to be considered to ensure that the changed data can meet the requirements and is compatible with the original system or tools.

[0130] Step S103: Perform version comparison on the first changed data to obtain the difference data;

[0131] Furthermore, in this embodiment of the invention, version comparison can be performed automatically to detect data differences and prevent data managers from overlooking changes.

[0132] Step S104: Fill the data change framework according to the difference data to obtain the first data change information, wherein the first data change information includes the data change task and the deadline for the data user to complete the data change task.

[0133] Furthermore, by populating the data change framework based on the discrepancies, the initial data change information can be obtained. After the change design and development are completed, testing and verification are required to ensure that the changed data can be correctly stored, retrieved, and used. Testing can include unit testing, integration testing, and system testing, while verification can include user acceptance testing.

[0134] Optionally, the data change requirement time can also be added to the first data change information, which is set by the data producer's data manager based on the data launch deadline.

[0135] This embodiment, through the above-described scheme, specifically generates a data change task based on the first changed data; designs a data change framework based on the data change task; performs version comparison on the first changed data to obtain difference data; and fills the data change framework with the difference data to obtain the first data change information. The first data change information includes the data change task and the deadline for the data user to complete the data change task. By accurately generating and publishing data change information, data users can reliably receive data change information, quickly respond to data change processing, and collaborate with data producers to complete data changes, thereby improving the consistency of changed data in the data chain.

[0136] Reference Figure 5 , Figure 5 for Figure 2 A schematic diagram illustrating the specific process for accepting the change response information in this embodiment. This embodiment is based on the above. Figure 2 In the illustrated embodiment, the step of accepting the change response information includes:

[0137] Step S301: Determine whether the change response information is the first change response information corresponding to the first data change information.

[0138] Step S302: If the change response information is the first change response information corresponding to the first data change information, then the acceptance is determined to be successful, and the first change response information is used to coordinate with the data user to launch the first changed data.

[0139] Step S303: If the change response information is not the first change response information corresponding to the first data change information, then the acceptance is determined to be unsuccessful, and the corresponding change incomplete information is recorded in the current business system.

[0140] Specifically, during the data acceptance testing phase, the data manager of the data producer can complete the acceptance of data changes based on the information collected and identified by the system. Once it is confirmed that the data change has been received and processed by the data link, the acceptance can be passed. Otherwise, the data acceptance is deemed to have failed, and the corresponding incomplete change information is recorded in the current business system to further determine the reasons and handling methods for the downstream data users not receiving the data changes.

[0141] Optionally, the specific steps for judgment include: (1) Clarifying the content and scope of the change: First, it is necessary to clarify the content and scope of the change and understand the affected systems and business areas in order to accurately determine whether the feedback information is related to the change. (2) Analyzing the content and source of the feedback information: Analyze the content and source of the feedback information to determine whether it is related to the change, and at the same time understand the claims or intentions of the feedback party. (3) Comparing historical data and information: Compare the feedback information with historical data and information to analyze whether there are any abnormalities or irregularities, thereby determining whether the feedback information indicates that feedback has occurred. (4) Comprehensive analysis and judgment: Comprehensively consider various factors, including the content and source of the feedback information, historical data and information, and other relevant information, to conduct comprehensive analysis and judgment in order to draw an accurate conclusion.

[0142] This embodiment, through the above scheme, specifically determines whether the change response information is the first change response information corresponding to the first data change information: if the change response information is the first change response information corresponding to the first data change information, then the acceptance is deemed successful, and the first change response information is used to coordinate with the data user to launch the first changed data; if the change response information is not the first change response information corresponding to the first data change information, then the acceptance is deemed unsuccessful, and the corresponding incomplete change information is recorded in the current business system. Through data acceptance, it is determined whether the data change has been received and processed by the data user, and then the data user receives the changed data and performs a collaborative launch operation to ensure the consistency of the data change.

[0143] Reference Figure 6 , Figure 6 This is a flowchart illustrating another exemplary embodiment of the data collaborative governance method of the present invention. The data collaborative governance method is applied to data users and includes:

[0144] Step A10: Receive the first data change information sent by the data producer;

[0145] Specifically, the data user receives the first data change information sent by the data producer, and can then confirm the authenticity and accuracy of the change information.

[0146] Optionally, after confirming the authenticity and accuracy of the changed information, if it is confirmed that the changed information is forged or incorrect, appropriate processing is required, such as correcting errors or triggering an alarm; if the changed information is confirmed to be authentic, a data change task can be automatically generated based on the received data change content and lineage, and the data manager can respond based on the task.

[0147] Step A20: Perform change response processing based on the first data change information to generate change response information;

[0148] Furthermore, upon receiving the first data change information, the data manager of the data user can assess the scope of the impact of the data change based on the upstream change information and the actual data usage, and confirm whether to accept the change.

[0149] Optionally, during the assessment of data changes, the nature of the data changes can be understood. Data changes may be caused by external or internal factors and may involve changes to relevant business data. Data users can confirm whether to accept the changes based on the actual situation and then generate change response information.

[0150] Step A30: Send the change response information to the data producer so that the data producer can accept the change response information and, based on the accepted change response information, collaborate with the data user to launch the first changed data online.

[0151] Furthermore, after receiving data change information, the data user assesses the scope of the data change's impact based on the upstream change information and actual data usage, confirms whether to accept the change, and provides feedback on the change response to the data producer.

[0152] Optionally, data users can also automatically identify and transmit change information to downstream data users based on the design changes and data lineage.

[0153] Optionally, after receiving the change response information from the data user, the data producer can accept the change response information and, based on the accepted change response information, collaborate with the data user to launch the first changed data online.

[0154] Optionally, during the data acceptance testing phase, the data producer completes the acceptance of data changes based on the information collected and identified by the system. Once it is confirmed that the data changes have been received and processed by the data link, the data changes can be accepted and then collaboratively deployed online. The collaborative deployment of upstream, midstream, and downstream data change parties is managed based on the deployment date to ensure the consistency and accuracy of data changes.

[0155] In this embodiment, the system receives first data change information from a data producer; performs change response processing based on the first data change information to generate change response information; and sends the change response information to the data producer for acceptance. Based on the accepted change response information, the data producer collaborates with the data user to deploy the first changed data. After recognizing the first changed data, the data producer processes it to obtain first data change information, which is then sent to the data user. The data user receives the data change information, performs change response processing, and returns change response information to the data producer. This ensures consistency between the data producer and the data user in data changes, improving data consistency along the data flow path, reducing business logic errors and system failures caused by data inconsistency, and thus reducing the occurrence of production accidents.

[0156] Reference Figure 7 , Figure 7 for Figure 6 A detailed flowchart of step A20 in the embodiment is shown below. Figure 7 As shown, in this embodiment, step A20 includes:

[0157] Step A201: Identify the scope of impact of the data change task in the first data change information;

[0158] Step A202: Determine whether the scope of influence includes the current business system;

[0159] Step A203: If the scope of impact includes the current business system, then execute the data change task within the deadline to obtain the second change data and generate the first change response information;

[0160] Step A204: If the scope of impact does not include the current business system, and / or the data change task is not executed within the deadline, then the second change response information is generated.

[0161] Optionally, the first data change information includes a data change task and a deadline for the data user to complete the data change task, and the change response information includes a first change response information or a second change response information.

[0162] Specifically, the data user identifies the scope of impact of the data change task in the first data change information, and determines whether the scope of impact includes the current business system through business information and data lifecycle information;

[0163] Optionally, if the scope of impact includes the current business system, the data change task is executed within the deadline. The data manager designs changes to the affected data, including changes to data structure, data content, data migration, and other change scenarios, records the design results, and further obtains the second changed data.

[0164] Optionally, if the scope of impact does not include the current business system, and / or the data change task is not executed within the deadline, then the second change response information is generated. Specifically, the second change response information includes the operation record of not completing the data change according to the data change information, and the reason for not completing the data change according to the data change information, so that the data producer can record and update the data lineage.

[0165] Optionally, the lifecycle of data lineage includes: (1) Defining data lineage: First, it is necessary to clearly define the data lineage, including which data are parent data, which data are child data, and the relationships between data. (2) Establishing a data lineage model: Based on the defined data lineage, establish a data lineage model to visualize the processes of data source, transmission, fusion, processing, and disappearance. (3) Identifying changes in data lineage: When changes occur in the processes of data source, transmission, fusion, processing, and disappearance, it is necessary to identify the changes in data lineage, including the reasons for the changes, the time, and the scope of impact. (4) Updating data lineage records: Based on the identified changes in data lineage, update the data lineage records in a timely manner, including modifying the relationships between data, updating the data source and destination, etc. (5) Verifying data lineage updates: After updating the data lineage, it is necessary to verify it to ensure the accuracy and consistency of the data. This may include testing the integrity, accuracy, and security of the data. (6) Publish a data lineage update report: After verification, a data lineage update report can be published so that relevant personnel can understand and respond to changes in data lineage in a timely manner.

[0166] Optionally, after the step of performing the data change task and obtaining the second changed data, the method further includes:

[0167] Based on the second changed data, generate second data change information;

[0168] Based on the data lineage corresponding to the second changed data, the second data change information is sent to the downstream data user so that the downstream data user can perform change response processing according to the second data change information, generate and return the corresponding change response information.

[0169] Specifically, after receiving data changes, the data user responds to the changes for the data they are responsible for and publishes the data changes to downstream data users based on data lineage, thereby achieving synchronized changes to related data in the data chain. During this process, data that needs to be transmitted can be processed and transformed according to data lineage to facilitate transmission and reception.

[0170] This embodiment, through the above-described scheme, specifically identifies the scope of influence of the data change task in the first data change information; determines whether the scope of influence includes the current business system; if the scope of influence includes the current business system, the data change task is executed within the deadline to obtain the second changed data, and the first change response information is generated; if the scope of influence does not include the current business system, and / or the data change task is not executed within the deadline, the second change response information is generated. By parsing the change information, the scope of influence of the data change task and the deadline for completing the data change task are obtained, enabling data users to quickly respond to data change processing based on the impact of the first data change information on themselves, and thus collaborate with data producers to complete data changes.

[0171] Furthermore, embodiments of the present invention also propose an apparatus, wherein the data collaborative governance apparatus includes:

[0172] The response module is used to respond to the detection of first changed data, process the first changed data, and obtain first data change information;

[0173] The sending module sends the first data change information to the corresponding data user based on the data lineage relationship corresponding to the first changed data, so that the data user can perform change response processing according to the first data change information, generate change response information and send it to the data producer;

[0174] The acceptance module accepts the change response information and, based on the accepted change response information, collaborates with the data user to put the first changed data online.

[0175] Reference Figure 8 , Figure 8 This is a schematic diagram of the overall process in an embodiment of the present invention, specifically including:

[0176] 1) Data Change Management: Manage changes to upstream business systems, including identifying data changes, managing data change processes, managing data change collaboration, accepting data changes, and coordinating the deployment of data changes.

[0177] Identify data changes: Dynamically tag downstream business systems to identify which systems use data and which do not, helping upstream business systems accurately identify whether they need to enter the data link change process.

[0178] Manage data change process: Provide data change process guidelines for data entering the data link change process, guiding data owners to complete the change processing by referring to the guidelines.

[0179] Managing Data Change Collaboration: Establish a data change impact analysis view, collect all published data change information, downstream response information to the changes, and key data information, so that data managers can keep track of the impact of data link changes and the progress of processing in real time, and resolve anomalies in a timely manner.

[0180] Acceptance of data changes: During the data acceptance testing phase, the data manager is supported in completing the acceptance of data changes based on the information collected and identified by the system. Once it is confirmed that the data changes have been received and processed by the data link, the acceptance is passed.

[0181] Collaborative deployment of data changes: Based on the deployment date, manage the collaborative deployment of data changes by upstream, midstream, and downstream parties to ensure the consistency and accuracy of data changes.

[0182] 2) Data Change Handling: Establish a data change handling tool for business systems, integrate it with data design, generate change element information, and support data managers to publish change information online to midstream data users.

[0183] Generate data change tasks: Based on the results of identifying data changes, generate data change tasks, and the data manager performs data change processing based on the tasks.

[0184] Data Tagging: Tagging data used downstream identifies which data is used by midstream and downstream entities, helping data managers accurately identify which data requires change processing and to publish changes to midstream and downstream entities. Tagging data with a wide impact or high importance supports data managers in strengthening the change processing of key data.

[0185] Design data changes: Establish a data change design tool to automatically generate change information based on the data design results.

[0186] Identify data discrepancies: Automatically compare different versions to identify data discrepancies and prevent data owners from overlooking changes.

[0187] Publishing data changes: Supports data owners to publish data changes to data users online, automatically identifies data users who use the changed data based on data lineage, and accurately conveys data changes.

[0188] 3) Data Change Response: The mid-to-downstream data managers receive data changes, process the changes for the data under their responsibility, and publish the data changes to the next downstream data users.

[0189] Generate data change pending tasks: Based on the received data change content and lineage, data change pending tasks are automatically generated, and the data owner starts processing based on the pending tasks.

[0190] Assessing Data Changes: Based on upstream change information and actual data usage, the data manager assesses the scope of impact of the changes to the data under their responsibility, confirms whether to accept the changes, and reports the assessment results back to the upstream change initiator.

[0191] Design Data Changes: The data manager designs changes to the affected data, including changes to data structure, data content, data migration, and other change scenarios, and records the design results.

[0192] Release data changes: Based on the design results of data changes and data lineage, automatically identify and transmit change information to downstream data users.

[0193] Collaborative deployment of data changes: The data manager completes the data change processing based on the upstream change requirements and collaborates with the upstream to complete the deployment of the changes.

[0194] 4) Data change tracking and management: Based on the processing and response status of pending tasks, change processing reports are automatically published, and managers can urge data responsible persons to complete the changes in a timely manner.

[0195] Generate data change reports: Collect information on change processing and response status along the data chain, generate data change reports, and visualize information such as the data person in charge, the number of tasks to be processed, and the importance level of data in the tasks.

[0196] Data producer tracking: Issue data change reports to data producer managers to encourage data managers to analyze, design, and release data changes in a timely manner.

[0197] Data User Tracking: Issue data change reports to data user managers to encourage data owners to promptly assess, respond to, and publish data changes.

[0198] 5) Data Change Type Definition: In this embodiment of the invention, data changes are classified into database-level, table-level, data content-level, file-level, and real-time data-level change types according to different scenarios, specifically including:

[0199] Database level: including database migration, adding / removing database shards, database shutdown, and database attribute changes;

[0200] Data table level: includes table structure, data download conditions, table migration, and table offline changes;

[0201] Data content level: includes code value, data format, data rules, and data modification changes;

[0202] File level: Includes file structure, file drop method, file attributes, warehouse entry information, file content, file directory migration, and file offline changes;

[0203] Real-time data level: including Topic structure, Topic migration, re-consumption of historical Topic data, Topic data content, Topic partition expansion, and Topic offline changes.

[0204] 6) Effect verification: The data collaborative governance method provided by this invention can be applied to various data link change scenarios.

[0205] This embodiment, through the aforementioned scheme, specifically employs a method for data change governance using change management, change processing, change response, and change tracking. It identifies whether data is used by other systems based on data lineage, uses pending tasks to connect all parties in the data link to carry out data changes, automatically identifies the scope of impact of changes on data users based on lineage, automatically generates data change content using development and design tools, and defines data change types. This achieves standardized end-to-end data change management, providing a universally applicable strategy to ensure the accuracy of data in the data link. For data producers, it clearly understands which changes will affect data users, accurately publishes change information, and verifies and accepts the responses of data users. For data users, it reliably receives change information from data producers, quickly responds to data change processing, and collaborates with data producers to complete data changes.

[0206] The principle and implementation process of data collaborative governance in this embodiment are explained in the above embodiments and will not be repeated here.

[0207] Furthermore, this embodiment of the invention also proposes a terminal device, which includes a memory, a processor, and a data collaborative governance program stored in the memory and executable on the processor. When the data collaborative governance program is executed by the processor, it implements the steps of the data collaborative governance method described above.

[0208] Since this data collaborative governance program employs all the technical solutions of all the aforementioned embodiments when executed by the processor, it possesses at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be elaborated upon here.

[0209] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a data collaborative governance program, which, when executed by a processor, implements the steps of the data collaborative governance method described above.

[0210] Since this data collaborative governance program employs all the technical solutions of all the aforementioned embodiments when executed by the processor, it possesses at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be elaborated upon here.

[0211] Compared to existing technologies, the data collaborative governance method, apparatus, terminal device, and storage medium proposed in this invention, in response to the identification of first changed data, process the first changed data to obtain first data change information; based on the data lineage corresponding to the first changed data, send the first data change information to the corresponding data user, so that the data user can perform change response processing according to the first data change information, generate change response information, and send it to the data producer; accept the change response information, and based on the accepted change response information, collaborate with the data user to launch the first changed data online. After the data producer identifies the first changed data, processes the first changed data to obtain first data change information and sends it to the data user, so that the data user receives the data change information and performs change response processing, and then returns change response information to the data producer, so that the data producer can collaborate with the data user to launch the first changed data online based on the accepted change response information. This ensures the consistency of data changes between the data producer and the data user, that is, improves the consistency of data in the flow link, reduces business logic errors and system failures caused by data inconsistency, and thus reduces the occurrence of production accidents.

[0212] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0213] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0214] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.

[0215] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A data collaborative governance method, characterized in that, The data collaborative governance method is applied to a data production party, and comprises the following steps: In response to identifying first change data, processing the first change data to obtain first data change information; Based on the data blood relationship corresponding to the first change data, the first data change information is sent to the corresponding data user, so that the data user performs change response processing according to the first data change information, generates change response information and sends it to the data production party; Accepting the change response information, and based on the change response information that passes the acceptance, the data user is cooperated to put online the first change data; The step of accepting the change response information comprises: Determine whether the change response information is the first change response information corresponding to the first data change information: If the change response information is the first change response information corresponding to the first data change information, it is determined that the acceptance passes, and the first change response information is used to cooperate the data user to put online the first change data; If the change response information is not the first change response information corresponding to the first data change information, it is determined that the acceptance does not pass, and the corresponding change incomplete information is recorded in the current business system.

2. The data co-governance method of claim 1, wherein, The step of responding to the identification of the first change data, processing the first change data to obtain the first data change information, further comprises: Obtain system data; Identify the key data associated with the data user in the system data; Tag processing is performed on the key data in the system data to obtain identification data; Identify whether the first change data exists in the identification data.

3. The data co-governance method of claim 1, wherein, The step of processing the first change data to obtain the first data change information comprises: According to the first change data, a data change to-do task is generated; Based on the data change to-do task, a data change framework is designed; Version comparison is performed on the first change data to obtain difference data; According to the difference data, the data change framework is filled to obtain the first data change information, wherein the first data change information comprises a data change task and a deadline for the data user to complete the data change task.

4. A data collaborative governance method, characterized in that, The data collaborative governance method is applied to a data production party, and comprises the following steps: Receive the first data change information sent by the data production party, wherein the first data change information is obtained by the data production party in response to identifying first change data, processing the first change data to obtain first data change information, and based on the data blood relationship corresponding to the first change data, the first data change information is sent to the corresponding data user; According to the first data change information, change response processing is performed to generate change response information; The change response information is sent to the data production party, so that the data production party accepts the change response information, and based on the change response information that passes the acceptance, the data user is cooperated to put online the first change data; The step of accepting the change response information by the data production party comprises: judging, by the data producer, whether the change response information is first change response information corresponding to the first data change information: if the change response information is the first change response information corresponding to the first data change information, determining, by the data producer, that the acceptance passes, and using the first change response information to coordinate the data consumer to online first changed data; if the change response information is not the first change response information corresponding to the first data change information, determining, by the data producer, that the acceptance fails, and recording corresponding change uncompleted information in the current business system.

5. The data co-governance method of claim 4, wherein, The first data change information includes a data change task and a deadline for completing the data change task by the data consumer, and the change response information includes first change response information or second change response information. The step of generating change response information according to the first data change information includes: identifying an influence range of the data change task in the first data change information; judging whether the influence range includes the current business system; if the influence range includes the current business system, executing the data change task within the deadline to obtain second changed data, and generating the first change response information; if the influence range does not include the current business system, and / or the data change task is not executed within the deadline, generating the second change response information.

6. The data co-governance method of claim 5, wherein, The step of executing the data change task to obtain second changed data further includes: generating second data change information according to the second changed data; based on a data blood relationship corresponding to the second changed data, sending the second data change information to a downstream data consumer, so that the downstream data consumer generates and returns corresponding change response information according to the second data change information.

7. A data co-governance apparatus, characterized by, The data collaborative governance device is applied to a data producer, and includes: a response module configured to, in response to identifying first changed data, process the first changed data to obtain first data change information; a sending module configured to, based on a data blood relationship corresponding to the first changed data, send the first data change information to a corresponding data consumer, so that the data consumer generates change response information according to the first data change information and sends the change response information to the data producer; an acceptance module configured to accept the change response information, and based on the acceptance passing, coordinate the data consumer to online the first changed data; The acceptance module is further configured to judge whether the change response information is first change response information corresponding to the first data change information: if the change response information is the first change response information corresponding to the first data change information, determining that the acceptance passes, and using the first change response information to coordinate the data consumer to online first changed data; If the change response information is not the first change response information corresponding to the first data change information, it is determined that the acceptance is not passed, and corresponding change uncompleted information is recorded in the current business system.

8. A terminal device, comprising: The terminal device comprises a memory, a processor, and a data collaboration governance program stored on the memory and executable on the processor, and the data collaboration governance program, when executed by the processor, implements the steps of the data collaboration governance method according to any one of claims 1-6.

9. A storage medium, characterized by The storage medium stores a data collaboration governance program, and the data collaboration governance program, when executed by the processor, implements the steps of the data collaboration governance method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Systems and methods for determining relationships among data elements

    US20180129699A1

  • Cognitive data outlier pre-check based on data lineage

    US20220350789A1