Data processing method and device, electronic equipment, storage medium and program product

Through the blood relationship analysis and change notification mechanism between the data consumer end and the data warehouse, the problem of timely notification of the consumer end when the data warehouse data is changed is solved, data accuracy and business stability are ensured, and timely monitoring and notification of the full-link data quality is achieved.

CN120523819APending Publication Date: 2025-08-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510646056.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When data changes in the data warehouse, the data consumer side cannot obtain notifications in a timely manner, resulting in data errors and unstable business services, and it is difficult for the existing technology to achieve full-link data quality assurance and information interoperability.

Method used

Through the data consumer end, the consumption data information is sent to the data warehouse, and blood analysis is carried out, the target data object of the affected data is determined, and a declaration request is initiated to the object to which it belongs. If it is passed, the change information is sent to the data consumer end when there is data change in the target data object, and a data change notification mechanism is established.

Benefits of technology

The data consumer side has achieved timely acquisition of data change information from upstream data warehouses, quickly responded to data impacts, ensured data accuracy and stable operation of business services, and improved the full-link data quality assurance capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523819A_ABST
    Figure CN120523819A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of data processing.The method comprises the steps that consumption data information sent by a data consumption end is obtained, and the consumption data information comprises affected data affected by data change of a data warehouse; performing consanguinity analysis on the affected data to obtain a consanguinity link of the affected data in the data warehouse; determining a target data object on which the influenced data depends based on the blood relationship link; initiating a declaration request to an object to which the target data object belongs, wherein the declaration request is used for declaring a data influence relationship between the target data object and the data consumption end; and if the declaration request passes, sending data change information to the data consumption end under the condition that the target data object has data change. According to the data change notification method and device, the problem of data change notification of the affected data consumption end when data change exists in the data warehouse can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art

[0002] In the current business environment, data consumers in business systems rely on data from data warehouses to provide corresponding business services. If data in the data warehouse is changed without timely notification to the data consumers, data errors may occur on these consumers. Therefore, how to promptly notify affected data consumers of data changes in the data warehouse has become a pressing issue. Summary of the Invention

[0003] In view of this, the present disclosure provides a data processing method, apparatus, electronic device, storage medium and program product to solve the problem of notifying affected data consumers of data changes when there are data changes in a data warehouse.

[0004] In a first aspect, the present disclosure provides a data processing method, applied to a data warehouse, comprising:

[0005] Obtaining consumption data information sent by a data consumer, wherein the consumption data information includes affected data affected by the data change in the data warehouse;

[0006] Performing lineage analysis on the affected data to obtain a lineage link of the affected data in the data warehouse;

[0007] Determining the target data object on which the affected data depends based on the lineage link;

[0008] Initiate a declaration request to the object to which the target data object belongs, wherein the declaration request is used to declare the data impact relationship between the target data object and the data consumer;

[0009] If the declaration request is approved, then when there is data change in the target data object, the data change information will be sent to the data consumer.

[0010] In a second aspect, the present disclosure provides another data processing method, which is applied to a data consumption end, and the method includes:

[0011] Obtaining consumption data information, wherein the consumption data information includes affected data affected by the data change in the data warehouse;

[0012] Sending the consumption data information to the data warehouse; wherein the data warehouse is used to perform a lineage analysis on the affected data, obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, the declaration request being used to declare the data impact relationship between the target data object and the data consumer end; the data warehouse is also used to, if the declaration request is passed, send data change information to the data consumer end in the event that there is a data change in the target data object.

[0013] In a third aspect, the present disclosure provides a data processing device, applied to a data warehouse, comprising:

[0014] An information acquisition module, configured to acquire consumption data information sent by a data consumer, wherein the consumption data information includes affected data affected by data changes in the data warehouse;

[0015] A data analysis module, configured to perform lineage analysis on the affected data to obtain lineage links of the affected data in the data warehouse;

[0016] A dependency determination module, configured to determine, based on the lineage link, a target data object on which the affected data depends;

[0017] A data reporting module, configured to initiate a reporting request to an object to which the target data object belongs, wherein the reporting request is used to report the data impact relationship between the target data object and the data consumption end;

[0018] The change notification module is used to send data change information to the data consumer if the declaration request is passed and there is data change in the target data object.

[0019] In a fourth aspect, the present disclosure provides another data processing device, applied to a data consumption end, the device comprising:

[0020] A data identification module is used to obtain consumption data information, wherein the consumption data information includes affected data affected by data changes in the data warehouse;

[0021] A data sending module is used to send the consumption data information to the data warehouse; wherein, the data warehouse is used to perform a lineage analysis on the affected data to obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, the declaration request being used to declare the data impact relationship between the target data object and the data consumer end; the data warehouse is also used to, if the declaration request is passed, send data change information to the data consumer end in the event that there is a data change in the target data object.

[0022] In a fifth aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data processing method of any of the above-mentioned embodiments by executing the computer instructions.

[0023] In a sixth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data processing method of any of the above-mentioned embodiments.

[0024] In a seventh aspect, the present disclosure provides a computer program product, comprising computer instructions, where the computer instructions are used to enable a computer to execute the data processing method of any of the above embodiments.

[0025] In the data processing method provided by the embodiment of the present disclosure, the data consumer sends the affected data affected by the data change in the data warehouse to the data warehouse. The data warehouse performs a lineage analysis on the affected data, and traces back the data objects in the data warehouse that have dependencies on the affected data to obtain a lineage link. Then, based on the lineage link, the target data object on which the affected data depends is determined, and a declaration request is initiated to the object to which the target data object belongs, so that the object to which it belongs is informed of the impact of its own business on the downstream data consumer. Furthermore, after the declaration request is passed, a data change notification mechanism is established between the target data object and the data consumer. In the case that there is a data change in the target data object, the affected data consumer is quickly determined, and the data change information is sent to the data consumer in a timely manner, so that the data consumer is informed of the data change information of the upstream data warehouse in a timely manner, and can quickly respond to the data impact that may be caused by the data change in the data warehouse, thereby ensuring the accuracy of the data and the stable operation of the business service.

[0026] The beneficial effects of the data processing device, electronic device, storage medium, and program product correspond to the beneficial effects of the data processing method and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 is an interactive schematic diagram of a related technology according to an embodiment of the present disclosure;

[0029] Figure 2 is a schematic diagram of an application scenario according to an embodiment of the present invention;

[0030] Figure 3 is a flow chart of a data processing method according to an embodiment of the present disclosure;

[0031] Figure 4 is a schematic diagram of a target data link according to an embodiment of the present disclosure;

[0032] Figure 5 is a schematic diagram of a first page according to an embodiment of the present disclosure;

[0033] Figure 6 is a flow chart of a data change notification mechanism according to an embodiment of the present disclosure;

[0034] Figure 7 is a flowchart of another data processing method according to an embodiment of the present disclosure;

[0035] Figure 8 is a schematic diagram of a process for identifying affected data according to an embodiment of the present disclosure;

[0036] Figure 9 This is a platform architecture diagram of a data warehouse and a data consumer according to an embodiment of the present disclosure;

[0037] Figure 10 is another platform architecture diagram of a data warehouse and a data consumer according to an embodiment of the present disclosure;

[0038] Figure 11 is a structural block diagram of a data processing device according to an embodiment of the present disclosure;

[0039] Figure 12 is a structural block diagram of another data processing device according to an embodiment of the present disclosure;

[0040] Figure 13 It is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0042] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0043] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0044] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0045] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0046] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0047] In the current business operating environment, the full-link data between data warehouses and data consumption ends has not yet been effectively connected, posing a significant challenge to the stable development and risk control of the business.

[0048] For data consumers, any changes to the data warehouse during business processes can have a direct and significant impact on downstream data consumers. Currently, data consumers face severe information asymmetry and a lack of quality assurance when relying on data from upstream data warehouses for changes and online operations. For example, during a marketing campaign, if the data warehouse changes the reward distribution logic without notifying the data consumer, it is very likely to result in over-rewarding users. This not only results in unnecessary capital loss for the company but can also trigger a series of chain reactions. Similarly, during the fund settlement process, if data errors or quality issues in the data warehouse lead to fund settlement errors, it will directly impact the company's financial situation and partnerships.

[0049] Therefore, data consumers urgently expect data warehouses to take on the responsibility of ensuring data quality throughout the entire chain. Regarding data accuracy, it is crucial to ensure that data flowing through the entire chain is error-free. Data latency is also a key issue. Data consumer businesses require high real-time performance, and any data delays can lead to delayed business decisions and missed market opportunities. For example, in real-time transaction monitoring scenarios, data delays can prevent companies from promptly detecting unusual transactions, thereby increasing financial risk. Therefore, it is necessary to establish an efficient notification mechanism to ensure that data consumers are immediately informed of data changes in the data warehouse, allowing them to quickly adjust their business strategies based on pre-set contingency plans and minimize the negative impact of data changes.

[0050] Data warehouses, on the other hand, face a major challenge: insufficient information about data usage scenarios at downstream data consumers. Because data warehouses are unaware of these scenarios, they often bear primary responsibility for online issues. For example, when a data consumer experiences a problem due to data issues, the data warehouse lacks clarity about the specific data usage and business logic at the consumer end, making it difficult to pinpoint the root cause and providing effective, targeted solutions. This often leaves data warehouses in a passive position when addressing issues, requiring significant time and effort to troubleshoot and making substantial progress difficult.

[0051] Therefore, data warehouses urgently need to clearly understand the flow and usage scenarios of data in order to anticipate potential risks in data usage and take effective preventive measures. Furthermore, in terms of data change management, data warehouses need to establish a comprehensive data change notification mechanism. When changes occur to data tasks, downstream data consumers are promptly notified so that they can prepare accordingly. Furthermore, monitoring efforts should be strengthened. By optimizing monitoring indicators and technical means, data quality and link status can be monitored in real time to ensure that problems are discovered and resolved as soon as they occur. This effectively enhances the data warehouse's support for data consumers and reduces overall business risk.

[0052] In summary, if Figure 1 As shown in the figure, the data in the data warehouse passes through the operational data storage (ODS) layer, the data detail (DWD) layer, and the application layer (APP) layer in sequence. The data consumer obtains the required data from the data warehouse. However, the data consumer has no data risk perception capability and can only perceive data quality problems in the upstream data warehouse when a data incident occurs, which is relatively passive. The data warehouse mainly relies on offline word of mouth and personal experience to sort out data and label tasks to determine which tasks in the data warehouse have data changes that will affect the data consumer. However, it is difficult to effectively identify the specific data links affected in the data consumer, making it difficult to achieve full-link data quality assurance and data monitoring.

[0053] In other words, data consumers operate in isolation from the data warehouse, and data change information between the two ends is not communicated. This makes it difficult for the data warehouse's data objects to understand downstream data usage scenarios. When data changes occur in the data warehouse, it's difficult to accurately assess the impact, leading to a wider impact, escalating data incidents, and even online failures.

[0054] As an optional application scenario of the embodiment of the present disclosure, Figure 2 As shown, Figure 2 This is an optional schematic diagram of an application scenario according to an embodiment of the present disclosure. The entire data processing scenario includes a data warehouse 1 and a data consumer 2. Data consumer 2 consumes data from data warehouse 1 to provide business services. This embodiment of the present disclosure establishes a data change notification mechanism between data warehouse 1 and data consumer 2, ensuring that any data changes in data warehouse 1 are promptly notified to the affected data consumer 2.

[0055] In view of this, according to an embodiment of the present disclosure, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0056] In this embodiment, a data processing method is provided, which can be used in a data warehouse. Figure 3 is a flow chart of a data processing method according to an embodiment of the present disclosure, such as Figure 3 As shown, the process includes the following steps:

[0057] Step S301: Obtain consumption data information sent by a data consumer, where the consumption data information includes affected data affected by data changes in a data warehouse.

[0058] In actual applications, data consumers can leverage the data processing and pattern recognition capabilities of large models and other recognition models to analyze and label the call interfaces and massive amounts of business data in the data consumer, identifying the call interfaces and business data that are significantly affected by data changes in the data warehouse as affected data. Alternatively, a manual approach can be used to manually identify the call interfaces and business data in the data consumer that are significantly affected by data changes in the data warehouse based on business experience, as affected data. Alternatively, the affected data identified by the recognition model can be manually verified and supplemented to ensure that the affected data is labeled comprehensively and accurately, providing a reliable basis for subsequent data change notifications.

[0059] Furthermore, the data consumer uses the field lineage tracing capability to identify the data affected by the data changes in the data warehouse. Then, using the recognition model and / or manual recognition method, the data consumer identifies the degree to which the data traced by the field lineage is affected by the data changes in the data warehouse, and the data with a higher degree of influence is regarded as the affected data.

[0060] It should be noted that the affected data includes at least one of the business data affected by the data change in the data warehouse and the interface information for calling the interface. The business data includes at least one of the table and the field, which can be adjusted according to actual conditions.

[0061] Step S302: perform lineage analysis on the affected data to obtain the lineage links of the affected data in the data warehouse.

[0062] Specifically, a lineage analysis is performed on the affected data to determine the upstream data objects of the affected data in the data warehouse, such as upstream data tables, upstream tasks, etc., to obtain the lineage link.

[0063] Step S303: Determine the target data object on which the affected data depends based on the lineage link.

[0064] Specifically, the previous data object of the affected data is determined based on the lineage link, and the previous data object is used as the target data object.

[0065] Optionally, the target data object may be an upstream data table of the affected data, or an upstream task of the affected data, etc., which is not limited here.

[0066] Step S304: Initiate a declaration request to the object to which the target data object belongs. The declaration request is used to declare the data impact relationship between the target data object and the data consumer.

[0067] Specifically, based on the object information of the target data object in the business process of the data warehouse, the object to which the target data object belongs is determined, and a declaration request is initiated to the object in the data governance platform of the data warehouse, thereby building an information interaction bridge between the upstream object and the downstream data consumer. This allows the object to which the target data object belongs to clearly understand the extent of the impact of its own business on the downstream data consumer, and encourages the upstream and downstream of the affected data to work together to prevent data risks caused by data changes, thereby ensuring stable business operation throughout the entire chain.

[0068] Step S305: If the declaration request is approved, then if there is data change in the target data object, the data change information is sent to the data consumer.

[0069] It can be understood that if the declaration request is approved, it means that the object to which the target data object belongs agrees to ensure the data quality of the affected data and implement task-level labeling for the target data object. Thus, based on the declaration record and the label information of the field lineage, data changes can be controllable, accurately notified, and strongly protected in the R&D process. Hierarchical and layered monitoring configuration and governance are performed for each data node of the affected data. Based on the declared field-level lineage link, the object to which the target data object belongs can view the affected downstream data consumer end in the event of data changes and quickly conduct a high-risk impact assessment. Furthermore, the data change information is promptly notified to the affected data consumer end. Therefore, business scenarios in which the data consumer end may be affected by data changes in the data warehouse can all use the data processing method provided in this embodiment to communicate data changes between the data consumer end and the data warehouse, ensuring that there are no major data accidents in the entire link.

[0070] The data processing method provided in this embodiment is that the data consumer sends the affected data affected by the data change in the data warehouse to the data warehouse. The data warehouse performs a lineage analysis on the affected data and traces back to the data objects in the data warehouse that have a dependency relationship with the affected data to obtain a lineage link. Then, based on the lineage link, the target data object on which the affected data depends is determined, and a declaration request is initiated to the object to which the target data object belongs, so that the object to which it belongs is informed of the impact of its own business on the downstream data consumer. Furthermore, after the declaration request is passed, a data change notification mechanism is established between the target data object and the data consumer. In the case that there is a data change in the target data object, the affected data consumer is quickly determined, and the data change information is sent to the data consumer in a timely manner, so that the data consumer is informed of the data change information of the upstream data warehouse in a timely manner, and can quickly respond to the data impact that may be caused by the data change in the data warehouse, thereby ensuring the accuracy of the data and the stable operation of the business service.

[0071] In some optional implementations, performing lineage analysis on the affected data in step S302 to obtain the lineage links of the affected data in the data warehouse includes:

[0072] Step a1: Acquire a first data link of a data consumption end.

[0073] Specifically, through the field-level lineage tracing capability, the data at the data consumer end is tracked to obtain the first data link at the data consumer end.

[0074] In actual applications, the data warehouse can use field-level lineage tracing capabilities to trace the data trace link of the data consumer to obtain the first data link. Alternatively, the data consumer can perform data tracing on its own to obtain the first data link and send it to the data warehouse.

[0075] Specifically, the data warehouse uses field-level lineage tracing capabilities to conduct in-depth cleaning and analysis of the calling relationships between data in the data warehouse to obtain the second data link of the data warehouse.

[0076] Step a2, determining a first link node in the first data link that accesses the data warehouse;

[0077] Specifically, see Figure 4The business service on the upstream server side (such as business service 1) writes data into the data warehouse. After multiple layers of processing in the ODS layer, DWD layer, and APP layer of the data warehouse, multiple data objects are obtained, such as data objects dataobject11 and data object12. The data consumer consumes the data of the data objects in the data warehouse to support one or more business services, such as business service 3 and business service 4. The data consumer mainly consumes data from the data warehouse in the following ways:

[0078] The first is that the data consumer directly queries the data tables in the data warehouse: the data consumer submits a Structured Query Language (SQL) query statement to the distributed SQL query engine to query the data in the data warehouse and consume the query results.

[0079] The second method is to import online data tables at the data consumer end: synchronize the data in the Hive data table in the data warehouse to the relational database management system (MySQL) at the data consumer end. The data consumer end consumes the data in the data warehouse by reading the data in the MySQL table.

[0080] The third type is that the data consumer provides data consumption services through the Application Programming Interface (API) platform: the data in the data warehouse provides data query capabilities through the API platform registration call interface, and the data consumer calls the registered call interface to consume the data in the data warehouse.

[0081] The fourth method is to import data from the data warehouse into the message queue (MQ), and the data consumer consumes MQ: the data in the data warehouse is synchronized to MQ, and the data consumer consumes MQ to consume the data in the data warehouse.

[0082] Therefore, by tracking the data tracking link of the data consumer end, the online data table queried by the data consumer end (such as MySQL, etc.), the consumed MQ information, and the called calling interface (that is, the interface of the API platform) can be obtained, and the link node corresponding to the online data table, MQ information, and calling interface in the first data link can be used as the first link node.

[0083] Step a3: Determine the data object called by the first link node in the data warehouse.

[0084] Specifically, by sorting out the interaction logic between the data consumer and the data warehouse, the terminal data table or task directly consumed by the data consumer in the data warehouse is accurately located to obtain the data object called by the first link node in the data warehouse.

[0085] Specifically, by monitoring data synchronization tasks and parsing their parameters, data objects in online data tables and MQ information sources are obtained. By parsing query SQL statements, the relationship between data warehouse data objects and online services / API platforms is obtained. The data objects called in the data warehouse by direct calls and call interfaces on the data consumer side are obtained, thereby obtaining the data objects called in the data warehouse by the first link node.

[0086] Specifically, the above step a3 includes:

[0087] Step a31 : determining candidate data objects for the data warehouse to provide data to the data consumer.

[0088] Step a32: Obtain a query statement for the candidate data object.

[0089] Specifically, when the data consumer consumes data from the data warehouse through the above four data consumption methods, it needs to use query statements to consume data. At this time, the query statements of the data consumer for the candidate data objects can be recorded.

[0090] It should be noted that the query statement is an SQL statement.

[0091] Step a33: parse the query statement to obtain the calling relationship between the candidate data object and the first link node.

[0092] Specifically, by parsing the query statement, the calling relationship between the data tables of the data warehouse and the online services and API platform provided by the data warehouse can be obtained, thereby determining the calling relationship between the called data object and the first link node.

[0093] Step a34: determining the data object called by the first link node in the data warehouse based on the calling relationship.

[0094] Specifically, the data object called by the first link node is determined from the candidate data objects based on the calling relationship.

[0095] Step a4: Taking the called data object as the starting point, construct the second data link of the data warehouse, and connect the first link node with the second link node corresponding to the called data object in the second data link to obtain the target data link.

[0096] It should be noted that the link nodes in the second data link correspond to the data objects in the data warehouse.

[0097] It can be understood that taking the data object called at the end of the data warehouse as the starting point can ensure a clear starting point for data lineage analysis and lay the foundation for subsequent data link tracking.

[0098] Specifically, see Figure 4, associate the corresponding first link node with the second link node, thereby aggregating the first data link and the second data link to obtain a data link between the data warehouse and the data consumer end from a full perspective, that is, the target data link.

[0099] Optionally, the data processing method of the present disclosure further includes: displaying the target data link in response to a display instruction for the target data link. Thus, it is possible to graphically allow relevant personnel to intuitively identify the upstream sources of the affected data, such as the data tables, tasks, and data flows associated with the affected data, thereby providing a reliable basis for tracing and optimizing data quality issues.

[0100] Step a5: Perform lineage analysis on the affected data based on the target data link to obtain a lineage link.

[0101] Specifically, the link where the affected data is located in the target data link is used as the lineage link of the affected data.

[0102] The data processing method provided in this embodiment obtains the first data link of the data consumer end and the second data link of the data warehouse, and then connects the first link node and the second link node that have a calling relationship in the first data link and the second data link. Therefore, it is possible to aggregate the data links between the data consumer end and the data warehouse to obtain the target data link, analyze the data calling situation of the affected data in the data warehouse from a full perspective, and then quickly determine the target data object that the affected data depends on in the data warehouse, thereby improving the analysis efficiency of the target data object.

[0103] In some optional implementations, the above step S303 includes:

[0104] Step b1: determine the leaf node associated with the data consumer in the bloodline link.

[0105] It can be understood that the leaf node is the last link node in the bloodline link that is connected to the data consumer end, that is, the second link node mentioned above.

[0106] Step b2: Use the data object corresponding to the leaf node as the target data object.

[0107] Specifically, the data table or task corresponding to the leaf node is used as the target data object.

[0108] It is worth noting that the data processing method disclosed in this disclosure initiates a declaration on the data governance platform based on the leaf nodes in the lineage chain. Relying on field-level lineage relationships, efficient transparent transmission of affected data between upstream and downstream is achieved. Therefore, it is possible to achieve a close integration of the declaration process and the transmission of data change information of the affected data, ensuring that data control of the affected data covers the entire business process and improving the timeliness and accuracy of data change processing.

[0109] The data processing method provided in this embodiment uses the leaf node associated with the data consumption end as the target data object, which can quickly trace the responsible party of the affected data in the data warehouse while ensuring the data quality bottom line of the affected data.

[0110] In some optional embodiments, the above-mentioned step S304 includes: displaying the affected data and the corresponding target data object on a first page; in response to a selection operation on the target data object, displaying a second page, the second page displaying a configuration item having at least one application information; in response to a configuration operation on the configuration item, generating a declaration request, and sending the declaration request to the belonging object.

[0111] For example, Figure 5 As shown, the data governance platform displays the first page, which displays the affected data (data1, data2), the target data object (object1) corresponding to data1, and the target data object (object2) corresponding to data2. If a user wishes to submit a request for data1, they can click object1. In response to the click on object1, the data governance platform displays the second page. By configuring the configuration items for the request information on the second page, a request can be initiated to the object to which object1 belongs.

[0112] In addition, the first page also displays the number of lineage links of all affected data, the number of affected data for which a data change notification mechanism has been established, the number of affected data waiting for approval, the number of affected data that have been approved for declaration, etc.

[0113] The data processing method provided in this embodiment displays the affected data and the corresponding target data objects on the first page. Therefore, it is convenient to intuitively view the affected data on the data consumer side, thereby selectively selecting the affected data for which data change services need to be provided and reporting to the object to which the target data object of the affected data belongs.

[0114] For example, see Figure 6As shown, data consumers bear the key responsibility for identifying data risk scenarios. Data consumers conduct a comprehensive review of their current business scenarios based on established impact assessment criteria. Through manual labeling and recognition model-based labeling, data consumption scenarios are tagged on the consumer's business data and call interfaces to identify data that is most significantly impacted by data changes in the data warehouse. Labeling data consumption scenarios requires comprehensive consideration of multiple factors, such as business complexity and transaction volume. Once the affected data is identified, if reporting is required, reporting is carried out based on the identified target data objects. The target data objects contain detailed information on the specific usage of the data in the business process and related parameters. Relevant personnel can flexibly choose to report individual affected data or select affected data in batches based on actual business needs. The standardized implementation of the reporting process ensures that affected data is promptly and accurately transmitted to subsequent approval processes.

[0115] In the approval process, first, the first-level approval is carried out by the relevant party. It is understandable that the relevant party has a deep understanding of the business data and related processes for which it is responsible, and can conduct a preliminary review of the declaration request based on business logic and data impact preferences. For example, it is necessary to determine whether the description of the data consumption scenario of the affected data reported is accurate, whether the reported processing plan complies with normal business operation specifications, etc. After the first-level approval is passed, the data quality assurance (QA) personnel will intervene to enter the second-level approval. From the perspective of quality control and compliance, the data QA will conduct a more rigorous and detailed review of the content in the declaration request. For example, it is necessary to check whether the declaration process is complete and compliant, and whether the risk prevention and control measures are appropriate and effective. Only after passing the two-level approval can the declaration request enter the next process. This multi-level approval mechanism can effectively ensure the scientificity and rigor of data change notification decisions.

[0116] After the application request is approved, the data warehouse's data governance platform accurately labels the target data objects (such as upstream tasks) associated with the affected data at the field level. Field-level lineage refers to the relationships and data flow between fields as data flows through various business processes. By labeling target data objects, the upstream target data objects of the affected data can be clearly identified.

[0117] To ensure the stable and efficient operation of the entire data chain, it is necessary to comprehensively improve data quality and timeliness monitoring. In terms of content quality, the accuracy and completeness of affected data transmission should be monitored in real time. For example, the affected data should be checked for errors or omissions. In terms of timeliness monitoring, the time interval between the request and the response of the affected data should be closely monitored to ensure that business operations can be completed within the specified time. Once the upstream target data object or data chain changes, such as data structure adjustments or interface protocol updates, the data warehouse can automatically trigger a notification mechanism to promptly and accurately notify data consumers of data change information. Data consumers can quickly adjust their business processes and related operations based on the data change information, effectively ensuring that the entire data chain can maintain high-quality and efficient operation in the face of various changes.

[0118] In some optional implementations, the data processing method of the present disclosure further includes:

[0119] Step c1: receiving a change message for affected data sent by a data consumer.

[0120] Step c2: Send the change message to the corresponding object.

[0121] It is worth noting that the data governance platform will use system messages as a notification method to promptly transmit change messages of affected data to the objects to which the target data objects belong.

[0122] For example, if the business service on the data server side involves adjusting the fields of a certain affected data, the object to which the target data object on which the affected data depends will be informed immediately through system messages so that the corresponding data processing and business adjustment preparations can be made in advance to ensure the stable operation of the entire data link.

[0123] In the data processing method provided by this embodiment, the data warehouse sends the change message for the affected data sent by the data consumer to the belonging object, thereby ensuring that the belonging object is promptly aware of the changes made by the downstream data consumer for the affected data and can promptly respond to the impact of the data changes on the affected data.

[0124] In this embodiment, another data processing method is provided, which can be used for data consumption terminals, such as servers, mobile phones, tablet computers, etc. Figure 7 is a flow chart of another data processing method according to an embodiment of the present invention. Figure 7 As shown, the process includes the following steps:

[0125] Step S701: Obtain consumption data information, which includes data affected by data changes in the data warehouse. Please refer to the above step S301 and will not be described in detail here.

[0126] Step S702, sending consumption data information to the data warehouse; wherein, the data warehouse is used to perform lineage analysis on the affected data, obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, the declaration request is used to declare the data impact relationship between the target data object and the data consumption end; the data warehouse is also used to send data change information to the data consumption end if the declaration request is passed and there is a data change in the target data object.

[0127] It should be noted that, for the relevant steps after the data warehouse receives the consumption data information, reference can be made to the above-mentioned data processing method for the data warehouse, which will not be elaborated here.

[0128] The data processing method provided in this embodiment is that the data consumer sends the affected data affected by the data change in the data warehouse to the data warehouse. The data warehouse performs a lineage analysis on the affected data and traces back to the data objects in the data warehouse that have a dependency relationship with the affected data to obtain a lineage link. Then, based on the lineage link, the target data object on which the affected data depends is determined, and a declaration request is initiated to the object to which the target data object belongs, so that the object to which it belongs is informed of the impact of its own business on the downstream data consumer. Furthermore, after the declaration request is passed, a data change notification mechanism is established between the target data object and the data consumer. In the case that there is a data change in the target data object, the affected data consumer is quickly determined, and the data change information is sent to the data consumer in a timely manner, so that the data consumer is informed of the data change information of the upstream data warehouse in a timely manner, and can quickly respond to the data impact that may be caused by the data change in the data warehouse, thereby ensuring the accuracy of the data and the stable operation of the business service.

[0129] In some optional implementations, the above step S701 includes:

[0130] Step d1: Acquire first data from a data warehouse.

[0131] It can be understood that the first data includes business data consumed in the data warehouse.

[0132] Step d2: analyzing the degree to which the first data is affected by the data change in the data warehouse based on the first recognition model to obtain a first recognition result of the first data.

[0133] Optionally, the first recognition model is a natural language model.

[0134] Specifically, the first recognition model is a large model. In addition, other natural language models, deep learning models, etc. can also be used, which are not limited here.

[0135] Specifically, the first data is input into the first recognition model, and the degree to which the first data is affected by the data change in the data warehouse is analyzed to obtain a first recognition result of the first data.

[0136] Step d3: determining the second data from all the first data based on the first recognition result.

[0137] Specifically, if the first recognition result indicates that there is a data risk in the first data, the first data is determined to be the second data.

[0138] Step d4: taking the second data as the affected data.

[0139] The data processing method provided in this embodiment utilizes the data processing capability of the first recognition model to perform deep semantic analysis on the first data, thereby being able to quickly analyze the extent to which the first data is affected by data changes in the data warehouse.

[0140] In some optional implementations, the above step S701 includes:

[0141] Step e1: Obtain interface information of a first calling interface facing the data warehouse.

[0142] Specifically, the interface information includes the interface name, description information, etc.

[0143] Step e2: Based on the second recognition model and the interface information of the first calling interface, the degree to which the first calling interface is affected by the data change in the data warehouse is analyzed to obtain a second recognition result of the first calling interface.

[0144] Optionally, the second recognition model is a natural language model.

[0145] Specifically, the second recognition model is a large model. In addition, other natural language models, deep learning models, etc. can also be used, which are not limited here.

[0146] Specifically, the interface information of the first calling interface is input into the second recognition model, and the degree to which the first calling interface is affected by the data change in the data warehouse is analyzed to obtain a second recognition result of the first calling interface.

[0147] Step e3: Based on the second identification result, determine the second calling interface from all the first calling interfaces.

[0148] Specifically, if the second identification result indicates that the first calling interface has a data risk, the first calling interface is determined as the second calling interface.

[0149] Step e4: Use the interface information of the second calling interface as the affected data.

[0150] It's worth noting that interface information, such as the interface name and description, often implies specific scenarios. For example, in a financial business scenario, when the interface name is "Refund," it intuitively reflects that the interface provides a refund service. When the interface name is "Payment," it is likely related to payment functionality. Therefore, the naming rules of the calling interface, based on business semantics, can provide a feasible path for the recognition model to identify the data consumption scenarios of the calling interface and determine the extent to which the calling interface is affected by data changes in the data warehouse. Therefore, with the help of the recognition model's powerful natural language processing capabilities, in-depth semantic analysis of interface information, such as the interface name and description, can be performed to effectively infer the data consumption scenarios of the calling interface and determine the extent to which the calling interface is affected by data changes in the data warehouse.

[0151] The data processing method provided in this embodiment utilizes the data processing capability of the second recognition model to perform deep semantic analysis on the interface information of the first calling interface. Thus, it is possible to quickly analyze the degree to which the first calling interface is affected by data changes in the data warehouse.

[0152] Specifically, see Figure 8 The data consumer retrieves local documents and identifies the intent of the call interface to determine the data and call interface related to the data warehouse. Using a recognition model, it analyzes the extent to which the data and call interface are affected by the data warehouse data change to determine the affected data. Then, based on the affected data, it pushes a message to the data warehouse to send information about the consumed data.

[0153] In some optional implementations, the data processing method applicable to the data consumer end of the present disclosure further includes:

[0154] Step f1: Obtain interface information of the third calling interface and the corresponding impact degree label.

[0155] Specifically, the impact labels can be obtained by manually labeling the interface information of some call interfaces (i.e., the third call interface) on the data consumer side. For example, the impact labels can be obtained by clearly labeling which call interfaces involve operations such as fund collection and payment, circulation, and settlement. Alternatively, a pre-trained model can be used to label the interface information of the third call interface to obtain the impact labels, which are not limited here.

[0156] Step f2: input the interface information of the third calling interface into a preset recognition model, analyze the degree to which the third calling interface is affected by the data change in the data warehouse, and obtain a third recognition result of the third calling interface.

[0157] Understandably, in the initial stage, the preset recognition model makes preliminary judgments on the interface information based on common language patterns and a small amount of prior knowledge, but its accuracy is limited.

[0158] Step f3: Based on the third recognition result and the impact degree label, adjust the parameters of the preset recognition model to obtain a second recognition model.

[0159] Understandably, inputting the interface information and impact label of the third call interface into the preset recognition model can guide the preset recognition model to learn the unique characteristics and patterns of the call interface's interface information, such as interface name and description. After multiple rounds of training and parameter optimization, the preset recognition model's recognition accuracy for call interfaces will be significantly improved, providing strong support for efficient management of call interfaces and risk prevention and control of data changes.

[0160] The data processing method provided in this embodiment uses the interface information of the third calling interface and the corresponding impact degree label to train the preset recognition model. Therefore, the preset recognition model can learn the unique characteristics of the interface information of the calling interface, thereby improving the recognition accuracy of the second recognition model finally trained for the degree of impact of the calling interface affected by data changes in the data warehouse.

[0161] In some optional implementations, the data processing method applicable to the data consumer end of the present disclosure further includes:

[0162] Step g1: In response to a data change operation on affected data, generating a change message for the affected data.

[0163] Step g2: Send the change message to the data warehouse, which is also used to send the change message to the corresponding object.

[0164] It should be noted that the relevant descriptions of the above steps g1 and g2 can be found in the relevant descriptions of the above steps c1 and c2, and will not be repeated here.

[0165] In the data processing method provided by this embodiment, the data warehouse sends the change message for the affected data sent by the data consumer to the belonging object, thereby ensuring that the belonging object is promptly aware of the changes made by the downstream data consumer for the affected data and can promptly respond to the impact of the data changes on the affected data.

[0166] As a specific application example, a first application is installed in the data warehouse, and the first application is used to execute the data processing method of the present disclosure applicable to the data warehouse. A second application is installed in the data consumption end, and the second application is used to execute the data processing method of the present disclosure applicable to the data consumption end. Figure 9 and Figure 10In the platform jointly constituted by the data warehouse and data consumption end of the present disclosure, the user-facing page of the platform provides page functions such as impact degree tagging, batch declaration, monitoring coverage, data change control, and data measurement. At the same time, it provides platform functions such as task declaration registration function, monitoring coverage function, data change control function, data measurement capability and lineage identification function. The data consumption end runs the second application to identify the tables, fields and call interfaces in the data consumption end that are affected by the data change of the data warehouse, obtain the affected data, and use this to send consumption data information to the data warehouse. The data warehouse runs the first application, using the field lineage capability of the data consumption end and the data warehouse to open the data link between the data consumption end and the data warehouse, obtain the lineage link of the affected data, and determine the target data objects that the affected data depends on in the data warehouse, such as data tables, upstream tasks, etc. Then, a declaration request is initiated to the object to which the target data object belongs. In the case of the declaration request, the data warehouse establishes a data change notification mechanism between the target data object and the data consumption end to promptly notify the other party when there is a data change related to the affected data on the target data object or the data consumption end. In addition, the data consumer can also monitor the number of tables, number of fields, and reporting rate in the affected data. The data warehouse can also manage the data change notification mechanism of the affected data through the transparent transmission rate of the affected data, the APP layer monitoring rate, the DWD layer monitoring rate, the DWS layer monitoring rate, etc., so as to understand the actual situation of data change processing between the data warehouse and the data consumer.

[0167] In this embodiment, a data processing device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0168] This embodiment provides a data processing device, which is applied to a data warehouse, such as Figure 11 As shown, the device includes:

[0169] The information acquisition module 1101 is used to acquire consumption data information sent by the data consumer, where the consumption data information includes the affected data affected by the data change in the data warehouse;

[0170] The data analysis module 1102 is used to perform lineage analysis on the affected data to obtain the lineage links of the affected data in the data warehouse;

[0171] A dependency determination module 1103 is configured to determine a target data object on which the affected data depends based on a lineage link;

[0172] The data declaration module 1104 is used to initiate a declaration request to the object to which the target data object belongs. The declaration request is used to declare the data impact relationship between the target data object and the data consumer.

[0173] The change notification module 1105 is used to send data change information to the data consumer if the declaration request is approved and there is data change in the target data object.

[0174] In some optional implementations, the data analysis module 1102 includes:

[0175] A link acquisition unit, configured to acquire a first data link from a data consumer;

[0176] A first analyzing unit, configured to determine a first link node in the first data link that accesses the data warehouse;

[0177] a second analyzing unit, configured to determine a data object called by the first link node in the data warehouse;

[0178] a node association unit, configured to construct a second data link of the data warehouse with the called data object as a starting point, and connect the first link node with the second link node corresponding to the called data object in the second data link to obtain a target data link;

[0179] The third analysis unit is used to perform lineage analysis on the affected data based on the target data link to obtain the lineage link.

[0180] In some optional embodiments, the second analysis unit includes:

[0181] The first analysis subunit is used to determine candidate data objects for the data warehouse to provide data to the data consumer;

[0182] A statement acquisition subunit is used to acquire query statements for candidate data objects;

[0183] A statement parsing subunit, configured to parse the query statement to obtain a call relationship between the candidate data object and the first link node;

[0184] The second analysis subunit is used to determine the data object called by the first link node in the data warehouse based on the calling relationship.

[0185] In some optional implementations, the dependency determination module 1103 includes:

[0186] A leaf node determination unit, used to determine the leaf node associated with the data consumer in the lineage link;

[0187] The dependency determination unit is used to take the data object corresponding to the leaf node as the target data object.

[0188] In some optional implementations, the data reporting module 1104 includes:

[0189] A first display unit, configured to display the affected data and the corresponding target data object on a first page;

[0190] A second display unit is configured to display a second page in response to a selection operation on the target data object, wherein the second page displays at least one configuration item of application information;

[0191] The data reporting unit is used to generate a reporting request in response to a configuration operation on a configuration item, and send the reporting request to the corresponding object.

[0192] In some optional embodiments, the data processing device of the present disclosure further includes:

[0193] The message receiving module is used to receive change messages for affected data sent by the data consumer;

[0194] The message forwarding module is used to send the change message to the corresponding object.

[0195] The data processing device provided by the embodiment of the present disclosure can execute the data processing method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. The data processing device provided by the embodiment of the present disclosure, the data warehouse performs a lineage analysis on the affected data of the data consumer end, to trace the target data object on which the affected data depends in the data warehouse, and initiates a declaration request to the object to which the target data object belongs, so that the object to which it belongs knows the impact of its own business on the downstream data consumer end. Furthermore, after the declaration request is passed, a data change notification mechanism is established between the target data object and the data consumer end. In the case that there is a data change in the target data object, the affected data consumer end is quickly determined, and the data change information is sent to the data consumer end in a timely manner, so that the data consumer end can respond to the data change information in a timely manner and adopt corresponding response strategies, thereby ensuring the accuracy of the data and the stable operation of the business service.

[0196] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0197] In this embodiment, a data processing device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0198] This embodiment provides a data processing device, which is applied to a data consumption end, such as Figure 12 As shown, the device includes:

[0199] The data identification module 1201 is used to obtain consumption data information, including affected data affected by data changes in the data warehouse;

[0200] The data sending module 1202 is used to send consumption data information to the data warehouse; wherein, the data warehouse is used to perform lineage analysis on the affected data, obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, and the declaration request is used to declare the data impact relationship between the target data object and the data consumption end; the data warehouse is also used to send data change information to the data consumption end if the declaration request is passed and there is a data change in the target data object.

[0201] In some optional implementations, the data identification module 1201 includes:

[0202] A first acquiring unit, configured to acquire first data from a data warehouse;

[0203] a first recognition unit, configured to analyze, based on a first recognition model, the degree to which the first data is affected by the data change in the data warehouse, and obtain a first recognition result of the first data;

[0204] a data screening unit, configured to determine second data from all first data based on the first recognition result;

[0205] The first determining unit is configured to use the second data as the affected data.

[0206] In some optional implementations, the data identification module 1201 includes:

[0207] A second acquiring unit is used to acquire interface information of a first calling interface facing the data warehouse;

[0208] A second identification unit is configured to analyze, based on the second identification model and the interface information of the first calling interface, the degree to which the first calling interface is affected by the data change in the data warehouse, and obtain a second identification result of the first calling interface;

[0209] An interface screening unit, configured to determine a second calling interface from among all first calling interfaces based on the second identification result;

[0210] The second determining unit is configured to use the interface information of the second calling interface as the affected data.

[0211] In some optional embodiments, the data processing device of the present disclosure further includes:

[0212] A sample acquisition module, used to obtain interface information of the third calling interface and a corresponding impact degree label;

[0213] a model prediction module, configured to input the interface information of the third calling interface into a preset recognition model, analyze the degree to which the third calling interface is affected by the data change in the data warehouse, and obtain a third recognition result of the third calling interface;

[0214] The parameter adjustment module is used to adjust the parameters of the preset recognition model based on the third recognition result and the influence degree label to obtain a second recognition model.

[0215] In some optional embodiments, the data processing device of the present disclosure further includes:

[0216] a change response module, configured to generate a change message for the affected data in response to a data change operation on the affected data;

[0217] The message sending module is used to send the change message to the data warehouse, and the data warehouse is also used to send the change message to the corresponding object.

[0218] The data processing device provided by the embodiment of the present disclosure can execute the data processing method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. In the data processing device provided by the embodiment of the present disclosure, the data consumer sends the affected data of the data warehouse to the data warehouse, so that the data warehouse performs a lineage analysis on the affected data, traces the target data object on which the affected data depends in the data warehouse, and initiates a declaration request to the object to which the target data object belongs, so that the object to which it belongs knows the impact of its own business on the downstream data consumer. Furthermore, after the declaration request is passed, the data warehouse builds a data change notification mechanism between the target data object and the data consumer. In the case that there is a data change in the target data object, the affected data consumer is quickly determined, and the data change information is sent to the data consumer in a timely manner, so that the data consumer responds to the data change information in a timely manner and adopts corresponding response strategies, thereby ensuring the accuracy of the data and the stable operation of the business service.

[0219] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0220] Figure 13 This is a structural block diagram of an electronic device provided in an embodiment of the present disclosure.

[0221] The following specific reference Figure 13, which shows a block diagram of the structure of the electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1302 or the program loaded from the memory 1308 into the random access memory (RAM) 1303. Various programs and data required for the operation of the electronic device are also stored in the RAM 1303. The processor 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0222] Typically, the following devices may be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 1309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 13 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0223] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1309, or installed from the memory 1308, or installed from the ROM 1302. When the computer program is executed by the processor 1301, the above-mentioned functions defined in the data processing method of the embodiment of the present disclosure are performed.

[0224] Figure 13 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0225] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the data processing method shown in the above embodiment is implemented.

[0226] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0227] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data processing method, characterized in that: Applied to a data warehouse, the method includes: Obtaining consumption data information sent by a data consumer, wherein the consumption data information includes affected data affected by the data change in the data warehouse; Performing lineage analysis on the affected data to obtain a lineage link of the affected data in the data warehouse; Determining, based on the lineage link, the target data object on which the affected data depends; Initiate a declaration request to the object to which the target data object belongs, wherein the declaration request is used to declare the data impact relationship between the target data object and the data consumer; If the declaration request is approved, then when there is data change in the target data object, the data change information will be sent to the data consumer.

2. The method according to claim 1, characterized in that The performing lineage analysis on the affected data to obtain the lineage link of the affected data in the data warehouse includes: Acquire a first data link of the data consumption end; Determining a first link node in the first data link that is connected to the data warehouse; determining a data object called by the first link node in the data warehouse; Taking the called data object as a starting point, constructing a second data link of the data warehouse, and connecting the first link node with the second link node corresponding to the called data object in the second data link to obtain a target data link; A lineage analysis is performed on the affected data based on the target data link to obtain the lineage link.

3. The method according to claim 2, characterized in that The determining the data object called by the first link node in the data warehouse includes: Determine candidate data objects for the data warehouse to provide data to the data consumer; Obtaining a query statement for the candidate data object; Parsing the query statement to obtain a calling relationship between the candidate data object and the first link node; The data object called by the first link node in the data warehouse is determined based on the calling relationship.

4. The method according to claim 1, wherein The determining, based on the lineage link, the target data object on which the affected data depends, includes: Determining a leaf node in the lineage link associated with the data consumer; The data object corresponding to the leaf node is used as the target data object.

5. The method according to claim 1, wherein The initiating a declaration request to the object to which the target data object belongs includes: Displaying the affected data and the corresponding target data object on a first page; In response to a selection operation on the target data object, displaying a second page, wherein the second page displays at least one configuration item of application information; In response to the configuration operation for the configuration item, the declaration request is generated, and the belonging declaration request is sent to the belonging object.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Receiving a change message for the affected data sent by the data consumer; The change message is sent to the belonging object.

7. A data processing method, characterized in that: Applied to the data consumption end, the method includes: Obtaining consumption data information, wherein the consumption data information includes affected data affected by the data change in the data warehouse; Sending the consumption data information to the data warehouse; wherein the data warehouse is used to perform a lineage analysis on the affected data, obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, the declaration request being used to declare the data impact relationship between the target data object and the data consumer end; the data warehouse is also used to, if the declaration request is passed, send data change information to the data consumer end in the event that there is a data change in the target data object.

8. The method according to claim 7, characterized in that The obtaining of consumption data information includes: Acquire first data from the data warehouse; analyzing, based on a first recognition model, the degree to which the first data is affected by the data change in the data warehouse, to obtain a first recognition result of the first data; Based on the first recognition result, determining second data from all the first data; The second data is used as the affected data.

9. The method according to claim 7, characterized in that The obtaining of consumption data information includes: Obtaining interface information of a first calling interface facing the data warehouse; Based on the second recognition model and the interface information of the first calling interface, analyzing the degree to which the first calling interface is affected by the data change in the data warehouse, to obtain a second recognition result of the first calling interface; Based on the second recognition result, determining a second calling interface from all the first calling interfaces; The interface information of the second calling interface is used as the affected data.

10. The method according to claim 9, characterized in that The method further comprises: Obtaining interface information of the third calling interface and the corresponding impact degree label; Inputting the interface information of the third calling interface into a preset recognition model, analyzing the degree to which the third calling interface is affected by the data change in the data warehouse, and obtaining a third recognition result of the third calling interface; Based on the third recognition result and the influence degree label, the parameters of the preset recognition model are adjusted to obtain the second recognition model.

11. The method according to claim 7, characterized in that The method further comprises: In response to the data change operation on the affected data, generating a change message for the affected data; The change message is sent to the data warehouse, and the data warehouse is further configured to send the change message to the belonging object.

12. A data processing device, characterized in that: Applied to a data warehouse, the device includes: An information acquisition module, configured to acquire consumption data information sent by a data consumer, wherein the consumption data information includes affected data affected by data changes in the data warehouse; A data analysis module, configured to perform lineage analysis on the affected data to obtain lineage links of the affected data in the data warehouse; A dependency determination module, configured to determine, based on the lineage link, a target data object on which the affected data depends; A data reporting module, configured to initiate a reporting request to an object to which the target data object belongs, wherein the reporting request is used to report the data impact relationship between the target data object and the data consumption end; The change notification module is used to send data change information to the data consumer if the declaration request is passed and there is data change in the target data object.

13. A data processing device, characterized in that: Applied to a data consumption end, the device includes: A data identification module is used to obtain consumption data information, wherein the consumption data information includes affected data affected by data changes in the data warehouse; A data sending module is used to send the consumption data information to the data warehouse; wherein, the data warehouse is used to perform a lineage analysis on the affected data to obtain the lineage link of the affected data in the data warehouse, and, based on the lineage link, determine the target data object on which the affected data depends, and initiate a declaration request to the object to which the target data object belongs, the declaration request being used to declare the data impact relationship between the target data object and the data consumer end; the data warehouse is also used to, if the declaration request is passed, send data change information to the data consumer end in the event that there is a data change in the target data object.

14. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 6, or the data processing method according to any one of claims 7 to 11, by executing the computer instructions.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 6, or to execute the data processing method according to any one of claims 7 to 11.

16. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the data processing method according to any one of claims 1 to 6, or to execute the data processing method according to any one of claims 7 to 11.