Data Processing Method, System, Electronic Device, and Storage Medium

By building a layered data management architecture and big data technology, the problems of low credibility of data statistics results, high demand delivery pressure, and poor user experience are solved, and data consistency and timeliness are improved, and data delivery quality and user experience are optimized.

CN116149823BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310184740.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-07-25
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

In the existing data management methods, there are problems such as low credibility in data statistics, high demand delivery pressure, and poor user experience. Especially in the fields of big data and autonomous driving, there is a lack of effective BI products, tools and self-service analysis tools, resulting in poor data statistics accuracy and user experience.

Method used

By building a new data management architecture, including data operation layer ODS, data detail layer DWD, data service layer DWS and application data layer ADS, we adopt layer scheduling and data fusion technology to standardize data call and management, introduce big data technologies such as SPARK SQL and Structured Streaming, establish a unified data access, data mart and visualization platform, and implement a unified data governance and point burial mechanism.

Benefits of technology

It improves the quality and efficiency of data delivery, improves user credibility and experience, ensures data consistency and timeliness, and optimizes the overall delivery quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116149823B_ABST
    Figure CN116149823B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data management method, apparatus, electronic device, and storage medium, which relate to the field of data processing, and particularly to fields such as big data and autonomous driving. The method includes: obtaining initial data from the Operational Data Store (ODS) based on data acquisition requirements; the initial data includes offline data and real-time data; based on the initial data, obtaining the first full amount of data corresponding to the first data scheduling level at the current moment in the Data Warehouse Detail (DWD), and the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment; the time span of the second data scheduling level is less than the time span of the first data scheduling level; according to the first full amount of data, the first incremental data, and a preset incremental generation rule, obtaining the first target full amount of data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS); controlling the Application Data Store (ADS) to call and output the first target full amount of data from the Data Warehouse Service (DWS).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and particularly to fields such as big data and autonomous driving, and provides a data processing method, system, electronic device, and storage medium. Background Art

[0002] The data platform faces many users and roles, including management, products, operations, and R & D, etc. Currently, the data platform generally combines various requirements together, resulting in mutual influence among various requirements and users. Under the overall architecture of the existing data platform, the following problems generally exist: (1) The credibility of data statistics results is not high: Overall, various requirements and users are not isolated, resulting in frequent iteration of indicators with low quality requirements affecting important indicators; unstable data sources, dirty data, and fluctuations in business data affect the accuracy of data statistics; there is no effective verification mechanism and process, and it is impossible to judge whether the statistical results are accurate; (2) The pressure of demand delivery is high: There is a lack of effective BI (Business Intelligence) products, tools, and data warehouses, resulting in strong dependence on R & D for demand iteration; there is no efficient self-service analysis tool provided for users, resulting in various requirements being squeezed into the data platform; there is no overall planning for the data indicator system, resulting in repeated development and frequent modification of many requirements; (3) The user experience is poor: The user interface lacks product and interaction intervention, the overall design style is relatively casual, and there are many user confusions, and users cannot understand and use the product well; there is a lack of user documentation and indicator explanations, and the statistical results of many indicators are relatively black-box, and users cannot effectively understand the true meaning of the indicators. Summary of the Invention

[0003] The technical problem to be solved by the present disclosure is to overcome the defects in the existing data management methods, such as low credibility of data statistics results, high pressure of demand delivery, and poor user experience, and provide a data processing method, system, electronic device, and storage medium.

[0004] The present disclosure solves the above technical problems through the following technical solutions:

[0005] According to one aspect of the present disclosure, a data management method is provided, and the data management method includes:

[0006] Obtaining initial data from the Operational Data Store (ODS) layer of data operation based on data acquisition requirements;

[0007] Wherein, the initial data includes offline data and real-time data;

[0008] Based on the initial data, obtaining the first full amount of data corresponding to the first data scheduling level at the current moment in the Data Warehouse Detail (DWD) layer, and the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment, where N is a positive integer;

[0009] Wherein, the first data scheduling level and the second data scheduling are used to respectively call the data belonging to the corresponding time span, and the time span of the second data scheduling level is less than that of the first data scheduling level;

[0010] According to the first full amount of data, the first incremental data, and a preset incremental generation rule, obtain the first target full amount of data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS;

[0011] Control the application data layer ADS to call and output the first target full amount of data from the data service layer DWS.

[0012] According to another aspect of the present disclosure, there is provided a data management device, the data management device includes:

[0013] An initial data acquisition module, configured to acquire initial data from the data operation layer ODS based on data acquisition requirements;

[0014] Wherein, the initial data includes offline data and real-time data;

[0015] A first full amount of data acquisition module, configured to acquire the first full amount of data corresponding to the first data scheduling level at the current moment in the data detail layer DWD based on the initial data;

[0016] A first incremental data acquisition module, configured to generate the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment in the data detail layer DWD based on the initial data, where N is a positive integer;

[0017] Wherein, the first data scheduling level and the second data scheduling are used to respectively call the data belonging to the corresponding time span, and the time span of the second data scheduling level is less than that of the first data scheduling level;

[0018] A first target full amount acquisition module, configured to obtain the first target full amount of data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS according to the first full amount of data, the first incremental data, and a preset incremental generation rule;

[0019] A data output module, configured to control the application data layer ADS to call and output the first target full amount of data from the data service layer DWS.

[0020] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0024] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.

[0025] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the above method when executed by a processor.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0028] Figure 1 It is a schematic diagram of the framework of an existing data management framework.

[0029] Figure 2 It is a flowchart of the data management method according to Embodiment 1 of the present disclosure.

[0030] Figure 3 It is a first schematic diagram of the data management framework according to Embodiment 1 of the present disclosure.

[0031] Figure 4 It is a schematic diagram of the fusion of new and old data in the data management method according to Embodiment 1 of the present disclosure.

[0032] Figure 5 It is a schematic diagram of the data scheduling logic in the data management method according to Embodiment 1 of the present disclosure.

[0033] Figure 6 It is a schematic diagram of a specific example of the fusion of new and old data in the data management method according to Embodiment 1 of the present disclosure.

[0034] Figure 7 It is a second schematic diagram of the data management framework according to Embodiment 1 of the present disclosure.

[0035] Figure 8 It is a schematic diagram of the modules of the data management device according to Embodiment 2 of the present disclosure.

[0036] Figure 9 It is a schematic diagram of the structure of an electronic device for implementing the data management method according to Embodiment 3 of the present disclosure. Detailed implementation manners

[0037] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0038] Embodiment 1

[0039] The existing data management architectures, such as Figure 1 shown, generally have the following problems:

[0040] (1) There is no unified mechanism and rule for defining index calibers

[0041] For the display design, index names, index calculation formulas, etc. of each data dashboard, they are all defined separately and there is no definition standard;

[0042] (2) There is no unified data warehouse modeling technical architecture

[0043] Data access tools: more than 40 scripts, single-machine operation, insufficient stability of the group cloud message middleware tool, and untimely follow-up of problems by operation and maintenance personnel, etc.;

[0044] Offline and real-time data warehouses: There is no unified and clear data warehouse layer construction, and there is no hierarchical call specification; different developers adopt their own cleaning rules and aggregate the same index from different layers;

[0045] Data calculation and storage: There is no big data calculation architecture, and all rely on Doris (distributed SQL database), and the superior performance and stability of Doris are seriously insufficient;

[0046] (3) There are no effective tools

[0047] Data lineage: There is no retrieval tool for each role to retrieve the definitions, cleaning links, etc. corresponding to each index, field, table, etc. of the index;

[0048] Data dictionary: There is no comprehensive index dictionary and no tool for automatically verifying whether there are changes or unentered indexes when going online;

[0049] (4) There is no perfect system monitoring means and system

[0050] Specifications: There are no data tracking specifications, requirement access specifications, requirement R & D specifications, etc.;

[0051] System monitoring means: There are no means such as automatically perceiving data consistency and data accuracy modules.

[0052] With the rapid development of the business, the business volume has increased from hundreds of thousands to millions. Based on the above data warehouse architecture, problems such as inconsistent metric calibers and calculation methods, metric errors, and increased page opening time have become increasingly prominent after the rapid increase and iteration of data platform reports.

[0053] To ensure data authority, platform availability, and performance, this disclosure constructs a new data management architecture that combines multiple guarantees according to business requirements and proposes a new data management solution.

[0054] Embodiment 1

[0055] As Figure 2 shown, the data management method of this embodiment includes:

[0056] S201. Obtain initial data from the Operational Data Store (ODS) layer of data operations based on data acquisition requirements;

[0057] Among them, the initial data includes offline data and real-time data;

[0058] S202. Based on the initial data, obtain the first full amount of data corresponding to the first data scheduling level at the current moment in the Data Warehouse Detail (DWD) layer, and the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment, where N is a positive integer;

[0059] Among them, the first data scheduling level and the second data scheduling level are used to separately call the data belonging to the corresponding time spans, and the time span of the second data scheduling level is less than that of the first data scheduling level;

[0060] Generally, the data within the most recent 2 days is taken. Of course, the data within the most recent 1 day or the most recent 3 days can also be taken; the value of N can be determined or adjusted according to the actual application situation.

[0061] Preferably, the first data scheduling level corresponds to data scheduling on a daily basis (or daily-level scheduling), and the second data scheduling level corresponds to data scheduling on a quarter-hour basis (or quarter-hour-level scheduling, corresponding to a 15-minute data slice).

[0062] S203. According to the first full amount of data, the first incremental data, and a preset incremental generation rule, obtain the first target full amount of data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer;

[0063] S204. Control the Application Data Store (ADS) layer to call and output the first target full amount of data from the Data Warehouse Service (DWS) layer.

[0064] Under the data management architecture of the data management solution in this embodiment, a complete data system construction facing decision-making, analysis, and operation can be completed, a unified data access, data engine, data mart, and visualization platform are constructed, and the overall delivery quality, delivery efficiency, and user experience are optimized.

[0065] In an implementable solution, before step S201, it further includes:

[0066] Based on different data classification criteria, several types of data source units are constructed;

[0067] Among them, in the autonomous driving scenario, several types of data source units include business data units, vehicle-end data units, cloud data units, road test data units, etc., that is, different types of source data are respectively stored in the corresponding units for overall management and subsequent calls, etc.

[0068] A preset data synchronization tool is used to synchronize the corresponding heterogeneous data in different types of data source units to obtain the first raw data after synchronization;

[0069] The first raw data is processed by a preset unified processing tool to obtain the second raw data that meets the preset normalization requirements;

[0070] Specifically, the preset unified processing tool includes a preset binary parsing tool and a preset transmission service tool that are processed in sequence; of course, it can also be other processing tools, as long as it can meet the corresponding data normalization processing requirements, which will not be elaborated here.

[0071] Synchronize the second raw data to the data operation layer ODS;

[0072] Specifically, step S201 includes:

[0073] Extract the initial data that meets the data acquisition requirements from the second raw data in the data operation layer ODS.

[0074] In this solution, after performing heterogeneous data synchronization processing and unified normalization processing on different types of data sources, they are introduced into the data operation layer ODS to obtain the initial data for subsequent direct calls, ensuring the accuracy and efficiency of subsequent data processing.

[0075] In an implementable solution, step S204 includes:

[0076] Obtain the second target full amount of data corresponding to the first data scheduling level in the data service layer DWS;

[0077] Control the application data layer ADS to call and output the first target full amount of data and the second target full amount of data from the data service layer DWS.

[0078] In this solution, the Application Data Layer (ADS) directly calls all the data from the Data Service Layer (DWS) (or Data Aggregation Layer), which restricts the Application Data Layer (ADS) to preferentially call the data in the common layer and does not allow the Application Data Layer (ADS) to reprocess the data from the Operational Data Store (ODS). This standardizes the data call specification of the data warehouse, effectively ensuring data consistency, the reliability of data storage and management, improving the quality and efficiency of data delivery, and enhancing user credibility at the same time.

[0079] In an implementable solution, step S203 includes:

[0080] Perform data fusion processing on the first all-data and the first incremental data to obtain the second all-data corresponding to the second data scheduling level in the Data Warehouse Detail Layer (DWD).

[0081] Obtain the third all-data corresponding to the first data scheduling level in the Data Service Layer (DWS).

[0082] Based on the second all-data and the third all-data, obtain the first target all-data corresponding to the second data scheduling level in the Data Service Layer (DWS) within the most recent N days.

[0083] In this solution, by fusing the first all-data at the day level in the Data Warehouse Detail Layer (DWD) and the first incremental data at the quarter level in the most recent 2 days, the second all-data at the quarter level in the most recent 2 days in the Data Warehouse Detail Layer (DWD) is obtained; for the Data Service Layer (DWS), based on the third all-data at the day level in this layer and the second all-data at the quarter level in the most recent 2 days in the Data Warehouse Detail Layer (DWD), the first target all-data at the quarter level in the most recent 2 days in the Data Service Layer (DWS) is obtained for output by the Application Data Layer (ADS), thus achieving the fusion and update of new and old data at the quarter level, ensuring data consistency, improving the quality and efficiency of data delivery, and enhancing user credibility at the same time.

[0084] In an implementable solution, the step of obtaining the first target all-data corresponding to the second data scheduling level in the Data Service Layer (DWS) within the most recent N days based on the second all-data and the third all-data specifically includes:

[0085] Compare the second all-data and the third all-data to obtain a comparison result.

[0086] Generate the actual incremental information corresponding to the Data Service Layer (DWS) based on the comparison result.

[0087] Based on the actual incremental information, determine the first target all-data corresponding to the second data scheduling level in the Data Service Layer (DWS) within the most recent N days.

[0088] In this solution, by comparing the second full - volume data and the third full - volume data, the actual incremental information of the changes is determined to identify which data are deleted data, which are newly added data, and which are changed data, etc., and the records of deletion, addition, and modification are found to ensure the accuracy of obtaining the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS, meeting the data management requirements.

[0089] In an implementable solution, the steps of comparing the second full - volume data and the third full - volume data to obtain the comparison result specifically include:

[0090] When there is the second full - volume data and there is no third full - volume data in the first data, the first marking information indicating that the first data is deleted data is generated;

[0091] When there is no second full - volume data and there is the third full - volume data in the second data, the second marking information indicating that the second data is newly added data is generated;

[0092] When there is the second full - volume data and there is the third full - volume data in the third data, the third marking information indicating that the third data is changed data is generated;

[0093] In this solution, based on the second full - volume data and the third full - volume data, the records of deletion, addition, and modification are found and corresponding marking processing is performed on them, which is convenient for distinction and management.

[0094] This data management method further includes:

[0095] Based on the first marking information, the second marking information, and the third marking information, an incremental table corresponding to the data service layer DWS is generated.

[0096] In this solution, the different marking information of the found deleted, newly added, and modified data is automatically used to generate an incremental table corresponding to the current data service layer DWS, which is accurate and intuitive. While improving the quality and efficiency of data delivery, it also enhances the user experience.

[0097] Specifically, in combination with Figure 3 and Figure 4 , for the data synchronization of the data operation layer ODS: the data of each business data source is synchronized from the business database to the OLAP ODS cluster (i.e., the data operation layer ODS) in real - time or offline via a unified data synchronization tool:

[0098] Data generation in the DWD (Data Detail Layer): Regularly generate dwd_df (daily full data in the data detail layer) and dwd_qi (quarterly incremental data in the data detail layer) in the OLAP ODS cluster respectively; if the data volume is small, dwd_qf (quarterly full data in the data detail layer) can be directly generated, and all subsequent processes will be simplified accordingly; while generating dwd_df and dwd_qi, be responsible for historical slice merging.

[0099] Full data generation in the DWS (Data Service Layer): Routinely export dwd_df to AFS, process it through SPARK and write it into UDW, and then routinely generate dws_df and import it into PALO; directly generate dws_df (daily full data in the data service layer) from dwd_df and write it into the ADS cluster (i.e., the Application Data Layer ADS) through the external table.

[0100] Incremental data generation in the DWS (Data Service Layer): Use dwd_qi and dws_df of T + a (a is a constant, for example, when a takes the value of 2 for the last 2 days) to generate dws_qf (quarterly full data in the data service layer, not stored in the database), and then combine it with dws_df to generate dws_qi (quarterly incremental data in the data service layer), and write it into the ADS cluster through the external table.

[0101] Data generation in the ADS (Application Data Layer): Obtain ads_qf (quarterly full data in the application data layer) based on dws_df and dws_qi.

[0102] Specifically, combined with Figure 5 the execution scheduling flowchart shown below, the generation rules corresponding to each of the above data layers are as follows:

[0103] dwd_qf = dwd_df(T + a) + dwd_qi(ad); a is a constant and can take the value of 2;

[0104] dws_qi = dwd_qf + dws_df(T + a)

[0105] ads_qf = dws_qf = dws_df(T + a) + dws_qi(ad)

[0106] In order to ensure the timeliness of generating slices for the quarterly full data after the change of irregular historical data, the following processing method is adopted:

[0107] For offline data, the DWD and DWS in the data detail layer and data service layer after daily full cleaning and aggregation are imported into the real-time OLAP Doris system in the data warehouse, and no 15 data shards are established to solve the problem of insufficient cluster capacity with large data volume;

[0108] For real-time data: The 15-minute slice only records the change records within 2 days. Records that do not exist are marked as invalid, newly added, or changed, and records with no changes are not recorded.

[0109] Specifically, retain the data change records for each period to solve the problem of asynchronous update timeliness of the entire ADS link at the bottom layer of each page, resulting in inconsistent page data. For the data service layer DWS with many dimensions and large amounts of data, only incremental data is stored in the OLAP Doris system to improve the overall efficiency of the data stream.

[0110] DWS temporary full volume = DWD full volume (T + 2) + DWD increment (changes in the last 2 days)

[0111] DWS increment = DWS full volume (T + 2) & DWS temporary full volume; among them, the generation process of the DWS increment table is as follows: Find the deleted, newly added, and modified records in DWS full volume (T + 2) and DWS temporary full volume to generate an effective DWS increment table. Among them:

[0112] Determine the data that belongs to the deleted: DWS full volume (T + 2) exists, DWS temporary full volume does not exist, and it is marked as invalid 0.

[0113] Determine the data that belongs to the newly added: DWS full volume (T + 2) does not exist, DWS temporary full volume exists, and it is marked as valid 1.

[0114] Determine the data that belongs to the changed: DWS full volume (T + 2) exists, DWS temporary full volume exists, take the data corresponding to the latest time of this record in the DWS temporary full volume, and mark it as valid 1.

[0115] Combined Figure 6 As shown, after the above data processing process, after the irregular historical data is changed, the quarter-level full volume data timeliness generates slices. The 15-minute slice only records the change records within 2 days. Records that do not exist are marked as invalid, newly added, or changed, and records with no changes are not recorded; that is, the offline and real-time solutions are combined to ensure the timeliness of obtaining the 15-minute (quarter) full volume data.

[0116] In an implementable solution, the data management method further includes:

[0117] Preset data governance rules;

[0118] Perform data governance operations on the first target full volume data and the second target full volume data output by the application data layer ADS based on the data governance rules.

[0119] Specifically, through the constructed data governance rules, the construction of data dictionaries and data lineage in data governance is completed; among them, automatic index verification is performed on the pages to be launched to determine whether they already exist or need to be newly added, etc.; the indexes include but are not limited to visualization, clear and unique definition, and retrievability; through data governance, the overall standardization and effectiveness of data management are ensured.

[0120] In an implementable solution, offline data and real-time data adopt hierarchical modeling in the data warehouse and meet the preset hierarchical call specifications.

[0121] Specifically, by using some big data technology solutions and tools, etc., strict and unified data warehouse hierarchical modeling of offline data and real-time data is completed, following the data warehouse hierarchical call specifications, to ensure the standardization, consistency, and efficiency of data usage and management.

[0122] In an implementable solution, the data management method further includes:

[0123] Preset a data execution mechanism;

[0124] Among them, the data execution mechanism includes a unified data logging mechanism and / or a quality assurance mechanism;

[0125] Based on the data execution mechanism, corresponding data processing processes are performed on the initial data.

[0126] In this solution, a unified data logging mechanism (standard output, review of logging points, etc.), a requirement access process, and a standardization mechanism (formulation and execution) and a quality assurance mechanism (offline self-testing and online regression, etc. solutions) are built. By constructing the data execution mechanism, the effective implementation of standardized data management is ensured.

[0127] Specifically, as Figure 7 shown, it is the data management architecture corresponding to the data management method of this embodiment. This data management architecture belongs to a big data architecture that combines multiple guarantees of building technology, standards, and mechanisms. The specific integration is as follows:

[0128] (1) Offline and real-time technologies for data warehouse modeling: Introduce big data technology solutions and tools such as SPARK SQL, Structured Streaming, Hadoop (SPARK SQL, Structured Streaming, and Hadoop are all data processing tools), and stream-batch integration. Complete strict and unified offline and real-time data warehouse hierarchical modeling, follow the data warehouse hierarchical call specification, and obtain a data hierarchical model. The data hierarchical model includes ODS: Source-attached data layer / Operational DataStore; DWD: Data Warehouse Detail; DWM: Data Warehouse Middle; DWS: Data Warehouse Service; DIM: Dimension table.

[0129] New technology switching for application layer pages: All pages are switched to the application data layer ADS layer after the unified data warehouse reconstruction. Among them, the application data layer ADS can only converge data from the data service layer DWS, and it is not allowed for the application data layer ADS to reprocess data from the operational data layer ODS. The highly aggregated data layer preferentially calls the data of the lightly aggregated data layer and tries to avoid fetching data from the data detail layer, that is, the data warehouse call specification is unified to ensure data consistency.

[0130] (2) Data governance: Complete the construction of data dictionaries and data lineage in data governance; Pages to be launched: Automatic indicator verification to determine whether they already exist or are newly added, etc.; Indicators: Visualization, clear and unique definition, and searchability.

[0131] (3) Data integration: Complete the construction of a heterogeneous data configuration synchronization tool in data integration;

[0132] (4) Data synchronization tool: An online data source and an offline data warehouse unified and configurable, full-volume and real-time, offline and online data synchronization tool; Encapsulation of a unified vehicle-side binary parsing module, introduction of a unified and stable transmission tool.

[0133] (5) Construction mechanism: Build a unified data logging mechanism, requirement access process and standardization mechanism, offline self-testing and online regression, etc. Among them, the unified data logging mechanism: Standardized output, review of logging; Requirement access process and standardization mechanism: Formulate and execute; Quality assurance mechanism: Offline self-testing plan, online regression, etc.

[0134] (6) Establish data standards: Data standards are the normative constraints formulated by an enterprise to ensure the consistency and accuracy of internal and external data usage and exchange; the goal of data standard management is to achieve the management of data integrity, effectiveness, consistency, standardization, openness, and sharing through unified data standard formulation and release, combined with means such as institutional constraints and system control, providing a management basis for data asset management.

[0135] It should be noted that, before a certain historical node, Figure 1 the original statistical caliber data under the old framework in Figure 7 After a certain historical node,

[0136] In this embodiment, a complete data processing flow should include the following tasks (see the table below). Each task generally corresponds to an SQL (Structured Query Language) file, and the core components of each task include: creating a physical table or an external table of the target cluster (create table), writing data (insert into xxx select...). The following table lists the composition of a complete data processing task:

[0137]

[0138]

[0139] In addition, based on the six evaluation criteria for data quality: consistency, accuracy, timeliness, compliance, stability, and completeness, a comprehensive evaluation of the data is carried out. The big data solution for the quarter-level fusion update of new and old data based on offline and real-time full-volume data ensures the consistency, timeliness, accuracy, compliance, stability, and completeness of the data.

[0140] The data management architecture under the data management solution in this embodiment can complete the construction of a complete data systemization for decision-making, analysis, and operation, build a unified data access, data engine, data mart, and visualization platform, and optimize the overall delivery quality, delivery efficiency, and user experience.

[0141] Embodiment 3

[0142] As Figure 8 shown, the data management device in this embodiment includes:

[0143] An initial data acquisition module 91, configured to acquire initial data from the data operation layer ODS based on data acquisition requirements;

[0144] Among them, the initial data includes offline data and real-time data;

[0145] The first full - volume data acquisition module 92 is used to acquire the first full - volume data corresponding to the first data scheduling level at the current moment in the data detail layer DWD based on the initial data;

[0146] The first incremental data acquisition module 93 is used to generate the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment in the data detail layer DWD based on the initial data, where N is a positive integer;

[0147] Among them, the first data scheduling level and the second data scheduling are used to respectively call the data belonging to the corresponding time span, and the time span of the second data scheduling level is less than that of the first data scheduling level;

[0148] Generally, the data within the most recent 2 days is taken. Of course, the data within the most recent 1 day or the most recent 3 days can also be taken; the value of N can be determined or adjusted according to the actual application situation.

[0149] Preferably, the first data scheduling level corresponds to data scheduling on a daily basis (or daily - level scheduling), and the second data scheduling level corresponds to data scheduling on a quarter - hour basis (or quarter - hour - level scheduling, corresponding to a 15 - minute data slice).

[0150] The first target full - volume acquisition module 94 is used to acquire the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS according to the first full - volume data, the first incremental data, and the preset incremental generation rule;

[0151] The data output module 95 is used to control the application data layer ADS to call and output the first target full - volume data from the data service layer DWS.

[0152] Under the data management scheme in this embodiment, the data management architecture can complete the construction of a complete data systemization for decision - making, analysis, and operation, build a unified data access, data engine, data mart, and visualization platform, and optimize the overall delivery quality, delivery efficiency, and user experience.

[0153] In an implementable solution, the data management device further includes:

[0154] The data source unit construction module is used to construct several types of data source units based on different data classification criteria;

[0155] Among them, in the autonomous driving scenario, several types of data source units include business data units, vehicle - end data units, cloud - end data units, road - test data units, etc., that is, different types of source data are respectively stored in the corresponding units for overall management and subsequent calls, etc.

[0156] The heterogeneous data synchronization processing module is used to synchronize the corresponding heterogeneous data in different types of data source units by using a preset data synchronization tool to obtain the first original data after synchronization;

[0157] The data unification processing module is used to process the first original data by using a preset unification processing tool to obtain the second original data that meets the preset normalization requirements;

[0158] The data synchronization module is used to synchronize the second original data to the data operation layer ODS;

[0159] The initial data acquisition module 91 is used to extract the initial data that meets the data acquisition requirements from the second original data in the data operation layer ODS.

[0160] In this solution, after heterogeneous data synchronization processing and unified normalization processing of different types of data sources, they are introduced into the data operation layer ODS to obtain the initial data for subsequent direct calls, ensuring the accuracy and efficiency of subsequent data processing.

[0161] In an implementable solution, the preset unified processing tool includes a preset binary parsing tool and a preset transmission service tool that are processed in sequence.

[0162] In an implementable solution, the data management device further includes:

[0163] The second target full amount acquisition module is used to acquire the second target full amount data corresponding to the first data scheduling level in the data service layer DWS;

[0164] The data output module 95 is further used to control the application data layer ADS to call and output the first target full amount data and the second target full amount data from the data service layer DWS.

[0165] In this solution, the application data layer ADS directly calls the full amount data in the data service layer DWS (or the data aggregation layer), that is, it restricts the application data layer ADS to preferentially call the data in the common layer and does not allow the application data layer ADS to reprocess the data from the data operation layer ODS, standardizes the data warehouse data call specification, thus effectively ensuring data consistency, ensuring the reliability of data storage and management, improving the quality and efficiency of data delivery, and enhancing the user's credibility.

[0166] In an implementable solution, the first target full amount acquisition module 94 includes:

[0167] The second full amount data acquisition unit is used to perform data fusion processing on the first full amount data and the first incremental data to obtain the second full amount data corresponding to the second data scheduling level in the acquisition data detail layer DWD;

[0168] The third full - volume data acquisition unit is used to acquire the third full - volume data corresponding to the first data scheduling level in the data service layer DWS;

[0169] The first target full - volume data acquisition unit is used to acquire the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS based on the second full - volume data and the third full - volume data.

[0170] In this solution, by fusing the first full - volume data at the day - level in the data detail layer DWD and the first incremental data at the quarter - hour level in the most recent 2 days, the second full - volume data at the quarter - hour level in the most recent 2 days in the data detail layer DWD is obtained; for the data service layer DWS, based on the third full - volume data at the day - level in this layer and the second full - volume data at the quarter - hour level in the most recent 2 days in the data detail layer DWD, the first target full - volume data at the quarter - hour level in the most recent 2 days in the data service layer DWS is obtained for output by the application data layer ADS, thus achieving the fusion and update of new and old data at the quarter - hour level, ensuring data consistency, improving the quality and efficiency of data delivery, and enhancing the credibility of users.

[0171] In an implementable solution, the first target full - volume data acquisition unit includes:

[0172] The comparison result acquisition subunit is used to compare the second full - volume data and the third full - volume data to obtain a comparison result;

[0173] The actual incremental information generation subunit is used to generate the actual incremental information corresponding to the data service layer DWS based on the comparison result;

[0174] The first target full - volume data acquisition subunit is used to determine the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS based on the actual incremental information.

[0175] In this solution, by comparing the second full - volume data and the third full - volume data to determine the actual incremental information that has changed, identifying which data are deleted data, which are new data, and which are data with changes, etc., and finding the deleted, new, and modified records, the accuracy of obtaining the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS is ensured to meet the data management requirements.

[0176] In an implementable solution, when there is the second full - volume data and there is no third full - volume data for the first data, the comparison result acquisition subunit is used to generate the first marking information indicating that the first data is deleted data;

[0177] When there is no second full - volume data but there is third full - volume data in the second data, second marker information indicating that the second data is new data is generated;

[0178] When there is second full - volume data and there is third full - volume data in the third data, third marker information indicating that the third data is changed data is generated;

[0179] In this solution, based on the second full - volume data and the third full - volume data, the deleted, new, and modified records are found and corresponding marker processing is performed on them, which is convenient for distinction and management.

[0180] The data management device further includes:

[0181] An incremental table generation module, configured to generate an incremental table corresponding to the data service layer DWS based on the first marker information, the second marker information, and the third marker information.

[0182] In this solution, different marker information of the found deleted, new, and modified data is automatically used to generate an incremental table corresponding to the current data service layer DWS, which is accurate and intuitive. While improving the quality and efficiency of data delivery, the user experience is also enhanced.

[0183] Specifically, in combination with Figure 3 and Figure 4 , for data synchronization in the data operation layer ODS: Data from each business data source is synchronized from the business database to the OLAP ODS cluster (i.e., the data operation layer ODS) in real - time or offline via a unified data synchronization tool:

[0184] Generation of data in the data detail layer DWD: In the OLAP ODS cluster, dwd_df (daily full - volume data in the data detail layer) and dwd_qi (quarter - hour - level incremental data in the data detail layer) are routinely generated respectively; if the data volume is small, dwd_qf (quarter - hour - level full - volume data in the data detail layer) can be directly generated, and all subsequent processes are correspondingly simplified; while generating dwd_df and dwd_qi, historical slice merging is also responsible for;

[0185] Generation of full - volume data in the data service layer DWS: dwd_df is routinely exported to AFS, processed by SPARK and written to UDW, and dws_df is routinely generated and imported into PALO; dwd_df directly routinely generates dws_df (daily full - volume data in the data service layer), and is written to the ADS cluster (i.e., the application data layer ADS) through an external table;

[0186] Incremental data generation of the Data Warehouse Service layer DWS: Use dwd_qi and dws_df of T + a (a is a constant, for example, when a takes the value of 2 for the most recent 2 days) to generate dws_qf (quarterly full-volume data of the Data Warehouse Service layer, not stored in the database), and then combine dws_df to generate dws_qi (quarterly incremental data of the Data Warehouse Service layer), which is written to the ADS cluster through the external table.

[0187] Data generation of the Application Data layer ADS: Obtain ads_qf (quarterly full-volume data of the Application Data layer) based on dws_df and dws_qi.

[0188] Specifically, combined with Figure 5 the execution scheduling flowchart shown below, the generation rules corresponding to each of the above data layers are as follows:

[0189] dwd_qf = dwd_df(T + a) + dwd_qi(ad); a is a constant and can take the value of 2;

[0190] dws_qi = dwd_qf + dws_df(T + a)

[0191] ads_qf = dws_qf = dws_df(T + a) + dws_qi(ad)

[0192] To ensure the generation of slices of the quarterly full-volume data in a timely manner after the change of irregular historical data, the following processing method is adopted:

[0193] For offline data, the data details layer DWD and the Data Warehouse Service layer DWS after daily full cleaning and aggregation are imported into the real-time OLAP Doris system, and no 15 data shards are established to solve the problem of insufficient cluster capacity due to large data volume;

[0194] For real-time data: The 15-minute slices only record the change records within 2 days: those that do not exist are marked as invalid, newly added, or changed, and those without changes are not recorded;

[0195] Specifically, retain the data change records of each period to solve the problem of inconsistent page data caused by the asynchronous timeliness of the full-link update of the underlying ADS of each page. For the Data Warehouse Service layer DWS with many dimensions and large data volume, only incremental data is stored in the OLAP Doris system to improve the overall efficiency of the data stream.

[0196] DWS temporary full volume = DWD full volume(T + 2) + DWD increment (changes in the most recent 2 days)

[0197] DWS Increment = DWS Full Volume (T+2) & DWS Temporary Full Volume; where the generation process of the DWS increment table is as follows: Find the deleted, newly added, and modified records from DWS Full Volume (T+2) and DWS Temporary Full Volume to generate an effective DWS increment table. Among them:

[0198] Determine the data that belongs to the deleted ones: DWS Full Volume (T+2) has it, DWS Temporary Full Volume doesn't have it, mark it as invalid 0.

[0199] Determine the data that belongs to the newly added ones: DWS Full Volume (T+2) doesn't have it, DWS Temporary Full Volume has it, mark it as valid 1.

[0200] Determine the data that belongs to the changed ones: DWS Full Volume (T+2) has it, DWS Temporary Full Volume has it, take the data corresponding to the latest time of this record in DWS Temporary Full Volume, and mark it as valid 1.

[0201] Combined Figure 6 As shown, after the above data processing process, after the irregular historical data is changed, the timeliness of the quarter-hourly full volume data generates slices. The 15-minute slices only record the changed records within 2 days, and the non-existent ones are marked as invalid, newly added, and changed ones, and the unchanged ones are not recorded; that is, the offline and real-time solutions are combined to ensure the timeliness of obtaining the 15-minute (quarter-hourly) full volume data.

[0202] In an implementable solution, the data management device further includes:

[0203] A data governance rule preset module, used to preset data governance rules;

[0204] A data governance processing module, used to perform data governance operations on the first target full volume data and the second target full volume data output by the application data layer ADS based on the data governance rules.

[0205] Specifically, through the constructed data governance rules, complete the construction of the data dictionary and lineage in data governance; among them, automatically verify the indicators of the page to be launched to judge whether it already exists or is newly added, etc.; the indicators include but are not limited to visualization, clear and unique definition, and retrievability; through data governance to ensure the overall standardization and effectiveness of data management.

[0206] In an implementable solution, the offline data and the real-time data adopt hierarchical modeling in the data warehouse and meet the preset hierarchical call specifications;

[0207] Specifically, adopt some big data technology solutions and tools, etc., to complete the strict and unified data warehouse hierarchical modeling of offline data and real-time data, and follow the data warehouse hierarchical call specifications to ensure the standardization, consistency, and efficiency of data use and management.

[0208] In an implementable solution, the data management device further includes:

[0209] A data execution mechanism construction module, configured to preset and construct a data execution mechanism;

[0210] Wherein, the data execution mechanism includes a unified data logging mechanism and / or a quality assurance mechanism;

[0211] A data processing module, configured to perform corresponding data processing processes on the initial data based on the data execution mechanism.

[0212] In this solution, a unified data logging mechanism (such as standardized output and audit of logging points), a requirement access process, and a standardization mechanism (formulation and execution) are established, as well as a quality assurance mechanism (such as offline self-testing and online regression). By constructing the data execution mechanism, the effective implementation of standardized data management is ensured.

[0213] It should be noted that the working principle of the data management device in this embodiment is the same as that of the data management method in Embodiment 1, so it will not be elaborated here.

[0214] Embodiment 3

[0215] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0216] Figure 9 The schematic block diagram of an example electronic device 900 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0217] As Figure 9 shown, the device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 902 or the computer program loaded from the storage unit 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0218] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as a keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as a disk, optical disc, etc.; and communication unit 909, such as a network card, modem, wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0219] Computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 901 executes the various methods and processes described above, such as the above-mentioned method. For example, in some embodiments, the above method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the above-described method can be executed. Alternatively, in other embodiments, computing unit 901 can be configured to execute the above method in any other suitable manner (e.g., by means of firmware).

[0220] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0221] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0222] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0223] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0224] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0225] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0226] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0227] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data management method, the data management method comprising: Obtaining initial data from the Operational Data Store (ODS) of the data operation layer based on data acquisition requirements; wherein the initial data includes offline data and real-time data; Based on the initial data, obtaining first full amount data corresponding to a first data scheduling level at the current moment in the Data Warehouse Detail (DWD) layer, and first incremental data corresponding to a second data scheduling level within the most recent N days before the current moment, where N is a positive integer; wherein the first data scheduling level and the second data scheduling level are used to respectively call data belonging to the corresponding time spans, and the time span of the second data scheduling level is less than the time span of the first data scheduling level; According to the first full amount data, the first incremental data, and a preset incremental generation rule, obtaining first target full amount data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer; Controlling the Application Data Store (ADS) layer to call and output the first target full amount data from the Data Warehouse Service (DWS) layer.

2. The data management method according to claim 1, the step of controlling the Application Data Store (ADS) layer to call and output the first target full amount data from the Data Warehouse Service (DWS) layer includes: Obtaining second target full amount data corresponding to the first data scheduling level in the Data Warehouse Service (DWS) layer; Controlling the Application Data Store (ADS) layer to call and output the first target full amount data and the second target full amount data from the Data Warehouse Service (DWS) layer.

3. The data management method according to claim 2, the step of obtaining first target full amount data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer according to the first full amount data, the first incremental data, and a preset incremental generation rule includes: Performing data fusion processing on the first full amount data and the first incremental data to obtain second full amount data corresponding to the second data scheduling level in the Data Warehouse Detail (DWD) layer; Obtaining third full amount data corresponding to the first data scheduling level in the Data Warehouse Service (DWS) layer; Based on the second full amount data and the third full amount data, obtaining the first target full amount data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer.

4. The data management method according to claim 3, the step of obtaining the first target full amount data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer based on the second full amount data and the third full amount data includes: Comparing the second full amount data and the third full amount data to obtain a comparison result; Generating actual incremental information corresponding to the Data Warehouse Service (DWS) layer based on the comparison result; Based on the actual incremental information, determining the first target full amount data corresponding to the second data scheduling level within the most recent N days in the Data Warehouse Service (DWS) layer.

5. The data management method according to claim 4, the step of comparing the second full amount data and the third full amount data to obtain a comparison result includes: When there is the first data with the second full - volume data present and the third full - volume data absent, then generate first marking information indicating that the first data is deleted data; When there is the second data with the second full - volume data absent and the third full - volume data present, then generate second marking information indicating that the second data is newly added data; When there is the third data with the second full - volume data present and the third full - volume data present, then generate third marking information indicating that the third data is changed data.

6. The data management method according to claim 5, wherein the data management method further comprises: Generate an incremental table corresponding to the data service layer DWS based on the first marking information, the second marking information, and the third marking information.

7. The data management method according to any one of claims 1 - 6, wherein the first data scheduling level corresponds to data scheduling on a daily basis, and the second data scheduling level corresponds to data scheduling on a quarterly basis.

8. The data management method according to claim 1, before the step of obtaining initial data from the data operation layer ODS based on data acquisition requirements, further comprising: Construct several types of data source units based on different data classification criteria; Use a preset data synchronization tool to synchronize heterogeneous data corresponding to different types of the data source units to obtain synchronized first - hand raw data; Process the first - hand raw data using a preset unified processing tool to obtain second - hand raw data that meets preset normalization requirements; Synchronize the second - hand raw data to the data operation layer ODS.

9. The data management method according to claim 8, the step of obtaining initial data from the data operation layer ODS based on data acquisition requirements, comprises: Extract the initial data that meets the data acquisition requirements from the raw data in the data operation layer ODS.

10. The data management method according to claim 8, wherein the preset unified processing tool comprises a preset binary parsing tool and a preset transmission service tool for sequential processing.

11. The data management method according to claim 2, wherein the data management method further comprises: Preset data governance rules; Perform data governance operations on the first target full - volume data and the second target full - volume data output by the application data layer ADS based on the data governance rules.

12. The data management method according to claim 1, wherein the offline data and the real - time data adopt hierarchical modeling in the data warehouse and meet preset hierarchical call specifications.

13. The data management method according to claim 1, wherein the data management method further comprises: Preset a data execution mechanism; Wherein, the data execution mechanism includes a unified data logging mechanism and / or a quality assurance mechanism; Perform corresponding data processing processes on the initial data based on the data execution mechanism.

14. A data management device, the data management device comprises: An initial data acquisition module, configured to obtain initial data from the data operation layer ODS based on data acquisition requirements; Wherein, the initial data includes offline data and real - time data; The first full - volume data acquisition module is used to acquire the first full - volume data corresponding to the first data scheduling level at the current moment in the data detail layer DWD based on the initial data; The first incremental data acquisition module is used to generate the first incremental data corresponding to the second data scheduling level within the most recent N days before the current moment in the data detail layer DWD based on the initial data, where N is a positive integer; Among them, the first data scheduling level and the second data scheduling are used to respectively call the data belonging to the corresponding time span, and the time span of the second data scheduling level is less than the time span of the first data scheduling level; The first target full - volume acquisition module is used to acquire the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS according to the first full - volume data, the first incremental data, and a preset incremental generation rule; The data output module is used to control the application data layer ADS to call and output the first target full - volume data from the data service layer DWS.

15. The data management device according to claim 14, wherein the data management device further comprises: The second target full - volume acquisition module is used to acquire the second target full - volume data corresponding to the first data scheduling level in the data service layer DWS; The data output module is further used to control the application data layer ADS to call and output the first target full - volume data and the second target full - volume data from the data service layer DWS.

16. The data management device according to claim 15, wherein the first target full - volume acquisition module comprises: The second full - volume data acquisition unit is used to perform data fusion processing on the first full - volume data and the first incremental data to obtain the second full - volume data corresponding to the second data scheduling level in the data detail layer DWD; The third full - volume data acquisition unit is used to acquire the third full - volume data corresponding to the first data scheduling level in the data service layer DWS; The first target full - volume data acquisition unit is used to acquire the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS based on the second full - volume data and the third full - volume data.

17. The data management device according to claim 16, wherein the first target full - volume data acquisition unit comprises: The comparison result acquisition subunit is used to compare the second full - volume data and the third full - volume data to obtain a comparison result; The actual incremental information generation subunit is used to generate the actual incremental information corresponding to the data service layer DWS based on the comparison result; The first target full - volume data acquisition subunit is used to determine the first target full - volume data corresponding to the second data scheduling level within the most recent N days in the data service layer DWS based on the actual incremental information.

18. The data management device according to claim 17, wherein the comparison result acquisition subunit is used to generate the first marker information indicating that the first data is deleted data when there is the second full - volume data and there is no third full - volume data in the first data; When there is no such second full - volume data and there is second data of the third full - volume data, second marker information indicating that the second data is new data is generated; When there is such second full - volume data and there is third data of the third full - volume data, third marker information indicating that the third data is changed data is generated.

19. The data management device according to claim 18, wherein the data management device further comprises: An incremental table generation module, configured to generate an incremental table corresponding to the data service layer DWS based on the first marker information, the second marker information, and the third marker information.

20. The data management device according to any one of claims 14 - 19, wherein the first data scheduling level corresponds to data scheduling on a daily basis, and the second data scheduling level corresponds to data scheduling on a quarter - hour basis.

21. The data management device according to claim 14, wherein the data management device further comprises: A data source unit construction module, configured to construct several types of data source units based on different data classification criteria; A heterogeneous data synchronization processing module, configured to use a preset data synchronization tool to synchronize corresponding heterogeneous data in different types of the data source units to obtain first original data after synchronization; A data unified processing module, configured to process the first original data using a preset unified processing tool to obtain second original data meeting preset normalization requirements; A data synchronization module, configured to synchronize the second original data to the data operation layer ODS.

22. The data management device according to claim 21, wherein the initial data acquisition module is configured to extract the initial data meeting the data acquisition requirements from the original data in the data operation layer ODS.

23. The data management device according to claim 21, wherein the preset unified processing tool comprises a preset binary parsing tool and a preset transmission service tool that are processed in sequence.

24. The data management device according to claim 15, wherein the data management device further comprises: A data governance rule preset module, configured to preset data governance rules; A data governance processing module, configured to perform data governance operations on the first target full - volume data and the second target full - volume data output by the application data layer ADS based on the data governance rules.

25. The data management device according to claim 14, wherein the offline data and the real - time data adopt hierarchical modeling in the data warehouse and meet preset hierarchical call specifications.

26. The data management device according to claim 14, wherein the data management device further comprises: A data execution mechanism construction module, configured to preset and construct a data execution mechanism; wherein the data execution mechanism includes a unified data logging mechanism and / or a quality assurance mechanism; A data processing module, configured to perform corresponding data processing processes on the initial data based on the data execution mechanism.

27. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-13.

29. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Lake and warehouse integration-based data processing system and method

    CN115599871A

  • Network operation data processing method and device, electronic equipment and medium

    CN115664992A