Big data job scheduling method and system based on field consanguinity
By building field blood relationship analysis, visualization and dependency optimization modules in big data analysis technology, and re-arrange online job flows, the problem of difficult to meet the timeliness requirements of online job flows is solved, and an efficient and transparent big data job scheduling system is realized.
Patent Information
- Application Number
- CN202510319932.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
In the existing big data analysis technology, the timeliness requirements of online job flow are difficult to meet, resulting in problems such as endless jobs or data loss.
By building a field blood relationship analysis module, a field blood relationship visualization module, and a job field blood relationship dependency analysis and optimization module, the online job flow is re-arranged, the data processing process is optimized, and resource consumption is reduced.
It realizes rapid and accurate acquisition of field blood ties, reduces online scheduled ODS jobs, improves job response speed, optimizes scheduling algorithms, reduces overload risk, and improves user experience.
Smart Images

Figure CN120144259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis. Specifically, it relates to an online job flow scheduling optimization based on field lineage in a business scenario, and particularly to a big data job scheduling method and system based on field lineage. Background Art
[0002] In the field of big data analysis, the offline and online job flow orchestration is mainly based on table-level lineage, which can meet most business scenarios. However, online jobs have strong timeliness requirements. That is, within a fixed scheduling period (such as 5 minutes, 10 minutes), if a job cannot be completed, the job flow in the next scheduling period will be skipped, resulting in data loss.
[0003] The existing invention patent with the publication number CN108694082A discloses a cross-domain job flow scheduling method and system, including: selecting a job flow scheduling cluster A in the collaborative scheduling network to receive the data processing service requirements sent by the application provider; performing job flow orchestration according to the logic of the data processing service requirements and dividing them into multiple data service processing blocks; allocating the multiple data service processing blocks to multiple job flow scheduling clusters in the collaborative scheduling network according to the job flow orchestration logic for processing; each of the multiple job flow scheduling clusters processes the corresponding allocated data service processing blocks and generates data; and outputting the generated data to a predetermined job flow scheduling cluster through a federated data channel and storing it in its corresponding database.
[0004] To optimize the above problems, it is now necessary to consider reorchestrating the online job flow from a finer-grained field-level lineage. Summary of the Invention
[0005] Aiming at the defects in the prior art, the present invention provides a big data job scheduling method and system based on field lineage.
[0006] According to the big data job scheduling method and system based on field lineage provided by the present invention, the solution is as follows:
[0007] In the first aspect, a big data job scheduling method based on field lineage is provided. The method includes:
[0008] Step S1: Build a field lineage parsing module, and through the field lineage parsing module, construct and maintain the complete transfer path of field lineage relationship data to ensure accurate tracking of the entire data processing process of each field from the source to the target.
[0009] Step S2: Based on the field blood relationship data provided by the field blood relationship parsing module, build a field blood relationship visualization module, and present the field blood relationship through the field blood relationship visualization module to enhance the transparency and understandability of data flow;
[0010] Step S3: Through the field blood relationship data, construct a job field blood relationship dependency parsing and optimization module to optimize the data processing flow and reduce resource consumption.
[0011] Preferably, the step S1 includes:
[0012] Step S1.1: Determine the field blood relationship parsing method; collect field blood relationship data during SQL runtime, and the execution engine is fixed;
[0013] Step S1.2: Deploy a log collection and processing tool; deploy Filebeat in the impala environment, use Filebeat to monitor and collect impala's blood relationship log events, and forward them to Logstash. Parse these blood relationship log events into metric fields through Logstash and store them in Elasticsearch.
[0014] Preferably, the step S2 includes:
[0015] Step S2.1: Perform data preparation and format conversion; that is, obtain the field blood relationship data for display from Elasticsearch; assemble the field blood relationship data into JSON format;
[0016] Step S2.2: Visually display the field blood relationship data assembled into JSON format, that is, use LineageX software to display the field blood relationship diagram.
[0017] Preferably, the step S3 includes:
[0018] Step S3.1: Analyze the existing job flow; check the table-level blood relationship of each detail table job, and extract the blood relationship dependency information between tables by checking the blood relationship of the tables;
[0019] Step S3.2: Query the foreign key blood relationship; query the blood relationship of foreign keys in the core SQL to determine which ODS tables are relied on to construct the foreign keys;
[0020] Step S3.3: Optimize the scheduling strategy; through the table-level blood relationship and the foreign key blood relationship, judge the ODS tables that appear in both the table-level blood relationship and the foreign key blood relationship, and schedule them once a day;
[0021] Re-arrange the construction process of the online detail fact table dwd_inc according to the field blood relationship, reduce the scheduling times of the ODS tables, and improve the efficiency.
[0022] In a second aspect, a big data job scheduling system based on field lineage is provided. The system includes:
[0023] Field lineage parsing module: Construct and maintain a complete flow path of field lineage relationship data to ensure accurate tracking of the entire data processing process of each field from the source to the target;
[0024] Field lineage visualization module: Based on the field lineage relationship data provided by the field lineage parsing module, present the field lineage relationship to enhance the transparency and understandability of data flow;
[0025] Job field lineage dependency parsing and optimization module: Optimize the data processing flow and reduce resource consumption through the field lineage relationship data.
[0026] Preferably, the field lineage parsing module includes:
[0027] Module M1.1: Determine the field lineage relationship parsing method; collect field lineage relationship data during SQL runtime, and the execution engine is fixed;
[0028] Module M1.2: Deploy a log collection and processing tool; deploy Filebeat in the impala environment, use Filebeat to monitor and collect impala's lineage log events, and forward them to Logstash. Parse these lineage log events into metric fields by Logstash and store them in Elasticsearch.
[0029] Preferably, the field lineage visualization module includes:
[0030] Module M2.1: Perform data preparation and format conversion; that is, obtain the field lineage relationship data for display from Elasticsearch; assemble the field lineage relationship data into JSON format;
[0031] Module M2.2: Visually display the field lineage relationship data assembled into JSON format, that is, use LineageX software to display the field lineage relationship diagram.
[0032] Preferably, the job field lineage dependency parsing and optimization module includes:
[0033] Module M3.1: Analyze the existing job flow; check the table-level lineage relationship of each detail table job, and extract the inter-table lineage dependency information by checking the lineage relationship of the tables;
[0034] Module M3.2: Query the foreign key lineage relationship; query the foreign key lineage relationship in the core SQL to determine which ODS tables are relied on to construct the foreign key;
[0035] Module M3.3: Scheduling Policy Optimization; By judging the ODS tables that appear in both the table-level lineage and the foreign-key lineage through table-level lineage and foreign-key lineage, and scheduling them once a day;
[0036] Rearrange the construction process of the online detailed fact table dwd_inc according to the field lineage, reduce the number of ODS table schedules, and improve efficiency.
[0037] Thirdly, a computer-readable storage medium storing a computer program is provided, and when the computer program is executed by a processor, the steps of the above-mentioned big data job scheduling method based on field lineage are implemented.
[0038] Fourthly, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the above-mentioned big data job scheduling method based on field lineage are implemented.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. Through the method provided by the present invention, users can quickly and accurately obtain the lineage of fields, and reduce the ODS jobs that need to be scheduled in online scheduling, effectively improving the timeliness of online jobs and enhancing the user experience;
[0041] 2. Through the technical means of optimizing job scheduling of the present invention, the response speed of jobs is improved, which is particularly important for an online data warehouse system that requires quick response; at the same time, the online scheduling algorithm is optimized, which is crucial for business processes that need to strictly comply with deadlines; effectively reduces the overload risk caused by the increase of business jobs, meets the demand of online short jobs first, and improves the user experience.
[0042] Other beneficial effects of the present invention will be elaborated in the specific implementation manners through the introduction of specific technical features and technical solutions. Those skilled in the art should be able to understand the beneficial technical effects brought by the technical features and technical solutions through these introductions. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0044] Figure 1 It is a flowchart of the present invention;
[0045] Figure 2 It is a diagram showing field lineage. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.
[0047] An embodiment of the present invention provides a big data job scheduling method based on field lineage. Referring to Figure 1 as shown, the method specifically includes the following content:
[0048] Step S1: Build a field lineage parsing module, and through the field lineage parsing module, construct and maintain a complete transfer path of field lineage relationship data to ensure accurate tracking of the entire data processing process of each field from the source to the target.
[0049] This step specifically includes:
[0050] Step S1.1: Determine the field lineage relationship parsing method; collect field lineage relationship data during SQL runtime, and the execution engine is fixed;
[0051] Step S1.2: Deploy a log collection and processing tool;
[0052] Specifically, deploy Filebeat (Filebeat is a lightweight log and file data collector) in the impala environment. Use the wget command to download directly on the machine where impala (impala is a massively parallel processing SQL query engine) is located. After completion, unzip it, and then enter the filebeat directory to create a startup configuration file and start it. Use Filebeat to monitor and collect the lineage log events of impala, and forward them to Logstash (Logstash is an open-source server-side data processing pipeline that can collect data from multiple sources, transform the data, and then send the data to the corresponding storage). Parse these lineage log events into metric fields through Logstash and store them in Elasticsearch (Elasticsearch is an open-source distributed RESTful search and analysis engine, an extensible data storage and vector database that can solve various emerging use cases).
[0053] Step S2: Based on the field lineage relationship data provided by the field lineage parsing module, build a field lineage visualization module, and present the field lineage relationship through the field lineage visualization module to enhance the transparency and understandability of data flow.
[0054] This step specifically includes:
[0055] Step S2.1: Perform data preparation and format conversion; that is, obtain the field lineage data for display from Elasticsearch. Since the collected data includes information such as which table a field belongs to and which tables' fields are processed to this field, this information can be processed into JSON format using a script.
[0056] Step S2.2: Visually display the field lineage data assembled in JSON format, that is, use LineageX software (an open-source software that can display the field lineage) to display the field lineage diagram.
[0057] Step S3: Build a job field lineage dependency parsing and optimization module through the field lineage data to optimize the data processing flow and reduce resource consumption.
[0058] This step specifically includes:
[0059] Step S3.1: Analyze the existing job flow; check the table-level lineage of each detail table job through the metadata information of the job, and by checking the table lineage, the inter-table lineage dependency information can be extracted.
[0060] Step S3.2: Query the foreign key lineage; query the foreign key lineage in the core SQL to determine which ODS (the full name of ODS is Operational Data Store, which is an "enterprise-wide, subject-oriented" data store for operational data. It is the layer closest to the data in the data source. After the data in the data source is extracted, cleaned, and transmitted, that is, after the so-called ETL process, it is loaded into this layer. The data in this layer is generally classified according to the classification method of the source business system.) tables are relied on to build the foreign key. By checking the foreign key relationships in the SQL, it can be known which tables do not need to be scheduled in each scheduling cycle.
[0061] Step S3.3: Optimize the scheduling strategy; in Step S3.1, the table lineage is obtained, and then in Step S3.2, the foreign key lineage is specifically extracted. It can be easily judged that for the ODS tables that appear in both the table-level lineage and the foreign key lineage, they are scheduled once a day.
[0062] Rearrange the construction process of the online detail fact table dwd_inc according to the field lineage (schedule the tables with foreign keys in the SQL daily, and do not need to be scheduled in each cycle), reduce the number of ODS table schedules, and improve efficiency.
[0063] The present invention also provides a big data job scheduling system based on field lineage. The big data job scheduling system based on field lineage can be implemented by executing the process steps of the big data job scheduling method based on field lineage. That is, those skilled in the art can understand the big data job scheduling method based on field lineage as a preferred implementation manner of the big data job scheduling system based on field lineage. The system specifically includes:
[0064] Field lineage parsing module: Build and maintain a complete flow path of field lineage relationship data to ensure accurate tracking of the entire data processing process of each field from the source to the target.
[0065] The field lineage parsing module includes:
[0066] Module M1.1: Determine the field lineage relationship parsing method; collect field lineage relationship data during SQL runtime, and the execution engine is fixed;
[0067] Module M1.2: Deploy a log collection and processing tool; specifically, deploy Filebeat (Filebeat is a lightweight log and file data collector) in the impala environment. Use the wget command to download directly on the machine where impala (impala is a massively parallel processing SQL query engine) is located. After completion, unzip it, and then enter the filebeat directory to create a startup configuration file and start it. Use Filebeat to monitor and collect impala's lineage log events and forward them to Logstash (Logstash is an open-source server-side data processing pipeline that can collect data from multiple sources, transform the data, and then send the data to the corresponding storage). Parse these lineage log events into metric fields through Logstash and store them in Elasticsearch (Elasticsearch is an open-source distributed RESTful search and analysis engine, an extensible data store, and a vector database that can solve various emerging use cases).
[0068] Field lineage visualization module: Based on the field lineage relationship data provided by the field lineage parsing module, present the field lineage relationship to enhance the transparency and understandability of data flow.
[0069] The field lineage visualization module includes:
[0070] Module M2.1: Perform data preparation and format conversion; that is, obtain the field lineage relationship data for display from Elasticsearch; since the collected data has information about which table a field belongs to and which table fields are processed into this field, these information can be processed into JSON format using a script;
[0071] Module M2.2: Visualize the field lineage data assembled into JSON format, that is, use LineageX software (an open-source software that can display field lineage) to display the field lineage graph.
[0072] Job Field Lineage Dependency Parsing and Optimization Module: Optimize the data processing flow and reduce resource consumption through field lineage data.
[0073] The job field lineage dependency parsing and optimization module includes:
[0074] Module M3.1: Analyze the existing job flow; check the table-level lineage of each detail table job through the metadata information of the job, check the lineage of the tables, and be able to extract the inter-table lineage dependency information;
[0075] Module M3.2: Query the foreign key lineage; query the foreign key lineage in the core SQL to determine which ODS (the full name of ODS is Operational Data Store, an operational data store. "Subject-oriented", the data operation layer, also called the ODS layer, is the layer closest to the data in the data source. After the data in the data source is extracted, cleaned, and transmitted, that is, after the so-called ETL, it is loaded into this layer. The data in this layer is generally classified according to the classification method of the source business system.) tables are relied on to build the foreign key, and checking the foreign key relationship in the SQL can know which tables do not need to be scheduled in each scheduling cycle;
[0076] Module M3.3: Optimize the scheduling strategy; in step S3.1, the table lineage is available, and then in step S3.2, the foreign key lineage is specifically extracted. It can be easily judged that for the ODS tables that appear in both the table-level lineage and the foreign key lineage, they are scheduled once a day;
[0077] Rearrange the construction process of dwd_inc according to the field lineage (schedule the foreign key tables in the SQL daily and do not need to be scheduled in each cycle), reduce the number of ODS table schedules, and improve efficiency.
[0078] Next, a more specific description of the present invention will be given.
[0079] The present invention provides a big data job scheduling system based on field lineage. The system includes three main modules: a field lineage parsing module, a field lineage visualization module, and a job field lineage dependency parsing and optimization module. The following is a specific description of each module and its role in the entire solution.
[0080] 1. Field Lineage Parsing Module:
[0081] Table-level blood relationship has extensive applications at the data warehouse level, such as offline / online job flow updates, job scheduling, dimension topology diagrams of dimension tables / detail fact tables, etc. There are currently two general methods for parsing field blood relationship:
[0082] Hard parsing of SQL: The advantage is that when there are few computing engines and tools and the syntax is compatible, the blood relationship of tables and fields can be directly parsed through tools; the disadvantage is that when the metadata changes, the field blood relationship may be inaccurate.
[0083] Collecting field blood relationship during SQL execution: The advantage is that the runtime state and information are the most accurate and there will be no SQL parsing syntax errors; the disadvantage is that parsing modules need to be developed for different engines, and the parsing speed must be fast enough to avoid affecting normal SQL execution.
[0084] Given that the execution engine of this big data warehouse is fixed as impala, the method of collecting blood relationship during SQL runtime is selected.
[0085] Specific implementation:
[0086] The lexical and syntax parsing of impala is developed through cup and flex. The source code for parsing field blood relationship is ColumnLineageGraph.java. However, impala does not provide an interface for querying field blood relationship. If you want to obtain the field blood relationship during SQL runtime, you can only parse the corresponding blood relationship logs during SQL execution. The specific steps are as follows:
[0087] 1) Deploy Filebeat: Deploy Filebeat in the impala environment.
[0088] 2) Monitor and collect logs: Filebeat monitors the field blood relationship logs of impala, collects log events, and forwards them to Logstash for indexing.
[0089] 3) Parse log data: Logstash parses these field blood relationship logs and extracts key information (such as table names, field names, operation types, etc.).
[0090] 4) Store data: Store the parsed data in Elasticsearch for subsequent querying and use.
[0091] 2. Field blood relationship visualization module:
[0092] Implementation steps:
[0093] 1) Obtain display information: Obtain information related to the field blood relationship to be displayed from Elasticsearch.
[0094] 2) Assembly data format: Assemble the lineage relationship into a specific JSON format.
[0095] 3) Visual display: Use LineageX software to display the field lineage graph, enabling users to intuitively understand the data flow path.
[0096] Field lineage display combines the previously obtained lineage relationship with LineageX Figure 2 to achieve the following functions: When the mouse cursor hovers over the target field, the source fields that make up the current field will be highlighted. Additionally, the arrangement between tables is automatically generated and does not require manual adjustment, making the field lineage relationship clear and straightforward.
[0097] 3. Job field lineage dependency parsing and optimization module:
[0098] Case analysis and optimization measures:
[0099] Taking the online detail fact table dwd_inc as an example, it is constructed every 10 minutes. Before constructing dwd_inc, all 16 ODS tables it depends on need to be fully extracted. However, not all ODS tables need to be re-extracted in each scheduling cycle. Some ODS tables are foreign keys (FKs, which are the primary keys of the dimension tables) in the core fields of the detail fact table. According to this FK, the corresponding dimension tables are associated, and the dimension tables are constructed once a day. Therefore, these ODS tables for constructing the FKs only need to be scheduled once a day, rather than in each cycle.
[0100] dwd_inc has 134 core fields, 40 of which are foreign keys. There are a total of 16 ODS tables it depends on, 5 of which are used to construct foreign keys. If the online dwd_inc construction is rearranged based on the field lineage relationship, the number of ODS tables that need to be scheduled in each cycle is 11, and the ODS job volume is reduced by 31%. The specific implementation plan is as follows:
[0101] 1) Update the job flow daily: When the online job flow is updated daily, check the table-level lineage relationship of each detail table job.
[0102] 2) Query the foreign key lineage relationship: Query the lineage relationship of the foreign keys in the core SQL of the job.
[0103] 3) Optimize the scheduling strategy: For ODS tables that appear in both the table-level lineage relationship and the foreign key lineage relationship, only need to be scheduled once a day, rather than a full extraction in each cycle.
[0104] Through the field lineage analysis module, the system can accurately track the entire process of data processing for each field; the field lineage visualization module enables these complex lineage relationships to be intuitively displayed; finally, the job field lineage dependency analysis and optimization module optimizes the data processing flow based on the detailed lineage information, significantly improving the scheduling efficiency and resource utilization rate. The three work together to form an efficient, transparent, and easy-to-maintain big data job scheduling system.
[0105] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as both software modules for implementing the method and the structures within the hardware component.
[0106] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A big data job scheduling method based on field lineage, characterized in that: include: Step S1: Build a field lineage analysis module, and use the field lineage analysis module to build and maintain the complete flow path of field lineage relationship data to ensure accurate tracking of the entire data processing process of each field from source to target; Step S2: Based on the field lineage relationship data provided by the field lineage analysis module, a field lineage visualization module is constructed to present the field lineage relationship through the field lineage visualization module, thereby enhancing the transparency and comprehensibility of data flow; Step S3: Build a job field lineage dependency analysis and optimization module through field lineage relationship data to optimize the data processing flow and reduce resource consumption.
2. The method for scheduling big data jobs based on field lineage according to claim 1, characterized in that: The step S1 comprises: Step S1.1: Determine the field lineage relationship analysis method; use SQL runtime to collect field lineage relationship data, and the execution engine is fixed; Step S1.2: Deploy log collection and processing tools; deploy Filebeat in the Impala environment, use Filebeat to monitor and collect Impala's lineage log events, and forward them to Logstash. Use Logstash to parse these lineage log events into indicator fields and store them in Elasticsearch.
3. The method for scheduling big data jobs based on field lineage according to claim 1, characterized in that: The step S2 comprises: Step S2.1: Perform data preparation and format conversion; that is, obtain field lineage relationship data for display from Elasticsearch; assemble the field lineage relationship data into JSON format; Step S2.2: Visualize the field lineage relationship data assembled into JSON format, that is, use LineageX software to display the field lineage relationship diagram.
4. The method for scheduling big data jobs based on field lineage according to claim 1, characterized in that: The step S3 comprises: Step S3.1: Analyze the existing job flow; check the table-level blood relationship of each detailed table job, and extract the blood dependency information between tables by checking the blood relationship of the table; Step S3.2: Query the blood relationship of foreign keys; query the blood relationship of foreign keys in the core SQL and determine which ODS tables are dependent on for building foreign keys; Step S3.3: Scheduling strategy optimization: through table-level kinship and foreign key kinship, determine the ODS table that appears in both table-level kinship and foreign key kinship, and schedule once a day; Rearrange the construction process of the online detail fact table dwd_inc according to the field relationship to reduce the number of ODS table scheduling times.
5. A big data job scheduling system based on field lineage, characterized in that: include: Field lineage analysis module: builds and maintains the complete flow path of field lineage data to ensure accurate tracking of the entire data processing process of each field from source to target; Field lineage visualization module: Based on the field lineage data provided by the field lineage analysis module, the field lineage relationship is presented to enhance the transparency and comprehensibility of data flow; Job field lineage dependency analysis and optimization module: optimizes data processing flow and reduces resource consumption through field lineage relationship data.
6. The big data job scheduling system based on field lineage according to claim 5 is characterized in that: The field lineage analysis module includes: Module M1.1: Determine the method for parsing field lineage relationships; use SQL runtime to collect field lineage relationship data, and the execution engine is fixed; Module M1.2: Deploy log collection and processing tools; deploy Filebeat in the Impala environment, use Filebeat to monitor and collect Impala's lineage log events, and forward them to Logstash. Use Logstash to parse these lineage log events into indicator fields and store them in Elasticsearch.
7. The big data job scheduling system based on field lineage according to claim 5 is characterized in that: The field lineage visualization module includes: Module M2.1: Perform data preparation and format conversion; that is, obtain field lineage relationship data for display from Elasticsearch; assemble the field lineage relationship data into JSON format; Module M2.2: Visualize the field lineage relationship data assembled in JSON format, that is, use LineageX software to display the field lineage relationship diagram.
8. The big data job scheduling system based on field lineage according to claim 5 is characterized in that: The job field lineage dependency analysis and optimization module includes: Module M3.1: Analyze the existing job flow; check the table-level relationship of each detailed table job, and extract the relationship dependency information between tables by checking the relationship of the tables; Module M3.2: Query the blood relationship of foreign keys; query the blood relationship of foreign keys in core SQL and determine which ODS tables are dependent on to build foreign keys; Module M3.3: Scheduling strategy optimization; through table-level kinship and foreign key kinship, determine the ODS table that appears in both table-level kinship and foreign key kinship, and schedule once a day; Rearrange the construction process of the online detail fact table dwd_inc according to the field relationship to reduce the number of ODS table scheduling times.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the big data job scheduling method based on field lineage described in any one of claims 1 to 4 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the big data job scheduling method based on field lineage described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Cross-domain job flow scheduling method and system
CN108694082A
Cited By
Blood relationship visual data construction method, device and equipment
CN120386817A