Monitoring Method, System and Related Devices for Digital Index Operations
The digital indicator job monitoring system addresses the lack of comprehensive monitoring in dual-active clusters by using SOD and MPP clusters with P9 and Ruban platforms to manage job dependencies and ensure high availability and efficient job transitions.
Patent Information
- Application Number
- CN202210004610.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-04
AI Technical Summary
The prior art cannot effectively manage the dependencies between multiple batch jobs in batch job management, and the monitoring is not comprehensive enough, especially when cross-platform cluster switching, it is difficult to achieve comprehensive and reliable job link monitoring.
By generating job link and job level information in the job management and control platform, combining the scheduling logs of the P9 scheduling platform and the Luban scheduling platform, the job operation status is determined, and job suspension and date settings are performed during cluster switching, and job monitoring information for digital indicators is generated.
It realizes comprehensive and reliable monitoring of batch jobs, especially when cluster switching, it can quickly locate exception points to ensure the continuity and reliability of the job link.
Smart Images

Figure CN114358595B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data, and in particular to a monitoring method, system and related devices for digital indicator operations. Background Art
[0002] At present, when enterprises manage batch jobs, they often combine multiple scheduling methods to form a batch job scheduling strategy to form an orderly management and control of batch jobs. These scheduling methods are as follows:
[0003] (1) Crontab command. In Linux, Unix, AIX and other systems, Crontab is used to set up periodic execution commands, controlling programs or scripts to run according to the set periodicity and time window. You can define attributes such as week, month, date, hour, minute, etc. In actual applications, Crontab is used to schedule batch jobs according to specified time points and cycles. However, since Crontab cannot manage the dependencies between multiple batch jobs, it is only used as a task triggering tool.
[0004] (2) Scheduling components of ETL products of data warehouse technology. ETL tools are increasingly used in IT system construction. Commonly used ETL tools include IBM's DataStage, Informatica's Powercenter, SAP's DataService, and open source kettle. In addition, there are now a wide variety of built-in tools for big data products, reporting products, and database products. In the actual system development process, the scheduling module of the ETL tool is mainly used to build task processes, implement process control of batch job operations, and meet the management of job dependencies.
[0005] Existing job monitoring only supports job dimension and job flow dimension monitoring, which is not comprehensive enough. Summary of the invention
[0006] In view of the above problems, the present invention provides a monitoring method, system and related devices for digital indicator operations that overcome the above problems or at least partially solve the above problems.
[0007] In a first aspect, a monitoring method for digital index operations is applied to a monitoring system for digital index operations, the system comprising at least: a SOD cluster, an MPP cluster and a job management and control platform, wherein the job management and control platform is equipped with a P9 scheduling platform and a Luban scheduling platform, the P9 scheduling platform is communicatively connected to the SOD cluster, and the Luban scheduling platform is communicatively connected to the MPP cluster;
[0008] The monitoring method comprises:
[0009] When it is detected that a cluster switch occurs, the job control platform generates job link and job hierarchy information of target digital metrics according to the database table information established in advance for each batch job, where the job link includes a plurality of the batch jobs, and the job hierarchy information characterizes the dependency hierarchy relationships between the batch jobs in the job link and other batch jobs, and the job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same;
[0010] For any of the job links, the following operations are all executed: the job control platform obtains the scheduling logs of the P9 scheduling platform for the batch jobs in the job link and obtains the scheduling logs of the Luban scheduling platform for the batch jobs in the job link according to the job identifiers;
[0011] The job control platform determines the job running status of the batch jobs in the job link respectively according to the scheduling logs, where the job running status characterizes the current running situation of the corresponding batch job;
[0012] The job control platform generates job monitoring information of the target digital metrics according to the job identifiers, job running status and job hierarchy information of the batch jobs in the job link.
[0013] Combined with the first aspect, in some optional embodiments, before it is detected that a cluster switch occurs, the method further includes:
[0014] If the SOD cluster fails and causes the job control platform to switch to the MPP cluster, the job control platform suspends each job flow running in the SOD cluster and records the current business date of each batch job that is suspended, where the job flow includes a plurality of the batch jobs;
[0015] For any of the suspended job flows, the following operation is all executed: the job control platform enables the corresponding job flow in the MPP cluster;
[0016] For any of the suspended batch jobs, the following operation is all executed: the job control platform sets the business date of the corresponding batch job in the MPP cluster to the current business date of the corresponding suspended batch job, so as to run the corresponding batch job in the MPP cluster when the current business date arrives.
[0017] Combined with the first aspect, in some optional embodiments, before it is detected that a cluster switch occurs, the method further includes:
[0018] If the MPP cluster fails and causes the job management and control platform to switch to the SOD cluster, the job management and control platform will suspend each job flow running in the MPP cluster and record the current business date of each batch job that is suspended, where each job flow includes multiple batch jobs;
[0019] For any suspended job flow, the following operations are performed: the job management and control platform enables the corresponding job flow in the SOD cluster;
[0020] For any suspended batch job, the following operations are performed: the job management and control platform sets the business date of the corresponding batch job in the SOD cluster to the current business date of the suspended batch job, so that when the current business date arrives, the corresponding batch job in the SOD cluster can be run.
[0021] Optionally, in some alternative embodiments, the job management and control platform generates job monitoring information of the target digital metrics according to the job identifiers, job running statuses, and job hierarchy information of each batch job in the job link, including:
[0022] The job management and control platform generates job monitoring information of the target digital metrics according to the job identifiers, job running statuses, the job hierarchy information, and the current business date of each batch job in the job link.
[0023] Combined with the first aspect, in some alternative embodiments, after generating the job link and job hierarchy information of the target digital metrics, the method further includes:
[0024] If an exception occurs during the running of a batch job, the job management and control platform determines the exception point according to the corresponding job hierarchy information.
[0025] In a second aspect, a monitoring system for digital metric jobs includes: an SOD cluster, an MPP cluster, and a job management and control platform, where the job management and control platform is equipped with a P9 scheduling platform and a Luban scheduling platform, the P9 scheduling platform is communicatively connected to the SOD cluster, and the Luban scheduling platform is communicatively connected to the MPP cluster;
[0026] The job management and control platform includes: a database table information unit, a log acquisition unit, a running status determination unit, and a monitoring information generation unit;
[0027] The database table information unit is used to generate the job link and job hierarchy information of the target digital metrics according to the database table information established in advance for each batch job when it is detected that a cluster switch occurs. Among them, the job link includes multiple batch jobs, and the job hierarchy information represents the dependency hierarchy relationship between each batch job in the job link and other batch jobs. The job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same;
[0028] The log acquisition unit is used to execute for any of the job links: obtain the scheduling logs of the P9 scheduling platform for each batch job in the job link and obtain the scheduling logs of the Luban scheduling platform for each batch job in the job link according to the job identifier;
[0029] The operation status determination unit is used to determine the job operation status of each batch job in the job link according to each of the scheduling logs, where the job operation status represents the current operation situation of the corresponding batch job;
[0030] The monitoring information generation unit is used to generate the job monitoring information of the target digital metrics according to the job identifiers, job operation statuses and job hierarchy information of each batch job in the job link.
[0031] Combined with the second aspect, in some optional embodiments, the job control platform further includes: a first suspension unit, a first enabling unit and a first date setting unit;
[0032] The first suspension unit is used to suspend each job flow running in the SOD cluster and record the current business date of each batch job suspended if the SOD cluster fails and causes the job control platform to switch to the MPP cluster before it is detected that a cluster switch occurs. Among them, the job flow includes multiple batch jobs;
[0033] The first enabling unit is used to execute for any of the suspended job flows: enable the corresponding job flow in the MPP cluster;
[0034] The first date setting unit is used to execute for any of the suspended batch jobs: set the business date of the corresponding batch job in the MPP cluster to the current business date of the corresponding suspended batch job, so as to run the corresponding batch job in the MPP cluster when the current business date arrives.
[0035] In combination with the second aspect, in some alternative embodiments, the job control platform further includes: a second suspension unit, a second enabling unit, and a second date setting unit;
[0036] The second suspension unit is configured to, before it is detected that a cluster switch occurs, if the MPP cluster fails and causes the job control platform to switch to the SOD cluster, suspend each job flow running in the MPP cluster, and record the current business date of each batch job that is suspended, where the job flow includes a plurality of the batch jobs;
[0037] The second enabling unit is configured to, for any of the suspended job flows, perform: enabling the corresponding job flow in the SOD cluster;
[0038] The second date setting unit is configured to, for any of the suspended batch jobs, perform: setting the business date of the corresponding batch job in the SOD cluster to the current business date of the correspondingly suspended batch job, so as to run the corresponding batch job in the SOD cluster when the current business date arrives.
[0039] In a third aspect, a computer-readable storage medium has a program stored thereon, and when the program is executed by a processor, the monitoring method of the digital metric job described in any one of the above is implemented.
[0040] In a fourth aspect, an electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the monitoring method of the digital metric job described in any one of the above.
[0041] With the above technical solutions, the monitoring method, system and related devices for digital index operations provided by the present invention can, when it is detected that a cluster switch occurs, generate job link and job hierarchy information of target digital indexes according to database table information pre-established for each batch job, where the job link includes multiple batch jobs, and the job hierarchy information represents the dependency hierarchy relationships between the batch jobs in the job link and other batch jobs, and the job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same; for any one of the job links, the following operations are performed: the job control platform obtains the scheduling logs of the batch jobs in the job link from the P9 scheduling platform according to the job identifiers, and obtains the scheduling logs of the batch jobs in the job link from the Luban scheduling platform; the job control platform determines the job running status of each batch job in the job link according to each scheduling log, where the job running status represents the current running situation of the corresponding batch job; the job control platform generates job monitoring information of the target digital indexes according to the job identifiers, job running status and job hierarchy information of the batch jobs in the job link. It can be seen from this that the present invention can monitor at the job link level for important batch jobs and data line scenarios (such as digital indexes), and the monitoring is relatively comprehensive and reliable.
[0042] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are hereinafter specifically exemplified. Brief Description of the Drawings
[0043] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0044] Figure 1 Shows a flowchart of a monitoring method for digital index operations provided by the present invention;
[0045] Figure 2 Shows a schematic diagram of a digital index definition table provided by the present invention;
[0046] Figure 3 Shows a schematic diagram of a job basic information table provided by the present invention;
[0047] Figure 4Shows a schematic diagram of a job dependency information table provided by the present invention;
[0048] Figure 5 Shows a schematic diagram of a table of all upstream dependency relationships of a job provided by the present invention;
[0049] Figure 6 Shows a schematic diagram of the scheduling log of a P9 scheduling platform provided by the present invention;
[0050] Figure 7 Shows a schematic diagram of the scheduling log of a Luban scheduling platform provided by the present invention;
[0051] Figure 8 Shows a schematic diagram of the structure of a monitoring system for digital index jobs provided by the present invention;
[0052] Figure 9 Shows a schematic diagram of the structure of an electronic device provided by the present invention. Detailed implementation manners
[0053] The most significant feature of digital transformation is to improve operational efficiency through digital applications, and ensuring the processing and output of digital indicators is crucial for digital transformation. Therefore, the financial industry has equipped high-priority job dual-active clusters for important batch processing and data lines of digital indicators, and realizes dual-database data sharing through dual-channel loading, which can quickly perform emergency switching.
[0054] For cross-platform job dual-active clusters, the job link of digital indicators is relatively complex, and how to better monitor jobs has become an urgent problem to be solved in this field.
[0055] Currently, when enterprises manage batch processing jobs, they often combine multiple scheduling methods to jointly form a batch processing job scheduling strategy to orderly control batch processing jobs. These scheduling methods are as follows:
[0056] (1) Crontab command. In systems such as linux, Unix, and AIX, Crontab is a command used to set periodic execution, and it controls programs or scripts to run according to the set periodicity and time window. Attributes such as week, month, date, hour, and minute can be defined. In practical applications, Crontab is used to schedule batch jobs at specified time points and periods. However, since Crontab cannot manage the dependencies between multiple batch processing jobs, it is only used as a task trigger tool.
[0057] (2) Scheduling components of ETL products. ETL tools are increasingly used in IT system construction. Commonly used ETL tools include IBM's DataStage, Informatica's Powercenter, SAP's DataService, and open source kettle. In addition, there are now a wide variety of built-in tools for big data products, reporting products, and database products. In the actual system development process, the scheduling module of the ETL tool is mainly used to build task processes, implement process control of batch job operations, and meet the management of job dependencies.
[0058] However, there are still several problems with the scheduling module of ETL tools, such as time trigger management, management of massive batch jobs, job monitoring and manual operations in operation and maintenance. ETL tools also have many inconveniences and cannot even meet many functions required for operation and maintenance. Therefore, they cannot be widely used as an overall enterprise scheduling platform.
[0059] (3) Self-developed scheduling programs. Due to the above problems, it is difficult for users to achieve good scheduling management and monitoring of batch jobs by using Crontab and ETL tools. In particular, in scenarios where the batch jobs are large in scale and the data flow between systems is complex, simple scheduling instructions and tools are even more difficult to handle. Therefore, many users develop scheduling tools based on their own needs. This type of scheduling tool is generally used for a certain project. There are also a few companies with relatively complete IT systems that choose to develop enterprise-level scheduling platforms themselves.
[0060] (4) Professional scheduling software. Professional scheduling software is designed for the control and scheduling of enterprise batch processing jobs. The functions of professional scheduling software integrate various user needs for batch processing job scheduling management, and can well implement job process management and time window management, and display monitoring views from the perspective of operation and maintenance.
[0061] The inventors have found that batch processing is different from online transactions. To address the high availability of batch processing services, regular batch processing cannot support active-active mode, but is deployed in master-slave or single-point mode. Therefore, the present invention proposes to equip batch processing jobs with high-priority active-active clusters, and realize dual database shunting and data supply through dual-path loading, so as to enable rapid emergency switching.
[0062] Subsequent terminology explanation:
[0063] Job: A program or running script that can be executed on the system, including the program and the parameter information required for the program to run. A job is the basic unit of execution and scheduling, representing an independent and executable function instance.
[0064] Job flow: A collection of one or more interdependent jobs with certain functions.
[0065] Job Group: For management needs, jobs (flows) are marked, and jobs (flows) under the same mark form a job group.
[0066] Component: Divided according to the new generation of CCB logic components, and each component is allowed to contain one or more job flows.
[0067] Digital Indicator: Define and decompose the business visual data requirements to form digital indicators. For example, "total deposit balance" is a digital indicator, which can be used to guide the personal customer peak season marketing business scenario.
[0068] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0069] The present invention provides a monitoring method for digital indicator jobs, which is applied to a monitoring system for digital indicator jobs. The system at least includes: an SOD cluster, an MPP cluster, and a job control platform. Among them, the job control platform is equipped with a P9 scheduling platform and a Luban scheduling platform. The P9 scheduling platform is communicatively connected to the SOD cluster, and the Luban scheduling platform is communicatively connected to the MPP cluster;
[0070] Optionally, the job control platform proposed by the present invention is composed of two parts: a job scheduling system and a job control system. The job scheduling system can centrally schedule and manage the batch jobs of each application system, and provide rich job scheduling modes and scheduling rules, including end-of-day batch, timed batch, cyclic batch, and online batch, etc. The job control system mainly includes modules such as job management, monitoring management, and data analysis, and realizes the unified definition and configuration management of jobs.
[0071] Optionally, the job scheduling system described in the present invention can be used as a data integration layer. The data of the data integration layer comes from each transaction system. Each transaction system transmits the data to the data cache layer of the data integration layer in the form of file push through a unified information exchange component. The source data area of the data integration layer stores the incremental files of the data cache layer in the database. Among them, the database table constructed by storing the incremental files is consistent with the database table structure of the data source, or consistent with the database table structure of the information exchange component. The dual-active cluster of the source data area includes an SOD cluster and an MPP cluster.
[0072] Optionally, the SOD cluster is: a cluster that constructs a source structure (close to the source system / component) data model based on the GP database.
[0073] The MPP cluster is: a data model cluster based on the MPP database to build a source-attached structure (close to the source system / component) on the cloud.
[0074] The P9 scheduling platform is: to centrally schedule and manage the batch jobs of the P9 application system.
[0075] The Luban scheduling platform is: to centrally schedule and manage the batch jobs of each application system.
[0076] Optionally, the same batch job can be set with the same job identifier in the P9 scheduling platform and the Luban scheduling platform, and different batch jobs can be set with different job identifiers; corresponding job flow identifiers can also be set for the job flow, and the naming methods of different job flow identifiers can be different. For example, in the P9 scheduling platform, the naming method for the job flow identifier can be: the English abbreviation of the application project + the source cluster name (such as SOD or MPP) + the name of the P9 scheduling platform. The present invention does not limit this.
[0077] Optionally, the same two sets of batch jobs can be configured with the same custom output condition name (for example: the English abbreviation of the application project + the job name). For example, job A depends on job B. The custom output condition name of job B will be passed as configuration information to job A, that is, the dependency information of job A will configure the custom output condition name of this job B.
[0078] As Figure 1 shown, the monitoring method includes: S100, S200, S300, and S400;
[0079] S100. When it is detected that a cluster switch occurs, the job control platform generates the job link and job hierarchy information of the target digital metrics according to the database table information established in advance for each batch job. Among them, the job link includes multiple batch jobs, and the job hierarchy information represents the dependency hierarchy relationship between each batch job in the job link and other batch jobs. The job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same;
[0080] Optionally, the present invention can implement a cross-platform cluster dual-loading mechanism, including batch jobs of different clusters accessing different storage units (including database GP and big data cloud MPP). Different job scheduling platforms (including the P9 scheduling platform and the Luban scheduling platform) are used to access different databases. Among them, the P9 scheduling platform is used when accessing database GP, and the Luban scheduling platform is used when accessing big data cloud MPP. The present invention does not limit this.
[0081] Optionally, for better monitoring of the subsequent operation link for digital metrics, the present invention can automatically pre-generate multiple database tables for each batch job, and the information of each database table records the relevant information of the corresponding batch job, collectively referred to as database table information. For example, as Figures 2 to 5 shown in the legend, where Figure 2 is the digital metric definition table, Figure 3 is the job basic information table, Figure 4 is the job dependency information table, Figure 5 is the table of all upstream dependency relationships of the job. The "fields" and "descriptions" in the figure are all optional implementation manners provided by the present invention, and the present invention does not limit this.
[0082] Optionally, based on the information Figures 2 to 5 shown, the operation link and operation hierarchy information of the target digital metric can be generated.
[0083] The operation hierarchy information obtains its pre-dependency job through the end job or a certain intermediate job. Its direct dependency is the first level, and the indirect dependency is the second level, and so on. For example, if job A depends on job B and job B depends on job C. Job B is the direct dependency of job A and is at the first level. Job C is the dependency of job A and is at the second level. Recursive loop can find all job dependency level relationships. Among them, each level of job dependency may include multiple other job information, and the present invention does not limit this.
[0084] Optionally, based on the current active job cluster (for example, the SOD cluster in the above-mentioned cluster dual loading is the active cluster), through the output job of the digital metric, the corresponding operation link is recursively found (obtained by associating through the job basic information table and the job dependency information table), and the present invention does not limit this.
[0085] Optionally, the present invention does not specifically limit the situation of cluster switching. For example, in combination with the Figure 1 shown implementation manner, in some optional implementation manners, before it is detected that the cluster switching occurs, the method further includes: steps 1.1, 1.2, and 1.3;
[0086] Step 1.1, if the SOD cluster fails and the job control platform switches to the MPP cluster, the job control platform suspends each job flow running in the SOD cluster and records the current business date of each batch job that is suspended, where the job flow includes multiple batch jobs;
[0087] Step 1.2, for any suspended job flow, execute: the job control platform enables the corresponding job flow in the MPP cluster;
[0088] Step 1.3: For any suspended batch job, the following operations are performed: the job control platform sets the business date of the corresponding batch job in the MPP cluster to the current business date of the corresponding suspended batch job, so that when the current business date arrives, the corresponding batch job in the MPP cluster can be run.
[0089] For another example, in combination with Figure 1 the embodiment shown, in some optional embodiments, before it is monitored that a cluster switch occurs, the method further includes: Step 2.1, Step 2.2, and Step 2.3;
[0090] Step 2.1: If the failure of the MPP cluster causes the job control platform to switch to the SOD cluster, the job control platform suspends each job flow running in the MPP cluster and records the current business date of each suspended batch job, where the job flow includes multiple batch jobs;
[0091] Step 2.2: For any suspended job flow, the following operations are performed: the job control platform enables the corresponding job flow in the SOD cluster;
[0092] Step 2.3: For any suspended batch job, the following operations are performed: the job control platform sets the business date of the corresponding batch job in the SOD cluster to the current business date of the corresponding suspended batch job, so that when the current business date arrives, the corresponding batch job in the SOD cluster can be run.
[0093] Optionally, in some optional embodiments, the job control platform generates job monitoring information of the target digital index according to the job identifier, job running status, and job hierarchy information of each batch job in the job link, including:
[0094] The job control platform generates job monitoring information of the target digital index according to the job identifier, job running status, the job hierarchy information, and the current business date of each batch job in the job link.
[0095] S200: For any job link, the following operations are performed: the job control platform obtains the scheduling logs of each batch job in the job link from the P9 scheduling platform according to the job identifier and obtains the scheduling logs of each batch job in the job link from the Luban scheduling platform;
[0096] Optionally, the present invention does not limit the information recorded in the scheduling logs of the P9 scheduling platform and the information recorded in the scheduling logs of the Luban scheduling platform. For example, Figure 6It is the scheduling log of the P9 scheduling platform. Figure 7 It is the scheduling log of the Luban scheduling platform, and the present invention does not limit this.
[0097] S300. The job control platform respectively determines the job running status of each batch job in the job link according to each of the scheduling logs, where the job running status represents the current running situation of the corresponding batch job.
[0098] Optionally, the job running status may include: unready, in execution, abnormal, normal and other states, and the present invention does not limit this.
[0099] S400. The job control platform generates job monitoring information of the target digital index according to the job identifier, job running status and job hierarchy information of each batch job in the job link.
[0100] Optionally, based on each of the scheduling logs, corresponding digital index job monitoring views can be generated according to information such as the cluster to which the job belongs, the job category, and whether it is an active cluster. Among them, the job monitoring view can display the following content: display the statistics of several states such as unready, in execution, abnormal, and normal of the entry job, intermediate job, and output job corresponding to the digital index; display the hierarchical relationship of dependencies of each job. For example, when clicking on the abnormal state, a list of jobs in the abnormal state can be displayed, including job name, job status, job execution time, etc. Clicking on the job name can further display the job hierarchy on which the abnormal job depends, and display information such as the job running status of each level. The abnormal state is displayed in red for warning, and the present invention does not limit this.
[0101] Optionally, the method for generating the digital index job monitoring view may include: counting the execution progress of important batch jobs and data line scenario (such as digital index) jobs (statistics of data in the states of not started, in execution, and abnormal). When a job execution is abnormal, the fault point can be quickly located according to the job hierarchy relationship, and the basic information and job status of each level of jobs on which the faulty job depends can be displayed. For example, in combination with Figure 1 In the embodiment shown, in some optional embodiments, after generating the job link and job hierarchy information of the target digital index, the method further includes:
[0102] If an abnormality occurs in the running of the batch job, the job control platform determines the abnormality point according to the corresponding job hierarchy information.
[0103] Optionally, at the first deployment, the present invention can use the SOD cluster as the active cluster. That is, the jobs corresponding to the SOD cluster are released and run normally, and the job flow where the jobs corresponding to the MPP cluster are located is suspended (or the job flow is predefined to be paused) to ensure that only the jobs of one cluster are running.
[0104] Optionally, when a cluster failure occurs and causes a cluster switch (when performing the switch operation, the job status needs to be updated while also updating whether the job is active). Based on the updated active job cluster (such as when switching from the SOD cluster to the MPP cluster during dual loading of the above-mentioned clusters, and the current MPP is the active cluster), a cross-platform digital metric job monitoring view is regenerated, and the present invention places no restrictions on this.
[0105] As Figure 8 As shown, the present invention provides a monitoring system for digital metric jobs, including: an SOD cluster 100, an MPP cluster 200, and a job management and control platform 300. Among them, the job management and control platform 300 is equipped with a P9 scheduling platform 310 and a Luban scheduling platform 320. The P9 scheduling platform is communicatively connected to the SOD cluster 100, and the Luban scheduling platform is communicatively connected to the MPP cluster 200;
[0106] The job management and control platform 300 includes: a database table information unit 311, a log acquisition unit 312, a running state determination unit 313, and a monitoring information generation unit 314;
[0107] The database table information unit 311 is used to generate the job link and job hierarchy information of the target digital metrics when it is detected that a cluster switch occurs, according to the database table information established in advance for each batch job. Among them, the job link includes multiple batch jobs, and the job hierarchy information represents the dependency hierarchy relationship between each batch job in the job link and other batch jobs. The job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same;
[0108] The log acquisition unit 312 is used to perform, for any of the job links: according to the job identifier, obtain the scheduling logs of each batch job of the job link from the P9 scheduling platform and obtain the scheduling logs of each batch job of the job link from the Luban scheduling platform;
[0109] The running state determination unit 313 is used to respectively determine the job running states of each batch job of the job link according to the scheduling logs, where the job running state represents the current running situation of the corresponding batch job;
[0110] The monitoring information generation unit 314 is used to generate the job monitoring information of the target digital metrics according to the job identifiers, job running states, and job hierarchy information of each batch job of the job link.
[0111] Combined with Figure 8In the illustrated embodiment, in some alternative embodiments, the job control platform 300 further includes: a first suspension unit, a first enabling unit, and a first date setting unit;
[0112] The first suspension unit is configured to suspend each job flow running in the SOD cluster 100 and record the current business date of each suspended batch job before it is detected that a cluster switch occurs and if the SOD cluster 100 fails, causing the job control platform 300 to switch to the MPP cluster 200, where the job flow includes a plurality of the batch jobs;
[0113] The first enabling unit is configured to, for any suspended job flow, execute: enabling the corresponding job flow in the MPP cluster 200;
[0114] The first date setting unit is configured to, for any suspended batch job, execute: setting the business date of the corresponding batch job in the MPP cluster 200 to the current business date of the corresponding suspended batch job, so as to run the corresponding batch job in the MPP cluster 200 when the current business date arrives.
[0115] Combined Figure 8 In the illustrated embodiment, in some alternative embodiments, the job control platform 300 further includes: a second suspension unit, a second enabling unit, and a second date setting unit;
[0116] The second suspension unit is configured to suspend each job flow running in the MPP cluster 200 and record the current business date of each suspended batch job before it is detected that a cluster switch occurs and if the MPP cluster 200 fails, causing the job control platform 300 to switch to the SOD cluster 100, where the job flow includes a plurality of the batch jobs;
[0117] The second enabling unit is configured to, for any suspended job flow, execute: enabling the corresponding job flow in the SOD cluster 100;
[0118] The second date setting unit is configured to, for any suspended batch job, execute: setting the business date of the corresponding batch job in the SOD cluster 100 to the current business date of the corresponding suspended batch job, so as to run the corresponding batch job in the SOD cluster 100 when the current business date arrives.
[0119] Optionally, in combination with any one of the above two embodiments, in some optional embodiments, the monitoring information generation unit 314 includes: a monitoring information generation subunit;
[0120] The monitoring information generation subunit is configured to generate job monitoring information of the target digital metrics according to the job identifiers, job running statuses, the job hierarchy information, and the current business date of the batch processing jobs in the job link.
[0121] Combined Figure 8 In combination with the embodiments shown, in some optional embodiments, the job control platform 300 further includes: an anomaly location unit;
[0122] The anomaly location unit is configured to, after generating the job link and job hierarchy information of the target digital metrics, if an anomaly occurs in the running of the batch processing job, determine the anomaly point according to the corresponding job hierarchy information.
[0123] The present invention provides a computer-readable storage medium having a program stored thereon, and the program, when executed by a processor, implements the monitoring method for digital metrics jobs described in any one of the above.
[0124] As Figure 9 shown, the present invention provides an electronic device 70, which includes at least one processor 701, at least one memory 702 connected to the processor 701, and a bus 703; wherein, the processor 701 and the memory 702 complete communication with each other through the bus 703; the processor 701 is configured to call program instructions in the memory 702 to execute the monitoring method for digital metrics jobs described in any one of the above.
[0125] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0126] Each embodiment in this specification is described in a relevant manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0127] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0128] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A monitoring method for digital index operations, characterized in that, A monitoring system applied to a digital metric operation, the system at least includes: an SOD cluster, an MPP cluster, and a job control platform. Among them, the SOD cluster is a source-attached structured data model cluster built based on a GP database, the MPP cluster is a cloud-based source-attached structured data model cluster built based on an MPP database, the job control platform is equipped with a P9 scheduling platform and a Luban scheduling platform. The P9 scheduling platform centrally schedules and manages batch jobs of the P9 application system, and the Luban scheduling platform centrally schedules and manages batch jobs of each application system. The P9 scheduling platform is communicatively connected to the SOD cluster, and the Luban scheduling platform is communicatively connected to the MPP cluster; The monitoring method includes: When it is detected that a cluster switch occurs, the job control platform generates job link and job hierarchy information of the target digital metric according to database table information pre-established for each batch job. Among them, the job link includes multiple batch jobs, and the job hierarchy information represents the dependency hierarchy relationship between each batch job in the job link and other batch jobs. The job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same; For any one of the job links, the following is executed: the job control platform obtains the scheduling logs of each batch job of the job link from the P9 scheduling platform according to the job identifier and obtains the scheduling logs of each batch job of the job link from the Luban scheduling platform; The job control platform determines the job running status of each batch job of the job link according to each of the scheduling logs, where the job running status represents the current running situation of the corresponding batch job; The job control platform generates job monitoring information of the target digital metric according to the job identifiers, the job running status, and the job hierarchy information of each batch job of the job link.
2. The method according to claim 1, wherein Before the step of "When it is detected that a cluster switch occurs", the method further includes: If the failure of the SOD cluster causes the job control platform to switch to the MPP cluster, the job control platform suspends each job flow running in the SOD cluster and records the current business date of each batch job suspended, where the job flow includes multiple batch jobs; For any one of the suspended job flows, the following is executed: the job control platform enables the corresponding job flow in the MPP cluster; For any one of the suspended batch jobs, the following is executed: the job control platform sets the business date of the corresponding batch job in the MPP cluster to the current business date of the corresponding suspended batch job, so that when the current business date arrives, the corresponding batch job in the MPP cluster can be run.
3. The method according to claim 1, wherein Before the step of "When it is detected that a cluster switch occurs", the method further includes: If the failure of the MPP cluster causes the job control platform to switch to the SOD cluster, the job control platform will suspend each job flow running in the MPP cluster and record the current business date of each batch job that is suspended, where the job flow includes multiple batch jobs; For any suspended job flow, the following operations are performed: the job control platform enables the corresponding job flow in the SOD cluster; For any suspended batch job, the following operations are performed: the job control platform sets the business date of the corresponding batch job in the SOD cluster to the current business date of the suspended batch job, so that when the current business date arrives, the corresponding batch job in the SOD cluster can be run.
4. The method according to claim 2 or 3, characterized in that The job control platform generates job monitoring information for the target digital metrics according to the job identifiers, job running status, and job hierarchy information of the batch jobs in the job link, including: The job control platform generates job monitoring information for the target digital metrics according to the job identifiers, job running status, job hierarchy information, and current business date of the batch jobs in the job link.
5. The method according to claim 1, characterized in that, After generating the job link and job hierarchy information for the target digital metrics, the method further includes: If an abnormality occurs during the running of a batch job, the job control platform determines the abnormality point according to the corresponding job hierarchy information.
6. A monitoring system for digital index operations, characterized in that, Including: An SOD cluster, an MPP cluster, and a job control platform, where the SOD cluster is a source-attached structure data model cluster built based on a GP database, the MPP cluster is a cloud source-attached structure data model cluster built based on an MPP database, the job control platform is equipped with a P9 scheduling platform and a Luban scheduling platform, the P9 scheduling platform centrally schedules and manages the batch jobs of the P9 application system, the Luban scheduling platform centrally schedules and manages the batch jobs of each application system, the P9 scheduling platform is communicatively connected to the SOD cluster, and the Luban scheduling platform is communicatively connected to the MPP cluster; The job control platform includes: a database table information unit, a log acquisition unit, a running status determination unit, and a monitoring information generation unit; The database table information unit is configured to, when detecting a cluster switch situation, generate the job link and job hierarchy information of the target digital metrics according to the database table information established in advance for each batch job, where the job link includes multiple batch jobs, the job hierarchy information represents the dependency hierarchy relationship between the batch jobs in the job link and other batch jobs, and the job identifiers configured for the same batch job in the P9 scheduling platform and the Luban scheduling platform are the same; The log acquisition unit is configured to, for any job link, perform the following operations: obtain the scheduling logs of the P9 scheduling platform for the batch jobs in the job link and obtain the scheduling logs of the Luban scheduling platform for the batch jobs in the job link according to the job identifiers; The running state determination unit is configured to respectively determine the job running states of the batch processing jobs of the job link according to each of the scheduling logs, where the job running state characterizes the current running situation of the corresponding batch processing job; The monitoring information generation unit is configured to generate job monitoring information of the target digital metrics according to the job identifiers, the job running states, and the job hierarchy information of the batch processing jobs of the job link.
7. The system according to claim 6, wherein The job control platform further includes: a first suspension unit, a first enabling unit, and a first date setting unit; The first suspension unit is configured to, before it is detected that a cluster switch occurs, if the SOD cluster fails and causes the job control platform to switch to the MPP cluster, suspend each job flow running in the SOD cluster, and record the current business date of each batch processing job that is suspended, where the job flow includes a plurality of the batch processing jobs; The first enabling unit is configured to, for any suspended job flow, execute: enabling the corresponding job flow in the MPP cluster; The first date setting unit is configured to, for any suspended batch processing job, execute: setting the business date of the corresponding batch processing job in the MPP cluster to the current business date of the corresponding suspended batch processing job, so as to run the corresponding batch processing job in the MPP cluster when the current business date arrives.
8. The system according to claim 6, wherein The job control platform further includes: a second suspension unit, a second enabling unit, and a second date setting unit; The second suspension unit is configured to, before it is detected that a cluster switch occurs, if the MPP cluster fails and causes the job control platform to switch to the SOD cluster, suspend each job flow running in the MPP cluster, and record the current business date of each batch processing job that is suspended, where the job flow includes a plurality of the batch processing jobs; The second enabling unit is configured to, for any suspended job flow, execute: enabling the corresponding job flow in the SOD cluster; The second date setting unit is configured to, for any suspended batch processing job, execute: setting the business date of the corresponding batch processing job in the SOD cluster to the current business date of the corresponding suspended batch processing job, so as to run the corresponding batch processing job in the SOD cluster when the current business date arrives.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the monitoring method for digital metric jobs according to any one of claims 1 to 5.
10. An electronic device, characterized in that, The electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the monitoring method for digital metric jobs according to any one of claims 1 to 5.
Citation Information
Patent Citations
Operation request method and device, electronic equipment and storage medium
CN109558446A
Monitoring method and device, computer device and storage medium
CN109873717A