Data Processing Timeliness Testing Method and Device

By obtaining and analyzing job runtime data and dependency data, identifying key jobs and matching them with the changed job sets, and conducting time tests, the problem that the existing technology cannot evaluate the overall time-saving risks of job changes in the R&D stage is solved, and efficient time-saving tests are achieved.

CN112907055BActive Publication Date: 2025-06-20INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110171584.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-08
Publication Date
2025-06-20
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

The existing technology cannot effectively evaluate the risk of operation changes in the R&D stage on the overall data processing timeliness, and traditional testing methods cannot cover the timeliness verification of all changed operations within a limited time.

Method used

By obtaining job run time data, job dependency data and change job sets, we determine the key jobs that affect the overall time of the job, and match them with the change job set to obtain the change key jobs. Perform a timeliness tests on changing key tasks to identify whether there is a timeliness risk.

Benefits of technology

It realizes the automatic identification of key operations that have time-saving risks, shortens the verification time period, and improves the time-saving testing efficiency of big data operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112907055B_ABST
    Figure CN112907055B_ABST
Patent Text Reader

Abstract

This specification provides a method and apparatus for testing the timeliness of data processing. Among them, the method includes: obtaining the job running time data, job dependency data, and changed job set of multiple jobs, where the multiple jobs include data processing of specified data according to a preset logic; determining the key jobs that affect the overall timeliness of the jobs based on the job running time data and job dependency data; matching the key jobs with the changed job set to obtain the changed key jobs; performing timeliness testing on the changed key jobs, and based on the timeliness test results, identifying whether there is a timeliness risk for the changed key jobs. The above method has a high degree of automation, can accurately locate the changed key jobs with timeliness risks, and make a judgment on the impact on the overall timeliness, shortening the verification time cycle and greatly improving the testing efficiency of big data jobs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular, to a method and device for testing data processing timeliness. Background Art

[0002] With the widespread rise of big data applications, the processing of massive data has brought new challenges, and the timeliness of data processing operations has been listed as an important indicator of the quality of big data applications. Due to the diversity and complexity of business types, the number of big data processing operations has been increasing continuously, and the dependency relationships among them have become increasingly complex. At the same time, with the implementation of the agile iterative R & D model, the changes in data processing logic and job dependency relationships have become increasingly frequent. The traditional testing methods cannot cover the timeliness verification of all changed operations within a limited time, and the impact on the overall timeliness cannot be accurately quantified. In the banking industry, there are often strict requirements for data processing timeliness, and serious consequences will occur if it is overdue.

[0003] Currently, there is a lack of a method for automatically identifying risks from the perspective of timeliness, and it is difficult to evaluate the risks brought by changes in the R & D stage to the overall timeliness.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this specification provide a method and device for testing data processing timeliness to solve the problem that the prior art cannot evaluate the risks brought by job changes in the R & D stage to the overall timeliness.

[0006] Embodiments of this specification provide a method for testing data processing timeliness, including: obtaining job running time data, job dependency relationship data, and a set of changed jobs for multiple jobs, where the multiple jobs include data processing of specified data according to a preset logic; determining key jobs that affect the overall timeliness of the multiple jobs according to the job running time data and the job dependency relationship data; matching the key jobs with the set of changed jobs to obtain changed key jobs; performing timeliness testing on the changed key jobs, and identifying whether there are timeliness risks for the changed key jobs according to the timeliness test results.

[0007] The embodiments of this specification also provide a data processing timeliness testing device, including: an acquisition module, configured to acquire the job running time data, job dependency data, and changed job set of multiple jobs, where the multiple jobs include data processing of specified data according to a preset logic; a determination module, configured to determine the critical jobs that affect the overall timeliness of the multiple jobs according to the job running time data and the job dependency data; a matching module, configured to match the critical jobs with the changed job set to obtain changed critical jobs; and an identification module, configured to perform timeliness testing on the changed critical jobs and identify whether there is a timeliness risk for the changed critical jobs according to the timeliness testing results.

[0008] The embodiments of this specification also provide a computer device, including a processor and a memory for storing processor-executable instructions, where when the processor executes the instructions, the steps of the data processing timeliness testing method described in any of the above embodiments are implemented.

[0009] The embodiments of this specification also provide a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of the data processing timeliness testing method described in any of the above embodiments are implemented.

[0010] In the embodiments of this specification, a data processing timeliness testing method is provided. The job running time data, job dependency data, and changed job set of multiple jobs can be acquired. According to the job running time data and the job dependency data, the critical jobs that affect the overall timeliness of the multiple jobs are determined. The critical jobs are matched with the changed job set to obtain changed critical jobs. Timeliness testing is performed on the changed critical jobs, and according to the timeliness testing results, it is identified whether there is a timeliness risk for the changed critical jobs. In the above solution, the critical jobs that affect the overall timeliness of the jobs can be determined, and they are matched with the changed job set to obtain changed critical jobs. By performing timeliness testing on the changed critical jobs, the timeliness testing scope is greatly reduced. Through timeliness verification of the timeliness testing results, the changed critical jobs with timeliness risks can be identified. The above method has a high degree of automation, can accurately locate the changed critical jobs with timeliness risks, and make a judgment on the impact on the overall timeliness, shortening the verification time cycle and greatly improving the timeliness testing efficiency of big data jobs. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. In the drawings:

[0012] Figure 1 The flowchart of the data processing timeliness testing method in an embodiment of this specification is shown;

[0013] Figure 2 Shows a schematic diagram of the AOE network in an embodiment of the present application;

[0014] Figure 3 Shows a flowchart for calculating the earliest start time of each job in an embodiment of the present application;

[0015] Figure 4 Shows a schematic diagram of the AOE network in an embodiment of the present application;

[0016] Figure 5 Shows a flowchart of the data processing timeliness test method in this specific embodiment;

[0017] Figure 6 Shows a schematic diagram of the data processing timeliness test device in an embodiment of this specification;

[0018] Figure 7 Shows a schematic structural diagram of a data processing timeliness test device based on the AOE network provided in an embodiment of the present application;

[0019] Figure 8 Shows a schematic diagram of the source data processing module in an embodiment of the present application;

[0020] Figure 9 Shows a schematic diagram of the big data job model construction module in an embodiment of the present application;

[0021] Figure 10 Shows a schematic diagram of the critical path construction module in an embodiment of the present application;

[0022] Figure 11 Shows a schematic diagram of the precise test module in an embodiment of the present application;

[0023] Figure 12 Shows a schematic diagram of the timeliness risk identification module in an embodiment of the present application;

[0024] Figure 13 Shows a schematic diagram of a computer device in an embodiment of this specification. Detailed implementation manners

[0025] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and thus implement the present application, rather than limiting the scope of the present application in any way. On the contrary, these embodiments are provided to make the disclosure of the present application more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0026] Those skilled in the art know that the embodiments of this specification can be implemented as a system, device, equipment, method, or computer program product. Therefore, the disclosure of this specification can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0027] The embodiments of this specification provide a method for testing the timeliness of data processing. Figure 1 The flowchart of the method for testing the timeliness of data processing in an embodiment of this specification is shown. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or non-creative labor. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of this specification and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, or even a distributed processing environment).

[0028] Specifically, as Figure 1 shown, the method for testing the timeliness of data processing provided by an embodiment of this specification may include the following steps:

[0029] Step S101, obtain the job running time data, job dependency data, and changed job set of multiple jobs.

[0030] The method in the embodiments of this specification can be applied to a job test server. The job test server can obtain the job running time data, job dependency data, and changed job set of multiple jobs. Among them, the multiple jobs include data processing of specified data according to a preset logic. For example, the multiple jobs can be multiple big data jobs required to complete one or more target tasks. Among them, the preset logic can be a data processing logic determined according to a preset business logic. The specified data can be business data of a bank, etc. The job running time data can be used to represent the time spent on executing a job. The job dependency data can be used to represent the dependency relationship between multiple jobs. The changed job set can be a set of jobs newly added or modified when the version is updated.

[0031] Step S102, determine the key jobs that affect the overall timeliness of the jobs according to the job running time data and job dependency data.

[0032] The job test server can determine the critical jobs that affect the overall timeliness of jobs based on the obtained job running time data and job dependency data. Specifically, the critical jobs that affect the overall timeliness of jobs can be the jobs on the critical path. Among them, the critical path can refer to the logical path with the longest delay from input to output. In big data jobs, the critical path can be the job execution path with the longest delay from the start of execution to the end of execution. According to the job running time data and job dependency data, the critical jobs that affect the overall timeliness of multiple jobs can be determined.

[0033] Step S103: Match the critical jobs with the set of changed jobs to obtain the changed critical jobs.

[0034] After determining the critical jobs, the critical jobs can be matched with the set of changed jobs to obtain the changed critical jobs. That is, the critical jobs existing in the set of changed jobs are used as the changed critical jobs. For example, if the critical jobs include Job A, Job C, Job D, and Job E, and the set of changed jobs includes Job C, Job E, Job F, and Job G, then the changed critical jobs are Job C and Job E.

[0035] Step S104: Conduct timeliness tests on the changed critical jobs, and based on the timeliness test results, identify whether there are timeliness risks for the changed critical jobs.

[0036] After obtaining the changed critical jobs, conduct timeliness tests on the changed critical jobs to obtain the timeliness test results. Among them, the timeliness test results can include the running timeliness data of the changed critical jobs. After obtaining the timeliness test results, it is possible to identify whether there are timeliness risks for the changed critical jobs based on the timeliness test results.

[0037] The method in the above embodiments can determine the critical jobs that affect the overall timeliness of jobs, match them with the set of changed jobs to obtain the changed critical jobs, greatly narrow the scope of timeliness tests by conducting timeliness tests on the changed critical jobs, and identify the changed critical jobs with timeliness risks through timeliness verification of the timeliness test results. The above method has a high degree of automation, can accurately locate the changed critical jobs with timeliness risks, make a judgment on the impact on the overall timeliness, shorten the verification time cycle, and greatly improve the timeliness test efficiency of big data jobs.

[0038] In some embodiments of this specification, obtaining the job running time data, job dependency data, and set of changed jobs of multiple jobs may include: collecting the job running time data of multiple jobs from the job operation and maintenance platform; performing lineage analysis on the data processing logic in the job scripts of multiple jobs to generate the job dependency data of multiple jobs; querying the job code submission records in the job version library to obtain the set of changed jobs of multiple jobs.

[0039] Specifically, the job running time data of multiple jobs can be collected from the job operation and maintenance platform in near real-time through continuous integration tasks, and the data can be formatted to unify the measurement unit of time. The job scripts of multiple jobs can be obtained from the repository, and the data processing logic in the job scripts can be analyzed for lineage to form the flow direction of the data stream, thereby automatically generating job dependency data. The changed job set can be obtained by querying the code commit records in the repository. In addition, the obtained job running time data, job dependency data, and changed job set can be saved to a database. The database type can be a relational database such as mysql or oracle. Through the above methods, the source data can be processed to obtain the required job running time data, job dependency data, and changed jobs.

[0040] In some embodiments of this specification, determining the key jobs that affect the overall timeliness of multiple jobs based on the job running time data and job dependency data may include: constructing an AOE network from multiple jobs according to the job running time data and job dependency data, where each job among the multiple jobs serves as a vertex of the AOE network, the job dependencies between jobs serve as directed edges of the AOE network, and the job running time of each job serves as the weight on the directed edge; based on the AOE network, determining the key jobs that affect the overall timeliness of multiple jobs.

[0041] The AOE network (Activity On Edge Network) is a commonly used weighted directed graph that can be used to estimate the shortest construction period of a project and which activities are critical to the progress of the project. Please refer to Figure 2 , which shows a schematic diagram of the AOE network in the embodiments of this application. In Figure 2 , the circles represent the vertices in the AOE network, the arrows represent the directed edges in the AOE network, and the values marked on the directed edges represent the weights of the directed edges. Among them, the vertices represent events, the directed edges represent activities, and the weights on the directed edges usually represent the duration of the activities. Only when all the activities represented by the directed edges entering a certain point have ended can the event represented by that vertex occur. Only after the event represented by a certain vertex occurs can the activities represented by the directed edges starting from that vertex begin. The time required to complete the entire project depends on the length of the longest path from the source point to the sink point, that is, the sum of the durations of all activities on this path. This path with the longest length is called the critical path, and the activities on the critical path are called critical activities. In Figure 2 , the critical path is A - B - C - H - J - K, and the critical activities include A, B, C, H, J, and K.

[0042] Based on the job running time data and job dependency data, multiple jobs can be constructed into an AOE network. Each job among the multiple jobs can be used as a vertex of the AOE network, the job dependencies between the multiple jobs can be used as directed edges of the AOE network, and the job running time of each job among the multiple jobs can be used as the weight on the directed edge. After obtaining the AOE network, based on the AOE network, the critical jobs that affect the overall timeliness of the jobs can be determined. By constructing multiple jobs into an AOE network, the critical jobs can be determined based on the AOE network, and the process is simple and efficient.

[0043] In some embodiments of the present specification, based on the AOE network, determining the critical jobs that affect the overall timeliness of the jobs may include: determining whether there is a cyclic graph in the AOE network; in the case where it is determined that there is no cyclic graph in the AOE network, based on the AOE network, determining the critical jobs that affect the overall timeliness of the jobs.

[0044] Specifically, after constructing the AOE network, it can be determined whether there is a cyclic graph in the AOE network according to the topological sorting algorithm. Among them, a cyclic graph refers to a link in the AOE network where the source point coincides with the sink point. In the case where it is determined that there is no cyclic graph in the AOE network, the critical jobs that affect the overall timeliness of the jobs are determined based on the AOE network. By the above method, the critical jobs are determined and subsequent steps are executed only when there is no cyclic graph in the AOE network, which can avoid wasting resources caused by executing subsequent steps when the AOE network does not conform to the data processing logic.

[0045] In some embodiments of the present specification, after determining whether there is a cyclic graph in the AOE network, it may further include: in the case where it is determined that there is a cyclic graph in the AOE network, determining that the multiple jobs do not conform to the data processing logic, and generating a notification message to be sent to the user.

[0046] Specifically, in the case where it is determined that there is a cyclic graph in the AOE network, it is determined that the multiple jobs do not conform to the data processing logic, and a notification message is generated and sent to the user. After the user receives the notification message, the user can troubleshoot and adjust the jobs based on the notification message. By the above method, the data processing logic of the jobs can be troubleshot in a timely manner, improving the accuracy and efficiency.

[0047] In some embodiments of this specification, based on the AOE network, determining the critical jobs that affect the overall timeliness of multiple jobs may include: performing a topological sort starting from the source point of the AOE network and calculating the earliest start time of each job among the multiple jobs; performing a topological sort starting from the sink point of the AOE network and calculating the latest start time of each job among the multiple jobs; determining whether the earliest start time and the latest start time of each job among the multiple jobs are equal; and determining the jobs with equal earliest start time and latest start time as the critical jobs that affect the overall timeliness of the jobs.

[0048] Specifically, according to the AOE network, a topological sort can be performed starting from the source point, the earliest start time of each job can be calculated, and added to the set of earliest start times of jobs. Please refer to Figure 3 , which shows the flowchart for calculating the earliest start time of each job in an embodiment of this application. As Figure 3 shown, calculating the earliest start time of a job may include the following steps: (1) Initialize the set of earliest start times of jobs, and uniformly assign 0 to the earliest start times of all jobs; (2) Start from the source point; (3) Perform a topological sort according to the AOE network; (4) Obtain the running time T i of the current job, that is, the job running time attribute of the current vertex; (5) Obtain the earliest start time ET i-1 of the previous job from the set; (6) Calculate the earliest start time of the current job: ET i = max(ET i-1 + T i ), if there are multiple previous jobs, take the maximum value of the calculation result as the earliest start time of this job; (7) Add the earliest start time ET i of the current job to the set of earliest start times of jobs; (8) Determine whether it is the sink point, if so, execute (9), otherwise return to (3) and perform a topological sort on the next vertex; (9) Obtain the set of earliest start times of all jobs.

[0049] According to the AOE network, a topological sort can be performed starting from the sink point, the latest start time of each job can be calculated, and added to the set of latest start times of jobs. The calculation of the latest start time of a job is the reverse process of the calculation of the earliest start time of a job. The steps for calculating the latest start time of a job can refer to the steps for calculating the earliest start time above. Among them, by reverse deduction based on the earliest start time, the calculation formula for the latest start time is LT i = min(LT i+1 - T i ).

[0050] For each job, obtain the earliest start time and the latest start time of the job from the set of the earliest start times of jobs and the set of the latest start times of jobs respectively. If the two are equal, then the job is a critical job. Since there are multiple combinations of the prerequisite paths for each job, there are multiple times to reach each job. Among them, the earliest start time is the time with the shortest duration to reach this point, and the latest start time is the time with the longest duration to reach this point. If the earliest start time and the latest start time of a vertex are equal, it means that this vertex has no extra slack time, indicating that this vertex is a critical activity. Please refer to Figure 4 , which shows a schematic diagram of the AOE network in the embodiments of the present application. In Figure 4 , the circles represent jobs in the AOE network. The characters on the left side of the circles represent the job numbers, the numerical values on the upper right side of the circles represent the earliest start times of the jobs, and the numerical values on the lower right side of the circles represent the latest start times of the jobs. That is, Figure 4 shows the AOE network with the earliest start times and the latest start times of the jobs updated. From Figure 4 , it can be conveniently determined that the critical jobs include: job V1, job V4, job V6, job V5, job V7, job V9, and job V10. Through the above method, topological sorting can be performed based on the AOE network to obtain the earliest start times and the latest start times of each job, so as to quickly determine the critical jobs.

[0051] In some embodiments of this specification, the timeliness test for changing critical jobs may include: obtaining the historical version and the current version of the changed critical job from the job version library; uploading the job scripts of the historical version and the current version to the directory of the scripts to be executed of the job scheduling server. Among them, the job scheduling server uses the job monitoring process to periodically scan the directory of the scripts to be executed. If a script to be executed is found, the job scheduling framework is run to execute the job script. After the job script is executed, the time consumed for executing the job script is determined as the timeliness result of the job corresponding to the job script and stored in the database, and the corresponding job script in the directory of the scripts to be executed is deleted.

[0052] Specifically, the historical version and the current version of the critical changed job can be obtained from the job version library. After respectively labeling version tags, the job scripts of the historical version and the current version of the critical changed job are uploaded to the directory of scripts to be executed of the job scheduling server. The job scheduling server can use the job monitoring process to regularly scan the directory of scripts to be executed of the job scheduling server. If a script to be executed is found, the job scheduling framework is automatically run to execute the job script. After the job execution is completed, the time consumed by the job running can be used as the aging result, incorporated into the aging memory, and the corresponding script in the directory of scripts to be executed is deleted. The aging memory can record the aging results of the jobs and save them to a database, and the database type can be a relational database such as mysql or oracle. Through the above method, the script to be verified is automatically deployed to the job scheduling server and then automatically executed by the job scheduling framework, and the results are saved to the database, which is fully automated, can improve efficiency, and save labor costs.

[0053] In some embodiments of this specification, according to the aging test results, identifying whether there is an aging risk for the critical changed job may include: obtaining the aging results of the historical version of the critical changed job and the aging results of the current version of the critical changed job from the database; obtaining the job running time threshold from the job version library; determining the path composed of critical jobs as the critical path and calculating the length of the critical path; determining the difference between the job running time threshold and the length of the critical path as the job running window time margin; and identifying whether there is an aging risk for the critical changed job according to the job running window time margin, the aging results of the historical version of the critical changed job, and the aging results of the current version of the critical changed job.

[0054] Specifically, the job running time threshold can be obtained from the version library. The aging results of the historical version of the critical changed job and the aging results of the current version of the critical changed job can be obtained from the database. The path composed of critical jobs can be determined as the critical path and the length of the critical path can be calculated. The difference between the job running time threshold and the length of the critical path is determined as the job running window time margin. The aging results of the historical version of the critical changed job and the aging results of the current version of the critical changed job are compared with the job running window time margin, so as to determine the impact on the overall aging and identify the critical changed jobs with aging risks. Through the above method, the impact of each critical changed job on the overall aging can be judged, so as to identify the critical changed jobs with aging risks.

[0055] In some embodiments of this specification, to identify whether there is a timeliness risk for a critical change job based on the time margin of the job running window, the timeliness results of the historical versions of the critical change job, and the timeliness results of the current version of the critical change job, it may include: determining whether the critical change job is a newly added job; in the case where it is determined that the critical change job is a newly added job, judging whether the timeliness result of the current version of the critical change job is greater than the time margin of the job running window; in the case where it is judged that the timeliness result of the current version of the critical change job is greater than the time margin of the job running window, determining the critical change job as a timeliness risk job.

[0056] Specifically, the change type of the critical change job can be judged. If the critical change job is a newly added job, then compare the timeliness result of the current version with the time margin of the job running window. If the timeliness result of the current version exceeds the time margin of the job running window, it indicates that all jobs on the critical path cannot be completed within the set job running window time window, and at the same time, the overall timeliness of the job will be extended. Then this job is identified as a timeliness risk job, otherwise there is no relevant risk. Through the above method, it can be determined whether the newly added critical change job extends the overall timeliness, and thus determine whether this job is a timeliness risk job.

[0057] In some embodiments of this specification, after determining whether the critical change job is a newly added job, it may further include: in the case where it is determined that the critical change job is not a newly added job, determining whether the timeliness result of the current version of the critical change job is greater than the timeliness result of the historical version of the critical change job; in the case where it is determined that the timeliness result of the current version of the critical change job is greater than the timeliness result of the historical version of the critical change job, determining whether the difference between the timeliness result of the current version of the critical change job and the timeliness result of the historical version of the critical change job is greater than the time margin of the job running window; in the case where it is determined that the difference between the timeliness result of the current version of the critical change job and the timeliness result of the historical version of the critical change job is greater than the time margin of the job running window, determining the critical change job as a timeliness risk job.

[0058] Specifically, in the case where the critical change job is not a newly added job, the critical change job is a modified job. The timeliness result of the current version can be compared with the timeliness result of the historical version. If the timeliness result of the current version exceeds the timeliness result of the historical version, it indicates that this modification will extend the overall timeliness of the job. In this case, compare the difference between the timeliness result of the current version and the timeliness result of the historical version with the time margin. If the former is greater than the latter, it indicates that all jobs on the critical path cannot be completed within the set time window, and this job is identified as a timeliness risk job, otherwise there is no relevant risk. Through the above method, it can be determined whether the modified critical change job extends the overall timeliness, and thus determine whether this job is a timeliness risk job.

[0059] In some embodiments of this specification, after determining the change-critical operation as a time-limit risk operation, the following may further be included: archiving the data related to the time-limit risk operation to a database; and / or, pushing the data related to the time-limit risk operation to a user.

[0060] Specifically, after determining the change-critical operation as a time-limit risk operation, the data related to the time-limit risk operation may be archived to a database for subsequent query and acquisition. After determining the change-critical operation as a time-limit risk operation, the data related to the time-limit risk operation may be pushed to a user. Among them, the data related to the time-limit risk operation may include at least one of the following: operation identifier, operation script, operation running time, and operation dependency relationship related to the operation, etc. By pushing the data related to the time-limit risk operation to a user, it is convenient for the user to quickly troubleshoot and handle.

[0061] The above method will be described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating the present application and does not constitute an improper limitation to the present application.

[0062] Please refer to Figure 5 , which shows the flowchart of the data processing time-limit test method in this specific embodiment. As Figure 5 shown, in this specific embodiment, the method includes the following steps:

[0063] Step 1, source data processing: Process the source data to obtain the operation running time, operation dependency relationship, and change operation set.

[0064] Step 2, big data job model construction: By mutually mapping the nodes and edges in the AOE network with big data job scheduling, a big data job model based on the AOE network is constructed.

[0065] Step 3, loop graph determination: Check whether there is a loop graph in the big data job model. If so, it indicates that it does not conform to the data processing logic, and relevant personnel will be automatically notified for troubleshooting.

[0066] Step 4, critical path construction: Use the critical path algorithm to perform path analysis on the big data job model to identify the critical path and critical operations that affect the overall time limit.

[0067] Step 5, job change set matching: Match the critical operations with the change operation set to obtain the change-critical operations.

[0068] Step 6, collection and deployment of change-critical operation scripts: Obtain the current version and historical version of the change-critical operation from the version library, label the version tags respectively, and then upload the job scripts of the two versions to the directory of the scripts to be executed of the job scheduling server.

[0069] Step 7, Change the Precision Test for Key Operations to Automation: The job monitoring process periodically scans the directory of scripts to be executed on the job scheduling server. If a script to be executed is found, the job scheduling framework is automatically run to execute the job script. After the execution is completed, the elapsed time of the job is recorded, and the corresponding script in the directory of scripts to be executed is deleted.

[0070] Step 8, Aging Risk Identification: Evaluate the running aging of the changed key operations, compare it with the time margin of the job running window, so as to determine the impact on the overall aging, and identify the changed key operations with aging risks.

[0071] Step 9, Automatically Push the Aging Risk Jobs to Relevant Personnel, and Suggest that Relevant Personnel Optimize the Aging of the Aging Risk Jobs or Adjust the Job Running Time Threshold.

[0072] The method in the above embodiments is highly automated and has been implemented for actual use, without relying on testers, and realizes the precision test of changed key operations unattended, quickly identifies aging risks, and effectively improves the test efficiency. Please refer to the following table to compare this solution with traditional manual testing, and it can be seen the beneficial effects that this solution can bring.

[0073] Table 1

[0074]

[0075] Based on the same inventive concept, an apparatus for testing data processing aging is also provided in the embodiments of this specification, as described in the following embodiments. Since the principle of the apparatus for testing data processing aging to solve problems is similar to that of the method for testing data processing aging, the implementation of the apparatus for testing data processing aging can refer to the implementation of the method for testing data processing aging, and the repeated parts will not be described again. Hereinafter, the term "unit" or "module" may be a combination of software and / or hardware that can implement a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and contemplated. Figure 6 is a structural block diagram of the apparatus for testing data processing aging in the embodiments of this specification, as Figure 6 shown, including: an acquisition module 601, a determination module 602, a matching module 603, and an identification module 604. The following describes this structure.

[0076] The acquisition module 601 is used to acquire the job running time data, job dependency data, and changed job set of multiple jobs, where the multiple jobs include data processing of specified data according to a preset logic.

[0077] The determination module 602 is configured to determine critical jobs that affect the overall timeliness of jobs among multiple jobs according to job running time data and job dependency data.

[0078] The matching module 603 is configured to match critical jobs with a set of changed jobs to obtain changed critical jobs.

[0079] The identification module 604 is configured to perform a timeliness test on the changed critical jobs, and identify whether there is a timeliness risk for the changed critical jobs according to the timeliness test results.

[0080] The following refers to Figures 7 to 12 A specific embodiment is used to illustrate the above device. However, it should be noted that this specific embodiment is only for better explaining the present application and does not constitute an improper limitation to the present application.

[0081] Please refer to Figure 7 , which shows a schematic structural diagram of the data processing timeliness test device in this specific embodiment. As Figure 7 shown, in this specific embodiment, the method includes the following steps:

[0082] Figure 7 It shows a schematic structural diagram of a data processing timeliness test device provided by an embodiment of the present application. As Figure 7 shown, the device may include: a source data processing module 701, a big data job model construction module 702, a critical path construction module 703, a precise test module 704, and a timeliness risk identification module 705. Each module will be described below.

[0083] The source data processing module 701 is configured to perform preprocessing on source data and incorporate it into a database, which can be divided into three parts. 1) Acquisition of job running time data, which collects job running time data from the operation and maintenance platform in near real time through continuous integration tasks; 2) Acquisition of job dependency data, which automatically generates job dependency data by performing blood relationship analysis on the data processing logic in job scripts; 3) Acquisition of a set of changed jobs by querying code commit records in the version library.

[0084] The big data job model construction module 702 is configured to construct a big data job model. Based on job running time data and job dependency data, a big data job model is constructed based on the AOE network, where each job is used as a vertex, the dependency relationship is used as a directed edge, and the running time is used as the weight on the edge.

[0085] The critical path construction module 703 is used to construct the critical path. The critical path algorithm is used to perform path analysis on the big data job model to identify the critical path and critical jobs that affect the overall timeliness. According to the AOE network diagram, topological sorting is performed starting from the source point and the sink point respectively, and the earliest start time and the latest start time of each job are calculated. If the earliest start time is equal to the latest start time, then the job is a critical job, and the path composed of critical jobs is the critical path.

[0086] The precise test module 704 is used to perform precise tests on the changed critical jobs. The changed job set is obtained from the version library and matched with the above-mentioned critical job set to automatically perform precise tests on the changed critical jobs. The historical version and the current version of the above-mentioned changed critical jobs are obtained from the version library, and the timeliness verification of the historical version and the current version is automatically performed through the job scheduling framework.

[0087] The timeliness risk identification module 705 is used to identify the changed jobs with timeliness risks. The precise test results are evaluated and compared with the time margin of the job running window (job running time threshold - critical path length) to identify the changed jobs with timeliness risks, thereby determining the impact on the overall timeliness and pushing the results to relevant personnel.

[0088] Please refer to Figure 8 , Figure 8 FIG. shows a schematic diagram of the source data processing module 800 in an embodiment of the present application. As Figure 8 shown, the source data processing module 800 may include: an information collector 801, a script parser 802, a version controller 803, and an information storage 804.

[0089] The information collector 801 can quasi-real-time collect job running time data from the operation and maintenance platform through a continuous integration task; format the data to unify the measurement unit of time.

[0090] The script parser 802 can obtain the job script from the version library, form the flow direction of the data stream through blood relationship analysis of the data processing logic in the job script, and thus automatically generate job dependency data.

[0091] The version controller 803 can obtain the changed job set by querying the code submission records in the version library.

[0092] The information storage 804 can save the job running time data obtained by the information collector 801, the job dependency data obtained by the script parser 802, and the changed job set obtained by the version controller 803 into a database. The database type can be a relational database such as mysql or oracle.

[0093] Please refer to Figure 9 ,Figure 9 It shows a schematic diagram of the big data job model construction module 900 in an embodiment of the present application. As Figure 9 shown, the big data job model construction module 900 may include: a point set builder 901, an edge set builder 902, an edge weight annotator 903, and a big data job model builder 904 based on the AOE network.

[0094] The point set builder 901 can obtain the information of all jobs from the information storage 804, including job names, applications to which the jobs belong, etc., and perform deduplication. Each job is defined as a vertex (Node), and each vertex contains three attributes: job name, subsequent job name, and job running time. After all vertices update the job name attribute according to the job information, they are added to the point set.

[0095] The edge set builder 902 can obtain all job dependencies from the information storage 804, and define each dependency as a link in a directed linked list (Link). Each link contains the current vertex and the subsequent vertex. After the directed linked list updates the subsequent job name attributes of all vertices according to the job dependencies, it is added to the edge set.

[0096] The edge weight annotator 903 can obtain the running times of all jobs from the information storage 804, and update the job running time attributes of all vertices accordingly.

[0097] The big data job model builder 904 based on the AOE network can construct an AOE network graph (Graph) according to the point set builder 901, the edge set builder 902, and the edge weight annotator 903. It judges whether the graph has a cycle according to the topological sorting algorithm. If so, it does not conform to the data processing logic, and automatically notifies relevant personnel for investigation.

[0098] Please refer to Figure 10 , Figure 10 It shows a schematic diagram of the critical path construction module 1000 in an embodiment of the present application. As Figure 10 shown, the critical path construction module 1000 may include a job earliest start time set builder 1001, a job latest start time set builder 1002, a critical job generator 1003, and a critical path generator 1004.

[0099] The earliest start time set builder 1001 of jobs can perform a topological sort starting from the source node according to the AOE network diagram, calculate the earliest start time of each job, and add it to the earliest start time set of jobs. Specifically, the earliest start time set builder 1001 of jobs can perform the following steps: (1) Initialize the earliest start time set of jobs, and uniformly assign 0 to the earliest start time of all jobs; (2) Start from the source node; (3) Perform a topological sort according to the AOE network; (4) Obtain the running time T i , that is, the job running time attribute of the current vertex; (5) Obtain the earliest start time ET i-1 of the previous job from the set; (6) Calculate the earliest start time of the current job: ET i = max(ET i-1 + T i ), if there are multiple previous jobs, take the maximum value of the calculation results as the earliest start time of this job; (7) Add the earliest start time ET i of the current job to the earliest start time set of jobs; (8) Determine whether it is a sink node, if so, execute (10), otherwise return to (3) and perform a topological sort on the next vertex; (10) Obtain the set of the earliest start times of all jobs.

[0100] The latest start time set builder 1002 of jobs can perform a topological sort starting from the sink node according to the AOE network diagram, calculate the latest start time of each job, and add it to the latest start time set of jobs. It is a reverse process compared to the earliest start time set builder 1001 of jobs.

[0101] The critical job generator 1003 can obtain the earliest start time and the latest start time of each job from the earliest start time set builder 1001 of jobs and the latest start time set builder 1002 of jobs respectively. If the two are equal, then this job is a critical job.

[0102] The critical path generator 1004 can obtain critical jobs from the critical job generator 1003, and the path composed of critical jobs is the critical path.

[0103] Please refer to Figure 11 , which shows a schematic diagram of the precision test module 1100 in an embodiment of the present application. As Figure 11 shown, the precision test module 1100 can include a change matcher 1101, a script collector 1102, a job scheduler 1103, and a timeliness memory 1114.

[0104] The change matcher 1101 can obtain the changed job set from the information memory 804, obtain critical jobs from the critical job generator 1003, and match the two to obtain changed critical jobs.

[0105] The script collector 1102 can obtain the historical version and the current version of the changed critical job from the repository. After marking the version tags respectively, it uploads the job scripts of the two versions to the directory of the scripts to be executed of the job scheduling server.

[0106] The job scheduler 1103 can use the job monitoring process to periodically scan the directory of the scripts to be executed of the job scheduling server. If there is a script to be executed, it automatically runs the job scheduling framework to execute the job script. After the execution is completed, it takes the running time as the timeliness data and incorporates it into the timeliness memory 1104, and deletes the corresponding script in the directory of the scripts to be executed.

[0107] The timeliness memory 1104 can record the job running time and save it to the database. The database type can be a relational database such as mysql or oracle.

[0108] Please refer to Figure 12 , which shows a schematic diagram of the timeliness risk identification module 1200 in the embodiments of the present application. As Figure 12 shown, the timeliness risk identification module 1200 can include a time margin calculator 1201, a timeliness risk identifier 1202, and a risk pusher 1203.

[0109] The time margin calculator 1201 can obtain the configuration of the job running time threshold from the repository, calculate the critical path length according to the critical path generator 1004, and subtract the two to calculate the time margin of the job running window.

[0110] The timeliness risk identifier 1202 can obtain the timeliness result from the timeliness memory 1104 and evaluate it, and compare it with the time margin of the job running window obtained by the time margin calculator 1201, so as to determine the impact on the overall timeliness and identify the changed jobs with timeliness risks. Specifically, the timeliness risk identifier 1202 can be used to perform the following steps: judge the job change type. If it is a newly added job, compare the current version timeliness result with the time margin. If the current version timeliness result exceeds the time margin, it indicates that all jobs on the critical path cannot be executed within the set time window, and at the same time, it will extend the overall timeliness of the job, and this job is identified as a timeliness risk job, otherwise there is no relevant risk. If it is a modified job, compare the current version timeliness result with the historical version timeliness result. If the current version timeliness result exceeds the historical version timeliness result, it indicates that this modification will extend the overall timeliness of the job. Then compare the difference between the current version timeliness result and the historical version timeliness result with the time margin. If the former is greater than the latter, it indicates that all jobs on the critical path cannot be executed within the set time window, and this job is identified as a timeliness risk job, otherwise there is no relevant risk. Archive the timeliness risk jobs into the database.

[0111] The risk pusher 1203 can automatically push the time-limit risk operations obtained by the time-limit risk identifier 1202 to relevant personnel via email, and suggest that the relevant personnel optimize the time-limit of the time-limit risk operations or adjust the operation running time threshold.

[0112] From the above description, it can be seen that the embodiments of this specification achieve the following technical effects: The key operations affecting the overall time-limit of the operations can be determined, and they are matched with the change operation set to obtain the changed key operations. By testing the changed key operations, the test scope is greatly reduced. Through the time-limit verification of the time-limit test results, the changed key operations with time-limit risks can be identified. The above method has a high degree of automation, can accurately locate the changed key operations with time-limit risks, and make a judgment on the impact on the overall time-limit, shortening the verification time cycle and greatly improving the test efficiency of big data operations.

[0113] The embodiments of this specification also provide a computer device, which can be specifically referred to Figure 13 to the schematic structural diagram of the computer device based on the data processing time-limit test method provided by the embodiments of this specification shown in the figure. The computer device can specifically include an input device 131, a processor 132, and a memory 133. Among them, the memory 133 is used to store the executable instructions of the processor. When the processor 132 executes the instructions, it implements the steps of the data processing time-limit test method described in any of the above embodiments.

[0114] In this embodiment, the input device can specifically be one of the main devices for information exchange between users and computer systems. The input device can include a keyboard, a mouse, a camera, a scanner, a light pen, a handwriting input board, a voice input device, etc.; the input device is used to input the original data and the programs for processing these data into the computer. The input device can also obtain and receive the data transmitted from other modules, units, and devices. The processor can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or a processor, a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. The memory can specifically be a memory device used to store information in modern information technology. The memory can include multiple levels. In a digital system, as long as it can store binary data, it can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as a RAM, a FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, a TF card, etc.

[0115] In this embodiment, the functions and effects specifically implemented by the computer device can be explained by comparison with other embodiments, and will not be elaborated here.

[0116] This specification embodiment also provides a computer storage medium based on a data processing timeliness test method. The computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the data processing timeliness test method described in any of the above embodiments are implemented.

[0117] In this embodiment, the above storage medium includes but is not limited to Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards stipulated by the communication protocol and is used for the interface of network connection communication.

[0118] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained by comparison with other embodiments, and will not be elaborated here.

[0119] Obviously, those skilled in the art should understand that the above modules or steps of the embodiments of this specification can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in the storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of this specification are not limited to any specific combination of hardware and software.

[0120] It should be understood that the above description is for illustrative purposes and not for limitation. Many embodiments and many applications other than the examples provided will be obvious to those skilled in the art. Therefore, the scope of this application should not be determined by the above description, but should be determined by the full scope of the foregoing claims and the equivalents of these claims.

[0121] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for testing data processing timeliness, characterized in that, Including: Obtain the job running time data, job dependency data, and changed job set of multiple jobs, where the multiple jobs include data processing of specified data according to a preset logic; the changed job set is a set of jobs newly added or modified when updating the version; Determine the critical jobs that affect the overall timeliness of the multiple jobs according to the job running time data and the job dependency data; the critical jobs are the jobs on the critical path, and the critical path is the logical path with the longest delay from input to output; Match the critical jobs with the changed job set, and regard the critical jobs existing in the changed job set as changed critical jobs; Conduct timeliness tests on the historical version and the current version of the changed critical jobs, and identify whether there is a timeliness risk for the changed critical jobs according to the timeliness test results.

2. The method according to claim 1, characterized in that, Obtain the job running time data, job dependency data, and changed job set of multiple jobs, including: Collect the job running time data of multiple jobs from the job operation and maintenance platform; Conduct lineage analysis on the data processing logic in the job scripts of the multiple jobs to generate the job dependency data of the multiple jobs; Query the job code submission records in the job version library to obtain the changed job set of multiple jobs.

3. The method according to claim 1, characterized in that, Determine the critical jobs that affect the overall timeliness of the multiple jobs according to the job running time data and the job dependency data, including: Construct an AOE network with the multiple jobs according to the job running time data and the job dependency data, where each job in the multiple jobs is used as a vertex of the AOE network, the job dependencies between jobs are used as directed edges of the AOE network, and the job running time of each job is used as the weight value on the directed edge; Based on the AOE network, determine the critical jobs that affect the overall timeliness of the multiple jobs.

4. The method according to claim 3, characterized in that, Based on the AOE network, determine the critical jobs that affect the overall timeliness of the multiple jobs, including: Determine whether there is a cyclic graph in the AOE network; In the case of determining that there is no cyclic graph in the AOE network, based on the AOE network, determine the critical jobs that affect the overall timeliness of the multiple jobs.

5. The method according to claim 4, characterized in that, After determining whether there is a cyclic graph in the AOE network, it further includes: In the case of determining that there is a cyclic graph in the AOE network, determine that the multiple jobs do not conform to the data processing logic, and generate a notification message to be sent to the user.

6. The method according to claim 3, characterized in that, Based on the AOE network, determine the critical jobs that affect the overall timeliness of the multiple jobs, including: Start from the source point of the AOE network for topological sorting, and calculate the earliest start time of each job in the multiple jobs; Start from the sink point of the AOE network for topological sorting, and calculate the latest start time of each job in the multiple jobs; Determine whether the earliest start time and the latest start time of each job in the multiple jobs are equal; Regard the jobs with equal earliest start time and latest start time as the critical jobs that affect the overall timeliness of the jobs.

7. The method according to claim 1, characterized in that, Conduct timeliness tests on the historical version and the current version of the critical change job, including: Obtain the historical version and the current version of the critical change job from the job version library; Upload the job scripts of the historical version and the current version to the directory of scripts to be executed on the job scheduling server. The job scheduling server uses a job monitoring process to periodically scan the directory of scripts to be executed. If a script to be executed is found, it runs the job scheduling framework to execute the job script. After the job script is executed, the time taken to execute the job script is determined as the timeliness result of the job corresponding to the job script and stored in the database, and the corresponding job script in the directory of scripts to be executed is deleted.

8. The method according to claim 7, characterized in that, According to the timeliness test results, identify whether there is a timeliness risk for the critical change job, including: Obtain the timeliness results of the historical version of the critical change job and the timeliness results of the current version of the critical change job from the database; Obtain the job running time threshold from the job version library; Determine the path composed of the critical jobs as the critical path and calculate the length of the critical path; Determine the difference between the job running time threshold and the length of the critical path as the job running window time margin; According to the job running window time margin, the timeliness results of the historical version of the critical change job, and the timeliness results of the current version of the critical change job, identify whether there is a timeliness risk for the critical change job.

9. The method according to claim 8, characterized in that,According to the job running window time margin, the timeliness results of the historical version of the critical change job, and the timeliness results of the current version of the critical change job, identify whether there is a timeliness risk for the critical change job, including: Determine whether the critical change job is a newly added job; In the case where it is determined that the critical change job is a newly added job, judge whether the timeliness result of the current version of the critical change job is greater than the job running window time margin; In the case where it is judged that the timeliness result of the current version of the critical change job is greater than the job running window time margin, determine the critical change job as a job with timeliness risk.

10. The method according to claim 9, wherein After determining whether the critical change job is a newly added job, it also includes: In the case where it is determined that the critical change job is not a newly added job, determine whether the timeliness result of the current version of the critical change job is greater than the timeliness result of the historical version of the critical change job; In the case where it is determined that the timeliness result of the current version of the critical change job is greater than the timeliness result of the historical version of the critical change job, determine whether the difference between the timeliness result of the current version of the critical change job and the timeliness result of the historical version of the critical change job is greater than the job running window time margin; In the case where it is determined that the difference between the timeliness result of the current version of the critical change job and the timeliness result of the historical version of the critical change job is greater than the job running window time margin, determine the critical change job as a job with timeliness risk.

11. The method according to claim 9 or 10, wherein After determining the critical change job as a job with timeliness risk, it also includes: Archive the data related to the time-limit risk operations into a database; and / or, Push the data related to the time-limit risk operations to the user.

12. A data processing timeliness testing device, wherein Comprising: An acquisition module, configured to acquire the operation running time data, operation dependency data, and changed operation set of multiple operations, wherein the multiple operations include data processing of specified data according to a preset logic; the changed operation set is a set of operations newly added or modified when updating the version; A determination module, configured to determine the critical operations that affect the overall timeliness of the multiple operations according to the operation running time data and the operation dependency data; the critical operations are the operations on the critical path, and the critical path is the logical path with the longest delay from input to output; A matching module, configured to match the critical operations with the changed operation set, and use the critical operations existing in the changed operation set as changed critical operations; An identification module, configured to perform timeliness tests on the historical version and the current version of the changed critical operations, and identify whether there is a timeliness risk for the changed critical operations according to the timeliness test results.

13. A computer device, wherein Comprising a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having computer instructions stored thereon, wherein When the instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Optimized task scheduling method and device based on data warehouse, equipment and medium

    CN111309712A

  • Automatic test monitoring method, device and equipment and storage medium

    CN111858352A