Data exchange scheduling method, system, electronic device, medium and program product
By defining preprocessing and dynamically generating job instances for to-process files, the problem that in the prior art generation of jobs cannot be adapted to actual changes is solved, efficient data exchange scheduling is achieved, and the system's fault tolerance and processing performance is improved.
Patent Information
- Application Number
- CN202111279303.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-10-29
Smart Images

Figure CN113918525B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a data exchange scheduling method, system, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] As the scale of enterprise business systems continues to expand, data exchange platforms help enterprises share data quickly and efficiently between systems, reduce data processing links, and improve data processing efficiency. For enterprise-level data exchange platforms that use files as the main object, there are massive files that need to be exchanged every day. The number, content, time, and frequency of files uploaded by file suppliers (referred to as upstream) are different every day, and each file involves multiple demanders (referred to as downstream). The files of each demander require different processing operations for processing, and there are execution order requirements between various operations. Data exchange scheduling matches the corresponding processing operations and operation execution links for each file that arrives at the platform, and then performs job scheduling processing.
[0003] At present, the processing method of the data exchange job scheduling system is to pre-generate jobs and job dependencies, and to detect the arrival of files to trigger the scheduling system to schedule related jobs. This method can solve the problem of reducing human intervention in the scheduling system, but there are at least two problems with the existing technology: 1) Jobs generated in advance cannot adapt to the actual changes in exchange data. Since the system involves many upstreams, if there is a file name and definition that is different from the upstream, the subsequent jobs will not be able to be called up. If the content and frequency of the upstream downloads change, the system will not be able to handle it. 2) Generating a full amount of jobs in advance will occupy a large amount of system resources and thus affect the availability of system services. Summary of the invention
[0004] To solve the problems existing in the prior art, the embodiments of the present disclosure provide a data exchange scheduling method, system, electronic device, computer-readable storage medium and computer program product to realize the conversion and scheduling of to-be-processed files to to-be-processed jobs.
[0005] The first aspect of the present disclosure provides a data exchange scheduling method, including: defining and preprocessing a file to be processed to generate job definition information and job dependency information of the file to be processed; instantiating the file to be processed according to the job definition information to obtain a job instance of the file to be processed, and generating dependency instance information of the job to be processed according to the job dependency information and the job instance; judging whether the job to be processed satisfies the dependency of a preceding job according to the dependency instance information, and if so, setting the job to be processed as a job to be scheduled and scheduling it.
[0006] Furthermore, the files to be processed are defined and preprocessed to generate job definition information and job dependency information of the files to be processed, including: obtaining basic information of multiple files to be processed; defining exchange requirements for the basic information of multiple files to be processed to generate exchange information; reading the exchange information within a preset time, and extracting keyword fields of the exchange information to generate a data definition file for the files to be processed; generating job definition information according to the processing method and job type of each data record in the exchange information, and generating job dependency information according to the execution order of each data record and job type in the exchange information.
[0007] Furthermore, according to the job definition information, before instantiating the file to be processed, the method also includes: adding the downloaded directory information in the data definition file to the detection directory table; detecting whether the directory information in the detection directory table generates a checklist verification file; if so, detecting whether the checklist verification file matches the data definition file; if so, adding the data definition file to the processing directory file, and entering the detected actual file information into the source data information file; if not, adding the data definition file to the abnormal processing directory file.
[0008] Furthermore, according to the job definition information, the file to be processed is instantiated to obtain a job instance of the file to be processed, including: matching the file to be processed recorded in the source data information file with the job definition information and the data definition file, and performing macro replacement to convert it into an instance job to obtain a job instance of the file to be processed.
[0009] Furthermore, based on the job dependency information and the job instance, the dependency instance information of the job to be processed is generated, including: based on the job name in the job instance file and the post-job name in the job dependency information, the dependency instance information of the job to be processed is generated.
[0010] Further, determining whether the job to be processed satisfies the dependency of the predecessor job includes: determining whether the job to be processed has a predecessor job dependency; if not, setting the job to be processed as a job to be scheduled; if so, determining whether the predecessor job has been processed; if the predecessor job has been processed, setting the job to be processed as a job to be scheduled.
[0011] Furthermore, the method further includes: determining whether the job to be scheduled is a zombie job, wherein a zombie job represents a job to be scheduled that has not been processed within a preset time; if so, resetting the zombie job to be scheduled.
[0012] Furthermore, the pending job is set as a pending job and scheduled, including: adding the pending job to the distribution job queue; taking out multiple pending jobs from the distribution job queue in turn and assigning them to the idle node addresses in the job client, and setting the status of the pending job to pending; submitting multiple pending jobs to the job client for job processing, and updating the job status of the pending job after processing.
[0013] Furthermore, the job status after the pending job is processed includes at least: processing, processing completed, or processing failed.
[0014] The second aspect of the present disclosure provides a data exchange scheduling system, including: a file preprocessing module, which is used to define preprocessing of files to be processed and generate job definition information and job dependency information of the files to be processed; a job instantiation module, which is used to instantiate the files to be processed according to the job definition information, obtain the job instance of the files to be processed, and generate dependency instance information of the jobs to be processed according to the job dependency information and the job instance; a job scheduling module, which is used to determine whether the jobs to be processed meet the dependencies of the previous jobs according to the dependency instance information, and if so, set the jobs to be processed as the jobs to be scheduled and schedule them.
[0015] Furthermore, the file preprocessing module is used to define preprocessing of the files to be processed and generate job definition information and job dependency information of the files to be processed, including: obtaining basic information of multiple files to be processed; defining exchange requirements for the basic information of multiple files to be processed and generating exchange information; reading the exchange information within a preset time, and extracting keyword fields of the exchange information to generate a data definition file for the files to be processed; generating job definition information according to the processing method and job type of each data record in the exchange information, and generating job dependency information according to the execution order of each data record and job type in the exchange information.
[0016] Furthermore, the system also includes: a file verification module, which is used to verify the files to be processed, including: adding the downloaded directory information in the data definition file to the detection directory table; checking whether a list verification file is generated for the directory information in the detection directory table; if so, checking whether the list verification file matches the data definition file; if they match, adding the data definition file to the processing directory file, and entering the actual file information detected into the source data information file; if they do not match, adding the data definition file to the abnormal processing directory file.
[0017] The third aspect of the present disclosure provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the data exchange scheduling method provided by the first aspect of the present disclosure is implemented.
[0018] The fourth aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the data exchange scheduling method provided by the first aspect of the present disclosure is implemented.
[0019] A fifth aspect of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the data exchange scheduling method provided by the first aspect of the present disclosure.
[0020] The present disclosure provides a data exchange scheduling method, system, electronic device, computer-readable storage medium and computer program product. The method and system are divided into two stages: pre-generation of job definitions and rapid generation of instance jobs from job definitions. The files to be processed are registered once for multiple times, which reduces the user usage cost. The method and system can dynamically generate jobs according to the actual files uploaded by the user, without limiting the time point and frequency of file upload, and has high fault tolerance. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] For a more complete understanding of the present disclosure and its advantages, reference will now be made to the following description taken in conjunction with the accompanying drawings, in which:
[0022] Figure 1 The following schematically shows an application scenario diagram of a data exchange scheduling method according to an embodiment of the present disclosure;
[0023] Figure 2 A flowchart of a data exchange scheduling method according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 3 A flowchart for defining preprocessing of a file to be processed according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 4 A flowchart of data file verification according to an embodiment of the present disclosure is schematically shown;
[0026] Figure 5 A flowchart of converting a job to be processed into a job to be scheduled according to an embodiment of the present disclosure is schematically shown;
[0027] Figure 6 A flowchart for scheduling a job to be scheduled according to an embodiment of the present disclosure is schematically shown;
[0028] Figure 7 A block diagram of a data exchange scheduling system according to an embodiment of the present disclosure is schematically shown;
[0029] Figure 8 A block diagram of a data exchange scheduling system according to another embodiment of the present disclosure is schematically shown;
[0030] Fig. 9 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0033] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0035] Some block diagrams and / or flow charts are shown in the accompanying drawings. It should be understood that some boxes or combinations thereof in the block diagrams and / or flow charts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that these instructions can create a device for implementing the functions / operations described in these block diagrams and / or flow charts when executed by the processor. The technology of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). In addition, the technology of the present disclosure can take the form of a computer program product on a computer-readable storage medium storing instructions, which can be used by an instruction execution system or in combination with an instruction execution system.
[0036] The disclosed embodiment provides a data exchange scheduling method that can be applied to a data exchange platform. When faced with changes in the number, content, time, and frequency of massive files, possible processing interruptions or problems that cannot be processed are avoided through dynamic job derivation. After decoupling the job derivation into two stages, most of the pressure of centralized job derivation is pre-placed in the first stage to generate job definitions. The second stage can quickly generate instance jobs according to the job definitions. At the same time, the file processing method fully releases the processing performance of the data exchange platform and avoids the platform pressure caused by job backlogs due to centralized time processing, so that the platform can process more data. The upstream and downstream data can complete the construction of the file exchange channel through simple definitions, and one definition can be used multiple times. The file can generate jobs according to the actual upstream download situation. Changes in the number, content, time, and frequency of files will not affect the generation of file jobs, which improves the fault tolerance of the system.
[0037] Figure 1 The exemplary system architecture 100 that can be applied to the data exchange scheduling method according to the embodiment of the present disclosure is schematically shown. It should be noted that: Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0038] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0039] Users (such as R&D personnel) can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as programming systems for various language software, test systems, web browser applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0040] The terminal devices 101, 102, 103 may be various electronic devices with display screens and supporting web browsing. The terminal devices 101, 102, 103 provide operating platforms for upstream users, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0041] The server 105 may be a server that provides various services, and may be a deployment server of the present method, such as a background management server (only an example) that provides support for applications used by users using the terminal devices 101, 102, and 103. The background management server may process files for received user requests, and feed back the processing results (such as job scheduling and processing, etc.) to the terminal device.
[0042] It should be noted that the data exchange scheduling method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the data exchange scheduling system provided in the embodiment of the present disclosure can generally be deployed in the server 105. The data exchange scheduling method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. The server cluster can be a data exchange platform server cluster, which can specifically include: a data preprocessor, a data detector, a job generator, a scheduling service end and a job client, etc. The data exchange platform server cluster will perform data pre-definition processing, instance data entry, job instantiation, job scheduling and processing. Correspondingly, the data exchange scheduling system provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0044] Figure 2 The flowchart of the data exchange scheduling method according to the embodiment of the present disclosure is schematically shown. Figure 2 As shown, the method includes: steps S201~S203.
[0045] In operation S201, definition preprocessing is performed on the file to be processed to generate job definition information and job dependency information of the file to be processed.
[0046] In operation S202, the file to be processed is instantiated according to the job definition information to obtain a job instance of the file to be processed, and dependent instance information of the job to be processed is generated according to the job dependency information and the job instance.
[0047] In operation S203, it is determined whether the job to be processed satisfies the dependency of the predecessor job according to the dependency instance information. If so, the job to be processed is set as a job to be scheduled and scheduled.
[0048] The following is a detailed description of an example flow of each step of the data exchange scheduling method of this embodiment.
[0049] In operation S201, definition preprocessing is performed on the file to be processed to generate job definition information and job dependency information of the file to be processed.
[0050] In the embodiments of the present disclosure, the data exchange scheduling method can be applied to electronic devices, which may include but are not limited to servers, server clusters, etc. Various application systems, data exchange scheduling systems, etc. can be deployed in the electronic devices, and all the files to be downloaded entered upstream are stored on the server. Among them, the files to be processed can be files in any format, including but not limited to .xls format files, etc.
[0051] According to the embodiments of the present disclosure, Figure 3 As shown, the file to be processed is defined and pre-processed to generate job definition information and job dependency information of the file to be processed, including: steps S301~S304.
[0052] In operation S301, basic information of a plurality of files to be processed is obtained.
[0053] In operation S302, basic information of a plurality of files to be processed is subjected to exchange requirement definition to generate exchange information.
[0054] Specifically, the basic information of all files to be downloaded recorded by the upstream on the data exchange platform is stored in the database of the data exchange platform, including but not limited to the file name, file field name (including the encoding type of the field) and file download information. Each file entered can be used by all downstreams. The downstream selects the file in the system as needed and defines the required processing method and receiving directory to complete the exchange requirement definition and form exchange information. Among them, the file name represents the unique primary key to distinguish the files to be processed, the file field name (including the encoding type of the field) refers to the file content structure information, and the file download information refers to the directory of the upstream downloaded data file.
[0055] In operation S303, the exchange information within a preset time is read, and the keyword field of the exchange information is extracted to generate a data definition file of the file to be processed.
[0056] In the embodiment of the present disclosure, this step can be processed by a data preprocessor, which reads the exchange information within the preset time and extracts the keyword field of the exchange information to generate a data definition file for the file to be processed. The preset time can be any time period set according to the needs, such as 1 hour, 2 hours, 6 hours, 24 hours, etc. The keyword field of the exchange information includes but is not limited to the file name, upstream identifier, downstream identifier, upstream download directory, downstream receiving directory, fields under the processing method, etc. The data definition file generated thereby also includes but is not limited to the file name, upstream identifier, downstream identifier, upstream download directory, downstream receiving directory and processing method.
[0057] In operation S304, job definition information is generated according to the processing mode and job type of each data record in the exchange information, and job dependency information is generated according to the execution order of each data record in the exchange information and the job type.
[0058] In the embodiment of the present disclosure, step S304 is executed by the data preprocessor, which specifically searches for the corresponding processing program command and job type according to the processing method of each data record, and performs job command definition splicing. The format of the job command definition splicing can be "processing command + # file source path # + # business date #" (#XXX# is a macro definition, which will be replaced with valid data in the instantiation stage), and generates job definition information. The job definition information can be saved in a job definition file in a preset format.
[0059] Furthermore, according to each data record in the exchange information and the execution order of the job type, each job name and the corresponding post-job name relationship record is used to generate job dependency information, which can be saved in a job dependency definition file in a preset format.
[0060] It should be noted that the file formats mentioned above (such as data definition files, job definition files, job dependency definition files, etc.) can be set according to actual needs, including but not limited to .xls format files, etc.
[0061] According to an embodiment of the present disclosure, after the to-be-processed file is pre-defined and processed, before the to-be-processed file is instantiated according to the job definition information, the method further includes a file verification process, which can be performed by a data detector, specifically including: Figure 4 As shown, the download directory information in the data definition file is added to the detection directory table; the directory information in the detection directory table is checked to see whether a checklist file (i.e., .CHK file) is generated; if so, after checking whether the checklist file is complete, check whether it matches the data definition file; if it matches, add the data definition file to the processing directory file, and enter the detected actual file information into the source data information file. If it does not match, add the data definition file to the abnormal processing directory file, and enter the error file into the data abnormality file. Specifically, the checklist file is checked to see whether the file already exists in the current directory and is of the same size, and the detected actual file information is entered into the source data information file, which at least includes the file name, file processing directory, processing type, business date, upstream name, downstream name, serial number, and job derivative status. Among them, the file processing directory refers to the directory where the file is currently located; the processing type refers to the way the file needs to be processed, such as 1, 2, 3, etc. The numbers 1, 2, and 3 represent different types of processing. For example, 1 can represent host decoding, etc., which can be set according to actual needs; the upstream name and the downstream name correspond to the upstream identifier and the downstream identifier respectively; the sequence number refers to the order of multiple downloads of the same file on the same business date, which solves the problem of uncertain upstream download frequency; the job derivation status includes but is not limited to waiting or scheduled.
[0062] In the embodiments of the present disclosure, a file verification process is set before the file instantiation process is performed to avoid the situation where the job generated in advance cannot adapt to the actual changed exchange data. Since the system involves more upstreams, if the file name and definition uploaded from the upstream are different, the subsequent jobs will not be able to be called up, so that the subsequent steps cannot be implemented normally.
[0063] In operation S202, the file to be processed is instantiated according to the job definition information to obtain a job instance of the file to be processed, and dependent instance information of the job to be processed is generated according to the job dependency information and the job instance.
[0064] In the embodiment of the present disclosure, the instantiation processing process of the file to be processed is executed by the job generator, specifically, the file to be processed is instantiated according to the job definition information to obtain a job instance of the file to be processed, including: matching the file to be processed recorded in the source data information file with the job definition information and the data definition file, and replacing the macro to convert it into an instance job to obtain a job instance of the file to be processed.
[0065] Specifically, the polling source data information file adds a label to the record to be processed and inserts it into the source data temporary file, searches the source data temporary file for the corresponding job in the job definition file according to the file name and upstream and downstream identifiers, replaces #file source path# in the job definition file with the file processing directory in the source data information file, and replaces #business date# with the corresponding business log to form a valid job command, sets the instantiated job information to a pending verification state, and obtains a job instance of the file to be processed. Among them, the job instance instantiates the job on behalf of the file to be processed, and generates a corresponding job instance file, and the information related to the job instance is saved in the job instance file.
[0066] Further, based on the job dependency information and the job instance, the dependency instance information of the job to be processed is generated, specifically including: based on the job name in the job instance file and the post-job name in the job dependency information, the dependency instance information of the job to be processed is generated. The dependency instance information includes but is not limited to the name of the predecessor job, the business date of the predecessor job, the serial number of the predecessor job, the name of the post-job, the business date of the post-job, the serial number of the post-job, and whether the dependency has been satisfied, etc. The dependency instance information can be saved in the dependency instance file.
[0067] In operation S203, it is determined whether the job to be processed satisfies the dependency of the predecessor job according to the dependency instance information. If so, the job to be processed is set as a job to be scheduled and scheduled.
[0068] In the embodiment of the present disclosure, the scheduling process of the pending job is performed by the scheduling server. Figure 5 As shown, specifically: determine whether the job to be processed satisfies the dependency of the predecessor job, specifically including: determine whether the job to be processed has a predecessor job dependency; if not, set the job to be processed as a job to be scheduled; if so, determine whether the predecessor job has been processed, and if the predecessor job has been processed, set the job to be processed as a job to be scheduled.
[0069] Specifically, in order to prevent the scheduled job from not being successfully processed after a certain period of time, the method provided by the present disclosure further includes: determining whether the scheduled job is a dead job, wherein a dead job represents a scheduled job that has not been processed within a preset time; if so, resetting the dead job to the scheduled job. For example, determining whether there is a dead job that has not been processed for more than 10 minutes among the scheduled jobs, the dead job is a job that has not been processed or has failed, resetting such a job to the scheduled job and rescheduling it.
[0070] like Figure 6 As shown, in the embodiment of the present disclosure, the to-be-scheduled job is scheduled and processed, specifically including: adding the to-be-scheduled job to the distribution job queue; taking out multiple to-be-scheduled jobs from the distribution job queue in turn and assigning them to the idle node address in the job client, and setting the status of the to-be-scheduled job to be processed; submitting multiple to-be-scheduled jobs to the job client for job processing, and updating the job status of the to-be-processed jobs after processing. Specifically, before assigning multiple to-be-scheduled jobs to the job client, it is necessary to regularly check the available status of the job client and update the library table according to the registered client resource table, and then obtain the available job client address and the number of available resources according to the library table to form a job client resource list, and select the available job client according to the job client resource list for job scheduling. Among them, the number of to-be-scheduled jobs assigned to the available job client is set according to the data of the idle node address in the job client, and the embodiment of the present disclosure does not limit this.
[0071] After receiving the assigned pending jobs, the job client executes the job processing command. After each pending job is processed, the status of the job in the library table is set to completed, and the idle number of the job client resources is increased by "1". At the same time, the job client will trigger the scheduling of dependent jobs, match the completed job name with the job dependency information file to check whether there is a subsequent job. If so, the dependency status of the subsequent job is set to satisfied. Among them, the job status after the pending job is processed includes at least processing, processing completed, or processing failed.
[0072] It should be noted that the preset time of 10 minutes and the job status after the pending job is processed in the above embodiment are only exemplary illustrations, and do not constitute a limitation of the embodiments of the present disclosure. In actual application, the unprocessed or failed jobs within the preset time may be pending jobs within 20 minutes or half an hour, and the job status may also be set according to actual needs.
[0073] The following will be combined with a specific embodiment to describe in detail the process of job definition and instantiation processing of the to-be-processed file provided by the embodiment of the present disclosure. It should be noted that the following content is only to help those skilled in the art understand the technical content of the present disclosure, but does not mean the limitation of the intermediate information generated by each processing process of the embodiment of the present disclosure and the applicable scenarios.
[0074] First, the upstream enters the basic information of all downstream files into the system. The specific content is as follows:
[0075] Table 1 Basic information of downloaded files
[0076]
[0077] In Table 1, the file structure information file at least includes information such as file field name, field type, field length, decoding type, length after decoding, and position after decoding. The downstream selects files in the system as needed based on all files entered by the upstream and defines the required processing method and receiving directory to complete the exchange requirement definition and form exchange information.
[0078] The data preprocessor reads the exchange information generated by the registration in the last hour, extracts the keyword field of the exchange information and generates a data definition file of the file to be processed. Following the above embodiment, the data definition file at least includes the following contents:
[0079] Table 2 Data definition file example
[0080]
[0081] Among them, the .CHK file generated by the downstream receiving directory is used for file verification before the job is instantiated, so as to be used by the platform for verification, so as to avoid triggering the next step of processing before the file is received. The generated data definition file can be represented by an xls file, and the various information in the file are column values. It should be noted that the above processing method is automatically set according to the actual content of the file to be processed, and different numerical values represent different processing methods. The processing method in the embodiment of the present disclosure is "1" representing host decoding, which is only an exemplary description and does not constitute a limitation of the embodiment of the present disclosure.
[0082] According to the processing method of each data record, the corresponding processing program command and job type are found, and the job command definition is spliced. The format of the job command definition splicing can be "processing command + # file source path # + # business date #" (#XXX# is a macro definition, which will be replaced with valid data in the instantiation stage) to generate job definition information. The corresponding job definition file includes at least the following content:
[0083] Table 3 Example of a job definition file
[0084]
[0085] According to the execution order of the job types, the relationship between each job name and the corresponding post-job name is recorded in the generated job dependency information. The job dependency information can be saved in a job dependency definition file in a preset format. This step represents the completion of the job pre-derivation stage.
[0086] After the upstream downloads the file to the data exchange platform, the checklist file (.CHK file) is also downloaded. The .CHK file records the name, size, business date (date when the data in the file is generated) and other information of the downloaded file for the platform to verify. The file verification step can be performed by the data detector, including the following: Figure 4 In the detection process shown, after the .CHK file is verified, the detected actual file information is entered into the source data information file. Following the above embodiment, the source data information file at least includes the following contents:
[0087] Table 4 Example of source data information file
[0088]
[0089] In the embodiments of the present disclosure, the source data information file records the actual file information of the file to be processed, which includes but is not limited to the contents shown in Table 4 above. The format of the source data information file can be set according to application requirements, for example, it can be a .xls format file, etc.
[0090] Next, the polling source data information file is used to add a label to the record to be processed and insert it into the source data temporary file. The source data temporary file is used to search for the corresponding job in the job definition file according to the file name and upstream and downstream identifiers. The #file source path# in the job definition file is replaced with the file processing directory in the source data information file, and the #business date# is replaced with the corresponding business log to form a valid job command. The instantiated job information is set to a pending verification state to obtain a job instance of the file to be processed, and a corresponding job instance file is generated. The above embodiment is used, and the job instance file at least includes the following contents:
[0091] Table 5 Example of job instance file
[0092]
[0093] Next, the job name in the job instance file is searched for a record with the same post-job name in the job dependency definition file, dependency instance information is generated, and the following job dependency instance file is formed accordingly. Following the above embodiment, the job dependency instance file includes at least the following contents:
[0094] Table 6 Example of job dependency instance file
[0095]
[0096] The source data information file with the above label is recorded and the job derivation status is updated to derivation, and the temporary file is deleted. In the embodiment of the present disclosure, the to-be-processed file is instantiated and then derived, and then the to-be-processed job is scheduled. The scheduling of the to-be-processed job can be performed by a scheduling server. A scheduling server is responsible for managing multiple job clients and processing them through multiple parallel threads. Different parallel threads process such as Figure 5 and Figure 6 The jobs shown are to be scheduled and assigned for submission and processing, which will not be described in detail here.
[0097] It should be noted that the contents of the basic information of the downloaded files, data definition files, job definition files, source data information files, job instance files and other files shown in the above embodiments are only examples. In other application scenarios and requirements, they can be replaced by other contents, and do not constitute a limitation of the embodiments of the present disclosure.
[0098] Figure 7 A block diagram of a data exchange scheduling system according to an embodiment of the present disclosure is schematically shown.
[0099] like Figure 7 As shown, the data exchange scheduling system 700 includes: a file preprocessing module 710, a job instantiation module 720 and a job scheduling module 730. The system 700 can be used to implement reference Figure 2 The described data exchange scheduling method.
[0100] The file preprocessing module 710 is used to perform definition preprocessing on the to-be-processed file and generate job definition information and job dependency information of the to-be-processed file. According to an embodiment of the present disclosure, the file preprocessing module 710 can be used to perform the above-mentioned Figure 2 The described step S201 will not be repeated here.
[0101] The job instantiation module 720 is used to instantiate the to-be-processed file according to the job definition information to obtain a job instance of the to-be-processed file, and to generate dependency instance information of the to-be-processed job according to the job dependency information and the job instance. According to an embodiment of the present disclosure, the job instantiation module 720 can be used to execute the above reference Figure 2 The described step S202 will not be repeated here.
[0102] The job scheduling module 730 is used to determine whether the pending job satisfies the preceding job dependency according to the dependency instance information. If so, the pending job is set as the pending job and is scheduled. According to an embodiment of the present disclosure, the job scheduling module 730 can be used to execute the above reference Figure 2 The described step S202 will not be repeated here.
[0103] According to an embodiment of the present disclosure, the file preprocessing module 710 is used to define preprocessing of files to be processed and generate job definition information and job dependency information of the files to be processed, including: obtaining basic information of multiple files to be processed; defining exchange requirements for the basic information of multiple files to be processed to generate exchange information; reading the exchange information within a preset time, and extracting keyword fields of the exchange information to generate a data definition file for the files to be processed; generating the job definition information according to the processing method and job type of each data record in the exchange information, and generating the job dependency information according to the execution order of each data record in the exchange information and the job type.
[0104] like Figure 8 As shown, the system 700 also includes: a file verification module 740, which is used to verify the file to be processed, including: adding the downloaded directory information in the data definition file to the detection directory table; checking whether a list verification file is generated for the directory information in the detection directory table; if so, checking whether the list verification file matches the data definition file; if so, adding the data definition file to the processing directory file, and entering the detected actual file information into the source data information file; if not, adding the data definition file to the abnormal processing directory file.
[0105] According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits, or at least part of the functions of any one of them can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as hardware circuits, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems on chips, systems on substrates, systems on packages, application specific integrated circuits (ASICs), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, submodules, units, and subunits can be at least partially implemented as computer program modules, and when the computer program modules are run, the corresponding functions can be performed.
[0106] For example, any multiple of the file preprocessing module 710, the job instantiation module 720, the job scheduling module 730, and the file inspection module 740 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the file preprocessing module 710, the job instantiation module 720, the job scheduling module 730, and the file inspection module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the file preprocessing module 710, the job instantiation module 720, the job scheduling module 730 and the file verification module 740 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0107] The data exchange scheduling method and system provided by the present disclosure can be used in the financial field or other fields. It should be noted that the data exchange scheduling method and system provided by the present disclosure can be used in the financial field, such as the scheduling processing after the files of various business systems in the financial field are converted into jobs, and can also be used in fields other than the financial field. The application field of the data exchange scheduling method and system provided by the present disclosure is not limited.
[0108] Fig. 9 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Fig. 9 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0109] like Fig. 9 As shown, the electronic device 900 described in this embodiment includes: a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 to a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0110] In RAM 903, various programs and data required for the operation of electronic device 900 are stored. Processor 901, ROM 902 and RAM 903 are connected to each other via bus 904. Processor 901 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 902 and / or RAM 903. It should be noted that the program can also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.
[0111] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage portion 908 as needed.
[0112] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0113] The embodiment of the present invention further provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the data exchange scheduling method according to the embodiment of the present disclosure is implemented.
[0114] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0115] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the data exchange scheduling method provided by the embodiment of the present disclosure.
[0116] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 901. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0117] In one embodiment, the computer program may be based on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0118] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0119] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0120] It should be noted that the functional modules in the various embodiments of the present disclosure may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention may essentially or in other words, the part that contributes to the prior art or all or part of the technical solution may be embodied in the form of a software product.
[0121] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0122] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0123] Although the present disclosure has been shown and described with reference to specific exemplary embodiments of the present disclosure, it should be understood by those skilled in the art that various changes in form and details may be made to the present disclosure without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure should not be limited to the above-mentioned embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims.
Claims
1. A data exchange scheduling method, It is characterized in that include: Perform definition preprocessing on the file to be processed, and generate job definition information and job dependency information of the file to be processed; The method of performing definition preprocessing on the files to be processed and generating the job definition information and job dependency information of the files to be processed includes: obtaining basic information of a plurality of files to be processed; defining exchange requirements for the basic information of the plurality of files to be processed and generating exchange information; reading the exchange information within a preset time and extracting keyword fields of the exchange information to generate a data definition file of the files to be processed; generating the job definition information according to the processing mode and job type of each data record in the exchange information, and generating the job dependency information according to the execution order of each data record in the exchange information and the job type; According to the job definition information, instantiate the file to be processed to obtain a job instance of the file to be processed, and generate dependent instance information of the job to be processed according to the job dependency information and the job instance; According to the dependency instance information, it is determined whether the job to be processed satisfies the predecessor job dependency. If so, the job to be processed is set as a job to be scheduled and scheduled.
2. The data exchange scheduling method according to claim 1, It is characterized in that Before instantiating the to-be-processed file according to the job definition information, the method further comprises: Add the download directory information in the data definition file to the detection directory table; Checking the directory information in the detection directory table to see whether a list verification file is generated; If so, checking whether the manifest verification file matches the data definition file; If there is a match, the data definition file is added to the processing directory file, and the actual file information detected is entered into the source data information file; if there is no match, the data definition file is added to the abnormal processing directory file.
3. The data exchange scheduling method according to claim 2, It is characterized in that The instantiating the file to be processed according to the job definition information to obtain a job instance of the file to be processed includes: The to-be-processed file recorded in the source data information file is matched with the job definition information and the data definition file, and macro replacement is performed to convert the file into an instance job, so as to obtain a job instance of the to-be-processed file.
4. The data exchange scheduling method according to claim 3, It is characterized in that The step of generating dependency instance information of a job to be processed according to the job dependency information and the job instance includes: The dependent instance information of the job to be processed is generated according to the job name in the job instance file and the post-job name in the job dependency information.
5. The data exchange scheduling method according to claim 1, It is characterized in that The determining whether the job to be processed satisfies the previous job dependency includes: Determine whether the job to be processed has a predecessor job dependency; If it does not exist, setting the pending job as the pending job; If so, determine whether the preceding job has been processed; if so, set the pending job as the pending job.
6. The data exchange scheduling method according to claim 5, It is characterized in that The method further includes: Determine whether the to-be-scheduled job is a dead job, wherein the dead job represents a to-be-scheduled job that has not been processed within a preset time; If so, reset the zombie job to the job to be scheduled.
7. The data exchange scheduling method according to claim 5, It is characterized in that The step of setting the pending job as a pending job and performing scheduling processing comprises: Add the to-be-scheduled job to the distribution job queue; Sequentially taking out a plurality of the to-be-scheduled jobs from the distribution job queue and assigning them to idle node addresses in the job client, and setting the status of the to-be-scheduled jobs to be processed; Submit the plurality of jobs to be scheduled to the job client for job processing, and update the job status of the jobs to be processed after the processing.
8. A data exchange scheduling system, It is characterized in that include: A file preprocessing module, used to perform definition preprocessing on the file to be processed, and generate job definition information and job dependency information of the file to be processed; The file preprocessing module is used to perform definition preprocessing on the files to be processed and generate job definition information and job dependency information of the files to be processed, including: obtaining basic information of multiple files to be processed; defining exchange requirements for the basic information of multiple files to be processed to generate exchange information; reading the exchange information within a preset time, and extracting keyword fields of the exchange information to generate a data definition file of the files to be processed; generating the job definition information according to the processing method and job type of each data record in the exchange information, and generating the job dependency information according to the execution order of each data record in the exchange information and the job type; A job instantiation module, used to instantiate the to-be-processed file according to the job definition information to obtain a job instance of the to-be-processed file, and to generate dependency instance information of the to-be-processed job according to the job dependency information and the job instance; The job scheduling module is used to determine whether the job to be processed satisfies the predecessor job dependency according to the dependency instance information, and if so, set the job to be processed as a job to be scheduled and perform scheduling processing.
9. The data exchange scheduling system according to claim 8, It is characterized in that The system also includes: The file verification module is used to verify the file to be processed, including: adding the downloaded directory information in the data definition file to the detection directory table; checking whether a list verification file is generated for the directory information in the detection directory table; if so, checking whether the list verification file matches the data definition file; if so, adding the data definition file to the processing directory file, and entering the detected actual file information into the source data information file; if not, adding the data definition file to the abnormal processing directory file.
10. An electronic device, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data exchange scheduling method as described in any one of claims 1 to 7 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the data exchange scheduling method according to any one of claims 1 to 7 is implemented.
12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the data exchange scheduling method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Job flow scheduling method and device, electronic equipment and storage medium
CN111930487A
Job scheduling method, apparatus and device, and medium
CN113419835A