Data Synchronization Method, Apparatus, Device, and Medium
By monitoring data changes and generating tasks to be executed, and using the configuration information of the data processing table and the output task table for data processing and output, the problem that enterprises need to preprocess after importing external data is solved, efficient data preprocessing and conversion are achieved, and response efficiency and user experience are improved.
Patent Information
- Application Number
- CN202211741863.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-12-30
AI Technical Summary
After an enterprise imports a large amount of data from an external system, it needs to preprocess it to adapt to different data formats of downstream applications, resulting in reduced response efficiency, especially in application scenarios for end-customer-oriented, affecting the user experience.
By monitoring data changes, generating and scheduling tasks to be executed, and using the configuration information in the data processing table and the output task table for data processing and output, efficient preprocessing and conversion of data is achieved.
The efficiency of data preprocessing conversion is improved, especially in the case of huge data volume and complex situations, the response efficiency and user experience are significantly improved through concurrent task execution.
Smart Images

Figure CN116028576B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data, and particularly to a data synchronization method, apparatus, device, medium, and program product. Background Art
[0002] In the daily operation of an enterprise, a large amount of data often needs to be imported from external systems. For example, in the risk management process of financial enterprises, data such as financial product information data, position data, and market data need to be imported from the outside. The amount of externally imported data may be large, and the source channels may also be numerous. As a result, the types and storage formats of the externally imported data may be diverse. At the same time, the data formats required by downstream applications in enterprise operations may also be diverse depending on different application scenarios and computing scenarios. Thus, the data imported from external systems usually cannot be directly used by downstream applications and need to be further processed (i.e., preprocessed) before use.
[0003] However, if the externally imported data is preprocessed every time when the downstream application needs to perform calculations, it is obvious that the response efficiency of the downstream application will be reduced. Especially when the downstream application provides services to end customers, the user experience will be seriously affected. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a data synchronization method, apparatus, device, medium, and program product that can efficiently preprocess externally imported data.
[0005] In the first aspect of the embodiments of the present disclosure, a data synchronization method is provided. The method includes: monitoring data changes in at least one storage location indicated by source data configuration information in a data processing table; when data changes in any first location among the at least one storage location are monitored, locating N records in the data processing table where the first source data configuration information indicating the first location is located, where N is an integer greater than or equal to 1; obtaining a data processing identifier and an output configuration identifier in each of the N records; wherein, in each record of the data processing table, a data processing logic corresponding to the data processing identifier is configured, and in each record of an output task table, output configuration information corresponding to the output configuration identifier is configured; generating a to-be-executed task based on the data processing identifier and the output configuration identifier in each record; scheduling the to-be-executed task, wherein when there is a target queue composed of prior tasks having the same output configuration identifier as the to-be-executed task, merging the to-be-executed task into the target queue to wait for scheduling; and when there is no such target queue, allocating the to-be-executed task to a new thread for scheduling and execution; when the to-be-executed task is scheduled, executing the to-be-executed task, wherein processing and converting the data in the first location according to a target data processing logic, where the target data processing logic is the data processing logic corresponding to the data processing identifier in the to-be-executed task; when the to-be-executed task is successfully executed, outputting the execution result of the to-be-executed task according to target output configuration information, where the target output configuration information is the output configuration information corresponding to the output configuration identifier in the to-be-executed task.
[0006] According to an embodiment of the present disclosure, executing the to-be-executed task includes: querying the output configuration information in the record where the output configuration identifier in the to-be-executed task is located from the output task table to obtain the target output configuration information; querying the data processing logic in the record where the data processing identifier in the to-be-executed task is located from the data processing table to obtain the target data processing logic; and reading the data to be processed from the first location based on the first source data configuration information and processing the data to be processed according to the target data processing logic to obtain the execution result.
[0007] According to an embodiment of the present disclosure, the output configuration identifiers and the output configuration information in different records of the output task table are all different; the source data configuration information in different records of the data processing table may be the same or different. Wherein, when the source data configuration information in different records of the data processing table is the same, the output configuration identifiers in these different records are different.
[0008] According to an embodiment of the present disclosure, the data processing identifiers in different records in the data processing table are different. Among them, when the source data configuration information in different records in the data processing table is the same, the data processing logics in these different records are also different.
[0009] According to an embodiment of the present disclosure, the data processing logic includes a processing script that can be hot-deployed. The processing of converting the data in the first position according to the target data processing logic includes: executing the processing script that implements the target data processing logic in a hot-deployed manner.
[0010] According to an embodiment of the present disclosure, the target queue is a segmented lock acquisition queue. Among them, when the to-be-executed task is scheduled, executing the to-be-executed task includes: after the to-be-executed task acquires the lock object allocated by the segmented lock acquisition queue, executing the to-be-executed task.
[0011] According to an embodiment of the present disclosure, the method further includes: querying the output configuration identifier of a prior task whose status is being executed or waiting to be executed from the task scheduling table; when there is an output configuration identifier that is the same as that of the to-be-executed task, determining that the target queue exists; otherwise, determining that the target queue does not exist. Among them, the task scheduling table records in real time the scheduling execution status of each task including the to-be-executed task.
[0012] According to an embodiment of the present disclosure, when the to-be-executed task is successfully executed, outputting the execution result of the to-be-executed task according to the target output configuration information includes: periodically scanning the task scheduling table; and when the status of the to-be-executed task is first scanned as successfully executed from the task scheduling table, outputting the execution result according to the target output configuration information.
[0013] According to an embodiment of the present disclosure, the method is applied to a distributed system. Generating a to-be-executed task based on the data processing identifier and the output configuration identifier in each record includes: a task initiation node in the distributed system initiating a synchronous task request based on the data processing identifier and the output configuration identifier in each record; a task scheduling node in the distributed system receiving the synchronous task request, generating the to-be-executed task based on the synchronous task request, and scheduling the to-be-executed task.
[0014] On the other hand, an embodiment of the present disclosure provides a data synchronization device. The device includes a data monitoring module, a task initiation module, a task scheduling module, and a task execution module. The data monitoring module is configured to monitor data changes in at least one storage location indicated by source data configuration information in a data processing table. The task initiation module is configured to: when data changes in any first location among the at least one storage location are monitored, locate N records in the data processing table where the first source data configuration information indicating the first location is located, where N is an integer greater than or equal to 1; and obtain a data processing identifier and an output configuration identifier in each of the N records; wherein, in each record of the data processing table, a data processing logic corresponding to the data processing identifier is configured, and in each record of the output task table, output configuration information corresponding to the output configuration identifier is configured. The task scheduling module is configured to: generate a to-be-executed task based on the data processing identifier and the output configuration identifier in each record; and schedule the to-be-executed task, wherein when there is a target queue composed of prior tasks having the same output configuration identifier as the to-be-executed task, merge the to-be-executed task into the target queue for waiting for scheduling, and when there is no such target queue, allocate the to-be-executed task to a new thread for scheduling and execution. The task execution module is configured to, when the to-be-executed task is scheduled, execute the to-be-executed task, including: processing and converting the data in the first location according to a target data processing logic, wherein the target data processing logic is the data processing logic corresponding to the data processing identifier in the to-be-executed task, and when the to-be-executed task is successfully executed, outputting the execution result of the to-be-executed task according to target output configuration information, wherein the target output configuration information is the output configuration information corresponding to the output configuration identifier in the to-be-executed task.
[0015] According to an embodiment of the present disclosure, the task execution module includes a data export module, a data conversion module, and a data import module. The data export module is configured to query the source data configuration information in the record where the data processing identifier in the to-be-executed task is located from the data processing table to obtain first source data configuration information, and read the data to be processed from the first location based on the first source data configuration information. The data conversion module is configured to query the data processing logic in the record where the data processing identifier in the to-be-executed task is located from the data processing table to obtain the target data processing logic, and process the data to be processed according to the target data processing logic. The data import module queries the output configuration information in the record where the output configuration identifier in the to-be-executed task is located from the output task table to obtain the target output configuration information, and outputs the execution result of the to-be-executed task according to the target output configuration information.
[0016] According to an embodiment of the present disclosure, the task scheduling module is further configured to: query an output configuration identifier of a prior task whose status is being executed or waiting to be executed from a task scheduling table; when there is an output configuration identifier that is the same as that of the to-be-executed task, determine that there is the target queue; otherwise, determine that there is no such target queue; wherein the task scheduling table records in real time the scheduling and execution status of each task including the to-be-executed task.
[0017] According to an embodiment of the present disclosure, the data import module is configured to periodically scan the task scheduling table; and when it is first scanned from the task scheduling table that the status of the to-be-executed task is execution success, output the execution result according to the target output configuration information.
[0018] A third aspect of the embodiments of the present disclosure provides an electronic device. The electronic device includes one or more processors and a memory. The memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method.
[0019] A fourth aspect of the embodiments of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above method.
[0020] A fifth aspect of the embodiments of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0021] The above one or more embodiments have the following advantages or beneficial effects: providing a complete set of data synchronization methods from data monitoring, to task initiation, task scheduling execution, and data output, realizing preprocessing in the data synchronization process, and improving the efficiency of data preprocessing conversion. Especially when the amount of preprocessed data is huge and the preprocessing situations are complex, concurrent task execution can effectively improve the data preprocessing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0023] Figure 1 Schematically shows an application scenario diagram of the data synchronization method and apparatus according to an embodiment of the present disclosure;
[0024] Figure 2 Schematically shows a block diagram of the data synchronization apparatus according to an embodiment of the present disclosure;
[0025] Figure 3 Schematically shows a flowchart of the data synchronization method according to an embodiment of the present disclosure;
[0026] Figure 4 Schematically shows the processing flow of a task to be processed waiting for scheduling in a data synchronization method according to an embodiment of the present disclosure;
[0027] Figure 5 Schematically shows the processing flow of outputting an execution result in a data synchronization method according to an embodiment of the present disclosure;
[0028] Figure 6 Schematically shows a partial processing flow chart in a data synchronization method according to another embodiment of the present disclosure;
[0029] Figure 7 Schematically shows the processing flow of data output in a data synchronization method according to another embodiment of the present disclosure; and
[0030] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the data synchronization method according to an embodiment of the present disclosure. Detailed implementation manners
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0032] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). The terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features.
[0035] Embodiments of the present disclosure provide a data synchronization method, apparatus, device, medium, and program product. The method of the embodiments of the present disclosure includes monitoring changes in source data to be synchronized, initiating a task when changes are monitored, and in a task scheduling manner, merging tasks with the same output configuration into the same queue for execution, while tasks with different output configurations can be executed concurrently. Then, after executing the tasks to process and convert the source data into intermediate data, it is transferred and stored according to the output configuration. In this way, a complete set of data synchronization methods from data monitoring, to task initiation, task scheduling execution, and data output is provided, realizing preprocessing during the data synchronization process and improving the efficiency of data preprocessing and conversion. Especially when the amount of preprocessed data is huge, concurrent task execution can effectively improve the efficiency of data preprocessing.
[0036] It should be noted that the data synchronization method and apparatus determined in the embodiments of the present disclosure can be used in the financial field and can also be used in any field other than the financial field. The present disclosure does not limit the application field.
[0037] Figure 1 An application scenario diagram of the data synchronization method and apparatus according to the embodiments of the present disclosure is schematically shown.
[0038] As Figure 1 shown, the application scenario 100 according to this embodiment may include a source storage system 101, a data synchronization apparatus 200, and a target storage system 102.
[0039] The source storage system 101 may receive and store data imported from at least one external system. The target storage system 102 may store and provide intermediate data used by at least one downstream application. The data synchronization apparatus 200 is disposed between the source storage system 101 and the target storage system 102. The composition of the data synchronization apparatus 200 may refer to the Figure 2 illustration.
[0040] The source storage system 101 and the target storage system 102 can be any one or more of a mySQL database, an Orical database, a shared file system, or a big data platform, and can include any one or more storage media (e.g., file type or database type).
[0041] The data synchronization device 200 can execute the data synchronization method of the embodiments of the present disclosure, read data from the source storage system 101, preprocess it, convert it into intermediate data that can be used by downstream applications, and then import it into the target storage system 102.
[0042] According to the embodiments of the present disclosure, during the data synchronization process, the data synchronization device 200 needs to use at least two tables: a data processing table and an output task table. Among them, the data processing table is used to specify the reading and preprocessing conversion methods for the data in the source storage system 101. The output task table is used to specify the output method of the preprocessed intermediate data. These two tables can be preset and can be dynamically adjusted during the data synchronization process.
[0043] Table 1 shows an output task table, and Table 2 shows a data processing table. Here, in combination with the examples in Table 1 and Table 2, the functions of the output task table and the data processing table and the relationship between the two are described as follows.
[0044] As shown in Table 1, the output task table may include an output configuration identifier and output configuration information. Among them, the output configuration information may include information such as output format, output location, type of target storage medium, description, etc. The output configuration information is used to indicate information such as the output location of the intermediate data, the output format (such as integer type, floating point type, or counting unit, etc.), and the type of target storage medium. In the application scenario 100, the output location is the storage address in the target storage system 102.
[0045] In the output task table, one output configuration identifier corresponds to one output configuration information, indicating the output method of one item of intermediate data.
[0046] The setting of the output task table can configure the output configuration identifier and output configuration information one by one according to the data or data tables required by the downstream application for calculation.
[0047] Table 1, Output Task Table
[0048]
[0049] Table 2 shows a data processing table. As shown in Table 2, each record in the data processing table includes a data processing identifier, source data configuration, data processing logic, and output configuration identifier, etc. Among them, the source data configuration is used to record the storage location of the data to be monitored in the source storage system 101, the source data structure, etc. The data processing logic is used to record the preprocessing conversion algorithm for the read data, etc. The value of the output configuration identifier comes from the output task table and is used to indicate that the processed data will be output according to the output configuration information corresponding to this output configuration identifier in the output task table. It can be seen that the output task table and the data processing table are corresponding and linked through the output task identifier.
[0050] Table 2, Data Processing Table
[0051]
[0052] In the application scenario 100, the data synchronization device 200 can monitor the changes of the corresponding data in the source storage system 101 according to the storage location indicated by the value in the source data configuration in Table 2. And when a change is monitored, the changed data is transformed according to the corresponding data processing logic in Table 2, and then the corresponding output configuration information is queried from Table 1 according to the output configuration identifier in Table 2, and then the transformed data is output to the target storage system 102 correspondingly.
[0053] By separately setting the output task table and the data processing table in the embodiments of the present disclosure, it is possible to decouple the processing transformation method and the output configuration of the data read from the same source data storage location, and at the same time, they are corresponding and linked through the output configuration identifier.
[0054] According to the embodiments of the present disclosure, the setting principles or logics of the output task table and the data processing table are different. Specifically, when setting the output task table, the output configuration identifiers and output configuration information (such as the output location) of different records are all different, that is, each output configuration identifier in the output task table indicates a kind of output configuration information, and different output configuration identifiers indicate different output configuration information. When setting the data processing table, the data processing identifiers in different records are different, but as long as at least one of the source data configuration, data processing logic, and output configuration identifier in different records is different.
[0055] Due to the association relationship and decoupling setting of the output task table and the data processing table, plus the difference in their setting logics, a complex corresponding relationship between data reading and output can be realized.
[0056] For example, referring to Table 2, the source data configuration information in different records in the data processing table can be the same or different. When the source data configuration information in different records in the data processing table is the same, the output configuration identifiers in these different records should be different. This means that data read from the same location in the source storage system 101 can have different output methods. For example, the same data can be preprocessed and then output to different types of media, or output to different files or data tables in the same medium, that is, to implement the one-to-many relationship between the input and output of the data synchronization device 200.
[0057] In some embodiments, when the source data configuration information in different records in Table 2 is the same, not only are the output configuration identifiers in these different records different, but the data processing logics can also be different. For example, for a product sales volume data obtained externally, one preprocessing requirement is to perform statistics on the sales volume data of the product, such as daily statistics or monthly statistics; another preprocessing requirement is to multiply the sales volume of the product by the unit price of the product to calculate the sales amount of the product. For the two different preprocessing requirements, two records are set in the data processing table. The values in the source data configuration of these two records are the same, but they correspond to different data processing identifiers, data processing logics, and output configuration identifiers. In this way, the same input of the data synchronization device 200, after different processes, corresponds to different outputs.
[0058] Also, for example, the output configuration identifiers in different records in the data processing table shown in Table 2 can also be the same or different. Among them, when the output configuration identifiers in different records in the data processing table are the same, the source data configuration information in different records in this data processing table should be different. This means that data read from different locations in the source storage system 101 can be preprocessed and then output in the same output manner. For example, the precious metal position data in the source storage system 101 can be from different external systems and stored in different locations in the source storage system 101. The data synchronization device 200 can read the precious metal position data stored in different locations in the source storage system 101, and after preprocessing, output it to the target location in the target storage system 102 according to the unified output configuration information. Thus, the many-to-one relationship from the input to the output of the data synchronization device 200 can be achieved.
[0059] It can be seen that by means of the independent setting and corresponding connection of the output task table and the data processing table, the data synchronization device 200 can achieve fine-grained and precise control over data preprocessing, input, and output.
[0060] Figure 2 Schematically shows a block diagram of the data synchronization device 200 according to an embodiment of the present disclosure.
[0061] Combined with Figure 1 andFigure 2 The data synchronization device 200 may include a data monitoring module 210, a task initiation module 220, a task scheduling module 230, and a task execution module 240. The task execution module 240 may be further divided into a data export module 241, a data conversion module 242, and a data import module 243.
[0062] Referring to the output task table shown in Table 1 and the data processing table shown in Table 2 above, the data synchronization device 200 is specifically described as follows.
[0063] The data monitoring module 210 is used to monitor data changes in at least one storage location indicated by the source data configuration information in the data processing table. In the application scenario 100, the at least one storage location is the storage location of the data to be monitored in the source storage system 101. The data monitoring module 210 is mainly responsible for monitoring the data to be synchronized and periodically scanning the data change information.
[0064] The task initiation module 220 is used to: when data changes in any first location among at least one storage location are monitored, locate N records in the data processing table where the first source data configuration information indicating the first location is located, and then obtain the data processing identifier and the output configuration identifier in each of the N records. Here, the "first location" is used to indicate the storage location of the data whose change is monitored in the source storage system 101.
[0065] In one embodiment, the task initiation module 220 may initiate a synchronization task according to the task configuration after receiving the data change message sent by the data monitoring module 210.
[0066] The task scheduling module 230 is used to: generate a task to be executed based on the data processing identifier and the output configuration identifier in each record, and then schedule the task to be executed.
[0067] According to an embodiment of the present disclosure, when performing task scheduling, different tasks can be distinguished as to whether they can be executed concurrently according to the output configuration identifier. Specifically, when there is a target queue composed of prior tasks having the same output configuration identifier as the task to be executed, the task to be executed is merged into the target queue and waits for scheduling. When there is no target queue, the task to be executed is assigned to a new thread for scheduling and execution.
[0068] In one embodiment, the task scheduling module 230 is responsible for scheduling the data synchronization task, including functions such as configuration query, lock control, generating a task execution object, concurrent processing, and task status monitoring.
[0069] The task execution module 240 is used to execute the to-be-executed task when the to-be-executed task is scheduled. The task execution module 240 can be divided into three types according to its functions: The data export module 241: exports the to-be-processed data from the source storage system 101 according to the configuration identifier transferred by the caller. The data processing module 242: calculates and / or converts the format of the to-be-processed data according to the output configuration identifier passed in by the caller. The data import module 243: determines the output configuration information according to the output configuration identifier passed in by the caller, and imports the processed data into the target storage system 102.
[0070] Specifically, the data export module 241 queries the source data configuration information from Table 2 according to the data processing identifier in the to-be-executed task, and then can read the to-be-processed data from the corresponding location in the source storage system 101 according to the queried source data configuration information.
[0071] The data conversion module 242 can query the data processing logic in the record where the data processing identifier in the to-be-executed task is located from Table 2 according to the data processing identifier in the to-be-executed task, and then process the to-be-processed data according to the data processing logic to obtain the execution result.
[0072] The data import module 243 can query the output configuration information in the record where the output configuration identifier in the to-be-executed task is located from Table 1 according to the output configuration identifier in the to-be-executed task, and then output the execution result of the to-be-executed task according to the output configuration information. For example, in the application scenario 100, the data import module 243 can import the execution result (i.e., the intermediate data after preprocessing and conversion) into the corresponding location in the target storage system 102 according to the output address in the obtained output configuration information.
[0073] It can be seen that after the present disclosure embodiment monitors the change of the corresponding data in the source storage system 101, it can concurrently execute the preprocessing tasks with different output methods through task scheduling. In the case where the amount of data imported by the external system is relatively large, there is a large amount of data to be preprocessed, and the situation is complex and cumbersome, the efficiency of data preprocessing and conversion can be greatly improved. Thus, it can efficiently preprocess and convert the externally imported data, and improve the usage efficiency and response efficiency of downstream applications.
[0074] It should be noted that the structure of the above data synchronization device 200 is merely exemplary. According to the embodiments of the present disclosure, any plurality of modules among the data monitoring module 210, the task initiation module 220, the task scheduling module 230, the data export module 241, the data conversion module 242, and the data import module 243 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to the embodiments of the present disclosure, at least one of the data monitoring module 210, the task initiation module 220, the task scheduling module 230, the data export module 241, the data conversion module 242, and the data import module 243 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware by integrating or packaging circuits, or can be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Or, at least one of the data monitoring module 210, the task initiation module 220, the task scheduling module 230, the data export module 241, the data conversion module 242, and the data import module 243 can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute the corresponding functions.
[0075] Next, based on Figure 1 the described scenarios and Figure 2 the data synchronization device shown, and in combination with the output task table shown in Table 1 above and the data processing table shown in Table 2, the data synchronization method of the embodiments of the present disclosure will be described in detail. It should be noted that the sequence numbers of the various operations in the following methods are only used as representations of the operations for description, and should not be regarded as indicating the execution sequence of the various operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0076] Figure 3 FIG. schematically shows a flowchart of a data synchronization method according to an embodiment of the present disclosure.
[0077] As Figure 3 shown, the data synchronization method according to this embodiment may include operation S310 to operation S390.
[0078] In operation S310, monitor data changes in at least one storage location indicated by the source data configuration information in the data processing table.
[0079] In one embodiment, the data monitoring module 210 may poll and scan the data processing table shown in Table 2 above, and monitor the corresponding data changes in the source storage system 101 according to the storage location indicated by the value in the source data configuration (i.e., source data configuration information) therein.
[0080] In another embodiment, the data range to be monitored may also be determined by filtering and outputting the task identifier from Table 1. For example, a flag indicating whether data output is allowed may be set for the output configuration identifier in Table 1. The data monitoring module 210 may first read the output configuration identifiers set to allow data output from Table 1, then find all records with the output configuration identifier in Table 2, and then monitor the data according to the information in the source data configuration in these records to check whether there is data change information. Or, for example, the data monitoring range may be determined by the calling party (e.g., downstream application) providing the output task identifier to be monitored. In this way, the synchronization of which data can be actively controlled.
[0081] In operation S320, when data changes in any first location among at least one storage location are monitored, N records where the first source data configuration information indicating the first location is located may be located from the data processing table, where N is an integer greater than or equal to 1.
[0082] In operation S330, the data processing identifier and the output configuration identifier in each of the N records are obtained. For example, the task initiation module 220 may query the data processing table and obtain the corresponding data processing identifier and output configuration identifier.
[0083] In operation S340, a to-be-executed task is generated based on the data processing identifier and the output configuration identifier in each record.
[0084] The generation of the to-be-executed task may be generated by either the task initiation module 220 or the task scheduling module 230.
[0085] In one embodiment, when the method of the embodiments of the present disclosure is applied to a distributed system, the task initiation node where the task initiation module 220 is located may generate a task initiation request based on the data processing identifier and the output configuration identifier in each record. Then, the task initiation request may be sent to the task scheduling module 230 on the task scheduling node by RPC calling the task scheduling module 230, so that the task scheduling module 230 generates a to-be-executed task for task scheduling and execution. In a distributed system, the computing functions of each node are highly specialized. Applying the method of the embodiments of the present disclosure to a distributed system can greatly improve the efficiency of data synchronization under the demand of a large amount of data.
[0086] As mentioned above, if the data monitoring module 210 first reads the data output configuration identifier from the output task table and then reads the source data configuration, data processing identifier, and output configuration identifier from the data processing table, it is possible that the output configuration identifier obtained in operation S330 is not allowed to be output or not called in the output task table. In this case, in operation S340, a task to be processed is not generated based on this output configuration identifier.
[0087] Next, the scheduling and execution of the tasks to be executed are carried out.
[0088] In operation S350, it is determined whether there is a target queue composed of prior tasks having the same output configuration identifier as the task to be executed. If so, operation S360 is executed; if not, operation S370 is executed.
[0089] In operation S360, when there is a target queue composed of prior tasks having the same output configuration identifier as the task to be executed, the task to be executed is merged into the target queue and waits for scheduling.
[0090] In operation S370, when there is no target queue, the task to be executed is assigned to a new thread for scheduling and execution.
[0091] According to an embodiment of the present disclosure, the queue can be a segmented lock receiving queue. The role of the lock is that a task can be executed only when the lock is obtained, and if the lock is not obtained, it waits. In this way, when there is a target queue, the tasks in the target queue can be sorted in the order of generation time and arranged in the segmented receiving queue, and the status is set to task waiting. When there is no target queue, a new lock object can be allocated to allow tasks to be executed concurrently.
[0092] In one embodiment, the task scheduling process of operations S350 to S370 can be executed by the task scheduling module 230.
[0093] Next, in operation S380, when the task to be executed is scheduled, the task to be executed is executed. When each task to be executed is executed, data is read from this first position in the source storage system 101, and then the data is processed and converted according to the data processing logic corresponding to the data processing identifier in the task to be executed.
[0094] Specifically, the task execution in operation S380 can be divided into two parts: exporting data from the source storage system 101 and processing and converting the exported data. In one embodiment, the task execution in operation S380 can be executed by the data export module 241 and the data conversion module 242.
[0095] For example, the task scheduling module 230 may schedule the tasks in the queue to the data export module 241 and the data conversion module 242. Then, the data export module 241 may query the source data configuration information in the record where the data processing identifier in the task to be executed is located from the data processing table, obtain the first source data configuration information, and read the data to be processed from the first location based on the first source data configuration information. Next, the data conversion module 242 queries the data processing logic in the record where the data processing identifier in the task to be executed is located from the data processing table, obtains the target data processing logic, and processes the data to be processed according to the target data processing logic.
[0096] According to an embodiment of the present disclosure, in the data processing table shown in Table 2, the data processing logic may include a processing script that can be hot-deployed. Thus, the data conversion module 242 may query the processing script in the record where the data processing identifier in the task to be executed is located from the data processing table, and then execute the processing script in a hot-deployed manner to implement the preprocessing conversion of the data. The processing script may be, for example, a Groovy script. Using a Groovy script can improve the flexibility of logic execution and support the random change and hot deployment of logic. Of course, the processing script may also be a script written in other interactive script logic languages.
[0097] Finally, in operation S390, when the task to be executed is successfully executed, the execution result of the task to be executed is output according to the target output configuration information, where the target output configuration information is the output configuration information corresponding to the output configuration identifier in the task to be executed.
[0098] In one embodiment, the data output in operation S390 may be performed by the data import module 243.
[0099] Specifically, the data import module 243 may query the output configuration information in the record where the output configuration identifier in the task to be executed is located from the output task table, obtain the target output configuration information, and output the execution result of the task to be executed according to the target output configuration information. For example, the data obtained by preprocessing and converting the task to be executed is imported into the target storage system 102 according to the output format and output address in the target output configuration information.
[0100] The data synchronization method according to the embodiments of the present disclosure performs data preprocessing during the data synchronization process. By means of concurrent execution of task scheduling, the data synchronization performance can be effectively improved. Moreover, during the task scheduling process, through the distinction of output configuration identifiers, conflicts do not occur between different output tasks during the concurrent execution of synchronization. According to the embodiments of the present disclosure, segment locks can be respectively set for the import processes of data flows with different output configuration identifiers to ensure no conflicts during the data synchronization process, realizing the efficiency and orderliness of data preprocessing and output.
[0101] According to an embodiment of the present disclosure, by configuring data processing logic in a data processing table, the data preprocessing transformation can achieve a fine-grained and precise effect, and the data processing logic is editable and configurable, and real-time hot deployment is also supported.
[0102] According to an embodiment of the present disclosure, a task scheduling table may be configured in the task scheduling module 230 to provide scheduling monitoring information for the task scheduling module 230. Table 3 illustrates a task scheduling table.
[0103] As shown in Table 3, the task scheduling table can record in real time the scheduling execution status of each task including the task to be executed. Every time the task scheduling module 230 generates a task, a record can be formed in the task scheduling table.
[0104] The input information in the task scheduling table is the information recorded according to the information passed in by the task initiation module 220, and may include information such as the output configuration identifier, data processing identifier, generation time, etc. of the task. In this way, in operation S350, scheduling can be performed according to the information in the task scheduling table.
[0105] Table 3 Task Scheduling Table
[0106]
[0107] Figure 4 Schematically shows the processing flow of the task to be processed waiting for scheduling in operation S350 in the data synchronization method according to an embodiment of the present disclosure.
[0108] As Figure 4 shown, according to this embodiment, operation S350 may include operations S351 to S354.
[0109] In operation S351, query the output configuration identifier of the prior task whose status is being executed or waiting to be executed from the task scheduling table.
[0110] In operation S352, determine whether there is an output configuration identifier that is the same as the task to be executed.
[0111] If so, in operation S353, determine that there is a target queue. Thus, the task to be executed can be grouped in the target queue.
[0112] If not, in operation S354, determine that there is no target queue. Then there is no task that outputs data to the same location as the task to be executed, so the task to be executed can be directly executed.
[0113] Since the task schedule table can record the scheduling execution status of each task, including the tasks to be executed, in real time, in some embodiments, the timing of importing data into the target storage system 102 can also be determined with the aid of the record information in the task schedule table in operation S390.
[0114] Figure 5 Schematically shows the processing flow of outputting the execution result in operation S390 of the data synchronization method according to an embodiment of the present disclosure.
[0115] As Figure 5 shown, according to this embodiment, operation S390 may include operation S391 and operation S392.
[0116] In operation S391, the task schedule table is scanned regularly.
[0117] In operation S392, when the status of the task to be executed is first scanned from the task schedule table as successfully executed, the execution result is output according to the target output configuration information.
[0118] Figure 6 Schematically shows a partial processing flow chart of the data synchronization method according to another embodiment of the present disclosure.
[0119] As Figure 6 shown, this partial processing flow may include step S61 to step S612. In combination with Figure 1 and Figure 2 , this partial processing flow is specifically the mutual cooperation processing flow among the data monitoring module 210, the task initiation module 220, the task scheduling module 230, the data export module 241, and the data conversion module 242 in the data synchronization device 200. The specific description is as follows.
[0120] In step S61, the data monitoring module 210 can regularly scan the change situation of the specified data configured in the data processing table. In this embodiment, the data monitoring module 210 can first read the output configuration identifier from the output task table, then find all records with the output configuration identifier in the data processing table, and then perform data monitoring according to the information in the source data configuration in these records.
[0121] Then in step S62, check whether there is data change information. If there is a change, read out the data flag of the changed data. If the changed data involves a data table, read out the flag of the changed data in this data table. Send it to the task initiation module 220 in combination with the output configuration identifier defined in the output task table.
[0122] Next in step S63, the task initiation module 220 combines information such as environment information and risk date, generates a synchronization task initiation request, and sends a new synchronization task request to the task scheduling module 230 based on the RPC protocol.
[0123] Then, the task scheduling module 230 executes steps S64 to S67.
[0124] In step S64, the task scheduling module 230 receives a task synchronization request, and according to the output configuration identifier therein, queries all output configuration information (such as parameters and contexts such as output location, output format, etc.) involved in the task from the output task table.
[0125] Meanwhile, in step S65, the task scheduling module 230 queries the task scheduling table, and according to the input information and task status in the table, checks whether there is a task with the same output location requirement that is being executed or waiting to be executed in the table.
[0126] If there is, then in step S66, the new task is classified into the segmented lock receiving queue where these tasks are located; if not, then in step S67, a new thread object is added, a new lock object is allocated, and the task is executed concurrently. The purpose of the synchronization measures and task allocation measures here is to improve the efficiency of data synchronization and the relative reliability of the results.
[0127] Next, the data export module 241 and the data conversion module 242 jointly execute steps S68 to S612.
[0128] After obtaining the lock of a task, the data export module 241 and the data conversion module 242 start in step S68 according to the settings in the output task table and the data processing table and the parameters passed in by the task initiation module 220 (including the output task identifier and the data processing identifier), and then in step S69, query the data processing table according to the data processing identifier, and then query the corresponding source data configuration information from the data processing table, and obtain the raw data to be synchronized according to the source data configuration information.
[0129] Then, in step S610, data conversion and format processing are performed through the raw data structure in the source data configuration queried during data processing and the groovy script recorded in the data processing field for executing data processing logic. The advantage of using the grooy script here is to improve the flexibility of logic execution and support arbitrary changes and hot deployment of logic.
[0130] Meanwhile, in step S611, the task status in the task scheduling table is updated after each step is completed.
[0131] And in step S612, the intermediate data after preprocessing and conversion is written into a file and uploaded to the shared file system.
[0132] Figure 7 Schematically shows the processing flow of data output in the data synchronization method according to another embodiment of the present disclosure.
[0133] As shown Figure 7 in combination with Figure 1 and Figure 2 , the processing flow for outputting data to the target storage system 102 may include steps S71 to S76. This processing flow may be executed by the data import module, and the data import module 243 is relatively independent compared to other modules in the data synchronization device 200.
[0134] In step S71, the data import module 243 periodically scans the execution status in the task scheduling table and makes a judgment in step S72.
[0135] If the task with the status of completed data conversion is scanned in step S72, then in step S73, according to the output task identifier in the input information field of the task scheduling table, the output configuration information corresponding to this output task identifier in the output task table is found, and then the data uploaded by the data conversion module 242 to the remote file system is imported into the corresponding target data layer in the target storage system 102. Then in step S74, the task status corresponding in the task scheduling table is set to successful.
[0136] If the task with the status of completed data conversion is not scanned in step S72, then in step S74, the data import module 243 checks whether the task execution time exceeds the timeout configured for each task in the output task table; if not, it returns to step S71 to continue scanning; if so, then in step S76, the task status in the task scheduling table is set to timeout.
[0137] In the embodiments of the present disclosure, the data import module 243 is also responsible for setting the task status of tasks that are in an unfinished state (i.e., except for the execution success and execution failure states) and whose time exceeds the timeout defined in the output task table to the timeout state. Correspondingly, each status change of the task is also recorded in real time in other information fields in the task scheduling table.
[0138] The embodiments of the present disclosure can help an enterprise preprocess and convert the data received from the outside and store it for backup to improve the response efficiency of downstream applications in enterprise operations. In one embodiment, when this method is applied to the risk management of a financial institution, in the inventor's work practice, it can effectively help the risk management system achieve the backup synchronization of position data for 56 products such as interest rate swaps, bond spot, bond forward, precious metal futures, foreign exchange options, foreign exchange spot, foreign exchange forward, and foreign exchange swap, and perform preprocessing operations for subsequent risk measurement work, saving a large amount of workload.
[0139] Figure 8 A block diagram of an electronic device suitable for implementing the data synchronization method according to the embodiments of the present disclosure is schematically shown.
[0140] As shown Figure 8As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0141] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 802 and / or the RAM 803. It should be noted that the program can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in the one or more memories.
[0142] According to an embodiment of the present disclosure, the electronic device 800 can further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 can further include one or more of the following components connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0143] The present disclosure also provides a computer-readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present disclosure is implemented.
[0144] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.
[0145] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiment of the present disclosure.
[0146] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0147] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0148] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0149] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0151] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0152] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A data synchronization method, comprising: monitoring data changes in at least one storage location indicated by source data configuration information in a data processing table; when data changes in any first location among the at least one storage location are monitored, locating N records in the data processing table where the first source data configuration information indicating the first location is located, where N is an integer greater than or equal to 1; obtaining a data processing identifier and an output configuration identifier in each of the N records; wherein, in each record of the data processing table, a data processing logic corresponding to the data processing identifier is configured, and in each record of the output task table, output configuration information corresponding to the output configuration identifier is configured; generating a to-be-executed task based on the data processing identifier and the output configuration identifier in each record; scheduling the to-be-executed task, wherein when there is a target queue composed of prior tasks having the same output configuration identifier as the to-be-executed task, merging the to-be-executed task into the target queue to wait for scheduling; and when there is no such target queue, allocating the to-be-executed task to a new thread for scheduling and execution; when the to-be-executed task is scheduled, executing the to-be-executed task; wherein, processing and converting the data in the first location according to a target data processing logic, where the target data processing logic is the data processing logic corresponding to the data processing identifier in the to-be-executed task; when the to-be-executed task is successfully executed, outputting the execution result of the to-be-executed task according to target output configuration information, where the target output configuration information is the output configuration information corresponding to the output configuration identifier in the to-be-executed task.
2. The method according to claim 1, wherein, the executing the to-be-executed task includes: querying the output configuration information in the record where the output configuration identifier in the to-be-executed task is located from the output task table to obtain the target output configuration information; querying the data processing logic in the record where the data processing identifier in the to-be-executed task is located from the data processing table to obtain the target data processing logic; reading the data to be processed from the first location based on the first source data configuration information and processing the data to be processed according to the target data processing logic to obtain the execution result.
3. The method according to claim 1 or 2, wherein, the output configuration identifiers and the output configuration information in different records of the output task table are all different; the source data configuration information in different records in the data processing table may be the same or different; wherein, when the source data configuration information in different records in the data processing table is the same, the output configuration identifiers in these different records are different.
4. The method according to claim 3, wherein, the data processing identifiers in different records in the data processing table are different; wherein, when the source data configuration information in different records in the data processing table is the same, the data processing logics in these different records are also different.
5. The method according to claim 1 or 2, wherein the data processing logic includes a processing script capable of hot deployment, and the data in the first location is processed and transformed according to the target data processing logic including: Executing, in a hot deployment manner, the processing script implementing the target data processing logic.
6. The method according to claim 1, wherein, the target queue is a segmented lock acquisition queue, and when the to-be-executed task is scheduled, executing the to-be-executed task includes: After the to-be-executed task acquires the lock object allocated by the segmented lock acquisition queue, executing the to-be-executed task.
7. The method according to claim 1, wherein, the method further includes: Querying the output configuration identifier of the prior task with a status of being executed or waiting to be executed from the task scheduling table; When there is an output configuration identifier identical to that of the to-be-executed task, determining that the target queue exists; otherwise determining that the target queue does not exist; wherein, the task scheduling table records in real time the scheduling execution status of each task including the to-be-executed task.
8. The method according to claim 7, wherein, when the to-be-executed task is successfully executed, outputting the execution result of the to-be-executed task according to the target output configuration information includes: Periodically scanning the task scheduling table; and When the status of the to-be-executed task is first scanned as successfully executed from the task scheduling table, outputting the execution result according to the target output configuration information.
9. The method according to claim 1, wherein, the method is applied to a distributed system, and generating a to-be-executed task based on the data processing identifier and output configuration identifier in each record includes: The task initiation node in the distributed system initiates a synchronous task request based on the data processing identifier and output configuration identifier in each record; The task scheduling node in the distributed system receives the synchronous task request, generates the to-be-executed task based on the synchronous task request, and schedules the to-be-executed task.
10. A data synchronization device, including a data monitoring module, a task initiation module, a task scheduling module, and a task execution module, wherein: The data monitoring module is used to monitor data changes in at least one storage location indicated by the source data configuration information in the data processing table; The task initiation module is used for: When monitoring data changes in any first location among the at least one storage location, locating N records in the data processing table where the first source data configuration information indicating the first location is located, where N is an integer greater than or equal to 1; and Obtaining the data processing identifier and output configuration identifier in each of the N records; wherein, in each record of the data processing table, data processing logic corresponding to the data processing identifier is configured, and in each record of the output task table, output configuration information corresponding to the output configuration identifier is configured; The task scheduling module is used for: Generating a to-be-executed task based on the data processing identifier and output configuration identifier in each record; and Schedule the to-be-executed task. When there is a target queue composed of prior tasks with the same output configuration identifier as the to-be-executed task, merge the to-be-executed task into the target queue to wait for scheduling; and when there is no such target queue, allocate the to-be-executed task to a new thread for scheduling and execution. The task execution module is used to execute the to-be-executed task when the to-be-executed task is scheduled, including: Process and convert the data in the first position according to the target data processing logic, where the target data processing logic is the data processing logic corresponding to the data processing identifier in the to-be-executed task; and When the to-be-executed task is successfully executed, output the execution result of the to-be-executed task according to the target output configuration information, where the target output configuration information is the output configuration information corresponding to the output configuration identifier in the to-be-executed task.
11. An electronic device, comprising: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and medium
CN114185932A
Task scheduling method and device, task processing method and device, electronic equipment and medium
CN115373822A