Task processing method and device, electronic equipment, medium and program product
By detecting changes in target data, automatically calculating the task backtracking period and optimizing the backtracking link, the problem of low task backtracking efficiency is solved and efficient and accurate data backtracking is achieved.
Patent Information
- Application Number
- CN202510935270.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-14
AI Technical Summary
The task backtracking process suffers from low backtracking efficiency and poor data accuracy due to manual operations, which is especially difficult to automate in multiple batches and complex backtracking links.
By detecting changes in target data, the matching backtracking link is determined, and the backtracking period of each task is automatically calculated based on the change cycle and task time parameters. Data is backtracked according to the backtracking link, and the backtracking link is optimized to avoid circular dependencies and data loss.
It realizes automated task backtracking without manual backtracking, improves backtracking efficiency and data accuracy, and is applicable to various task types and scenarios.
Smart Images

Figure CN120780436A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to the technical field of data processing, and especially to a task processing method and device, electronic equipment, medium and program product. BACKGROUND
[0002] With the continuous development of computer technology, the data link between the upper application and the lower data architecture is also more and more complex. The upstream application modification and business change need to repair the lower interface, data structure or data table, and thus the tasks involved in the entire data link need to be traced back.
[0003] However, the task backtracking process involves manual backtracking, and the backtracking efficiency is low. SUMMARY
[0004] The present disclosure provides a task processing method and device, electronic equipment, medium and program product.
[0005] According to an aspect of the present disclosure, a task processing method is provided, comprising: in response to detecting that target data changes, determining a first backtracking link matched with the target data, the first backtracking link comprising a plurality of first tasks having a dependency relationship; determining a backtracking period of each of the plurality of first tasks based on a change period of the target data and a time parameter of the first tasks; and backtracking data of the first tasks within the backtracking period according to the first backtracking link to obtain a backtracking result.
[0006] According to another aspect of the present disclosure, a task processing device is provided, comprising: a first determining module configured to determine a first backtracking link matched with target data in response to detecting that the target data changes, the first backtracking link comprising a plurality of first tasks having a dependency relationship; a second determining module configured to determine a backtracking period of each of the plurality of first tasks based on a change period of the target data and a time parameter of the first tasks; and a backtracking module configured to backtrack data of the first tasks within the backtracking period according to the first backtracking link to obtain a backtracking result.
[0007] According to another aspect of the present disclosure, an electronic equipment is provided, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as above.
[0008] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method as above.
[0009] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method as described above.
[0010] It should be understood that the matters described herein are intended to be illustrative rather than limiting. For example, while the application is illustrated and described in relation to a task processing method and device, the application is not intended to be limited to only these implementations. Rather, the scope of the application is to be interpreted only limited by the claims. Various changes and modifications to the implementations described herein will be apparent to those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification. In the drawings:
[0012] Figure 1 An application scenario diagram is schematically shown according to an embodiment of the present disclosure, which can be applied to a task processing method and device;
[0013] Figure 2 A flowchart of a task processing method is schematically shown according to an embodiment of the present disclosure;
[0014] Figure 3 A scenario diagram is schematically shown according to an embodiment of the present disclosure, which is a period mapping relationship matching a first task in a first backtracking link;
[0015] Figure 4 An application scenario diagram is schematically shown according to an embodiment of the present disclosure, which is a determination of a backtracking period of a first task;
[0016] Figure 5 A flowchart is schematically shown according to an embodiment of the present disclosure, which is a determination of a first backtracking link;
[0017] Figure 6A An application scenario diagram is schematically shown according to an embodiment of the present disclosure, which is an optimization of a second task in a second backtracking link to obtain a first backtracking link;
[0018] Figure 6B An application scenario diagram is schematically shown according to another embodiment of the present disclosure, which is an optimization of a second task in a second backtracking link to obtain a first backtracking link;
[0019] Figure 6C An application scenario diagram is schematically shown according to another embodiment of the present disclosure, which is an optimization of a second task in a second backtracking link to obtain a first backtracking link;
[0020] Figure 7 A scenario diagram is schematically shown according to an embodiment of the present disclosure, which is a determination of a second backtracking link;
[0021] Figure 8 A block diagram of a system architecture is schematically shown according to an embodiment of the present disclosure, which applies a task processing method;
[0022] Figure 9 a structural block diagram of a task processing apparatus is shown schematically according to an embodiment of the present disclosure; and
[0023] Figure 10 a block diagram of an electronic device suitable for implementing a task processing method is shown schematically according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in a descriptive sense only. Thus, it will be apparent to one of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0025] In the task backtracking process, on the one hand, due to the date offset of the multiple task call data tables, the multiple tasks cannot be triggered for backtracking through a unified date, thereby resulting in the need for manual backtracking of the tasks, and the backtracking efficiency is low. In addition, when the task backtracking involves multiple batches, the task backtracking between the multiple batches will also have a date offset, thereby resulting in the inability to automatically trigger between the batches, and the backtracking efficiency is low. On the other hand, due to the complexity of the tasks and data tables involved in the entire backtracking link, some special scenarios also need to be manually operated for task backtracking, not only the backtracking efficiency is low, but also the backtracking data is prone to errors, affecting the data accuracy.
[0026] Therefore, an embodiment of the present disclosure provides a task processing method, comprising: in response to detecting that target data changes, determining a first backtracking link matched with the target data, the first backtracking link comprising multiple first tasks having a dependency relationship; determining a backtracking period of each of the multiple first tasks based on a change period of the target data and a time parameter of the first tasks; and backtracking data of the first tasks within the backtracking period according to the first backtracking link to obtain a backtracking result. By adopting the above-mentioned embodiment, the present disclosure can at least partially solve the technical problem of low task backtracking efficiency, and achieve the technical effect of improving the backtracking efficiency.
[0027] Figure 1 a block diagram of an electronic device suitable for implementing a task processing method is shown schematically according to an embodiment of the present disclosure.
[0028] It should be noted that, Figure 1The examples shown are merely examples of application scenarios in which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the application scenario in which the task processing method and apparatus can be applied may include a terminal device, but the terminal device can implement the task processing method and apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0029] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0031] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0032] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.
[0033] The server can be a cloud server, also known as a cloud computing server or cloud host. It is a hosting product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server can also be a distributed system server or a server integrated with blockchain.
[0034] It should be noted that the task processing method provided in the embodiments of the present disclosure can also be executed by the server 105. Correspondingly, the task processing apparatus provided in the embodiments of the present disclosure can be arranged in the server 105. The task processing method provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105. Correspondingly, the task processing apparatus provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105. Alternatively, the task processing method provided in the embodiments of the present disclosure can also be executed by the terminal device 101, 102, or 103. Correspondingly, the task processing apparatus provided in the embodiments of the present disclosure can be arranged in the terminal device 101, 102, or 103.
[0035] For example, the user can trigger the detection of the changed target data through the interaction with the interaction interface of the terminal device 101, 102, or 103. The terminal device 101, 102, or 103 sends a triggering instruction to the server 105. In response to detecting that the target data has changed, the server 105 determines a first rollback link matched with the target data, the first rollback link including a plurality of first tasks that exist in a dependency relationship; determines a rollback period of each of the plurality of first tasks based on a change period of the target data and a time parameter of the first task; and performs rollback on data of the first task within the rollback period according to the first rollback link to obtain a rollback result.
[0036] It should be understood that, Figure 1 The number of terminal devices, networks and servers in the above description is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks and servers.
[0037] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of the user's personal information comply with the relevant legal regulations, necessary security measures are taken, and do not violate public order and good customs.
[0038] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before the user's personal information is acquired or collected.
[0039] Figure 2 A flowchart of a task processing method according to an embodiment of the present disclosure is schematically shown. As shown in Figure 2 The task processing method 200 includes operations S210-S230.
[0040] In operation S210, in response to detecting that the target data has changed, a first rollback link matching the target data is determined, the first rollback link including a plurality of first tasks having a dependency relationship.
[0041] In operation S220, based on a change period of the target data and a time parameter of the first tasks, a rollback period of each of the first tasks is determined.
[0042] In operation S230, data of the first tasks within the rollback period is rolled back according to the first rollback link, to obtain a rollback result.
[0043] The target data refers to data that has changed, which can be all data in a data table, data in a field of a data table, or a change in the data structure of a data table, such as adding or deleting a field. The change period refers to the time period during which the target data has changed.
[0044] For example, the data of field 1 has changed between January 31 and February 15, so the change period is from January 31 to February 15. For a data table with added or deleted fields, the change period can be the entire life cycle of the data involved in the data table. For example, a data table includes data between January 1 and February 15. If field A is added or deleted, the data between January 1 and February 15 for field A will be added or deleted, and the change period is from January 1 to February 15.
[0045] The first task includes a task that directly or indirectly calls the target data, where direct calling refers to a task that directly uses the target data, and indirect calling refers to a task that uses data determined based on the target data. Indirect calling can be transmitted through the dependency relationship between the first tasks.
[0046] Dependency generally refers to the interrelation and constraint relationship between components, data, and systems. In embodiments of the present disclosure, the dependency relationship between the plurality of first tasks refers to the association of data between the plurality of first tasks. Through the dependency relationship between the plurality of first tasks, the plurality of first tasks can be combined into a first rollback link.
[0047] For example, the first rollback link includes two link branches, first task A→first task B→first task C and first task A→first task D, first task A is a task that directly calls the target data, second task B and D have a dependency relationship with first task A, and second task C has a dependency relationship with second task B. The above first rollback link is a first rollback link matching the target data, that is, the first rollback link includes a plurality of first tasks affected by the change of the target data and the order of the influence.
[0048] In one embodiment, the target data that changes can be detected by a timing or polling instruction on a certain data table. In another embodiment, the target data that changes can also be determined through interaction with the user. For example, the user inputs the target data that changes and its change period through interaction with the terminal device, and the server receives the target data and the change period after detecting the user input.
[0049] The time parameters include a plurality of time parameters related to the first task, such as a task execution period of the first task, a data period of the first task calling data, and the like.
[0050] The backtracking period refers to a data period that needs to be backtracked in the first task. For example, the first task calls data table 1 and data table 2, data table 1 needs to backtrack data from February 1 to February 2, and data table 2 needs to backtrack data from February 1 to February 5, if the data in data table 1 is the target data, the backtracking period of the first task is from February 1 to February 2.
[0051] Since a plurality of first tasks can call the target data at a plurality of times or call data generated based on the target data, it is difficult to backtrack the plurality of first tasks by a unified time. However, there is a calling relationship between the first task and the target data, there is a dependency relationship between the plurality of first tasks, and the change period of the target data also affects the data of the plurality of first tasks, which can be determined by the first backtracking link. Therefore, the backtracking period of each first task can be determined based on the change period and the time parameters of the first task.
[0052] The change of the target data not only affects the target data itself, but also chain-affectedly affects the first tasks that directly or indirectly call the target data, and the data transmission between the first tasks is limited by the dependency relationship. For example, for the first task A→the first task B→the first task C in the first backtracking link, when backtracking, the data of the first task A in its backtracking period needs to be backtracked first, and then the data of the first task B in its backtracking period needs to be backtracked; and then the data of the first task C in its backtracking period needs to be backtracked, so as to ensure the correct transmission of data in the entire first backtracking link.
[0053] Therefore, when backtracking the task, the data of each first task in its backtracking period is backtracked according to the dependency relationship between the plurality of first tasks in the first backtracking link, and the data of each first task obtained after backtracking is the backtracking result.
[0054] In the embodiments of the present disclosure, since the first rollback link matching the target data is determined in response to detecting that the target data changes, the plurality of first tasks affected by the target data can be accurately determined; further, based on the change period of the target data and the time parameter of the first task, the respective rollback period of each first task is automatically determined, and the data of the first task within the rollback period is rolled back according to the first rollback link to obtain the rollback result, so that the data rollback of each first task within the rollback period can be realized without manual rollback, and the rollback efficiency is improved.
[0055] According to the embodiments of the present disclosure, for the operation S220, the rollback period of the first task is determined based on the change period of the target data and the time parameter of the first task, including: determining a period mapping relationship matching the first task based on the task type of the first task; and determining the rollback period of the first task based on the period mapping relationship matching the first task, the change period and the time parameter.
[0056] The period mapping relationship refers to the mapping relationship between the change period, the time parameter and the rollback period. For example, the period mapping relationship can be a mapping relationship obtained by performing mathematical transformation on the change period and the time parameter to obtain the rollback period.
[0057] The first rollback link usually includes first tasks of multiple task types, and the multiple task types have different calling times or calling time ranges for the same data. Therefore, for the multiple task types, the rollback period of the first task can be determined through multiple period mapping relationships.
[0058] In one embodiment, the multiple period mapping relationships can use the same time parameter, and different mathematical transformations are performed on the time parameter and the change period to obtain the rollback period. Alternatively, the multiple period mapping relationships can also use different time parameters, and data changes are performed on the change period and the different time parameters to obtain the rollback period.
[0059] In the embodiments of the present disclosure, the period mapping relationship matching the first task is determined through the task type of the first task in the first rollback link, and the rollback period of the first task is determined based on the period mapping relationship matching the first task, the change period and the time parameter. Therefore, the embodiments of the present disclosure can realize the rollback period calculation of the first task based on the same change period, without the need for the user to manually perform the rollback operation, thereby improving the rollback efficiency; in addition, through the period mapping relationship matching the first task, the rollback period calculation of the first task of multiple task types is also supported, thereby expanding the applicable scenarios of the multiple task types in the rollback scenario.
[0060] According to an embodiment of the present disclosure, based on a task type of the first task, a period mapping relationship matched with the first task is determined, including: in a case where the task type is a first type, a first period mapping relationship is determined, the first type representing a task including a preset data set, the preset data set including a data set in which latest data covers historical data; in a case where the task type is a second type, a second period mapping relationship is determined, the second type representing a task including the preset data set and having a recovery cover data operation; and in a case where the task type is a third type, a third period mapping relationship is determined, the third type representing a task not including the preset data set.
[0061] The preset data set refers to a data set in which latest data covers historical data for data of a same field or a same data table. The data cover in the preset data set can be performed according to a preset rule or periodically. For example, the preset data set can include a dynamic partition table, which is used to dynamically create partitions according to time dimensions such as days, months or quarters, that is, the partitions are periodically covered according to day, month or quarter, so that when data of a specific time period is queried, the corresponding partition can be quickly located.
[0062] The first task of the first type, that is, the task including the preset data set, for example, can be a routine task including the preset data set. Because historical data in the preset data set can be covered to cause data loss, there can be the first task of the second type for performing a recovery operation on the covered data. The first task of the second type can be a backtracking task including the dynamic partition table. In addition, the first task of the third type can be a normal routine task (not including the preset data set). The routine task is executed once periodically, such as a task executed once a day; and the backtracking task is a task that needs to be executed multiple times in a short time.
[0063] In an embodiment, the first backtracking link can simultaneously include the first task of at least one type. For example, the first backtracking link can simultaneously include the first task of the second type and the third type to avoid data loss in the preset data set. Alternatively, for a routine task having special requirements, the first backtracking link can further include the first type to meet the special requirements of the first task of the first type, such as supporting backtracking of the first task of the first type.
[0064] In an embodiment of the present disclosure, by determining the first period mapping relationship, the second period mapping relationship and the third period mapping relationship respectively matched with the first type, the second type and the third type, backtracking of the special first task including the preset data set is supported, which helps to meet the backtracking requirements of users in various scenarios and improves the backtracking efficiency of the special task.
[0065] Figure 3A scenario diagram illustrating a period mapping relationship matched with a first task in a first backtracking link according to an embodiment of the present disclosure is shown schematically. As shown in the scenario 300, the first backtracking link includes a plurality of first tasks with a dependency relationship, such as a first task 1 301, a first task 2 302, a first task 3 303, and so on, wherein the first task 2 and the first task 3 both have a dependency relationship with the first task 1. For the first task 3 303, the task type thereof can be a first type, a second type, or a third type, whereby a first period mapping relationship 3031, a second period mapping relationship 3032, or a third period mapping relationship 3033 matched with the first task 3 303 can be determined. In the first backtracking link, the first task 1 301, the first task 2 302, and so on can be any task type. Figure 3 As shown in the scenario 300, the first backtracking link includes a plurality of first tasks with a dependency relationship, such as a first task 1 301, a first task 2 302, a first task 3 303, and so on, wherein the first task 2 and the first task 3 both have a dependency relationship with the first task 1. For the first task 3 303, the task type thereof can be a first type, a second type, or a third type, whereby a first period mapping relationship 3031, a second period mapping relationship 3032, or a third period mapping relationship 3033 matched with the first task 3 303 can be determined. In the first backtracking link, the first task 1 301, the first task 2 302, and so on can be any task type.
[0066] According to an embodiment of the present disclosure, the time parameter includes a first start parameter and a first end parameter for backtracking data, and the change period includes a second start parameter and a second end parameter.
[0067] Based on the period mapping relationship, the change period, and the time parameter, a backtracking time period of the first task is determined, including one of the following: according to the first period mapping relationship, determining a third end parameter of the backtracking time period based on the second end parameter and the first start parameter, and a third start parameter of the backtracking time period being the same as the second start parameter; according to the second period mapping relationship, determining the third start parameter based on the second start parameter and the first start parameter, and determining the third end parameter based on the second end parameter and the first start parameter; and according to the third period mapping relationship, determining the third start parameter based on the second start parameter and the first end parameter, and determining the third end parameter based on the second end parameter and the first start parameter.
[0068] The first start parameter and the first end parameter of the first task can be an offset start time and an offset end time compared with data at a certain moment of target data. The second start parameter and the second end parameter of the change period can be a start change time and an end change time of the target data. The third start parameter and the third end parameter of the backtracking time period are a start time and an end time of data that the first task needs to backtrack.
[0069] For example, for a first task of the first type, the backtracking time period of the first task is: the second start parameter~(the second end parameter-the first start parameter), the first period mapping relationship is that the third start parameter=the second start parameter, and the third end parameter=the second end parameter-the first start parameter.
[0070] For the first task of the second type, the backtracking period of the first task is: (second start parameter + first start parameter) ~ (second end parameter - first start parameter), and the second period mapping relationship can be: third start parameter = second start parameter + first start parameter, and third end parameter = second end parameter - first start parameter.
[0071] For the first task of the third type, the backtracking period of the first task is: (second start parameter - first end parameter) ~ (second end parameter - first start parameter), and the third period mapping relationship can be: third start parameter = second start parameter - first end parameter, and third end parameter = second end parameter - first start parameter.
[0072] In another embodiment, for the first task of the second type, if the time corresponding to the third end parameter exceeds the current time, the current time can be taken as the third end parameter to avoid the recovery operation of invalid data.
[0073] In the embodiments of the present disclosure, the first period mapping relationship, the second period mapping relationship and the third period mapping relationship can be automatically calculated by a predetermined time function, such as the date_add() function.
[0074] For example, the following two first tasks in Table 1 are taken as examples for illustration. The identification of the first task task1 is 12345, and the identification of task2 is 11111; the identification corresponding to the backtracking task field of task1 is 12345, that is, the backtracking task of task1 is task1 itself, and the dependency table does not include the dynamic partition table, and thus the task type of task1 is the third type; similarly, the identification of the backtracking task of task2 is 22222, which is different from 11111, and the dependency table includes the dynamic partition table turing.turing_table_x, and thus the task type of task2 is the second task. The dependency table is the data table to which the target data called by the first task belongs.
[0075] If the target data is the data in the turing.turing_table_b table, and the change period is 0510-0513, that is, May 10-13, then the backtracking period of task 12345 is 0515-0523, May 15 = May 10 - (-5 days), and May 23 = May 13 - (-10 days).
[0076] If the target data is data in the turing.turing_table_x table, the change period is still 0510-0513, and the backtracking period of task 22222 is 0410-0513, April 10 = May 10 + (-30 days), and May 13 - (-30 days) = June 12, which is greater than the current May 13, so the third end parameter of the backtracking period is taken as the current May 13.
[0077] Table 1
[0078]
[0079] In addition, if task2 is a downstream task of task3, the target data is data in table X, and the data in turing.turing_table_x is calculated using the data in table X, the backtracking period of backtracking task 22222 can also be calculated in the above manner, which will not be repeated here.
[0080] In other embodiments, a new field "is backfill" of a Boolean type can be used to mark whether the task type of the first task is a backtracking task. If the field value is true, it is a backtracking task; if the field value is false, it is a routine task.
[0081] In embodiments of the present disclosure, the backtracking period of the first task can be automatically calculated according to the time parameter and the change period by the first period mapping relationship, the second period mapping relationship, and the third period mapping relationship, without manual backtracking, thereby improving the task backtracking efficiency.
[0082] Figure 4 An application scenario diagram for determining the backtracking period of the first task according to embodiments of the present disclosure is schematically shown. As shown in FIG. 4, Figure 4 For the first task 401, a matching first period mapping relationship 4011, a second period mapping relationship 4012, or a third period mapping relationship 4013 can be determined according to the task type of the first task 401. Then, based on the first period mapping relationship 4011, the second period mapping relationship 4012, or the third period mapping relationship 4013, based on the time parameter 402 of the first task 401 and the change period 403 of the target data, the backtracking period 404 of the first task can be determined.
[0083] For example, the target data can be business-side data of a certain search engine app or forum chat app, such as the number of active users, the number of new users, the channel source for logging into the app, traffic, etc. The above target data can be stored by a separate data table, and can also be stored as a field of a business data table.
[0084] If the business changes, the upstream node re-counts the active user quantity from May 10 to May 13, and can pass the active user quantity as the changed data to the downstream node. After the downstream node receives the active user quantity, the active user quantity from May 10 to May 13 can be regarded as target data, and May 10 to May 13 is the change period of the target data. The first backtracking link matched with the target data includes a first task A for calculating the active user change rate and a first task B for counting the active user quantity. The two first tasks respectively need to update the data table storing the active user change rate and the data table storing the total active user quantity of consecutive days by using the active user quantity.
[0085] For the data table related to the first task A, when counting the active user change rate of each month, the data of May 10 is associated with the data of April 10, and the data table (which is a preset data set) is covered once every 30 days. However, since the change on May 10 affects the change on April 10, it is necessary to restore the data covering operation of the data table of the first task A and restore the data of April 10. Thus, the first task A belongs to the first type. The first task B only counts the active user quantity, and the mechanism of the data table does not need the latest data to cover the historical data. Therefore, the first task B is of the third type. If there is a special requirement, the first task A can be manually marked without performing the data restoring and covering operation. In this case, the first task A is of the first type.
[0086] The first start parameter and the first end parameter of the first task A of the first type are -30 and 0, respectively, and the first start parameter and the first end parameter of the second task B of the third type are -10 and -5, respectively. For the first tasks A and B, based on the respective time parameters and the respective corresponding period mapping relationship and the change period (May 10 to May 13), the backtracking period of each first task can be determined. For example, the backtracking period of the first task A is 0410-0513, and the backtracking period of the first task B is 0515-0523 (see the above description for specific calculation).
[0087] Since the first tasks A and B have different time parameters, the backtracking periods of the first tasks A and B cannot be determined by the unified May 10 to May 13, and the efficiency of manually triggering the backtracking of the first tasks A and B is too low. Through the above embodiment, multiple first tasks matched with the target data can be determined from the target data, the backtracking periods of the respective first tasks can be determined based on the change period of the target data and the time parameters, and data backtracking is performed. The whole process does not need the user to manually trigger the backtracking operation, thereby achieving the technical effect of improving the backtracking efficiency.
[0088] According to an embodiment of the present disclosure, the time parameter further comprises a task execution parameter; and the method further comprises: in a case where the task execution parameter meets a predetermined time condition, adjusting the third start parameter and the third end parameter of the backtracking time period.
[0089] In this embodiment, the task execution parameter can comprise a task execution time, such as the task timing time in Table 1 above. Since different systems have different timing times for executing tasks, and the time for dividing the data of the day and the next day can be different, the task execution parameter can be compared with the predetermined time condition to determine whether to adjust the backtracking time period.
[0090] For example, the preset time condition can be that the task execution parameter is greater than a predetermined time parameter, such as 22:00:00. If the task execution parameter 23:00:00 is greater than the predetermined time parameter 22:00:00, the third start parameter and the third end parameter can be reduced by one, such as adjusting the 4 / 10-5 / 13 determined above to 4 / 9-5 / 12; otherwise, the third start parameter and the third end parameter are not adjusted.
[0091] In an embodiment of the present disclosure, by adjusting the third start parameter and the third end parameter of the backtracking time period in a case where the task execution parameter meets the predetermined time condition, the determination of the backtracking time period can meet the actual data recovery situation, thereby ensuring the accuracy of the backtracking data.
[0092] Figure 5 A flowchart for determining a first backtracking link according to an embodiment of the present disclosure is schematically shown. As shown in Figure 5 Embodiment 500 for determining a first backtracking link includes operations S511-S513, which can be one specific embodiment of operation S210 described above.
[0093] In operation S511, in response to detecting a change in the target data table, a plurality of second tasks for calling the target data are determined from the database.
[0094] In operation S512, based on the dependency relationship between the plurality of second tasks, a second backtracking link is determined.
[0095] In operation S513, the second tasks in the second backtracking link are optimized to obtain a first backtracking link.
[0096] The database can include code files of a plurality of tasks, and by parsing the code files in the database, the calling relationship between the plurality of tasks and the target data can be determined. The second task is a task that directly or indirectly calls the target data. It should be noted that the first task and the second task are only used to distinguish the operations in different stages, and the first task and the second task can be the same or different.
[0097] Similarly, the dependency relationship between the plurality of second tasks can be determined by analyzing the code files of the tasks; based on the dependency relationship between the plurality of second tasks, an initial second backtracking link can be composed.
[0098] In some embodiments, the first backtracking link and the second backtracking link can be represented by a directed acyclic graph (DAG), in which the nodes are tasks and the directed edges are the dependency relationships between the tasks.
[0099] Since the second tasks involved in the second backtracking link are more, and the dependency relationship between the second tasks can cause the backtracking tasks to enter a loop, so that the backtracking cannot be normally performed; or, the second backtracking link includes second tasks that do not actually need to be backtracked; or, the first type of second tasks in the second backtracking link can cause data loss, therefore, the second tasks in the second backtracking link are optimized by operation S513 to obtain the first backtracking link. The optimization of the second tasks includes but is not limited to: deleting, replacing, and merging the second tasks.
[0100] In embodiments of the present disclosure, in response to detecting that the target data changes, a plurality of second tasks for calling the target data are determined from the database; based on the dependency relationship between the plurality of second tasks, a second backtracking link is determined; and the second tasks in the second backtracking link are optimized to obtain a first backtracking link. Since the second tasks in the second backtracking link are optimized after the preliminary second backtracking link is constructed based on the dependency relationship, the dependency relationship between the plurality of first tasks in the first backtracking link is correct, and the backtracking link is correct, so that correct data can be backtracked based on the first backtracking link.
[0101] In some embodiments, for operation S513, the optimization of the second tasks in the second backtracking link to obtain the first backtracking link includes: based on the dependency relationship between the plurality of second tasks, determining a second task pair with a circular dependency relationship from the second backtracking link, the two second tasks in the second task pair including the same data table; based on the second task pair, adjusting the second task in the second task pair that is located downstream of the second backtracking link; and deleting the second task in the second task pair that is located upstream of the second backtracking link to obtain the first backtracking link.
[0102] By detecting whether there are two second tasks including the same data table in the second backtracking link, it is determined whether the above two second tasks are a second task pair with a circular dependency relationship.
[0103] In some embodiments, different second tasks in the second backtracking link can call data of different partitions of the same data table (it can be understood that the data of different partitions are all affected by the target data), since the different second tasks all call the same data table, the backtracking operation in the second backtracking link can circulate between the above two second tasks, causing the task backtracking to fail to proceed normally. Therefore, if the two second tasks include the same data table, the two second tasks are referred to as a second task pair with a circular dependency relationship.
[0104] It can be understood that in the second backtracking link, the two second tasks in the second task pair have a preceding and following relationship, that is, one is upstream and one is downstream.
[0105] Since the two second tasks in the second task pair both include the same data table, the above two second tasks can be merged. When merging, considering that the second task downstream of the second backtracking link can involve other data, to ensure the dependency relationship between the data, the second task upstream of the second backtracking link in the second task pair can be merged into the second task downstream of the second backtracking link. Then, the second task upstream of the second backtracking link can be deleted from the second backtracking link.
[0106] In the embodiments of the present disclosure, by determining the second task pair with a circular dependency relationship from the second backtracking link based on the dependency relationship between the plurality of second tasks, the two second tasks in the second task pair include the same data table; based on the second task pair, adjusting the second task downstream of the second backtracking link in the second task pair; deleting the second task upstream of the second backtracking link in the second task pair, obtaining the first backtracking link, thereby, the backtracking into the loop of the first backtracking link caused by the circular dependency relationship can be avoided, and the accuracy of the backtracking result is ensured.
[0107] Figure 6A An application scenario diagram of obtaining the first backtracking link by optimizing the second task in the second backtracking link according to an embodiment of the present disclosure is schematically shown. As shown in Figure 6A In the embodiment 600A, the second backtracking link includes task 1→task 3→task 4→task 2, wherein task 1 and task 2 respectively include the same data table, such as table A, and task 1 and task 2 call the data of partition 1 and partition 2 of table A respectively. Therefore, the actual dependency relationship of the second backtracking link is actually task 1→task 3→task 4→task 2→task 3, that is, after the data of table A of task 2 is backtracked, task 3 will be executed based on the dependency relationship of table A in “task 1→task 3”, thereby causing a loop. By merging task 1 into task 2, the optimization of the second backtracking link is realized, and the first backtracking link obtained thereby does not have a circular dependency relationship.
[0108] According to an embodiment of the present disclosure, for operation S513, the second tasks in the second rollback link are optimized to obtain the first rollback link, including: determining the second tasks of the first type from the second rollback link; and replacing the second tasks of the first type with the second tasks of the second type to obtain the first rollback link, the second tasks of the second type being created based on the second tasks of the first type.
[0109] Data coverage exists in the first type of routine tasks, which may cause data loss. To avoid possible data loss, at least one second task of the first type can be determined according to the task label of the second task in the second rollback link; and the second task of the first type is replaced with the second task of the second type to obtain the first rollback link.
[0110] The second task of the second type can be created based on the second task of the first type, and an operation of recovering the coverage data is added between the second tasks of the first type and the second type.
[0111] Figure 6B An application scenario diagram of optimizing the second tasks in the second rollback link to obtain the first rollback link according to another embodiment of the present disclosure is schematically shown. As shown in Figure 6B The second rollback link in embodiment 600B includes task 1→task 3→task 4→task 2, task 3 is a routine task including a dynamic partition table, that is, task 3 is a second task of the first type, and thus, the optimization of the second rollback link is implemented by replacing task 3 with a rollback task including a dynamic partition table, such as task 5, to obtain the first rollback link.
[0112] In the embodiments of the present disclosure, by replacing the second task of the first type with the second task of the second type to obtain the first rollback link, the problem of the rollback result caused by data loss can be avoided, and the accuracy of the rollback result is ensured.
[0113] In this embodiment, the second tasks in the second rollback link are all tasks of the first type or the third type, that is, all are routine tasks; and all or part of the second tasks of the first type in the second rollback link can be replaced with the second tasks of the second type in the above manner.
[0114] According to an embodiment of the present disclosure, for operation S513, the second tasks in the second rollback link are optimized to obtain the first rollback link, including: obtaining a rollback whitelist matched with target data, the target data including unchanged fields associated with the target data, the rollback whitelist including tasks for using the unchanged fields; and deleting the second tasks belonging to the rollback whitelist from the second rollback link to obtain the first rollback link.
[0115] In an embodiment of the present disclosure, the target data of the change can be the data of a certain field in a data table, however, the subsequent second task can only call the unchanged field in the data table, and thus the unchanged field in the data table is the unchanged field associated with the target data. Since the subsequent second task can only call the unchanged field in the data table, the subsequent second task actually does not need to be traced back. Thus, for each target data of the change, there is a dynamic whitelist for backtracking.
[0116] In an embodiment of the present disclosure, for each second task in each second backtracking link, the first backtracking link is obtained by matching the second task with the tasks in the backtracking whitelist and deleting the second task belonging to the backtracking whitelist from the second backtracking link.
[0117] In another embodiment, the backtracking whitelist further includes a custom task. The custom task can be a second task determined based on expert experience or based on user interaction.
[0118] Figure 6C An application scenario diagram of optimizing the second task in the second backtracking link to obtain the first backtracking link according to another embodiment of the present disclosure is schematically shown. As shown in Figure 6C The second backtracking link of the embodiment 600C includes task 1→task 3→task 4→task 2, wherein task 4 is a task in the backtracking whitelist, and thus task 4 can be deleted from the second backtracking link, and the first backtracking link is obtained based on the relationship between task 4 and the upstream task 3 and the relationship between task 4 and the downstream task 2. In the first backtracking link, task 1 and task 3 are directly connected.
[0119] In an embodiment of the present disclosure, the first backtracking link is obtained by deleting the second task belonging to the backtracking whitelist from the second backtracking link, which can reduce the amount of tasks for backtracking on the basis of ensuring the accuracy of the backtracking result, and improve the backtracking efficiency.
[0120] According to an embodiment of the present disclosure, the second backtracking link obtained in operation S512 can be a link after the repeated link branches are removed.
[0121] In an embodiment, a search algorithm can be used to remove the repeated link branches, such as a Breadth-First Search (BFS) algorithm, and thus the second backtracking link obtained is the minimum execution plan for performing task backtracking.
[0122] Figure 7 A scenario diagram for determining the second backtracking link according to an embodiment of the present disclosure is schematically shown. As shown in Figure 7As shown in the embodiment 700, based on the dependency relationship between the second tasks, two groups of backtracking branches, backtracking branch 701 and backtracking branch 702, can be obtained. The backtracking branch 701 includes task 1→task 3→task 5, task 2→task 3→task 5, task 2→task 4, and the backtracking branch 702 includes task 6→task 3→task 5, task 6→task 7→task 4. As can be seen, task 1, task 2 and task 6 all have a dependency relationship with “task 3→task 5”, and the second backtracking link 703 can be obtained by removing the duplicate branches.
[0123] In this embodiment, the second tasks in the second backtracking link can have different levels. For example, the second task that calls external data has a lower level, and the second task that finally outputs data to the outside has a higher level. As shown in the embodiment 700, Figure 7 As shown, task 1, task 2 and task 6 are level 0, task 3 and task 7 are level 1, and task 5 and task 4 are level 2. The lower the level, the shallower the dependency depth of the task on the target data; on the contrary, the higher the level, the deeper the dependency depth of the task on the target data.
[0124] After the second backtracking link is optimized, the level between the tasks in the first backtracking link is usually unchanged.
[0125] According to an embodiment of the present disclosure, for operation S230, the data of the first task within the backtracking period is backtracked according to the first backtracking link to obtain a backtracking result, including: according to the dependency relationship between the plurality of first tasks in the first backtracking link, adding the task instance of the first task to at least one task queue, the at least one task queue corresponding to at least one priority; and executing the task instance in the at least one task queue to backtrace the data of the first task within the backtracking period to obtain the backtracking result.
[0126] When backtracking the tasks, the dependency relationship between the first tasks in the first backtracking link is used for backtracking. If there is a dependency relationship “task 1→task 3” between task 1 and task 3, the data in task 1 is backtracked first, and then the data in task 3 is backtracked. Therefore, when the plurality of first tasks are added to the task queue, it is necessary to ensure that the first task located upstream of the first backtracking link is executed before the first task located downstream of the first backtracking link.
[0127] For ease of understanding, the following will take the first task belonging to the same level in the first backtracking link generated according to the dependency relationship as an example to explain how to add the task instance of the first task to the task queue.
[0128] The first task can implement the backtracking of the data in the backtracking period through a task instance. In an embodiment, the backtracking period can include a plurality of backtracking sub-periods of a predetermined length, such as one day, whereby a single first task can have one or more task instances, each of which is used to backtrack the data in the backtracking sub-period. Alternatively, the first task can backtrack the data of a plurality of fields through a plurality of task instances.
[0129] According to an embodiment of the present disclosure, there can be one or more task queues for executing task instances, and the task instance of each first task can be added to one or more task queues.
[0130] For example, there can be at least one priority corresponding to at least one task queue, and there can be a plurality of priorities among the plurality of first tasks, and the task instances of the first tasks can be added to the task queues of the at least one priority according to the priorities of the first tasks.
[0131] For another example, there can also be a plurality of priorities among the plurality of task instances of the first tasks, such as a higher priority of the task instance of the backtracking sub-period in front than a priority of the task instance of the backtracking sub-period behind. The task instances with the plurality of priorities in each first task can also be added to a plurality of task queues.
[0132] For another example, for the task instances or the first tasks without priority order, the task instances of the first tasks can be sequentially added to the task queues according to the order of the first tasks in the first backtracking link, and the task queues sequentially execute the task instances based on the First In First Out (FIFO) principle. Alternatively, the FIFO and the priority strategy can also be combined, the task instances of the first tasks are first added to the corresponding task queues according to the priorities to implement the execution of the task instances with higher priorities first, and then, for the task instances without priority difference, the task instances are executed based on the FIFO strategy of the task queues.
[0133] In addition, the first task can also have a concurrency number parameter and / or a batch parameter, the concurrency number parameter is used to indicate the number of task instances of the first task executed simultaneously, and the batch parameter is used to indicate the batch of the plurality of first tasks executed simultaneously or the batch of the plurality of task instances in the first task executed simultaneously.
[0134] In an embodiment, for the task instances determined according to the backtracking sub-period, there is a time dependency among the plurality of task instances, and therefore, the concurrency number parameter of the first task is 1. For the task instances of the plurality of first tasks at the same level, the batch parameters can be the same, that is, the task instances of the above-mentioned plurality of first tasks can be executed in the same batch.
[0135] In the embodiments of the present disclosure, by executing the task instances of each first task with the task queue with multiple priorities according to the dependency relationship of the multiple first tasks in the first backtracking link, the first task or task instance with higher priority can be guaranteed to be backtracked first, and the resources of the task queue can be prevented from being preempted by the first task or task instance with lower priority. Thus, the timeliness of task backtracking can be guaranteed in the embodiments of the present disclosure.
[0136] In one embodiment, the number of task instances executed by each task queue or all task queues can also be limited. For example, the number of task instances executed by each task queue or all task queues at the same time is limited to a target threshold, such as 30. In addition, the number of task instances that occupy more CPU resources can be further limited by a token bucket algorithm to ensure that the target threshold of the number of task instances executed by each task queue or all task queues at the same time can dynamically change while balancing the CPU resources.
[0137] In another embodiment, each task queue can also correspond to an executable time period of executable task instances. The task instances of the first task are added to at least one task queue based on the executable time period of each task queue. Specifically, the task queue that can execute task instances at the current time can be dynamically determined according to the executable time period of each task queue, and then the task instances of each first task are added to the task queue that can execute task instances.
[0138] In still another embodiment, each task queue can also correspond to a queue blacklist including the identifiers of the first tasks that are not executed by the task queue to balance the resources of each task queue. For example, the queue blacklist can be configured in the form of a cron expression.
[0139] In still another embodiment, each task queue stores the data of the backtracking of each task instance to the cluster, such as a Redis cluster, so as to avoid the loss of the backtracked data in the case of interruption of the execution of the task instances by the task queue. In this embodiment, a checkpoint can be created according to a predetermined strategy, so that in the case of interruption of the execution of the task instances by the task queue, whether there is a previously saved checkpoint file is detected, and if there is a saved checkpoint file, the execution can be continued from the last interruption instead of re-executing all the task instances.
[0140] Figure 8 A block diagram of a system architecture applying the task processing method according to the embodiments of the present disclosure is schematically shown.
[0141] As Figure 8As shown, the system framework 800 includes bottom-level operators, architecture changes, data detection, and application support from bottom to top. The bottom-level operators can include operators for data extraction (Extract), transformation (Transform), and loading (Load), referred to as etl operators; or structured query language (sql) operators for various operations on databases. The bottom-level operators are used to implement development tasks in routine production scenarios and to implement start-back operations in backtracking scenarios.
[0142] Architecture changes refer to changes from architecture 1 to architecture 2, which can be implemented by performing data import operations in routine scenarios, thereby achieving uniformity of the architecture in migration scenarios. For example, files in a distributed file system can be migrated to a database, a data warehouse, or a columnar repository, such as migration of asf to a sql database, a palo data warehouse, or a ClickHouse database; the data architecture in a data warehouse can also be changed, such as optimization of the architecture of a palo data warehouse; or data under the turing framework can be migrated to the ClickHouse database framework.
[0143] Data detection is mainly routine detection of data in routine scenarios, such as detecting the difference in data volume before and after the target data changes, and in the case where the data volume fluctuation is less than a fluctuation threshold (such as less than 5%), the backtracking result of the target data is considered correct by default. Alternatively, the difference in each field before and after the change in the target data is detected, and a visual field change heat map is generated for the user to view. Alternatively, the time at which the backtracking result is obtained is announced. In addition, if the data volume fluctuation in the backtracking scenario is greater than or equal to the fluctuation threshold, real-time backtracking announcement is performed, and the backtracking result is not passed; and a backtracking blocking operation is automatically triggered, pushing the alarm information of the context traceability and the reason analysis of the backtracking result not passing. The reason analysis can be: missing dependencies / insufficient resources / syntax error classification alarms, etc. In addition, data detection also supports obtaining real-time logs to display the real-time backtracking progress and display the occupation of key indicators. The key indicators include the total number of tasks, the success rate, the average execution time, the resource utilization, etc.
[0144] The application support can support multiple data applications, and the data applications can include a Turing framework-based data development governance platform (Turing Data Studio, TDA), a visual data analysis platform (Turing Data Analysis, UBI), and the like. In a routine scenario, multiple data processing operations of the data applications can be supported, and in a backtracking scenario, platform announcements are also supported, such as automatically calculating a downstream influence link (such as a first backtracking link and a second backtracking link) after a target data change, generating a visual DAG graph, and the like. The application also supports outputting an announcement template adapted to the TDA / UBI platform to show a user a change summary, an influence range, an operation suggestion, a background description, a data volatility rate, and the like of the target data. The application also supports intelligent customer service connection, and solves problems existing in a backtracking result in real time through a Frequently Asked Questions (FAQ) database, a discussion group, and the like. In addition, the application also supports calculating an amount of data influenced by the backtracked data in the entire system, such as counting the amount of data in a partition dimension of a table.
[0145] In an embodiment of the present disclosure, when a task instance is executed, through mechanisms such as queue priority, concurrent quantity limitation, a blacklist, and a whitelist, the balance between system stability and flexibility can be achieved in the process of ensuring backtracking stability.
[0146] Figure 9 A structural block diagram of a task processing apparatus according to an embodiment of the present disclosure is schematically shown.
[0147] As shown in Figure 9 The task processing apparatus 900 includes a first determination module 910, a second determination module 920, and a backtracking module 930.
[0148] The first determination module 910 is configured to, in response to detecting a change in target data, determine a first backtracking link matched with the target data, the first backtracking link including a plurality of first tasks having a dependency relationship.
[0149] The second determination module 920 is configured to determine a backtracking time period of each of the plurality of first tasks based on a change period of the target data and a time parameter of the first task.
[0150] The backtracking module 930 is configured to backtrack data of the first task within the backtracking time period according to the first backtracking link, to obtain a backtracking result.
[0151] According to an embodiment of the present disclosure, the second determination module 920 includes a first determination submodule and a second determination submodule.
[0152] The first determination submodule is configured to determine a period mapping relationship matched with the first task based on a task type of the first task.
[0153] The second determining submodule is configured to determine a backtracking time period of the first task based on the period mapping relationship matched with the first task, the change period, and the time parameter.
[0154] According to an embodiment of the present disclosure, the first determining submodule comprises a first determining unit, a second determining unit, and a third determining unit.
[0155] The first determining unit is configured to determine the first period mapping relationship in a case where the task type is a first type, the first type representing a task comprising a preset data set, the preset data set comprising a data set in which latest data covers historical data.
[0156] The second determining unit is configured to determine the second period mapping relationship in a case where the task type is a second type, the second type representing a task comprising a preset data set and having a recovery covering data operation.
[0157] The third determining unit is configured to determine the third period mapping relationship in a case where the task type is a third type, the third type representing a task not comprising a preset data set.
[0158] According to an embodiment of the present disclosure, the time parameter comprises a first start parameter and a first end parameter for backtracking data, and the change period comprises a second start parameter and a second end parameter.
[0159] The second determining submodule comprises one of the following: a fourth determining unit, a fifth determining unit, and a sixth determining unit.
[0160] The fourth determining unit is configured to determine a third end parameter of the backtracking time period based on the second end parameter and the first start parameter according to the first period mapping relationship, the third start parameter of the backtracking time period being the same as the second start parameter.
[0161] The fifth determining unit is configured to determine the third start parameter based on the second start parameter and the first start parameter, and determine the third end parameter based on the second end parameter and the first start parameter according to the second period mapping relationship.
[0162] The sixth determining unit is configured to determine the third start parameter based on the second start parameter and the first end parameter, and determine the third end parameter based on the second end parameter and the first start parameter according to the third period mapping relationship.
[0163] According to an embodiment of the present disclosure, the time parameter further comprises a task execution parameter, and the second determining submodule further comprises a seventh determining unit configured to adjust the third start parameter and the third end parameter of the backtracking time period in a case where the task execution parameter meets a predetermined time condition.
[0164] According to an embodiment of the present disclosure, the first determining module comprises an obtaining sub-module, a third determining sub-module and an optimization module.
[0165] The obtaining sub-module is configured to determine, from the database, a plurality of second tasks for invoking the target data in response to detecting that the target data has changed.
[0166] The third determining sub-module is configured to determine a second backtracking chain based on a dependency relationship between the plurality of second tasks.
[0167] The optimization module is configured to optimize the second tasks in the second backtracking chain to obtain a first backtracking chain.
[0168] According to an embodiment of the present disclosure, the optimization module comprises a task pair determining unit, an adjusting unit and a first deleting unit.
[0169] The task pair determining unit is configured to determine, from the second backtracking chain, a second task pair having a circular dependency relationship based on a dependency relationship between the plurality of second tasks, two second tasks in the second task pair comprising a same data table.
[0170] The adjusting unit is configured to adjust a second task located downstream of the second backtracking chain in the second task pair based on the second task pair.
[0171] The first deleting unit is configured to delete a second task located upstream of the second backtracking chain in the second task pair to obtain the first backtracking chain.
[0172] According to an embodiment of the present disclosure, the optimization module comprises an eighth determining unit and a replacing unit.
[0173] The eighth determining unit is configured to determine a first type of second task from the second backtracking chain.
[0174] The replacing unit is configured to replace the first type of second task with a second type of second task to obtain the first backtracking chain, the second type of second task being created based on the first type of second task.
[0175] According to an embodiment of the present disclosure, the optimization module comprises an obtaining unit and a second deleting unit.
[0176] The obtaining unit is configured to obtain a backtracking white list matched with the target data, a target data table comprising an unchanged field associated with the target data, and the backtracking white list comprising a task for using the unchanged field.
[0177] The second deleting unit is configured to delete a second task belonging to the backtracking white list from the second backtracking chain to obtain the first backtracking chain.
[0178] According to an embodiment of the present disclosure, the backtracking module comprises a queue processing sub-module and an execution sub-module.
[0179] The queue processing submodule is configured to add the task instance of the first task into at least one task queue according to the dependency relationship among the plurality of first tasks in the first backtracking link, the at least one task queue corresponding to at least one priority.
[0180] The execution submodule is configured to execute the task instance in the at least one task queue to backtrack the data of the first task within the backtracking period to obtain a backtracking result.
[0181] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0182] According to an embodiment of the present disclosure, an electronic device comprises at least one processor and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0183] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method as described above.
[0184] According to an embodiment of the present disclosure, a computer program product comprises a computer program, and the computer program is used to implement the method as described above when executed by a processor.
[0185] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0186] According to an embodiment of the present disclosure, an electronic device comprises at least one processor and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0187] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method as described above.
[0188] According to an embodiment of the present disclosure, a computer program product comprises a computer program, and the computer program is used to implement the method as described above when executed by a processor.
[0189] Figure 10 A block diagram of an electronic device suitable for implementing the task processing method according to an embodiment of the present disclosure is schematically shown.
[0190] Electronic device is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0191] like Figure 10 As shown, electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of electronic device 1000 may also be stored in RAM 1003. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.
[0192] Multiple components in electronic device 1000 are connected to an input / output (I / O) interface 1005, including an input unit 1006, such as a keyboard and mouse; an output unit 1007, such as various types of displays and speakers; a storage unit 1008, such as a magnetic disk and optical disk; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0193] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the task processing method. For example, in some embodiments, the task processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the task processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the task processing method by any other suitable means, such as by means of firmware.
[0194] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0195] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0196] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0197] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0198] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0199] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0200] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.
[0201] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A task processing method, comprising: In response to detecting a change in target data, determining a first backtracking link matching the target data, the first backtracking link including a plurality of first tasks having a dependency relationship; determining a backtracking period for each of the plurality of first tasks based on a change cycle of the target data and a time parameter of the first task; as well as According to the first backtracking link, the data of the first task within the backtracking period is backtracked to obtain a backtracking result.
2. The method according to claim 1, wherein The determining the backtracking period of the first task based on the change cycle and the time parameter of the first task includes: determining, based on the task type of the first task, a period mapping relationship matching the first task; and A backtracking period of the first task is determined based on a period mapping relationship matching the first task, the change period, and the time parameter.
3. The method according to claim 2, wherein: The determining, based on the task type of the first task, a period mapping relationship matching the first task includes: determining a first period mapping relationship when the task type is a first type, the first type representing a task including a preset data set, the preset data set including a data set that overwrites historical data with the latest data; In a case where the task type is a second type, determining a second period mapping relationship, where the second type represents a task including the preset data set and having an operation of recovering overwritten data; In a case where the task type is a third type, a third period mapping relationship is determined, where the third type represents a task that does not include the preset data set.
4. The method according to claim 3, wherein: The time parameter includes a first start parameter and a first end parameter for backtracking data, and the change period includes a second start parameter and a second end parameter. Determining the backtracking period of the first task based on the period mapping relationship, the change period, and the time parameter includes one of the following: determining, according to the first period mapping relationship, a third end parameter of the traceback period based on the second end parameter and the first start parameter, wherein the third start parameter of the traceback period is the same as the second start parameter; According to the second period mapping relationship, determining the third start parameter based on the second start parameter and the first start parameter, and determining the third end parameter based on the second end parameter and the first start parameter; According to the third period mapping relationship, the third start parameter is determined based on the second start parameter and the first end parameter, and the third end parameter is determined based on the second end parameter and the first start parameter.
5. The method according to claim 4, wherein The time parameter also includes a task execution parameter; the method further includes: When the task execution parameter satisfies a predetermined time condition, the third start parameter and the third end parameter of the backtracking period are adjusted.
6. The method according to any one of claims 2 to 5, wherein In response to detecting a change in the target data table, determining a first backtracking link matching the target data includes: In response to detecting a change in the target data, determining a plurality of second tasks for calling the target data from a database; Determining a second backtracking link based on the dependency relationship between the plurality of second tasks; as well as The second task in the second backtracking link is optimized to obtain the first backtracking link.
7. The method according to claim 6, wherein: The optimizing the second task in the second backtracking link to obtain the first backtracking link includes: Based on the dependency relationships among the plurality of second tasks, determining a second task pair having a circular dependency relationship from the second backtracking link, wherein two second tasks in the second task pair include a same data table; Based on the second task pair, adjusting a second task in the second task pair that is located downstream of the second backtracking link; and The second task located upstream of the second backtracking link in the second task pair is deleted to obtain the first backtracking link.
8. The method according to claim 6 or 7, wherein: The optimizing the second task in the second backtracking link to obtain the first backtracking link includes: determining a second task of the first type from the second backtracking link; and The first type of second task is replaced by a second type of second task to obtain the first backtracking link, where the second type of second task is created based on the first type of second task.
9. The method according to any one of claims 6 to 8, wherein The optimizing the second task in the second backtracking link to obtain the first backtracking link includes: Acquire a retroactive whitelist matching the target data, wherein the target data table includes unchanged fields associated with the target data, and the retroactive whitelist includes tasks for using the unchanged fields; and The second task belonging to the backtrace whitelist is deleted from the second backtrace link to obtain the first backtrace link.
10. The method according to any one of claims 1 to 9, wherein The step of backtracking the data of the first task within the backtracking period according to the first backtracking link to obtain a backtracking result includes: According to the dependency relationship between the plurality of first tasks in the first backtracking link, adding the task instance of the first task to at least one task queue, the at least one task queue corresponding to at least one priority; and At least one task instance in the task queue is executed to backtrack data of the first task within the backtracking period to obtain a backtracking result.
11. A task processing device comprising: a first determining module configured to determine, in response to detecting a change in a target data table, a first backtracking link matching the target data, the first backtracking link including a plurality of first tasks having a dependency relationship; a second determining module, configured to determine respective backtracking periods of the plurality of first tasks based on a change cycle of the target data and a time parameter of the first task; as well as A backtracking module is configured to backtrack the data of the first task within the backtracking period according to the first backtracking link to obtain a backtracking result.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 10.