Data backtracking method, device, electronic device and storage medium
By generating target topological relationships and coordinating the computing environment, and dynamically adjusting the computing environment of task instances, the problems of low data backtracking calculation efficiency and large resource overhead in the existing technology are solved, and efficient data quality optimization is achieved.
Patent Information
- Application Number
- CN202210299742.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-03-22
AI Technical Summary
In the prior art, the calculation efficiency of batch data backtracking of a single data source across time intervals is low, while the calculation efficiency of undifferentiated parallel data backtracking of multiple data sources consumes a lot of computing resources, and cannot effectively solve the problem of deterioration in data quality caused by abnormal data in big data.
By generating the target topological relationship, based on the first dependency and second dependency of the data to be traced, the computing environment is coordinated and the computing engine is called for task calculations, the computing environment of each task instance is dynamically adjusted, and the data traceback process is optimized.
It improves data backtracking computing efficiency, reduces computing resource overhead, and realizes effective optimization of data quality.
Smart Images

Figure CN114691658B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, further to the field of big data, and particularly to a data backtracking method, device, electronic device, and storage medium. Background Art
[0002] In today's big data era, the explosive growth of enterprise big data can lead to practical issues such as data timeliness, data security, and data quality. When managing data quality, due to the complex production chains and system relationships of big data, when a piece of data in the data is abnormal, it can affect its associated upstream and downstream data and systems, resulting in poor data quality. Data quality is a top strategic priority for enterprise organizations, and therefore requires optimization.
[0003] Related technologies typically use methods to assess data quality, such as batch data backtracking across time intervals from a single data source or parallel data backtracking from multiple data sources without differentiation. However, batch data backtracking across time intervals from a single data source is computationally inefficient, while parallel data backtracking from multiple data sources without differentiation requires significant computational resources.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The present disclosure provides a data backtracking method, device, electronic device, and storage medium.
[0006] According to one aspect of the present disclosure, a data backtracking method is provided, comprising: generating a target topological relationship based on a first dependency relationship and a second dependency relationship of data to be backtracked, wherein the data to be backtracked comprises: a plurality of data objects, each of the plurality of data objects comprises: a plurality of data units, the first dependency relationship is used to describe the dependency relationship between data units of different data objects, the second dependency relationship is used to describe the reference time dependency relationship between different data objects, and the target topological relationship is used to determine a plurality of task instances to be calculated; coordinating a corresponding computing environment for each of the plurality of task instances to be calculated according to a preset concurrency; calling a computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process.
[0007] According to another aspect of the present disclosure, a data backtracking device is provided, including: an analysis module for generating a target topology relationship based on a first dependency relationship and a second dependency relationship of the data to be backtracked, wherein the data to be backtracked includes: multiple data objects, each of the multiple data objects includes: multiple data units, the first dependency relationship is used to describe the dependency relationship between data units of different data objects, the second dependency relationship is used to describe the reference time dependency relationship between different data objects, and the target topology relationship is used to determine multiple task instances to be calculated; a collaboration module for coordinating a corresponding computing environment for each task instance in the multiple task instances to be calculated according to a preset concurrency; a computing module for calling a computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a computing result, wherein the computing result is used to adjust the processing progress of the data backtracking process.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data backtracking method proposed in the present disclosure.
[0009] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the data backtracking method proposed in the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When a processor executes the computer program, the data backtracking method proposed in the present disclosure is executed.
[0011] In the embodiments of the present disclosure, the task instances required for calculation of the data to be backtracked and the topological relationship between the task instances are obtained through the field-level lineage information of the data to be backtracked and the benchmark time dependency model, and then the corresponding computing environment is coordinated for each of the multiple task instances to be calculated according to the preset concurrency, and finally the computing engine corresponding to the computing environment is called to perform task calculation to obtain the calculation result, thereby achieving the purpose of dynamically adjusting the computing environment of each task instance corresponding to the data to be backtracked to perform calculations to complete data backtracking, and realizing the technical effect of improving the efficiency of data backtracking calculations and reducing the resource overhead of data backtracking calculations, thereby solving the technical problems of low computing efficiency and high computing resource overhead of data backtracking methods in related technologies.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0014] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data backtracking method according to an embodiment of the present disclosure;
[0015] Figure 2 is a flow chart of a data backtracking method according to an embodiment of the present disclosure;
[0016] Figure 3 is a flowchart of an optional example topology view construction according to an embodiment of the present disclosure;
[0017] Figure 4 This is a structural diagram of the topological relationship between different data tables at an optional field level according to an embodiment of the present disclosure;
[0018] Figure 5 This is a schematic diagram of an optional calculation of the time partition for backtracking the vertex data of the field to be backtracked in the full link according to an embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram of an optional task instance topology view according to an embodiment of the present disclosure;
[0020] Figure 7 It is a structural block diagram of a data backtracking device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] In related technologies, a method of batch data backtracking across time intervals from a single data source can be used to process data quality. This involves optimizing the single data source itself, sequentially identifying the downstream data sources affected by the single data source, and batch processing other data across data intervals to improve data quality. However, the computational efficiency of the method of batch data backtracking across time intervals using a single data source is low. When an enterprise's data link has a topological depth of 5 to 10 layers or even more, it cannot meet the processing requirements for complete repair of data across the entire link. In response to this, a method of indiscriminate parallel data backtracking for multiple data sources is also provided in related technologies. This method simultaneously focuses on multiple data sources in an associated system and uses indiscriminate parallel backtracking to speed up the overall data backtracking of the data link, thereby completing data backtracking. Although this method improves the efficiency of data processing, it has a technical problem of high computing resource overhead due to the inability to accurately identify the affected data range.
[0024] According to an embodiment of the present disclosure, a data backtracking method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0025] The method embodiments provided in the embodiments of the present disclosure can be executed in a mobile terminal, a computer terminal or a similar electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein. Figure 1The present invention shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data backtracking method.
[0026] like Figure 1 As shown, the computer terminal 100 includes a computing unit 101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 102 or a computer program loaded from a storage unit 108 into a random access memory (RAM) 103. Various programs and data required for the operation of the computer terminal 100 can also be stored in the RAM 103. The computing unit 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.
[0027] Multiple components in the computer terminal 100 are connected to the I / O interface 105, including an input unit 106, such as a keyboard, a mouse, etc.; an output unit 107, such as various types of displays, speakers, etc.; a storage unit 108, such as a magnetic disk, an optical disk, etc.; and a communication unit 109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 109 allows the computer terminal 100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0028] The computing unit 101 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 101 performs the data backtracking method described herein. For example, in some embodiments, the data backtracking method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer terminal 100 via the ROM 102 and / or the communication unit 109. When the computer program is loaded into the RAM 103 and executed by the computing unit 101, one or more steps of the data backtracking method described herein can be performed. Alternatively, in other embodiments, the computing unit 101 can be configured to perform the data backtracking method by any other appropriate means (e.g., by means of firmware).
[0029] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0030] It should be noted that, in some optional embodiments, the above Figure 1 The electronic device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the electronic device described above.
[0031] Under the above operating environment, the present disclosure provides the following Figure 2 The data backtracking method shown in Figure 1 The computer terminal or similar electronic device shown is used for execution. Figure 2 This is a flow chart of a data backtracking method provided according to an embodiment of the present disclosure. Figure 2 As shown, the method may include the following steps:
[0032] Step S21: generating a target topology relationship based on a first dependency relationship and a second dependency relationship of the data to be backtracked, wherein the data to be backtracked includes: a plurality of data objects, each of the plurality of data objects includes: a plurality of data units, the first dependency relationship is used to describe the dependency relationship between data units of different data objects, the second dependency relationship is used to describe the reference time dependency relationship between different data objects, and the target topology relationship is used to determine a plurality of task instances to be calculated;
[0033] The aforementioned backtracking data may be data that needs to be searched forward according to optimal conditions to reach a target. The aforementioned data objects may be complex information representations understood by software, including external entities, things, roles, organizational units, etc. The aforementioned data units may be the basic units of network information transmission. The aforementioned topological relationships may be relationships between various spatial data that satisfy the principles of topological geometry.
[0034] Step S22: Coordinate a corresponding computing environment for each of the multiple task instances to be calculated according to a preset concurrency;
[0035] The concurrency can be the number of users that can interact with the server at a given time. The computing environment can be built on an open network infrastructure, integrating and utilizing distributed autonomous resources to provide a harmonious, secure, and transparent integrated service environment for end users or application systems.
[0036] For example, a locking mechanism of a relational data management system is used to coordinate different computing engines for the data to be calculated. Specifically, in a relational data management system, the locking mechanism of the relational data management system can coordinate different computing engines for the data to be calculated according to the level of concurrency. For example, when the concurrency is relatively low, the first storage engine can be used, and deadlock will not occur during the calculation process when using this engine. Deadlock can be a blocking phenomenon caused by two or more processes competing for resources or communicating with each other during execution.
[0037] Step S23: calling the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process.
[0038] The computing engine mentioned above may be a computer program specifically for processing data.
[0039] According to the above steps S21 to S23 of the present disclosure, the time partitions to which all the vertex data of the fields to be backtracked should be backtracked are inferred through the field-level lineage information of the data to be backtracked and the benchmark time dependency model, and then the task instances required to be calculated for the backtracked data and the topological relationships between the task instances are obtained. Then, the corresponding computing environment is coordinated for each of the multiple task instances to be calculated according to the preset concurrency, and finally the computing engine corresponding to the computing environment is called to perform task calculation on the task instance corresponding to the computing environment to obtain the calculation result. In this way, the purpose of completing data backtracking is achieved by dynamically adjusting the computing environment corresponding to each task instance through the topological relationship between the task instances required to be calculated for the backtracked data, and the technical effect of improving the data backtracking computing efficiency and reducing the data backtracking computing resource overhead is achieved, thereby solving the technical problems of low computing efficiency and high computing resource overhead of the data backtracking method in the related technology.
[0040] In one optional embodiment, vertices are constructed using metadata of data units of different data objects, and edges are constructed using dependencies between data units of different data objects to obtain a first dependency relationship. A second dependency relationship is then obtained using a base time offset and a dependency time step span between different data objects. The target topology relationship is then generated using the first and second dependency relationships. The metadata is used to describe the structure, semantics, purpose, and usage of the information.
[0041] For example, a credit system processes overdue information for different users. To view overdue information for different users at different time points, the data backtracking mechanism of the credit system can be a full-link field-level data backtracking mechanism based on lineage relationships. First, the initial lineage is generated through the data lineage model. Then, the benchmark time dependency model of the data table is used to infer the time partition to which the data of all the vertex data of the field to be backtracked in the entire link should be backtracked. Secondly, the analysis module traverses all the vertices of the field to be backtracked, marking the out-degrees of the edges involved in all the vertices of the field to be backtracked, and marking the data tables pointed to by the fields. Finally, based on the number of benchmark times, the out-degrees are converted into an equal number of vertices. Then, based on the association between the data table and the task, the target topological relationship is generated. In addition, the full-link field-level data backtracking mechanism based on lineage relationships is completed by the analysis module, the coordination module, and the execution module. Among them, the operation of the analysis module is based on the data lineage model.
[0042] The data backtracking method of the above embodiment is further introduced below.
[0043] As an optional implementation, in step S21, generating a target topological relationship based on the first dependency relationship and the second dependency relationship of the data to be backtracked may include the following method steps:
[0044] S211: traverse each data unit associated with the first dependency relationship in the backtracking data to obtain a first traversal result;
[0045] S212: Mark the out-degrees of the edges involved in the traversed data units based on the first traversal result, and mark the data objects pointed to by the traversed data units to obtain a marking result;
[0046] S213, determining the amount of reference time using the second dependency relationship;
[0047] S214: Convert the out-degree into vertices with equivalent data based on the number of benchmark times and the marking result to obtain a conversion result.
[0048] S215 . Generate a target topological relationship based on the transformation result and a preset association relationship, wherein the preset association relationship is used to describe the association relationship between the multiple data objects and the multiple task instances.
[0049] The traversal can be performed by sequentially visiting each node in the tree (or graph) along a search route. The operations performed on the visited nodes depend on the specific application problem and can include checking or updating the node's value. The out-degree can be the number of outgoing edges from a vertex in a directed graph. The benchmark time can be a custom time point based on project requirements.
[0050] Figure 3 This is a flowchart of an optional example topology view constructed according to an embodiment of the present disclosure, such as Figure 3 As shown, the initial lineage is first generated through the data lineage model. Then, the benchmark time dependency model of the data table is used to infer the time partitions to which all vertex data of the fields to be backtracked in the entire link should be backtracked. Secondly, the analysis module traverses all vertices of the fields to be backtracked, marks the out-degrees of the edges involved in the vertices of all the fields to be backtracked, and marks the data tables pointed to by the fields. Finally, according to the number of benchmark times, the out-degrees are converted into the same number of vertices. Then, based on the association between the data table and the task, an instance topology view can be constructed.
[0051] The data lineage model can be a collection of metadata used to describe data dependencies. The data lineage model can include a table level and a field level. The table level can describe the dependencies between upstream and downstream data tables using a directed acyclic graph (DAG) structure. Each DAG vertex represents a specific table, and each DAG edge describes the dependencies between data tables. The field level can be a more granular level than the table level. Vertices in the DAG can describe field metadata, while edges describe both field metadata and field dependencies. Field vertex information also contains pointers to associated table information, enabling traceability of the relationship with the data table. The reference time dependency model can be expressed as a two-tuple {offset, step}, where offset represents the reference time offset and step represents the step size of the dependency time. For example, B relies on A{0,1} means that B relies on A's data for the current day's time partition, while B relies on A{-2,2} means that B relies on all of A's data for the past two days' time partitions.
[0052] Figure 4 This is a structural diagram of the topological relationship between different data tables at an optional field level according to an embodiment of the present disclosure, such as Figure 4As shown, data table A can be represented by Table A. Table A includes fields A1, A2, A3, and A4. Table B includes fields B1, B2, and B3. Field A1 can be associated with both field B1 and field C1. In addition, Table B relies on the time-partitioned data of Table A for the past day.
[0053] Figure 5 It is a schematic diagram of an optional calculation of the time partitions to which the vertex data of the field to be traced back in the entire link should be traced back according to an embodiment of the present disclosure. Specifically, in the process of calculating the time partitions to which all the vertex data of the field to be traced back in the entire link should be traced back through the reference time dependency model of the data table, the graph traversal is first started from the starting vertex (i.e., the original data to be traced back), and then the reference time input parameter information carried by the starting vertex is used, wherein the reference time input parameter information can be a time point customized by the relevant business system, and finally, the time partitions to which the vertex data of the field to be traced back in the entire link should be traced back are calculated in sequence during the traversal process. For example, if Table D should trace back the partition data with the reference time input parameter information of 20220102, and Table E depends on all the data of the time partitions of Table D in the past two days, it can be calculated that Table E should trace back the partition data with the reference time of 20220103 and the reference time of 20220104.
[0054] It's important to note that after the above traversal is complete, the analysis module can filter out tables that have no indirect association with the backtracking field. For affected tables, the out-degree can be converted to an equal number of vertices based on the number of benchmark times. Based on the association between the data table and the task instance, a task instance topology view can be constructed.
[0055] Figure 6 is a schematic diagram of an optional task instance topology view according to an embodiment of the present disclosure, such as Figure 6 As shown, Tables A, B, C, D, and E are affected. Therefore, after the traversal is complete, the out-degrees can be converted to an equal number of vertices based on the number of benchmark times. Then, based on the associations between Tables A, B, C, D, and E and the task instances, a task instance topology view can be constructed. Optionally, since Table F has no indirect association with the backtracking field, the analysis module can filter out Table F after the traversal is complete.
[0056] As an optional implementation, in step S22, coordinating a corresponding computing environment for each of the multiple task instances to be calculated according to the preset concurrency may include the following method steps:
[0057] S221, traverse the target topological relationship to obtain a second traversal result;
[0058] S222: Fill the element information contained in the execution queue according to the second traversal result to obtain a filling result, wherein the execution queue is used to perform timing control on the target topological relationship;
[0059] S223. Coordinate a corresponding computing environment for each task instance in the plurality of task instances to be calculated according to the preset concurrency and the filling result.
[0060] The above-mentioned element information may be the name of the data table to be backtracked, shard information, computing task name, benchmark time, required resource amount, etc.
[0061] Still taking a credit system that processes overdue information of different users as an example, the coordination module in the data backtracking mechanism of the credit system can first perform a breadth-first traversal of the task instance topological relationship by calling the design pattern provided by the analysis module, and add the traversed vertex information to the execution queue until all traversals are completed to obtain the filling result. Then, the coordination module consumes the elements of the execution queue in turn according to the preset concurrency. Finally, the coordination module calls the execution module to open up a computing node with the same concurrency. At the same time, the coordination module will establish a connection with the execution module. When it is detected that the computing node of the execution module is idle, the computing node will be removed from the execution queue and submitted to the idle computing node. In this way, the coordination of the computing environment of each task instance is completed to reduce the overhead of data backtracking on computing resources.
[0062] As an optional implementation, in step S22, coordinating a corresponding computing environment for each of the multiple task instances to be calculated according to the preset concurrency may further include:
[0063] Configure a fault-tolerant mode, where the fault-tolerant mode is used to determine a fault-tolerant handling method for calculation anomalies in response to calculation anomalies in some task instances during task calculation for multiple task instances to be calculated.
[0064] The above-mentioned configuration fault tolerance mode can be one of the three modes: fast failure mode, automatic failure recovery mode, and fail-safe mode. The fast failure mode can be a mode that requires that the backtracking of the entire link must be successful. If a single task exception is encountered, the full link backtracking is terminated, the backtracking result is set to failure, and the intermediate data with side effects generated during the task backtracking process is deleted, and the task is restored to an unexecuted state. The automatic failure recovery mode can be a mode that considers the local failure of individual tasks in the link due to abnormal runtime environment. If a single task exception is encountered, other non-strongly dependent nodes in the queue are submitted to the execution module first. When other nodes are completed, the failed node is submitted. The fail-safe mode can be a mode that hopes that more data in the link will be backtracked and recovered. This mode will ignore the failed task node and its downstream, and give priority to tasks that run successfully.
[0065] As an optional implementation, in step S23, calling a computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment, and obtaining the calculation result may include the following method steps:
[0066] S231. Obtain target parameter information, where the target parameter information includes: preset concurrency and fault tolerance mode;
[0067] S232. Based on the target parameter information, call the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result.
[0068] The computing engine mentioned above may be a Spark computing engine, a Flink computing engine, a MapReduce computing engine, etc., wherein the execution module may decide which computing engine to use based on the submitted target parameter information.
[0069] Still taking a credit system processing overdue information of different users as an example, when the execution module in the credit system's data backtracking mechanism receives a request command from the coordination module, it first starts the corresponding computing node, and then distributes the task instance to the corresponding computing environment for task calculation based on the submitted target parameter information to obtain the calculation results. In this way, the computing environment corresponding to each task instance is dynamically adjusted to improve the accuracy of the calculation results.
[0070] As an optional implementation, in step S23, calling a computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment, and obtaining the calculation result may further include the following method steps:
[0071] S233, detecting whether the task status of the task instance corresponding to the computing environment is normal, and whether the calculation result meets the preset conditions;
[0072] S234: In response to the task status of the task instance corresponding to the computing environment being normal and the computing result meeting the preset conditions, the computing result is returned.
[0073] The above-mentioned preset conditions can be conditions set by the quality inspection system of the third party. The above-mentioned calculation results can be the path and execution status of the task instance output, where the execution status can be either success or failure.
[0074] Still taking a credit system processing overdue information of different users as an example, after the calculation engine completes the result calculation, the execution module in the credit system's data backtracking mechanism can first confirm whether the task status is normal, and then verify whether the data of the result path meets the expected conditions. When the running result meets the expected conditions, it is marked in the data storage system, where the mark is usually an agreed mark. Finally, the path and execution status of the task instance are returned to the coordination module, thereby completing the return of the calculation result.
[0075] As an optional implementation, in step S232, based on the target parameter information, calling the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment, and obtaining the calculation result may also include:
[0076] During the task calculation process for multiple task instances to be calculated, in response to calculation exceptions occurring in some task instances, exception information is reported, wherein the exception information is used to determine whether to trigger a fault-tolerant mode.
[0077] Still taking a credit system processing overdue information of different users as an example, when the execution module in the credit system's data backtracking mechanism performs task calculations on multiple task instances to be calculated, when some task instances have calculation anomalies, the coordination module in the data backtracking mechanism will determine whether to trigger the corresponding fault-tolerant mode, thereby improving the computational efficiency of the data backtracking method.
[0078] It should be noted that the causes of computing anomalies may be computing resource congestion causing the task to be stuck for a long time, excessive storage cluster load, etc., where the stuck state may be the freezing of the computing program used for the computing task.
[0079] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0081] The present disclosure also provides a data backtracking device for implementing the above-mentioned embodiments and preferred embodiments, which will not be repeated hereafter. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0082] Figure 7 is a structural block diagram of a data backtracking device according to one embodiment of the present disclosure, such as Figure 7 As shown, the data backtracking device 700 includes: an analysis module 701 , a collaboration module 702 , and a calculation module 703 .
[0083] The analysis module 701 is used to generate a target topological relationship based on the first dependency and the second dependency of the data to be backtracked, wherein the data to be backtracked includes: multiple data objects, each of the multiple data objects includes: multiple data units, the first dependency is used to describe the dependency between data units of different data objects, the second dependency is used to describe the reference time dependency between different data objects, and the target topological relationship is used to determine multiple task instances to be calculated; the collaboration module 702 is used to coordinate the corresponding computing environment for each task instance in the multiple task instances to be calculated according to the preset concurrency; the computing module 703 is used to call the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process.
[0084] Optionally, the analysis module 701 is further configured to: construct vertices using metadata of data units of different data objects, and construct edges using dependency relationships between data units of different data objects, to obtain a first dependency relationship.
[0085] Optionally, the analysis module 701 is further configured to obtain a second dependency relationship by utilizing a reference time offset and a step span of a dependency time between different data objects.
[0086] Optionally, the analysis module 701 is also used to: traverse each data unit associated with the first dependency relationship in the backtracked data to obtain a first traversal result, mark the out-degree of the edge involved in the traversed data unit based on the first traversal result, and mark the data object pointed to by the traversed data unit to obtain a marking result, use the second dependency relationship to determine the number of reference times, and convert the out-degree into a vertex of equivalent data through the number of reference times and the marking result to obtain a conversion result, thereby generating a target topological relationship based on the conversion result and the preset association relationship.
[0087] Optionally, the collaborative module 701 is also used to: traverse the target topological relationship to obtain a second traversal result, fill in the element information contained in the execution queue based on the second traversal result, and obtain a filling result, thereby coordinating the corresponding computing environment for each task instance in the multiple task instances to be calculated according to the preset concurrency and filling result.
[0088] Optionally, the collaborative module 701 is further used to configure a fault-tolerant mode, wherein the fault-tolerant mode is used to determine a fault-tolerant processing method for calculation anomalies in response to calculation anomalies in some task instances during task calculation for multiple task instances to be calculated.
[0089] Optionally, the computing module 701 is also used to: obtain target parameter information, wherein the target parameter information includes: preset concurrency and fault tolerance mode; based on the target parameter information, call the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain the calculation result.
[0090] Optionally, the computing module 701 is also used to detect whether the task status of the task instance corresponding to the computing environment is normal and whether the calculation result meets the preset conditions. When the task status of the task instance corresponding to the computing environment is normal and the calculation result meets the preset conditions, the calculation result is returned.
[0091] Optionally, the calculation module 701 is further used to: during the task calculation process for multiple task instances to be calculated, when calculation exceptions occur in some task instances, report exception information, wherein the exception information is used to determine whether to trigger the fault tolerance mode.
[0092] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0093] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including a memory and at least one processor, wherein the memory stores computer instructions, and the processor is configured to execute the computer instructions to perform the steps in the above method embodiment.
[0094] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0095] Optionally, in the present disclosure, the processor may be configured to execute the following steps through a computer program:
[0096] Step S1: generating a target topology relationship based on a first dependency relationship and a second dependency relationship of the data to be backtracked, wherein the data to be backtracked includes: a plurality of data objects, each of the plurality of data objects includes: a plurality of data units, the first dependency relationship is used to describe the dependency relationship between data units of different data objects, the second dependency relationship is used to describe the reference time dependency relationship between different data objects, and the target topology relationship is used to determine a plurality of task instances to be calculated;
[0097] Step S2, coordinating a corresponding computing environment for each of the multiple task instances to be calculated according to a preset concurrency;
[0098] Step S3: calling the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process.
[0099] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0100] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are configured to execute the steps of the above method embodiment when running.
[0101] Optionally, in this embodiment, the non-transitory computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0102] Step S1: generating a target topology relationship based on a first dependency relationship and a second dependency relationship of the data to be backtracked, wherein the data to be backtracked includes: a plurality of data objects, each of the plurality of data objects includes: a plurality of data units, the first dependency relationship is used to describe the dependency relationship between data units of different data objects, the second dependency relationship is used to describe the reference time dependency relationship between different data objects, and the target topology relationship is used to determine a plurality of task instances to be calculated;
[0103] Step S2, coordinating a corresponding computing environment for each of the multiple task instances to be calculated according to a preset concurrency;
[0104] Step S3: calling the computing engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment to obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process.
[0105] Alternatively, in this embodiment, the non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any suitable combination of the above. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0106] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product. The program code for implementing the data backtracking method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0107] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.
[0108] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0109] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0110] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0111] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0113] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.
Claims
1. A data backtracking method, comprising: Generate a target topology relationship based on a first dependency relationship and a second dependency relationship of data to be backtracked, wherein the data to be backtracked includes: a plurality of data objects, each of the plurality of data objects includes: a plurality of data units, the first dependency relationship is used to describe a dependency relationship between data units of different data objects, the second dependency relationship is used to describe a reference time dependency relationship between different data objects, and the target topology relationship is used to determine a plurality of task instances to be calculated; Coordinate a corresponding computing environment for each of the plurality of task instances to be calculated according to a preset concurrency; Invoking a computing engine corresponding to the computing environment to perform task computing on the task instance corresponding to the computing environment to obtain a computing result, wherein the computing result is used to adjust the processing progress of the data backtracking process; Among them, based on the first dependency and the second dependency of the data to be backtracked, generating the target topological relationship includes: traversing each data unit associated with the first dependency in the data to be backtracked to obtain a first traversal result; marking the out-degree of the edge involved in the traversed data unit based on the first traversal result, and marking the data object pointed to by the traversed data unit to obtain a marking result; using the second dependency to determine the number of reference times; through the number of reference times and the marking result, converting the out-degree into vertices of equivalent data to obtain a conversion result; generating the target topological relationship based on the conversion result and a preset association relationship, wherein the preset association relationship is used to describe the association relationship between the multiple data objects and the multiple task instances.
2. The data backtracking method according to claim 1, wherein: The data backtracking method further includes: Vertices are constructed using meta-information of data units of different data objects, and edges are constructed using dependency relationships between data units of different data objects to obtain the first dependency relationship.
3. The data backtracking method according to claim 1, wherein: The data backtracking method further includes: The second dependency relationship is obtained by using the reference time offset and the step span of the dependency time between different data objects.
4. The data backtracking method according to claim 1, wherein: Coordinating a corresponding computing environment for each of the multiple task instances to be calculated according to the preset concurrency includes: Traversing the target topological relationship to obtain a second traversal result; Filling the element information contained in the execution queue according to the second traversal result to obtain a filling result, wherein the execution queue is used to perform timing control on the target topological relationship; A corresponding computing environment is coordinated for each task instance in the plurality of task instances to be calculated according to the preset concurrency and the filling result.
5. The data backtracking method according to claim 4, wherein: The data backtracking method further includes: A fault-tolerant mode is configured, wherein the fault-tolerant mode is used to determine a fault-tolerant processing method for calculation anomalies in response to calculation anomalies occurring in some task instances during the task calculation process for the multiple task instances to be calculated.
6. The data backtracking method according to claim 5, wherein: Calling a computing engine corresponding to the computing environment to perform task computing on the task instance corresponding to the computing environment, and obtaining the computing result includes: Acquire target parameter information, wherein the target parameter information includes: the preset concurrency and the fault tolerance mode; Based on the target parameter information, a computing engine corresponding to the computing environment is called to perform task calculation on the task instance corresponding to the computing environment to obtain the calculation result.
7. The data backtracking method according to claim 6, wherein: The data backtracking method further includes: Detecting whether the task status of the task instance corresponding to the computing environment is normal and whether the computing result meets the preset conditions; In response to the task status of the task instance corresponding to the computing environment being normal and the computing result meeting the preset condition, the computing result is returned.
8. The data backtracking method according to claim 6, wherein: The data backtracking method further includes: During the task calculation process for the plurality of task instances to be calculated, in response to calculation anomalies occurring in some task instances, anomaly information is reported, wherein the anomaly information is used to determine whether to trigger the fault-tolerant mode.
9. A data backtracking device, comprising: an analysis module, configured to generate a target topology relationship based on a first dependency relationship and a second dependency relationship of data to be backtracked, wherein the data to be backtracked includes: a plurality of data objects, each of the plurality of data objects includes: a plurality of data units, the first dependency relationship is used to describe a dependency relationship between data units of different data objects, the second dependency relationship is used to describe a base time dependency relationship between different data objects, and the target topology relationship is used to determine a plurality of task instances to be calculated; A coordination module, configured to coordinate a corresponding computing environment for each of the plurality of task instances to be calculated according to a preset concurrency; A calculation module, configured to call a calculation engine corresponding to the computing environment to perform task calculation on the task instance corresponding to the computing environment, and obtain a calculation result, wherein the calculation result is used to adjust the processing progress of the data backtracking process; Among them, the analysis module is also used to: traverse each data unit associated with the first dependency relationship in the data to be backtracked to obtain a first traversal result; mark the out-degree of the edge involved in the traversed data unit based on the first traversal result, and mark the data object pointed to by the traversed data unit to obtain a marking result; use the second dependency relationship to determine the number of reference times; convert the out-degree into a vertex of equivalent data through the number of reference times and the marking result to obtain a conversion result; generate the target topological relationship based on the conversion result and a preset association relationship, wherein the preset association relationship is used to describe the association relationship between the multiple data objects and the multiple task instances.
10. The device according to claim 9, wherein The analysis module is also used for: Vertices are constructed using meta-information of data units of different data objects, and edges are constructed using dependency relationships between data units of different data objects to obtain the first dependency relationship.
11. The device according to claim 9, wherein The analysis module is also used for: The second dependency relationship is obtained by using the reference time offset and the step span of the dependency time between different data objects.
12. The device according to claim 9, wherein The collaboration module is also used to: Traversing the target topological relationship to obtain a second traversal result; Filling the element information contained in the execution queue according to the second traversal result to obtain a filling result, wherein the execution queue is used to perform timing control on the target topological relationship; A corresponding computing environment is coordinated for each task instance in the plurality of task instances to be calculated according to the preset concurrency and the filling result.
13. The device according to claim 12, wherein The collaboration module is also used to: A fault-tolerant mode is configured, wherein the fault-tolerant mode is used to determine a fault-tolerant processing method for calculation anomalies in response to calculation anomalies occurring in some task instances during the task calculation process for the multiple task instances to be calculated.
14. The device according to claim 13, wherein The calculation module is also used for: Acquire target parameter information, wherein the target parameter information includes: the preset concurrency and the fault tolerance mode; Based on the target parameter information, a computing engine corresponding to the computing environment is called to perform task calculation on the task instance corresponding to the computing environment to obtain the calculation result.
15. The device according to claim 14, wherein The calculation module is also used for: Detecting whether the task status of the task instance corresponding to the computing environment is normal and whether the computing result meets the preset conditions; In response to the task status of the task instance corresponding to the computing environment being normal and the computing result meeting the preset condition, the computing result is returned.
16. The device according to claim 14, wherein The calculation module is also used for: During the task calculation process for the plurality of task instances to be calculated, in response to calculation anomalies occurring in some task instances, anomaly information is reported, wherein the anomaly information is used to determine whether to trigger the fault-tolerant mode.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the data tracing method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the data tracing method according to any one of claims 1 to 8.
19. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data backtracking method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Flow processing operation scheduling method and system for dynamically adjusting task allocation
CN107580023A
Data topology generation method and device, electronic equipment and storage medium
CN113656407A