Computer-implemented method for executing process having plurality of work steps

By defining and sorting work steps, and using computing units to execute in parallel, the complexity and type safety problems of data preprocessing in complex software systems are solved, and code clarity and efficient data processing are achieved.

CN120122922APending Publication Date: 2025-06-10ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411786993.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-12-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle data preprocessing in complex software systems, especially when database data comes from multiple sources and requires nonlinear interactions, resulting in difficult code complexity, type safety, and difficulty in testing and inspection.

Method used

A method is proposed to simplify software development and processing complex data by defining a process with multiple working steps, check whether the input can be generated by the process, sort the working steps, instantiate and execute the working steps, and use the computing unit to execute the working steps in parallel to simplify software development and process complex data.

Benefits of technology

It simplifies data preprocessing in complex software systems, improves code clarity and type safety, reduces error risks and test difficulties, and is suitable for processing multi-source database data and other complex data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122922A_ABST
    Figure CN120122922A_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for executing a process having a plurality of work steps using a plurality of computing units, where the method comprises the following steps:-defining (S10) a process having a plurality of work steps, where each work step requires one or more defined inputs, and wherein each working step produces one or more defined outputs, wherein the outputs of some working steps are used as inputs for the other working steps; checking (S12) whether all required inputs can be produced by the process; -if the check indicates that all required inputs can be produced by the process, sorting (S14) the working steps corresponding to dependencies derived by the inputs and the outputs; -instantiating (S16) the sorted work steps; and-performing (S18), by the one or more computing units, the instantiated work steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the processing of data, in particular to the preprocessing of database data. Background Art

[0002] Modern software products are usually complex and process different types of data. Execution usually requires many different components, such as libraries. However, libraries are not always available in a suitable programming language. In other cases, libraries are not compatible with each other for other reasons.

[0003] The above components must interact with each other in a non-linear manner. Whether the software is generated by a program or by object-oriented programming, the software usually contains code that is difficult to maintain. If other functions are implemented in this situation, the obscurity of the software will increase with its increasing complexity.

[0004] The interfaces of software components become increasingly opaque, difficult to find, difficult to use, or difficult to change in terms of increasingly complex code. In addition, ensuring type safety between components is a tedious and often overlooked task. Both aspects lead to a high risk of unexpected behavior and errors.

[0005] When a project reaches a certain level of maturity, the next step is usually to expand the application to many other products and other application scenarios. Due to the above complexity and the difficulty of maintaining the involved pipeline, this task is often particularly laborious.

[0006] For complex and often intertwined functions, it becomes increasingly difficult to correctly test and inspect the software. As a result, the testing is no longer carried out in correspondence with the actual requirements.

[0007] The above problems particularly occur in research projects with growing and changing requirements, especially projects involving artificial intelligence (AI), data science, etc. Summary of the Invention

[0008] Therefore, the task on which the present invention is based is to propose a method by which software for complex processes can be constructed simply and quickly.

[0009] This task is solved by the subject matter of the independent claims.

[0010] According to a first aspect of the present invention, this task is solved by a computer-implemented method for executing a process with multiple working steps using one or more computing units, wherein the method comprises the following steps:

[0011] - Defining a process with multiple working steps,

[0012] Each of these work steps requires one or more defined inputs, and each of these work steps produces one or more defined outputs,

[0013] wherein the outputs of some of the work steps are used as inputs for other work steps;

[0014] - Check whether all required inputs can be produced by the process;

[0015] - If the check shows that all required inputs can be produced by the process, then sort the work steps according to the dependencies derived from the inputs and outputs;

[0016] - Instantiate the sorted work steps; and

[0017] - Execute the instantiated work steps by one or more computing units.

[0018] The process to be executed includes multiple work steps. Each work step here has tasks that contribute to the overall goal of the process. These tasks can be different. For example, one task can be to determine new values from pre-given data, in particular by calculation. Another task can include converting the same input data into a different format.

[0019] These tasks can be diverse and essentially include arbitrary code to perform these tasks. Importantly, each work step is provided with defined inputs, and each work step produces defined outputs. The outputs must be used as inputs for other work steps or as outputs of the entire process.

[0020] In the example already mentioned, the output of a work step can include an array with the generated values. Thus, the output of the task of the last work step or the entire process can be, for example, to create a table with the input values and the generated values.

[0021] The foregoing examples are intended to illustrate the basic idea of the present invention in a simple manner. In practice, the processes are much more complex, especially when they include preprocessing training data to train machine learning algorithms.

[0022] The process can also include the following work steps, which include different tasks such that different programming languages are more or less suitable for performing these work steps. As a result, the work steps can be defined by code written in a suitable language respectively. As long as the inputs and outputs are clearly defined, different software packages can work together.

[0023] If a process with all work steps is defined and it is set for each work step what inputs it requires and what outputs it produces, then check the process.

[0024] During the check, it is determined whether suitable inputs can be generated for each working step by the process. This can also include providing the input by reading or receiving input data from a storage medium.

[0025] An exemplary method can include, for example, three working steps f, g, and h. The first working step f reads data x from a storage medium, in particular from a hard disk memory or a working memory. The output is f(x). The second working step g requires the output of the first working step f as input and produces the output g(f(x)). The third working step h requires the outputs of the other two working steps as input and thereby produces the output h(g(f(x)), f(x)).

[0026] During the check, it is now tested whether suitable outputs can be used as inputs for each working step. For example, if one of the working steps from the above example requires an input called k(x), the check will show that not all inputs can be generated by the process, since only f, g, and h are produced. In this case, an output can be generated for the developer that indicates that the input required for the working step is not defined. The developer must then define further working steps to produce the output k(x) from the data x.

[0027] In the next step, the working steps are sorted so that the working steps can be executed logically. Here, the working steps that produce further inputs used as outputs must be executed before the working steps that use the outputs of other working steps as inputs.

[0028] Returning to the above example, the working steps f, g, and h can, for example, be linearly arranged one after the other in that order. The working step f only requires the input data x. The working step g requires the output of the working step f and must therefore be executed after the working step f. The working step h requires the outputs of the working steps f and g and must therefore be executed after these working steps.

[0029] This example may seem trivial. However, it illustrates the method according to the invention. If the process includes more working steps, for example dozens or even hundreds, the dependencies can become unclear and difficult to manage.

[0030] If the order of the working steps is determined, the working steps can be instantiated and then executed.

[0031] "Instantiating" is a term in software development that involves the process of creating class instances. In object-oriented programming languages, a class represents a blueprint or template for an object, and an instance is a concrete implementation of that class.

[0032] Instantiation involves creating a concrete object (instance) based on a class - based definition. This process allocates memory space for the object and initializes the object according to the attributes and methods defined in the class. Instantiation is a fundamental step in object - oriented programming and enables the use of classes as modules for structured software.

[0033] If multiple programs are used for this process, these programs can be combined, for example, as sub - programs or grouped in a program library. The programs can be executed by a superior script corresponding to a determined order, if necessary, in different runtime environments, in different programming languages, and / or on different hardware components.

[0034] In particular, the present invention itself can be implemented as a program. The user pre - specifies the process to be executed by defining work steps with inputs and outputs. In this way, software developers are supported when creating complex processes that include multiple work steps. The creation is simplified and thus clearer. The present invention thus solves its task.

[0035] In one embodiment, the process is pre - processing database data, where the database data originates from more than one source.

[0036] Database data particularly includes data from multiple databases that should be processed together. This can include, for example, merging data from different databases and / or cascading data sets.

[0037] It may be the case that transformations must be performed first in order to process data from different sources. For example, a table must be transposed or inverted. Other examples include the conversion, encoding, or decoding of data.

[0038] Due to different data formats and database structures, the process of accessing data from multiple databases can be particularly complex. This embodiment can advantageously structure these processes, thereby making them simpler and clearer. Thus, for example, individual work steps can be defined for each database, which first coordinate the data, i.e., convert the data into the same format, and then merge or process the data together.

[0039] In one embodiment, the process is defined in a configuration file.

[0040] The process and all its work steps can be stored in a configuration file. Thereby, regardless of its complexity, the process can be configured centrally and clearly. In particular, the expected inputs and possible outputs of the individual work steps can also be stored in the configuration file.

[0041] In one embodiment, environment variables are used to execute work steps.

[0042] An environment variable is a variable that can be input by a user when executing a work step. Thereby, a process can react to changes that a user wants to make when a specific event occurs. For example, the event can be an error message caused by one of the work steps or an intermediate result defined by the work step.

[0043] Thereby, the user obtains the possibility to react to the event without having to adapt a configuration file and recompile the program to execute the work step. This makes the proposed method more flexible.

[0044] In one embodiment, sorting the work steps includes topological sorting.

[0045] "Topological sorting" is a concept from graph theory that is applied in computer science and especially in software development and particularly within the scope of this embodiment. This sorting enables a linear arrangement of nodes in a directed acyclic graph (DAG), where the edges represent directed connections between the nodes.

[0046] The main condition for topological sorting is that for each directed edge in the DAG from node A to node B, it should be ensured that A is located before B in the order. In other words, this sorting respects the direction of the edges and ensures that every dependency in the DAG is taken into account.

[0047] Topological sorting can have various applications, especially in the fields of building and task scheduling in software development. For example, when tasks are interdependent, topological sorting can be used to determine the order of these tasks to ensure correct execution. For example, an algorithm for topological sorting can be implemented by means of depth-first search (DFS).

[0048] In one embodiment, sorting the work steps includes parallelizing the work steps, where the work steps are executed simultaneously if the dependencies of the inputs and outputs of the respective work steps allow.

[0049] This embodiment can especially be executed together with topological sorting, where topological sorting is performed before the parallelization of the work steps. Topological sorting can especially be used to timely find work steps that cannot be completed, such as work steps that cannot be completed due to circular dependencies.

[0050] In this embodiment, work steps that are independent of each other are executed in parallel, which presupposes an architecture that performs parallel work of computing units. For example, two work steps f and g using the same input x can be executed separately and in parallel with each other by different computing units, especially by different cores of a multi-core processor or by different computers within a cluster.

[0051] Through parallelization, work steps can be carried out simultaneously, thereby reducing the total duration required for the process to be executed. The degree of parallelization can also be associated with the physical possibilities of the available computing units. For example, if the execution system includes four computing cores, at most four work steps can be arranged in parallel and then these four work steps are executed.

[0052] In one embodiment, the process is a process for processing health data, wherein the input of one or more work steps includes personal patient data, insurance data, treatment information, and / or diagnosis, and wherein the output of at least one work step includes a message to the employer, a prescription for a drug, and / or a bill.

[0053] Doctors, health insurance companies, and, if necessary, employers process various personal data of patients. Data protection should always be considered here, so that only the data actually required for the actual activities is provided to each person. For example, for an employer, what the patient's diagnosis is does not matter. All that matters is whether the patient can work.

[0054] In addition, there are no standards that doctors, insurance companies, or generally software providers have to follow. This results in a separate process that can be established in each clinic, which processes patient data in its own way.

[0055] Therefore, the process to be carried out according to the present invention in this embodiment may include work steps of combining patient data from different sources, reducing, sending, or converting patient data from different sources for the purpose of data protection.

[0056] In one embodiment, the process is a process for processing logistics data, wherein the input of one or more work steps includes the location of goods, information about the transportation of goods (especially in a warehouse), receipt information, or goods inventory, and wherein the output of at least one work step includes an inventory and / or a list with the location of goods, especially the location of goods within a warehouse.

[0057] Logistics data can be complex at multiple levels, making it difficult to process. For example, one possibility is to locate and take inventory of goods within a logistics complex. Modern logistics centers have a high throughput of goods, where the sender and the recipient can include different goods formats. For example, goods from Asia may have different descriptions compared to goods produced in Europe for the European market.

[0058] To process the data, the following process can be established, where each work step takes into account different countries of origin and regions of origin, types of goods, possible hazard warnings and / or certificates, and tax information and / or destination descriptions.

[0059] In one embodiment, the process is an accounting process, wherein the input of one or more work steps includes information about a transfer, in particular the transfer amount, the recipient, the sender, and / or the currency of the transfer, and wherein the output of at least one work step includes an account balance.

[0060] Institutions with a large number of booking processes, in particular trading companies, banks, or enterprises with a large number of employees, must be able to track and assign booking types. However, booking types can be logically associated with different metadata and have different formats. Therefore, the assignment and / or analysis of data can be very time-consuming.

[0061] The process to be executed can include a work step of sorting the data and further processing the data corresponding to its sorting.

[0062] In one embodiment, the process is a process for data transmission, in particular a process for the flow of data, wherein the input of one or more work steps includes the flow history of the user profile, the recommended user hardware, the used user hardware, the geographical location of the user, and / or the user's selection, and wherein the output of at least one work step includes a media stream and / or a recommendation for a media stream.

[0063] For example, the area requirements and the permission status can be adapted to the stream by using the process to be executed. In particular for media content that is fully or partially locally restricted, it can thus be adapted to the media content. Furthermore, creating consumption recommendations based on personal data can depend on the target group and / or the area where the user is located. In this embodiment, any framework conditions for stream transmission can be considered.

[0064] On the other hand, the invention relates to a computer program having program code for performing the method as described above when the computer program is executed on a computer.

[0065] On the other hand, the invention relates to a computer-readable data carrier having program code of a computer program for performing the method as described above when the computer program is executed on a computer.

[0066] On the other hand, the invention relates to a system for using one or more computing units to execute a process having a plurality of work steps, wherein the system is configured to perform the method as described above.

[0067] Thus, a method for using one or more computing units to execute a process having a plurality of work steps, a computer program having program code, a computer-readable data carrier, and a system having a plurality of computing units are generally described.

[0068] The described designs and extensions can be combined with each other arbitrarily.

[0069] Further possible designs, extensions and implementations of the present invention also include combinations of features of the present invention that are not explicitly mentioned, either previously or below, in the description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The drawings should provide a further understanding of the embodiments of the present invention. The drawings illustrate embodiments and are used in conjunction with the description to explain the principles and concepts of the present invention.

[0071] Other embodiments and many of the mentioned advantages can be derived with reference to the drawings. The elements shown in the drawings are not necessarily drawn to scale with respect to each other.

[0072] Figure 1 Schematically shows the flow of the proposed method according to one embodiment;

[0073] Figure 2 Shows five topologically sorted work steps; and

[0074] Figure 3 Shows a directed acyclic graph (DAG) of a method with seven steps.

[0075] In the figures of the drawings, unless otherwise stated, the same reference numerals denote the same or functionally identical elements, components or assemblies. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] Figure 1 Schematically shows the flow of a method for performing a process with multiple work steps.

[0077] The method starts in step S10 by defining the process. To this end, all work steps are described, and it is set which inputs are required for each work step and what outputs the work steps produce. The outputs of some work steps are used as inputs for other work steps, from which the dependencies of the work steps on each other result. For example, the definition of the process can be saved in a configuration file.

[0078] In step S12, it is checked whether the input of each work step is available through the process. If a work step requires an input that cannot be produced or obtained within the process, the check fails. In this case, the method must return to step S10 to adapt the definition of the process.

[0079] If the check is successful, the work steps can be sorted in the next step S14. The sorting brings the work steps into the order in which these work steps are to be executed. Since the work steps are interdependent, the work steps cannot be executed in an arbitrary order. In principle, the work steps whose output is used as input by other work steps should be executed before these other work steps.

[0080] In step S16, the work steps are instantiated. This can include, for example, reserving working memory, compiling code, or other steps necessary for generally executing the program code.

[0081] Then in step S18, the work steps are executed corresponding to the set order. Once all the work steps have been executed, the process ends like this.

[0082] Figure 2 and Figure 3 shows two ways how the work steps can be ordered.

[0083] Figure 2 Topological sorting is shown in. The process described here includes five work steps. In the first work step, data x is read in. In the second work step f, the read-in data x is processed into output f(x). The output does not have to be a mathematical function of the input data x.

[0084] For example, x can be a table with information, and f(x) describes the transposed table.

[0085] The next two work steps g and h equally use the output f(x). Work steps g and h produce outputs g(f) and h(f), where f(x) is used as the input respectively.

[0086] Due to the dependence on the same input, work steps g and h can be arranged side by side with each other. However, in the case of topological sorting, a one-dimensional order of the work steps is produced, so that g or h will be executed before the other respectively. In the example shown, g is arranged before h. However, this order can also be reversed.

[0087] In the final work step k, output k(g,h) is produced, for which the outputs g(f) and h(f) are used as inputs respectively. Due to this dependence, k must be executed after work steps g and h. The process is completed with the output of k(g,h). In one implementation, the result can also be stored or sent as a data packet for a specific task (for example, for controlling a machine or a facility).

[0088] Figure 3 A directed acyclic graph (DAG) showing another exemplary process is shown. The process includes seven work steps f, g, h, k, l, m, and p.

[0089] The first three work steps f, g, and h process the read-in data x into outputs f(x), g(x), and h(x) respectively. Therefore, these three work steps are not dependent on each other and can be executed in parallel. If the execution system has three or more computing units, work steps f, g, and h can be executed in parallel as shown here. Even if only two computing units are available, two computing steps can be executed in parallel, which also saves time.

[0090] Work step k requires the outputs of work steps f and g as inputs. Therefore, work step k can only be executed after work steps f and g are completed. Since work step k does not depend on the output of work step h, work step k does not have to wait for work step h to be completed.

[0091] Work step l requires the outputs f(x) and k(f,g) as inputs. Therefore, work step l can only be executed after work step k is completed. Thus, work step l directly depends on work steps f and k and also indirectly depends on work step g. Work step l also does not depend on work steps h or m, so work step l can be executed independently of work steps h or m.

[0092] Work step m uses the output h(x) as an input and thus must be downstream of work step h. Work step m does not require further inputs, so work step m can, for example, be executed directly after work step h is completed. Work step m can also be executed in parallel with work step k or l, provided that the execution system has a sufficient number of computing units.

[0093] The last work step p uses the outputs l(f,k), k(f,g), and m(h) and thus directly or indirectly depends on all the other work steps. Therefore, work step p must be executed at the end of the process after all the other work steps are completed.

Claims

1. A computer-implemented method for executing a process having a plurality of working steps using a plurality of computing units, wherein the method comprises the following steps: - defining (S10) a process having a plurality of working steps, wherein each work step requires one or more defined inputs, and wherein each work step produces one or more defined outputs, The output of some of these work steps is used as input to other work steps; - checking (S12) whether all required inputs can be generated by the process; - if the check shows that all required inputs can be generated by the process, the work steps are ordered according to the dependencies resulting from the inputs and the outputs (S14); - instantiating (S16) the sequenced work steps; and - executing ( S18 ) the instantiated working steps by one or more computing units, The sequencing of the work steps ( S14 ) includes parallelizing the work steps, wherein the work steps are executed simultaneously if dependencies of inputs of the respective work steps permit. 2 . The computer-implemented method of claim 1 , wherein the process is pre-processing database data, wherein the database data originates from more than one source.

3. The computer-implemented method of any one of the preceding claims, wherein the process is defined in a configuration file.

4. A computer-implemented method according to any one of the preceding claims, wherein the working step is performed (S18) using environment variables.

5. The computer-implemented method according to any of the preceding claims, wherein the ordering (S14) of the work steps comprises a topological ordering.

6. A computer-implemented method according to any of the preceding claims, wherein the process is a process for processing logistics data, wherein the input of one or more work steps comprises the location of goods, information about the transportation of goods, in particular in a warehouse, receipt information or goods inventory, and wherein the output of at least one work step comprises an inventory and / or a list with the location of goods, in particular the location of goods within a warehouse.

7. A computer program having a program code for performing the method according to any of the preceding claims when the computer program is executed on a computer.

8. A computer-readable data carrier having a program code of a computer program for carrying out the method according to any one of claims 1 to 6 when the computer program is executed on a computer. 9 . A system for executing a process having a plurality of working steps using one or more computing units, wherein the system is designed to execute the method according to claim 1 .