Unsupervised data processing

The unsupervised multiple-workflow scheduling system addresses the challenges of explicit mappings and scripts in conventional data processing by using a resource registry for autonomous data processing services, enhancing scalability and fault tolerance.

WO2025136403A1PCT designated stage expired Publication Date: 2025-06-26HITACHI VANTARA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2023/085612
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Conventional data processing systems require explicit mappings and scripts for managing metadata processing, which can lead to errors, are labor-intensive, and make it challenging to manage large numbers of data sets.

Method used

The system employs unsupervised multiple-workflow scheduling (UMWS) that manages metadata processing without explicit mappings or scripts, using a resource registry to maintain references to data resources and update their statuses, allowing data processing services to operate autonomously and independently.

Benefits of technology

This approach simplifies the orchestration of metadata processing, enables recovery from service failures, and facilitates load balancing by allowing additional services to be initiated or idle services to be shut down, without the need for centralized scheduling or pre-planned mappings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2023085612_26062025_PF_FP_ABST
    Figure US2023085612_26062025_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, a first data processing service executing on a computing device accesses data resources maintained in a storage. The first data processing service adds, to a list of resource references, respective references to respective data resources of the data resources maintained in the storage. A second data processing service executing on the computing device searches the list of resource references for respective references having the first data processing service ID and a first status indicator. The second data processing service processes a first batch of the respective data resources corresponding to the respective references located via the searching performed by the second data processing service. In addition, the second data processing service updates the references to the respective data resources processed by the second data processing service to include a second data processing service ID and the first status indicator.
Need to check novelty before this filing date? Find Prior Art

Description

UNSUPERVISED DATA PROCESSINGTECHNICAL FIELD

[0001] This disclosure relates to the technical field of processing data to extract other data, such as in systems that store large amounts of data.BACKGROUND

[0002] Metadata discovery is a process that employs automated tools for data processing to discover information about data sources in data sets. Metadata discovery can provide mappings between the data sources and one or more metadata registries. The process of metadata discovery may be performed with many thousands of data sources (or even millions of data sources, such as in the case of emails, other types of electronic messages, and so forth).

[0003] Conventional tools used for scheduling and orchestrating jobs and workflows for data processing typically use an explicit mapping of the data sources to data processing services, such as through the use of DAGs (directed acyclic graphs), scripts, or the like, for providing the explicit mapping. However, providing an explicit mapping or script for every job may be undesirable as explicit mappings and scripts can introduce errors, and can be labor intensive, as new mappings / scripts or modifications of old mappings / scripts are required for new jobs. Furthermore, metadata discovery may typically be performed as batch processing, and each data processing service may need to be operable and compatible with each data source. This can make management of the processing of a large number of data sets challenging.SUMMARYIn some implementations, a first data processing service executing on a computing device accesses data resources maintained in a storage. The first data processing service adds, to a list of resource references, respective references to respective data resources of the data resources maintained in the storage. A second data processing service executing on the computing device searches the list of resource references for respective references having the first data processing service ID and a first status indicator. The second data processing service processes a first batch of the respective data resources corresponding to the respective references located via the searching performed by the second data processing service. In addition, the second data processing service updates the references to the respective data resources processed by thesecond data processing service to include a second data processing service ID and the first status indicator.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The detailed description is set forth with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different figures indicates similar or identical items or features.

[0005] FIG. 1 illustrates an example architecture of a computing system configured for performing unsupervised data processing according to some implementations.

[0006] FIG. 2 illustrates an example resource reference data structure of a portion of a resource reference according to some implementations.

[0007] FIG. 3 is a flow diagram illustrating an example process performed by a data processing service according to some implementations.

[0008] FIG. 4 illustrates an example block diagram of a framework that includes a flow diagram of an example process for data processing according to some implementations.

[0009] FIG. 5 illustrates an example hardware and logical configuration of a computing system according to some implementations.DESCRIPTION OF THE EMBODIMENTS

[0010] Some implementations herein are directed to techniques and arrangements for unsupervised multiple-workflow scheduling (UMWS) that may be applied for managing metadata processing without utilizing an explicit mapping of resources to the data processing services, and without the use DAGs or scripts for workflow management. For instance, implementations herein may simplify orchestration of metadata processing by removing explicit descriptions of work flow from the processing management. Additionally, examples herein enable recovery after failure of a data processing service, such as by enabling rerunning of the data processing service and cleaning up the status of a failed data processing service. Further, load balancing is simplified by some examples herein. For example, one or more additional data processing services may be initiated as needed. Alternatively, one or more currently active data processing services may be shut down, if the one or more data processing services are determined to be idle.

[0011] In some examples, the data resources herein may include any datasets that are accessible via direct or indirect references to physical storage locations of those datasets. Forinstance, a direct reference may include a uniform resource identifier (URI) while an indirect reference may include a pointer to a direct reference. Some more specific, nonlimiting examples of data resources may include databases, tables in databases, columns / rows in tables, comma-separated-value files, unstructured documents such as PDF, PPT, DOC, image files, electronic messages, other types of files, other types of data objects, and so forth.

[0012] Additionally, a plurality of data processing services are employed herein for performing the processing of the data resources. The data processing services may be, or may include, programs that are configured to perform processing of identified data resources, and different data processing services are configured to each perform different data processing functions. The data processing services are further configured to perform self-scheduling, and do not rely on an overall scheduling manager or a previously defined schedule. To perform the self-scheduling function, the data processing services herein may include some features or capabilities that are designed to support resource registry management while the data processing services are performing data processing. For instance, each data processing service may be configured to update the statuses of the data resources that it is processing, as well as updating other properties of the data resource. In some examples, some of the capabilities of the data processing services may be implemented via annotations, or the like, that are supported in a variety of programming languages. Examples, of the capabilities that may be implemented in at least some of the data processing services herein may include the ability to acquire a batch of one or more data resources for processing, and the ability to update the status of the resource processing of the particular resource in a reference to the data resource maintained in the resource registry.

[0013] In addition, the data processing services herein may be idempotent such that a data processing service may be applied to the same data multiple times and consistently reach the same result. This capability of the data processing services herein relates to recovery of a data processing service from failure. For instance, recovery may include rerunning the failed data processing service against the same data resource. Because the data processing services herein are idempotent, the recovery execution of the data processing service produces the same result as would have been expected to be produced had the execution of the earlier data processing service not failed.

[0014] The examples herein may employ a resource registry (RR) that may be a data structure that provides a list of references to respective data resources (e.g., one listed reference per data resource) that are loaded or otherwise currently targeted for processing. As mentioned above, the data resources may be of different types and may be processed by differentrespective data processing services that are configured to process the different respective data resource types. Each data resource referenced in the resource registry may include at least some of the following associated information, such as, a resource identifier (ID), a resource location, a data processing service ID, a data processing service instance UUID, a resource type, and a resource status.

[0015] For example, the “resource ID” may be an identifier that is unique or otherwise individually distinguishable within the resource registry, such as an SHA- 1 hash of some or all of the corresponding content of the resource. As an alternative to SHA-1 any other type of hashing, fingerprinting, or identification technology may be employed for producing an individually distinguishable identifier. Furthermore, the “resource location” may include a resource location, address, pointer, or other reference to the location of the data resource, e.g., an indication of how to access the data resource. The “data processing service ID” may be an identifier of the data processing service that performs a specific current step of the overall processing of the data. Additionally, the “data processing service instance UUID” may be an identifier that is individually distinguishable or otherwise unique within the system that identifies the instance of the data processing service that performs processing of the current resource. For example, multiple instances of the same processing data processing service may be executed simultaneously, and those instances may be distinguished from each other in some examples herein by the data processing service instance UUID. Further, the resource type may indicate the data type of the data resource, examples of which were discussed above. In addition, the “status” of the resource processing may be one of several possible statuses such as “waiting”, “locked”, “in-process”, “failed”, and / or “completed”.

[0016] For discussion purposes, some example implementations are described in the environment of one or more computing systems that perform unsupervised data processing, such as for metadata discovery. However, implementations herein are not limited to the particular examples provided, and may be extended to other types of computing system architectures, other types of storage environments, other types of data and data processing, and so forth, as will be apparent to those of skill in the art in light of the disclosure herein.

[0017] FIG. 1 illustrates an example architecture of a computing system 100 configured for performing unsupervised data processing according to some implementations. In this example, one or more data processing computing devices 102 may communicate with one or more resource computing devices 104 over one or more networks 106. For instance, the resource computing devices 104 may include a resource computing device 104-1 having data resources 108-1, and a resource computing device 104-2 having data resources 108-2.Alternatively, in other examples, the data resources 108 may be stored at or by the data processing computing device(s) 102, such as in a local storage device, or the like. In some examples, one or more of the resource computing devices 104 may include a network storage that stores data, such as object data or any other type of data that may be subject to metadata discovery. In some examples, such a network storage may be provided by one or more commercial cloud-based storage providers, such as AMAZON®, MICROSOFT®, IBM®, GOOGLE®, HITACHI VANTARA®, or the like, and may typically, but not necessarily, be located at a location that is remote from the data processing computing device(s) 102. Additionally, as an alternative to public cloud storage, one or more private storage systems (cloud or local) may be provided as one or more of the resource computing devices 104. Alternatively, in yet other examples, some or all of the data resources 108 may be stored at or by the data processing computing device(s) 102, such as in a local storage device (not shown in FIG. 1), e.g., a local drive, a local storage array, a local network attached storage (NAS), other type of local storage system, or the like. For instance, in some examples, the resource computing devices 104 might not be included. Thus, the data resources 108 may be stored in any combination of public, private, and / or local storage.

[0018] The one or more networks 106 may include any suitable network, including a wide area network (WAN), such as the Internet; a local area network (LAN), such as an intranet; a wireless network, such as a cellular network, a local wireless network, such as Wi-Fi, and / or short-range wireless communications, such as BLUETOOTH®; a wired network including Libre Channel, fiber optics, Ethernet, or any other such network, a direct wired connection, or any combination thereof. Accordingly, the one or more networks 106 may include both wired and / or wireless communication technologies. Components used for such communications can depend at least in part upon the type of network, the environment selected, or both. Protocols for communicating over such networks are well known and will not be discussed herein in detail. Accordingly, the data processing computing devices 102 and the resource computing devices 104 are able to communicate over the one or more networks 106 using wired or wireless connections, and combinations thereof.

[0019] Additional details of example hardware configurations of the computing system 100 are discussed below with respect to EIG. 5. Eurthermore, the computing system 100 is not limited to the hardware configurations described and illustrated in this disclosure, but may include any suitable or desired hardware configuration able to access stored data and perform the data processing functions described herein. In addition, the hardware configuration at one of the data processing computing devices 102 may be different from that at another one of thedata processing computing devices 102 in some cases. Further, the hardware configuration at one of the resource computing devices 104 may be different from that at another one of the resource computing devices 104 in some cases.

[0020] In some examples, the data processing computing devices 102 are configured to perform data processing, such for discovering metadata associated with a plurality of data resources 108. The data processing computing devices 102 may execute a plurality of data processing services 114 that each are able to autonomously operate to select data resources and perform processing of the selected data resources, such as for accessing data resources, determining metadata associated with the selected data resources, or for performing other data processing functions. In the examples herein, the processing of data resources to extract metadata may typically include performing multiple steps, such as: (1) collecting available metadata, e.g., database URIs, metadata provided by a database (e.g., tables, columns, etc.), file locations, file types, file sizes, and so on; (2) ingesting data resources to discover internal metadata, e.g., basic statistics for numerical data and calculating fingerprints using different hashing techniques; and (3) analyzing ingested data sets of data resources, such as by applying Regex (regular expression) analysis or other analysis techniques.

[0021] The computing system 100 provides an arrangement for managing and processing a large number of data resources for metadata discovery, or the like, that does not rely on explicit mapping or scripting of data processing services to the data resources. The computing system 100 is scalable to perform processing of very large amounts of heterogeneous data resources of different types, which may include structured and unstructured data having different formats and configurations, such as databases, parts of databases, documents, images, electronic messages, other types of data objects, and so forth. Further, in the computing system 100 there is no need to assign data resources to specific data processing services, or vice versa.

[0022] In the example of FIG. 1, a resource registry (RR) 110 is provided by the data processing computing device(s) 102. Alternatively, in other examples, the resource registry 110 may be provided by a separate computing device (not shown in FIG. 1). The resource registry 110 may be, or may include, a list of respective references 112 that refer respective ones of the data resources 108 that are currently targeted for processing. As mentioned above, the data resources 108 may be of various different resource types having different data structures, formats, configurations, or the like. The different resource types may be processed by one or more different respective data processing services 114 that each are configured to perform different processing operations based on the resource type of respective data resource upon which they are configured to operate.

[0023] The example of FIG. 1 includes a reference 112- 1 to a first data resource, a reference 112-2 to a second data resource, ..., and a reference 112-N to an Nth data resource. Further, in this example, each reference 112 to a data resource 108 may include at least a resource location 116, a resource type 118, a data processing service ID 120, and a resource status 122. For instance, the resource location 116 may indicate a path, a URI, a URL, a pointer to one of these, and / or other location information for accessing the corresponding data resource to perform processing. Additionally, the resource type 118 may indicate the type of the data resource. As mentioned above, some nonlimiting examples of data resource types 118 may include databases, tables in databases, columns / rows in tables, comma-separated-value files, unstructured documents, such as PDF, PPT, DOC, image files, electronic messages, other types of files or data objects, and so forth.

[0024] Further, the data processing service ID 120 associated with a reference 112 may identify the last data processing service 114 that worked on the data resource 108 to which the reference 112 corresponds or, in the case that the data resource 108 is currently being worked on, the ID of the data processing service 114 that is performing a current step of the overall processing of the data resource 108. For example, some of the data processing services 114 may be configured to only work on their designated data resource types after another data processing service 114 has already worked on the data resource. Consequently, the data processing service ID 120 associated with a particular reference 112 can help a particular data processing service 114 determine, in part, whether it should start processing of the corresponding data resource 108.

[0025] In addition, the resource status 122 may indicate a current status of the processing of the corresponding data resource 108. Examples of the status of the data resource processing may be one of several possible statuses such as “waiting”, “locked”, “in-process”, “failed”, or “completed”. For instance, “waiting” may indicate that the data resource is ready for further processing and that processing has not yet been completed. “Locked” may indicate that a data processing service 114 has included the data resource 108 in a batch of one or more data resources identified for further processing by that data processing service 114. “In process” may indicate that a data processing service 114 is currently processing the data resource 108. “Failed” may indicate that the data processing failed and that the data processing service 114 that failed may need to be restarted. “Completed” may indicate that data processing for the particular data resource 108 is complete and no further processing is expected.

[0026] The data resources 108 referenced in the resource registry may, in some cases, include some additional information that may be associated with the data resources 108 andthe corresponding data reference 112, such as, a resource ID and a data processing service instance UUID (not shown in FIG. 1). For example, as mentioned above, the “resource ID” may be an identifier that is unique or otherwise individually distinguishable within the resource registry, such as an SHA- 1 hash of some or all of the corresponding content of the resource, some or all of the content of metadata of the data resource, or the like. As an alternative to SHA-1 any other type of hashing, fingerprinting, or identification technology may be employed for creating the resource IDs herein. Additionally, the “data processing service instance UUID” may be an identifier that is individually distinguishable or otherwise unique within the system 100, and that identifies a particular instance of the data processing service 114 that performs processing of the current data resource 108. For example, multiple instances of the same data processing service 114 may be executed concurrently or otherwise contemporaneously, and those different instances may be distinguished from each other in some examples herein by the data processing service instance UUID.

[0027] As mentioned above, the data resources 108 may include a variety of different types of data and may be accessed via direct or indirect location references to physical locations of the data resources. Direct locations may include paths, URIs, URLs, and so forth. For example, a database connection string such as "server=127.0.0.1;uid=root;pwd=12345;database=test” may be provided as a direct reference to a database that is one of the data resources 108. As another example, an identifier in a lookup table that points to a record with a connection string may be an indirect reference to the same data resource 108.

[0028] The data processing services 114 may be configured to perform various different data processing operations depending on the resource type 118 for which they are configured. Additionally, depending on the amount of data resources 108, multiple instances of the same data processing service 114 may operate concurrently with each other. In the example of FIG. 1, three different data processing services 114 are illustrated including a first data processing service 114-1, a second data processing service 114-2, and a third data processing service 114-3. In other examples, there may be fewer, more, or many more data processing services 114, depending on the number of different resource types that require different processing steps. In the illustrated example, for discussion purposes, suppose that the first data processing service 114-1 is a metadata gathering service; second data processing service 114- 2 is a data ingest processing service; and data processing service 114-3 is a Regex data analysis service.

[0029] As illustrated in FIG. 1, the first data processing service 114-1 may perform steps such as reading basic metadata from external data resources such as at the first resourcecomputing device(s) 104-1 and / or the second resource computing device(s) 104-2. For each data resource read, the first data processing service 114-1 may create a reference 112 in the resource registry 110. As mentioned above, each created reference 112 may include a resource location 116, a resource type 118, a data processing service ID 120, and a resource status 122 of the respective data resource 108. After the reference 112 has been created for a particular data resource 108, the first data processing service 114-1 may set the data processing service ID to its own identifier, which suppose, in this example, is “DPS_ID1”, and may set the resource status 122 of the data resource to “waiting”, which indicates that the data resource 108 is waiting for further processing to be performed. Additionally, in some examples, the data processing service 114-1 may perform additional processing such as depending on the resource type, the system configuration, or the like. In some examples, the first data processing service 114-1 may be configured to operate periodically and / or in response to receiving a trigger communication from another program, such as directly, or via an application programming interface (API).

[0030] In addition, the second data processing service 114-2 may subsequently perform a search of references 112 in the resource registry 110 for references 112 that include a data processing service ID 120 corresponding to DPS_ID1, and having a status of “waiting”. Further, in some examples, the search criteria may also include a specified resource type 118 that corresponds to a data type that the second data processing service 114-2 is configured to process. In the illustrated example, suppose that the second data processing service 114-2 is configured to process data resources of resource type “XYZ”. In other examples, however, the resource type may be implied based on the specified data processing service ID. Thus, in these examples, it is not necessary to specify the resource type 118 as one of the search criteria.

[0031] When the second data processing service 114-2 has located one or more references 112 that match the search criteria, the second data processing service 114-2 may create a batch of one or more data resources 108, such as by changing the resource status 122 in each of the located references 112 from “waiting” to “locked”. This can prevent other instances of the second data processing service 114-2 from attempting to perform processing on the same data resources 108 that are currently selected by the instant second data processing service 114-2.

[0032] In some examples, the second data processing service 114-2 may be configured to perform the search periodically. For example, if no references 112 are located by the search, the second data processing service 114-2 may remain idle for a specified interval, and may repeat the search following the specified interval.

[0033] Following creation of the batch of one or more data resources 108, the second data processing service 114-2 may begin processing of the respective data resources 108 in the created batch of data resources 108. In this example, suppose that the processing includes ingesting the metadata of the respective data resources 108 into a data structure, database, or the like. Furthermore, when the second data processing service 114-2 begins processing a particular data resource 108 in the current batch, the second data processing service 114-2 may change the resource status 122 of the particular data resource 108 from “locked” to “in-process” so that the current status of the particular data resource 108 is accurately indicated by the reference 112 corresponding to the particular data resource 108 that is being processed. Additionally, the second data processing service 114-2 may change the data processing service ID 120 associated with the reference 112 in the resource registry 110 from the ID of the first data processing service to the ID of the second data processing service, e.g., “DPS_ID2”. Following completion of the processing of the particular data resource 108, the second data processing service 114-2 may again change the status 122 of the resource in the corresponding reference 112 to another appropriate status such as “waiting”, “completed”, or “failed”. For example, if there is still more processing to be performed for the data resource 108, the status may be indicated to be “waiting” for the further processing. Alternatively, if no further processing is needed, then the status may be changed to “completed”. As yet another alternative, if the processing failed, then the resource status 122 may be changed to “failed” to indicate that the processing by the second data processing service 114-2 needs to be repeated. In the illustrated example suppose that the status is set to “waiting” to indicate that additional processing of the data resource still needs to be performed.

[0034] In the illustrated example, suppose that the third data processing service 114-3 subsequently performs a search of the resource registry 110 for references 112 having a data processing service ID of “DPS_ID2” and a resource status of “waiting”. Additionally, in some examples, the search criteria may also include a resource type, such as resource type “XYZ”. However, in some examples herein it is not necessary to include the resource type as one of the search criteria since the resource type may be implied based on the identifier of the prior data processing service. For example, suppose that the third data processing service 114-3 is configured to process data resources 108 having resource type “XYZ” and that have already been ingested by the second data processing service 114-2. Since the second data processing service 114-2 is configured to only process the resource type XYZ, then any reference having the data processing service identifier “DPS-ID2” may be assumed to be of the resource type XYZ, and it is not necessary to include the reference type 118 as one of the search criteria.

[0035] When one or more references 112 that match the search have been located by the third data processing service 114-3, the third data processing service 114-3 may create a batch of the one or more data resources 108 that match the search by changing the resource status 122 from “waiting” to “locked”, thereby indicating that those data resources are part of a batch that will be processed by that instance of the third of data processing service 114-3. In some examples, the third data processing service 114-3 may be configured to perform the search periodically.

[0036] After the batch of one or more data resources 108 has been created, the third data processing service 114-3 may begin processing the data resources 108 in the batch. In this example, suppose that the third data processing service 114-3 is configured to perform Regex analysis processing on the ingested data that was previously ingested by the second data processing service 114-2. For instance, as is known in the art, Regex analysis may include identifying regular expressions in text data using a string searching algorithm to analyze portions of text that may have been stored for each of the data resources 108 to identify text that may be informative about the corresponding data resource 108. When the third data processing service 114-3 selects a particular data resource 108 in the batch for processing, the third data processing service 114-3 may change the status 122 in the reference 112 for that resource 108 from “locked” to “in-process”. In addition, the third data processing service 114- 3 may change the data processing service ID 120 in the reference 112 for that resource 108 from “DPS_ID2” to “DPS_ID3” (i.e., its own identifier). Further, when the third data processing service 114-3 has completed processing, if the processing was successful and there is no additional processing to be performed in the workflow for that data resource 108, the third data processing service 114-3 may change the resource status 122 in the reference 112 for that resource 108 from “in-process” to “completed”. Alternatively, if there is still additional processing to be performed in the workflow for that data resource 108, the status is set to “waiting”. Additionally, if the processing failed, then the status is set to “failure”.

[0037] The data processing services 114 herein may each operate autonomously and independently of the other data processing services. As one example, each data processing service 114 may periodically search the resource registry 110 for data resources that are ready for processing by the particular data processing service based on the resource type 118, the data processing service ID 120, and the data resource status 122. A batch of one or more resources 108 that are identified for processing may be processed by the data processing service 114 independently from the other data processing services 114. Accordingly, the data processingservices 114 herein do not require any centralized scheduler, or pre-planned mapping or script, to complete processing of large amounts of data resources 108.

[0038] Each data processing service 114 may be an executable instance of a program that is specifically configured to perform processing on one or more particular data resource types that have achieved a particular stage of processing. Furthermore, as discussed above, each of the data processing services 114 may operate in an idempotent manner such that the data processing service 114 will output the same result consistently when processing the same data. This feature enables recovery of data and processing functionality when one of the data processing services 114 fails. For instance, following a failure of one of the data processing services 114, the system 100 may rerun the same data processing service 114, such as by initiating a new instance of the data processing service 114 for processing the same data resource 108 that was subject to the failure by the prior data processing service 114. Because the data processing services 114 herein are idempotent, the rerun of the processing will produce the same result as would have been expected for the previously run processing that failed.

[0039] Additionally, the updating of the status of the respective data resources 108 and their corresponding references 112 in the resource registry 110 enables data resources 108 for processing to be added to the resource registry 110 while one or more of the data processing services 114 are currently working on processing a previously created batch of data resources 108. Furthermore, different data processing services 114 may need different amounts of time to complete their respective processing steps. To accommodate for this, a different number of instances of different data processing services 114 may be initiated in some examples herein for the same data types or the like, such as when there are a large number of data resources of a particular data type, and an executing data processing service 114 satisfies a threshold number of data resources 108 when creating a batch for processing. In some examples, reaching the threshold may trigger the instantiation of another instance of the same data processing service 114.

[0040] When new data resources are added during the processing herein, these resources can be registered in the resource registry 110, and can be made available for processing (i.e., status = “waiting”). For example, in the case of collecting database metadata, some examples may start with only one resource (i.e., the database itself). Then, when the data processing service 114 gets to database tables, these tables may represent new data resources 108 that should be processed in a next step, and may each be listed using separate reference 112 in the resource registry 110. Similarly, the columns inside each of the tables may also be treated as new data resources 108. The columns also may each be listed as separate data resources 108using separate references 112 in the resource registry 110. Consequently, in some examples herein, the data processing services 114 may generate new data resources as output, and these new data resources 108 may be registered and be made available for processing as well.

[0041] The computing system 100 described with respect to FIG. 1 solves the problems with the conventional techniques discussed above and without using an explicit description of a workflow or assignments of data resources to the specific data processing services 114. Thus, examples herein manage resource references by generating a centralized resource registry as a single place for maintaining references to all data resources and managing resource status. For instance, references to all resources upon which processing is to be performed may be added to the resource registry.

[0042] In the examples herein, each resource reference 112 in the resource registry 110 provides information about resource location 116, resource type 118, last data processing service ID 120, and status 122 of the processing of the given resource. Further, the data processing services 114 are configured to independently select, from the resource registry 110, resources that are available for processing by particular ones of the data processing services 114. For example, the data processing services 114 may select certain resources from the resource registry for forming a batch of resources for processing based on the resource status 122, the resource type 118, and information of the most recent data processing service 114 that performed operations on the certain data resources. Upon completion of processing a data resource, the data processing service 114 that performed the processing may update the resource reference in the resource registry 110 by changing preceding data processing service ID to its own, and by updating the resource status.

[0043] In some examples, following a failed processing by one of the data processing services 114, the data processing computing device(s) 102 may execute a three step recovery procedure to resume execution of the failed processing. The steps of the recovery procedure may be performed independently of each other, e.g., sequentially in series, or concurrently in parallel. In particular, the computing device 102 may restart the failed data processing service 114 by instantiating a new instance of the data processing service 114. Further, the computing device 102 may release previously locked resource references by changing the status from “locked” to “waiting” in the corresponding reference 112. Alternatively, for resources with a status that is indicated to be in process, the computing device 102 may rollback the data processing service instance UUID to a blank value, rollback the data processing service ID to the data processing service ID of the preceding data processing service 114 in the workflow, and rollback the resource status to “waiting”.

[0044] Additionally, the computing device 102 may roll back any changes to the data resource or other data made by the failed data processing service 114. The newly instantiated data processing service 114 will subsequently add the data resource to the next batch that it creates for processing and re-execute the processing of the data resource. In some examples, a data processing management program (not shown in FIG. 1) or other suitable program may be executed on the data processing computing device(s) 102 to perform the above discussed recovery procedure. In addition, the data processing management program or other suitable program may be executed to perform simple and unambiguous application maintenance, as discussed additionally below.

[0045] FIG. 2 illustrates an example registry reference data structure 200 of a portion of a registry reference 112 according to some implementations. In this example, the registry reference data structure 200 includes a “resource ID” 202, the “data processing service ID” 120, a “data processing service UUID” 204 and the “resource status” 122. As discussed above with respect to FIG. 1, additional information that may be included in the registry reference data structure 200 may include the resource location 116, the resource type 118 (not shown in FIG. 2). Furthermore, other examples, other types of information (not shown in FIG. 2) related to a data resource may also be included in the reference data structure 200.

[0046] In this example, the changes to the status of the data resource 108, as discussed above with respect to FIG. 1 for the operations performed by the second data processing service 114-2 are reflected in the changes illustrated from 210-216. For example, as indicated at 210, following the search performed by the second data processing service 114-2, suppose that the second data processing service 114-2 locates the reference 112 in this example. The resource status 122 is currently indicated to be “waiting”, and the data processing service ID is “DPS_ID1” i.e., the identifier of the first data processing service 114-1. Furthermore, in this example, the data processing service instance UUID 204 is blank, having been removed by the first data processing service 114-1 following completion of processing by that service.

[0047] At 212, the second data processing service 114-2 changes the resource status 122 to “locked” and enters its own data processing service instance UUID at 204. As mentioned above, the data processing service instance UUID 204 enables a determination of which instance of the second data processing service 114-2 is performing processing on the particular data resource, such as in the case that the data processing fails or the like. As discussed above with respect to FIG. 1, following changing of the resource status 122 to “locked”, the second data processing service 114-2 may create a batch of one or more data resources for processing.

[0048] At 214, following creation of the batch, when the second data processing service 114-2 begins processing the data resource corresponding to this registry reference 112, the second data processing service 114-2 changes the resource status 122 from “locked” to “in- process”. may add its own data processing service ID 120 (i.e., “DPS-ID2”) to the data structure 200 in place of the identifier of the first data processing service 114-1.

[0049] At 216, when the second data processing service 114-2 has completed processing of the data resource corresponding to this registry reference 112, the second data processing service 114-2 may remove its own data processing service instance UUID 204 from the data structure 200, and may change the resource status 122 from “in-process” to “waiting”. Accordingly, the reference is now ready to be discovered by the third data processing service 114-3, as discussed above with respect to FIG. 1.

[0050] FIGS. 3-4 include flow diagrams illustrating example processes according to some implementations. The processes are illustrated as collections of blocks in logical flow diagrams, which represent sequences of operations, some or all of which may be implemented in hardware, software or a combination thereof. In the context of software, the blocks may represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, program the processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like, that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as a limitation. Any number of the described blocks can be combined in any order and / or in parallel to implement the process, or alternative processes, and not all of the blocks need be executed. For discussion purposes, the processes are described with reference to the environments, frameworks, and systems described in the examples herein, although the processes may be implemented in a wide variety of other environments, frameworks, and systems.

[0051] FIG. 3 is a flow diagram illustrating an example process 300 performed by a data processing service 114 according to some implementations. In some cases, the process 300 may be executed at least in part by one or more of the data processing computing devices 102, such as in the computing system 100 discussed above with respect to FIG. 1. For example, each of the data processing services 114 discussed above with respect to FIGS. 1 and 2 may be configured to perform the same or a similar sequence of operations for operating independently of one another for unsupervised data processing.

[0052] At 302, the computing device may start the data processing service. As one example, a data processing management program executing on the computing device may starteach of the data processing services installed on the computing device, or may select certain data processing services to start, such as depending on the data types that are intended to be processed.

[0053] At 304, the data processing service executing on the computing device may perform a search of the resource registry to attempt to locate references to resources based on a specified data processing service ID, a resource type, and that has a status of “waiting”. Furthermore, such as in the case that there is an order of data processing service IDs for particular workflows, the resource type may not need to be included in the search.

[0054] At 306, the data processing service executing on the computing device may determine whether any references were located. If so, the process goes to 310. If not, the process goes to 308.

[0055] At 308, when no references were located by the search, the data processing service executing on the computing device may wait for a period of time and then return to 304 to conduct an next search for references.

[0056] At 310, when one or more references are located by the search, the data processing service executing on the computing device may create a batch for processing by setting the status for the located resources to “locked”, and by setting the “data processing service instance UUID” to the UUID of the current data processing service.

[0057] At 312, the data processing service executing on the computing device may select a resource to the current data processing service ID, and set the resource status of the selected resources to in-process.

[0058] At 314 the data processing service executing on the computing device may perform processing on the selected data resource. Numerous different types of data processing will be apparent to those of skill in the art having the benefit of the disclosure herein. Accordingly, implementations herein are not limited to any particular type of data processing for the selected resources.

[0059] At 316, following processing, the data processing service executing on the computing device may remove data processing service instance UUID from the reference data structure, and set the status to “waiting” (or “completed” if it is the last data processing service in the workflow).

[0060] At 318, the data processing service executing on the computing device may determine whether all of the data resources in the batch have been processed. If so, the process goes to 304 to conduct a new search. If not, the process goes to 312 to select a next resource from the batch for processing. Thus, the data processing services 114 herein do not rely onscheduling, but instead are configured to operate autonomously and independently of the other data processing services 114, and may run until all of the relevant data resources have been processed. The overall data processing of the data resources may be determined to be complete when all of the data processing services 114 are in an idle state.

[0061] Furthermore, if there are multiple data processing jobs to be executed, then all of the data processing services for all of these jobs can be executed at the same time. Initially, only the source data processing services would be able to perform any processing in order to populate the resource registry. Thereafter though, the other data processing services would be able to begin processing of the data resources based on the contents of the resource registry. However, there is no need for a supervisory program or the like for making assignments of the data processing services 114, since the data processing services 114 are configured to operate periodically and autonomously for obtaining relevant data resources for processing.

[0062] In addition, while the workflow is not used for scheduling the data processing services 114 herein, the last data processing service 114 in each workflow may still be configured, as the last data processing service in a workflow, to set the status of a data resource to “completed”. When a data processing service 114 sets the status of a resource to “completed”, then no further processing is needed for that data resource. Accordingly, the data processing management program or other suitable program may perform a garbage collectionlike process of removing the references of “completed” data resources from the resource registry and may remove any associated data from local storage. Furthermore, in some examples, rather than setting the status to “completed”, the last data processing service 114 in the workflow may simply perform the function of removing the reference from the resource registry and removing any associated data from the local storage.

[0063] Furthermore, some examples herein may treat the resource registry 110 similarly to a queue for managing the number of instances of each data processing service 114 that are executing at any particular time in the computing system 100. For example, the resource registry may be considered to be similar to a queue in that it contains a plurality of references to data resources that are awaiting processing by one or more of the data processing services 114. Accordingly, the average wait time of a reference may be compared with the number of references in the resource registry and the rate at which resources are being added to the resource registry by source data processing services. Additionally, different data processing services in the same workflow may require different amounts of time to complete their processing. Accordingly, a basic comparison may be applied that links the number of references in the resource registry with the frequency at which resources are being added to theresource registry and the average processing time of the resources by given data processing services, and may be used to estimate how many instances of each data processing service in each workflow should be executed to optimize the size of the resource registry and the total processing time.

[0064] FIG. 4 illustrates an example block diagram of a framework 400 that includes a flow diagram of an example process for data processing according to some implementations. For example, the framework 400 may include a software container 401 for containing a data processing service 114-a and a data processing service 114-b. For example, the software container 401 may be a DOCKER container or any other suitable type of software container that enables the data processing services 114-a and 114-b to execute as intended. Alternatively, in other examples, the data processing services 114-a and 114-b may each have their own software container 401.

[0065] As one example, the process of FIG. 4 may be applied for determining metadata from unstructured data resources, such as emails having attachments. For instance, email attachments may have a variety of different formats and a variety of file sizes. Accordingly, the framework 400 of FIG. 4 is configured with the data processing service 114-a for extracting attachments from the emails, and the data processing service 114-b for processing the extracted attachments.

[0066] At 402, the computing device may receive or otherwise access a plurality of emails at least some of which may have attachments.

[0067] At 404, the data processing service 114-a executing on the computing device may process the emails, such as by accessing the emails and determining information about the attachments for use in generating resource references for the email attachments.

[0068] At 406, the data processing service 114-a executing on the computing device may extract the attachments from the emails.

[0069] At 408, the data processing service 114-a executing on the computing device may store the attachments to a storage 407, which may be a local storage device in some examples, or any other suitable data storage.

[0070] At 410, the data processing service 114-a executing on the computing device may send resource references corresponding to the respective attachments to a database 409 or other suitable data structure. For example, the database 409 may maintain a resource registry 411, which may include a list of the resource references that reference the attachments extracted from the emails, and that includes a processing status of the respective attachments and a preceding data processing service ID (e.g., the ID of the data processing service 114-a).

[0071] At 412, the data processing service 114-b executing on the computing device may perform a search of the resource registry references to create batches of the attachments for processing. For example, the data processing service 114-b may locate references to the attachments that meet a search criteria that includes a processing status of “waiting” and the data processing service ID for the data processing service 114-a.

[0072] At 414 the data processing service 114-b executing on the computing device read the attachments from the storage 407. For example, based on the created batch, the data processing service 114-the may retrieve the attachments included in the batch for processing.

[0073] At 416, the data processing service 114-b executing on the computing device may process the attachments, such as by extracting text from the attachments.

[0074] At 418, the data processing service 114-b executing on the computing device may extract addresses from the text extracted from the attachments.

[0075] At 420, the data processing service 114-b executing on the computing device may save the extracted address metadata in a database collection. For example, suppose that the database includes an address mated metadata collection 421. The data processing service one- B may save the extracted address metadata for each attachment to metadata 424 for all extracted addresses grouped by email. Furthermore, while a specific example of metadata extraction is described above, imitations herein are not limited to such and numerous variations will be apparent to those of skill in the art having the benefit of the disclosure herein.

[0076] The example processes described herein are only examples of processes provided for discussion purposes. Numerous other variations will be apparent to those of skill in the art in light of the disclosure herein. Additionally, while the disclosure herein sets forth several examples of suitable frameworks, architectures and environments for executing the processes, implementations herein are not limited to the particular examples shown and discussed. Furthermore, this disclosure provides various example implementations, as described and as illustrated in the drawings. However, this disclosure is not limited to the implementations described and illustrated herein, but can extend to other implementations, as would be known or as would become known to those skilled in the art.

[0077] FIG. 5 illustrates an example hardware and logical configuration of a computing system 500 according to some implementations. In some examples, the computing system 500 may correspond to the computing system 100 discussed above with respect to FIG. 1. The computing system 500 includes one or more data processing computing devices 102 that may be, or that may correspond to, the one or more data processing computing devices 102 discussed above with respect to FIG. 1. The data processing computing devices 102 are able tocommunicate with one or more data source computing devices 104, such as through the one or more networks 106 discussed above, and / or are otherwise coupled to one or more storage devices 502 for accessing the data resources 108.

[0078] In some examples, the data processing computing devices 102 may include one or more servers or other types of computing devices that may be embodied in any number of ways. For instance, in the case of a server, the programs, services, applications, other functional components, and at least a portion of data storage may be implemented on at least one server, such as in a plurality of servers, a server farm, a data center, a cloud-hosted computing service, and so forth, although other computer architectures may additionally or alternatively be used. In the illustrated example, each data processing computing device 102 may include, or may have associated therewith, one or more processors 510, one or more computer-readable media 512, and one or more communication interfaces 514.

[0079] Each processor 510 may be a single processing unit or a number of processing units, and may include single or multiple computing units, or multiple processing cores. The processor(s) 510 can be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, system-on-chip processors, digital signal processors, state machines, logic circuitries, graphics processors, and / or any devices that manipulate signals based on operational instructions. As one example, the processor(s) 510 may include one or more hardware processors and / or logic circuits of any suitable type specifically programmed or configured to execute the algorithms and processes described herein. The processor(s) 510 may be configured to fetch and execute computer-readable instructions stored in the computer- readable media 512, which may be executed to program the processor(s) 510 to perform the functions described herein.

[0080] The computer-readable media 512 may include volatile and nonvolatile memory and / or removable and non-removable media implemented in any type of technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. For example, the computer-readable media 512 may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, optical storage, solid state storage, magnetic tape, magnetic disk storage, storage arrays, network attached storage, storage area networks, cloud storage, or any other medium that can be used to store the desired information and that can be accessed by a computing device. Depending on the configuration of the data processing computing devices 102, the computer-readable media 512 may be a tangible non-transitory medium to the extent that, when mentioned, non-transitory computer- readable media exclude media such as energy, carrier signals, electromagnetic waves, and / orsignals per se. In some cases, the computer-readable media 512 may be included in the data processing computing devices 102, while in other examples, the computer-readable media 512 may be partially separate from the data processing computing devices 102.

[0081] The computer-readable media 512 may be used to store any number of functional components that are executable by the processor(s) 510. In many implementations, these functional components comprise instructions or programs that may be executed by the processor(s) 510 and that, when executed, specifically program the processor(s) 510 to perform the actions attributed herein to the data processing computing device(s) 102. Functional components stored in the computer-readable media 512 may include the data processing services 114, such as 114-1, 114-2, 114-3, ...., and so forth. In addition, the functional components may include a data processing (DP) management program 518 that may be configured, in some examples, to start and stop instances of the data processing services 114, execute recovery from processing failures, perform garbage collection as discussed above, perform load balancing, as discussed above, and so forth. In some cases, the functional components may be stored in a storage portion of the computer-readable media 512, loaded into a local memory portion of the computer-readable media 512, and executed by the one or more processors 510.

[0082] In addition, the computer-readable media 512 may store data and data structures used for performing the functions and services described herein. For example, the computer- readable media 512 may store a metadata data structure 520 such as a metadata database that may include stored metadata 522 discovered by the computing system 500. Additional data and data structures stored in the computer readable media may include the resource registry 110, which may contain a plurality of resource references 112.

[0083] The data processing computing device 102 may also include or maintain other functional components and data in the computer readable media 512, which may include programs, drivers, etc., and the data used or generated by the functional components. Further, the data processing computing device 102 may include many other logical, programmatic, and physical components, of which those described above are merely examples that are related to the discussion herein.

[0084] The communication interface(s) 514 may include one or more interfaces and hardware components for enabling communication with various other devices, such as over the one or more network(s) 106. For example, the communication interface(s) 514 may enable communication through one or more of a LAN, the Internet, cable networks, cellular networks, wireless networks (e.g., Wi-Fi) and wired networks (e.g., Fibre Channel, fiber optic, Ethernet),direct connections, as well as close-range communications such as BLUETOOTH®, and the like, as additionally enumerated elsewhere herein.

[0085] In some examples, the data source computing devices 104 may include one or more storage computing devices 530, which may include one or more servers or any other suitable computing device, such as any of the examples discussed above with respect to the data processing computing device(s) 102. The storage computing device(s) 530 may each include one or more processors 532, one or more computer-readable media 534, and one or more communication interfaces 536. For example, the processors 532 may correspond to any of the examples discussed above with respect to the processors 510, the computer-readable media 534 may correspond to any of the examples discussed above with respect to the computer- readable media 512, and the communication interfaces 536 may correspond to any of the examples discussed above with respect to the communication interfaces 514.

[0086] In addition, the computer-readable media 534 may include a storage program 538 as a functional component executed by the one or more processors 532 for managing the storage of data on a storage 540 included in the data source computing devices 104. The storage 540 may include one or more controllers 542 associated with the storage 540 for storing data on one or more local storage devices. For instance, the controller 542 may control the storage devices 502, such as for configuring the storage devices 502 in a RAID configuration, an erasure coded configuration, and / or any other suitable storage configuration. The storage devices 502 may be any type of storage device, such as hard disk drives, solid state drives, optical drives, magnetic tape, combinations thereof, and so forth. Additionally, while several examples of computing systems have been described herein, numerous other systems able to implement the distributed object storage and replication techniques herein will be apparent to those of skill in the art having the benefit of the disclosure herein.

[0087] Various instructions, methods, and techniques described herein may be considered in the general context of computer-executable instructions, such as computer programs and applications stored on computer-readable media, and executed by the processor(s) herein. Generally, the terms program and application may be used interchangeably, and may include instructions, routines, scripts, modules, objects, components, data structures, executable code, etc., for performing particular tasks or implementing particular data types. These programs, applications, and the like, may be executed as native code or may be downloaded and executed, such as in a virtual machine or other just-in-time compilation execution environment. Typically, the functionality of the programs and applications may be combined or distributed as desired in various implementations. An implementation of these programs, applications, andtechniques may be stored on computer storage media or transmitted across some form of communication media.

[0088] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

CLAIMS1. A system comprising: one or more computing devices configured by executable instructions to perform operations comprising: accessing, by a first data processing service executable by the one or more computing devices, data resources maintained in a storage; adding, by the first data processing service, to a list of resource references, respective references to respective data resources of the data resources maintained in the storage, wherein the respective references added to the list of resource references include a first status indicator and a first data processing service identifier (ID); searching, by a second data processing service executable by the one or more computing devices, the list of resource references for the respective references having the first data processing service ID and the first status indicator; processing, by the second data processing service, a first batch of one or more of the respective data resources corresponding to one or more respective references located via the searching by the second data processing service; and updating, by the second data processing service, the one or more references to the one or more respective data resources processed by the second data processing service to include a second data processing service ID and the first status indicator.

2. The system as recited in claim 1, the operations further comprising searching, by a third data processing service executable by the one or more computing devices, the list of resource references for references having the second data processing service ID and the first status indicator; processing, by the third data processing service, a second batch of one or more of the respective data resources corresponding to respective references located via the searching by the third data processing service to determine metadata associated with the one or more respective data resources; and updating, by the third data processing service, the respective references to the one or more data resources processed by the third data processing service, the updating the respective references by the third data processing service including updating the respective references to include a third data processing service ID and one of a second status indicator or the first status indicator.

3. The system as recited in claim 2, wherein: the first status indicator indicates that the respective data resources are waiting for further processing; and the second status indicator indicates that processing of the respective resources is completed.

4. The system as recited in claim 2, wherein searching, by the third data processing service, list of resource references for data resources having the second data processing service ID and the first status indicator further comprises searching list of resource references for data resources having the second data processing service ID, the first status indicator, and a resource type that corresponds to the third data processing service.

5. The system as recited in claim 2, wherein the second status indicator indicates that the processing by the third data processing service failed, the operations further comprising: starting a new instance of the third data processing service; and changing the respective reference to the respective data resource to include the first status indicator and the second data processing service ID.

6. The system as recited in claim 1, the operations further comprising: prior to performing the processing by the second data processing service, changing, by the second data processing service, the first status indicator to a second status indicator that indicates that the respective resources have been selected as part of the first batch of one or more data resources for processing by the second data processing service.

7. The system as recited in claim 6, the operations further comprising: changing, by the second data processing service, the second status indicator to a third status indicator based on the second data processing service performing the processing; and changing, by the second data processing service, the third status indicator to one of the first status indicator or to a fourth status indicator when the second data processing service has completed the processing.

8. The system as recited in claim 7, wherein: the first status indicator indicates that the corresponding respective data resource is waiting for additional processing; the second status indicator indicates that the corresponding respective data resource has been selected for inclusion in the first batch; the third status indicator indicates that the corresponding respective data resource is being processed by the second data processing service; and the fourth status indicator indicates that the processing of the corresponding respective data resource is one of completed or failed.

9. The system as recited in claim 1, wherein searching, by the second data processing service, the resource registry for data resources having the first data processing service ID the first status indicator further comprises searching the resource registry for data resources having the first data processing service ID, the first status indicator, and a resource type that corresponds to the second data processing service.

10. The system as recited in claim 1, the processing by the second data processing service including at least one of reading the one or more respective data resources from a storage or determining information associated with the one or more respective data resources.

11. The system as recited in claim 1, the operations further comprising, based at least on a comparison of a number references included in the list of resource references and an average processing time for the corresponding data resources, starting an additional instance of the second data processing service for execution on the one or more computing devices.

12. A method comprising: accessing, by a first data processing service executing on a computing device, data resources maintained in a storage; adding, by the first data processing service, to a list of resource references, respective references to respective data resources of the data resources maintained in the storage, wherein the respective references added to the list of resource references include a first status indicator and a first data processing service identifier (ID);searching, by a second data processing service executing on the computing device, the list of resource references for the respective references having the first data processing service ID and the first status indicator; processing, by the second data processing service, a first batch of one or more of the respective data resources corresponding to one or more respective references located via the searching by the second data processing service; and updating, by the second data processing service, the one or more references to the one or more respective data resources processed by the second data processing service to include a second data processing service ID and the first status indicator.

13. The method as recited in claim 12, further comprising: searching, by a third data processing service executing on the computing device, the list of resource references for references having the second data processing service ID and the first status indicator; processing, by the third data processing service, a second batch of one or more of the respective data resources corresponding to respective references located via the searching by the third data processing service to determine metadata associated with the one or more respective data resources; and updating, by the third data processing service, the respective references to the one or more data resources processed by the third data processing service, the updating the respective references by the third data processing service including updating the respective references to include a third data processing service ID and one of a second status indicator or the first status indicator.

14. One or more non-transitory computer-readable media storing one or more programs executable by a computing device to configure the computing device to perform operations comprising: accessing, by a first data processing service executing on the computing device, data resources maintained in a storage; adding, by the first data processing service, to a list of resource references, respective references to respective data resources of the data resources maintained in the storage, wherein the respective references added to the list of resource references include a first status indicator and a first data processing service identifier (ID); 1searching, by a second data processing service executing on the computing device, the list of resource references for the respective references having the first data processing service ID and the first status indicator; processing, by the second data processing service, a first batch of one or more of the respective data resources corresponding to one or more respective references located via the searching by the second data processing service; and updating, by the second data processing service, the one or more references to the one or more respective data resources processed by the second data processing service to include a second data processing service ID and the first status indicator.

15. The one or more non-transitory computer-readable media as recited in claim 14, the operations further comprising: searching, by a third data processing service executing on the computing device, the list of resource references for references having the second data processing service ID and the first status indicator; processing, by the third data processing service, a second batch of one or more of the respective data resources corresponding to respective references located via the searching by the third data processing service to determine metadata associated with the one or more respective data resources; and updating, by the third data processing service, the respective references to the one or more data resources processed by the third data processing service, the updating the respective references by the third data processing service including updating the respective references to include a third data processing service ID and one of a second status indicator or the first status indicator.

Citation Information

Patent Citations

  • Scheduling software jobs having dependencies

    US20200073727A1

  • Techniques and Architectures for Providing an Extract-Once Framework Across Multiple Data Sources

    US20220092048A1

  • Heterogeneous data platform

    US20220414105A1