System and method for processing accident data through data pipeline
By processing incident data through a data pipeline system, efficient real-time analysis and storage of incident data were achieved, solving the problem of low incident handling efficiency after system changes and improving the response speed and data processing capabilities of the IT system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FIDELITY INFORMATION SERVICES LLC
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to efficiently process incident data after system changes, leading to downtime and inefficiency. Furthermore, the increased complexity of IT systems places a heavy burden on IT teams, making it difficult to quickly identify and resolve the root causes of incidents.
Accident data is processed through a data pipeline system, including collection points, front-end processors, data storage systems, processing platforms, and data sinks. Machine learning modules are used to analyze accident data, enabling real-time data processing and optimized storage, and supporting the classification and transmission of multi-format data.
It improved the efficiency of incident data processing, reduced downtime, reduced the burden on the IT team, enabled the rapid identification and resolution of incidents, and mitigated the company's economic and reputational losses.
Smart Images

Figure CN121925633A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This patent application claims the benefit of priority to U.S. Nonprovisional Application No. 18 / 478,106, filed on September 29, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The various embodiments disclosed herein generally relate to processing incident data, and more specifically to processing incident data via a data pipeline. Background Technology
[0004] Any change to a system carries some degree of risk, namely that the system may no longer function as expected. Furthermore, even if system performance is not immediately affected, changes may cause problems later, and it may take a significant amount of time and resources to determine what caused the change in system performance.
[0005] For example, in software, deploying, refactoring, or releasing software code carries different types of associated risks, depending on the code being changed. Unclear details about vulnerabilities or risks in a deployed codebase can increase the risk of system outages. Technological shifts are significant for any product and present both substantial risks and opportunities for software companies.
[0006] Downtime and / or incidents incur financial losses for the company in service level agreement (SLA) expenditures, but more importantly, they waste employee time through rework and can negatively impact the company's reputation among its customers. The highest costs are attributed to defects entering production, including cascading effects and direct costs to all downstream teams. Furthermore, even after modifications have been deployed, incident teams may waste time determining what caused the change in system performance.
[0007] Information technology (IT) operations, such as executing change requests, can carry varying levels of risk and impact. In large IT organizations, change-related incidents can account for 70% to 80% of critical incidents, placing a significant burden on IT teams. Modern IT architectures are becoming increasingly complex. Resolving recurring incidents across large systems within an IT environment often involves dispersed staff and systems, as well as separate work orders and time-staggered solutions, resulting in significant inefficiencies in large IT organizations. Furthermore, many IT systems can only handle specific forms of data, contributing to further inefficiencies.
[0008] This disclosure aims to overcome one or more of the challenges mentioned above. Summary of the Invention
[0009] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative, and not intended to limit the disclosed embodiments as claimed.
[0010] In some aspects, the technology described herein relates to a method for processing data through a data pipeline, the method being executed by one or more processors and comprising: receiving data from one or more data sources by a collection point configured to perform at least one of extracting, transforming, or loading the data; transmitting the data from the collection point to a front-door processor configured to process the data; transmitting the processed data from the front-door processor to a data storage system configured to store the processed data; transmitting the processed data from the front-door processor to a processing platform, the processed data transmitted from the front-door processor including data already classified by the front-door processor, the processing platform being configured to apply one or more real-time processing techniques including filtering the processed data; transmitting the processed data from the processing platform to one or more data sinks, each of the one or more data sinks being configured to provide short-term storage of the processed data in an optimized format; and outputting the processed data to an artificial intelligence module.
[0011] In some respects, the techniques described herein relate to a method in which one or more data sources include data from cloud-based environments and / or internal systems.
[0012] In some respects, the techniques described herein relate to a method in which, when data is received from a cloud-based environment, the data is transmitted to a secondary collection point configured to perform additional processing on the data prior to its receipt at the collection point.
[0013] In some respects, the techniques described herein relate to a method in which data from one or more data sources includes at least one of the following: incident data, alarm data, or change data.
[0014] In some respects, the techniques described herein relate to a method in which data from one or more data sources includes data in multiple formats.
[0015] In some respects, the techniques described herein relate to a method in which data from one or more data sources undergoes a format change during data reception.
[0016] In some respects, the techniques described herein relate to a method in which the processing of data by a front-door processor may include: classifying the data into multiple client categories to form multiple datasets associated with the respective client categories, wherein the multiple datasets are stored separately in a data storage system.
[0017] In some respects, the techniques described herein relate to a method in which transferring processed data from a processing platform to one or more data destination layers includes transferring multiple datasets to multiple data destination layers based on associated corresponding client categories.
[0018] In some respects, the technology described herein relates to a method that may further include: determining that data is no longer being received by the collection point; and, upon determining that the data is no longer being received by the collection point, transferring the processed data from the data storage system to the processing platform.
[0019] In some respects, the techniques described herein relate to a method in which processed data transmitted from a front-door processor to a processing platform includes both stream-processed data and batch-processed data.
[0020] In some respects, the techniques described herein relate to a method that further includes: transferring processed data from one or more data sinks to one or more machine learning systems.
[0021] In some aspects, the technology described herein relates to a system for a data pipeline, the system comprising: a memory storing processor-readable instructions; and at least one processor configured to access the memory and execute the processor-readable instructions to perform operations including: receiving data from one or more data sources by a collection point configured to perform at least one of extracting, transforming, or loading the data; transferring the data from the collection point to a front-door processor configured to process the data; transferring the processed data from the front-door processor to a data storage system configured to store the processed data; transferring the processed data from the front-door processor to a processing platform, the processed data including data already classified by the front-door processor, the processing platform being configured to apply one or more real-time processing techniques including filtering the processed data; and transferring the processed data from the processing platform to one or more data sink layers, each of the one or more data sink layers being configured to provide short-term storage of the processed data in an optimized format, and outputting the processed data to an artificial intelligence module.
[0022] In some respects, the techniques described herein relate to a system in which one or more data sources include data from a cloud-based environment and / or an internal system.
[0023] In some respects, the techniques described herein relate to a system in which, when data is received from a cloud-based environment, the data is transmitted to a secondary collection point configured to perform additional processing on the data prior to its receipt at the collection point.
[0024] In some respects, the technology described herein relates to a system in which data from one or more data sources includes at least one of the following: incident data, alarm data, or change data.
[0025] In some respects, the techniques described herein relate to a system in which data from one or more data sources includes data in multiple formats.
[0026] In some respects, the techniques described herein relate to a system in which data from one or more data sources undergoes a format change during data reception.
[0027] In some respects, the technology described herein relates to a system in which the processing of data by a front-door processor includes classifying the data into multiple client categories, thereby forming multiple datasets associated with the respective client categories, wherein the multiple datasets are stored separately in a data storage system.
[0028] In some respects, the techniques described herein relate to a system in which the transfer of processed data from a processing platform to one or more data destination layers includes transferring multiple datasets to multiple data destination layers based on associated corresponding client categories.
[0029] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing processor-readable instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving data from one or more data sources by a collection point configured to perform at least one of extracting, transforming, or loading the data; transferring the data from the collection point to a front-door processor configured to process the data; transferring the processed data from the front-door processor to a data storage system configured to store the processed data; transferring the processed data from the front-door processor to a processing platform, the processed data including data already classified by the front-door processor, the processing platform being configured to apply one or more real-time processing techniques including filtering the processed data; and transferring the processed data from the processing platform to one or more data sinks, each of the one or more data sinks being configured to provide short-term storage of the processed data in an optimized format, and outputting the processed data to an artificial intelligence module.
[0030] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description which follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and obtained by means of the elements and combinations particularly pointed out in the appended claims.
[0031] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative, and not intended to limit the disclosed embodiments as claimed. Attached Figure Description
[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate a variety of exemplary embodiments and, together with the description, serve to explain the principles of the disclosed embodiments.
[0033] Figure 1 An exemplary system overview for data pipelines for data transmission and aggregation according to one or more embodiments is described.
[0034] Figure 2 A flowchart is depicted for a method of receiving and processing data using a data pipeline, according to one or more embodiments.
[0035] Figure 3 A flowchart is depicted for a method of processing data via a data pipeline according to one or more embodiments.
[0036] Figure 4 Specific implementations of general-purpose computer systems capable of performing the techniques presented herein are illustrated. Detailed Implementation
[0037] The various embodiments disclosed herein generally relate to processing incident data, and more specifically to processing incident data via a data pipeline.
[0038] The subject matter of this disclosure will now be described more fully with reference to the accompanying drawings, which illustrate specific exemplary embodiments by way of example. The embodiments or specific implementations described herein as “exemplary” should not be construed as superior to, for example, other embodiments or specific implementations; rather, they are intended to reflect or indicate that the embodiments are “example” embodiments. The subject matter can be embodied in a variety of different forms, and therefore the covered or claimed subject matter is intended to be construed as not being limited to any of the exemplary embodiments set forth herein; exemplary embodiments are provided only as illustrative. Similarly, the claimed or covered subject matter is intended to be within a reasonably broad scope. Among other things, for example, the subject matter can be embodied as a method, apparatus, component, or system. Therefore, embodiments can take the form of, for example, hardware, software, firmware, or any combination thereof (other than software itself). Therefore, the following detailed description is not intended to be understood in a limiting sense.
[0039] Throughout the specification and claims, terms may have nuanced meanings as presented or implied in the context, beyond their expressly stated meanings. Similarly, the phrase "in one embodiment" as used herein does not necessarily refer to the same embodiment, and the phrase "in another embodiment" as used herein does not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include combinations of all or part of exemplary embodiments.
[0040] The terminology used below may be interpreted in its broadest and most reasonable manner, even when used in conjunction with a detailed description of certain specific instances of this disclosure. Indeed, some terms may even be emphasized below; however, any term intended to be interpreted in any limiting manner will be so publicly and specifically defined in this Detailed Description section.
[0041] For example, software companies strive to avoid downtime caused by incidents resulting from software or hardware component upgrades or team member changes. The system described herein can be configured to analyze and / or process event data from IT systems. The system described herein can, for example, receive event data streams within a specific time period. The system can further receive batches of event data. This event data can be further described as Information Technology (IT) event data. Event data can include, but is not limited to: (1) incidents, (2) alarms, (3) change data, (4) problems; and / or (5) anomalies.
[0042] An incident can be any event that may disrupt the operation, service, or functionality of a system or cause the loss of such operation, service, or functionality. Incidents can be manually reported by customers or employees, automatically logged by internal systems, or otherwise captured. Incidents may occur due to factors such as hardware failure, software failure, software defect, human error, and / or cyberattacks. For example, deploying, refactoring, or releasing software code may result in an incident. For example, an incident may be detected during downtime or performance changes. Incidents can include characteristics, where incident characteristics can refer to the quality or features associated with the incident. For example, incident characteristics may include, but are not limited to, the severity of the incident, the urgency of the incident, the complexity of the incident, the scope of the incident, the cause of the incident and / or what configurables correspond to the incident (e.g., which systems / platforms / products are affected by the incident), how the incident is described in free-form text, which business segments are affected, which categories / subcategories are affected, and / or the assignment group to which the incident belongs.
[0043] Change data can refer to information describing modifications made to data within a system or database. Change data can track changes that occur over one or more time periods. Problem data can refer to any data that causes problems or hinders the normal operation of the system. Anomaly data can refer to data that indicates that the system deviates from standards or normal operation.
[0044] Event data can further include the entities affected by the event and their corresponding relationships. Event data can be associated with one or more Configurable Items (CIs). A Configurable Item (CI) can refer to a component of the system, which can be identified as a self-contained unit for change control and identification. For example, a specific application, service, product, or server can be defined by a CI.
[0045] An alert can refer to a notification of an event to the system or user. An alert can include a set of events representing deviations from the normal behavior of the system. For example, an alert can include metadata, including a short field description that includes no text fields (e.g., a summary of the alert), first occurrence, timestamp, alert key, etc. Understanding the different types of alerts within the system from various perspectives can help in resolving incidents.
[0046] One or more embodiments disclosed herein can aggregate and transmit data to alleviate the burden on companies to identify and resolve incidents. In the context of this disclosure, an incident can be any change to the system, such as downtime or performance changes. Incidents can be manually reported by customers or employees, automatically logged by internal systems, or otherwise captured. One or more embodiments can provide IT management, governance, and operations with solutions for identifying and resolving incidents and having a continuous, dynamic impact. One or more embodiments can be extended to clients and users with services and software connected to the systems described herein. One or more embodiments can provide data-agnostic tools to ingest, process, and analyze large volumes of data. One or more embodiments can provide data pipelines (e.g., software platforms) configured to receive data from data sources, transmit and process data, and provide processed data to one or more data sinks. One or more embodiments can allow aggregation, correlation, and parsing options by ingesting, storing, and processing data input. Data input can, for example, come from enterprise-level and business tools and correspond to incident-related data. One or more embodiments can allow a wide variety of data processing to identify correlations, similarities, and root causes, and to recommend corrective actions based on received data and user feedback mechanisms.
[0047] One or more embodiments can utilize a combination of open-source software solutions to collect third-party and system-level data via a collection point. The data can then be passed from the collection point to a front-door processor, where it can then be passed to a data storage system for long-term storage and retrieval. The front-door processor can further pass the data to a processing platform, where the raw data can be aggregated and preprocessed. The data can then be passed to one or more data sinks, where it can be retrieved by one or more machine learning systems configured to evaluate the data using machine learning algorithms, including but not limited to natural language processing, graph embedding, association rule modeling, and anomaly detection. Based on user needs, the machine learning systems can then provide output via an Application Programming Interface (API), which can then trigger automation, update logging systems, or provide user insights via a presentation layer.
[0048] Figure 1 An exemplary system overview for data pipelines for data transmission and aggregation according to one or more embodiments is described. Data pipeline system 100 may be a platform with multiple interconnected components. Data pipeline system 100 may include one or more servers, intelligent networked devices, computing devices, components, and corresponding software for aggregating and processing data.
[0049] like Figure 1 As shown, the data pipeline system 100 may include a data source 101, a collection point 120, a secondary collection point 110, a front-door processor 140, a data storage device 150, a processing platform 160, a data destination layer 170, a data destination layer 171, and an artificial intelligence module 180.
[0050] Data source 101 may include internal data 103 and third-party data 199. Internal data 103 may be a data source directly linked to data pipeline system 100. Third-party data 199 may be an external data source connected to data pipeline system 100, as will be described in more detail below.
[0051] Both internal data 103 and third-party data 199 of data source 101 can include incident data 102. Incident data 102 can include incident reports, where information for each incident is provided as one or more of the following: incident number, closure date / time, category, closure code, closure comment, long description, short description, root cause, or assignment group. Incident data 102 can include incident reports, where information for each incident is provided as one or more of the following: issue keywords, description, summary, tags, issue type, fix version, environment, author, or comments. Incident data 102 can include incident reports, where information for each incident is provided as one or more of the following: filename, script name, script type, script description, display identifier, message, submitter type, submitter link, attributes, file changes, or branch information. Incident data 102 can include one or more of the following: real-time data, market data, performance data, historical data, utilization data, infrastructure data, or security data. These are merely examples of information that can be used as data, and this disclosure is not limited to these examples.
[0052] Incident data 102 can be automatically generated by monitoring tools that generate alerts and incident data to provide notifications of high-risk actions and failures in the IT environment, and can be generated as work orders. Incident data may include metadata such as text fields, identifiers, and timestamps.
[0053] Internal data 103 may be stored in a relational database that includes an incident table. The incident table may be provided as one or more tables and may include, for example, one or more of the following: issues, tasks, risk conditions, incidents, or changes. The relational database may be stored in the cloud. The relational database may be encrypted and connected to a gateway. The relational database may send periodic updates to and receive periodic updates from the cloud. The cloud may be a remote cloud service, a local service, or any combination thereof. The cloud may include a gateway connected to a processing API configured to deliver data to collection point 120 or secondary collection point 110. The incident table may include incident data 102.
[0054] The data pipeline system 100 may include third-party data 199 generated and maintained by third-party data producers. Third-party data producers may generate incident data 102 from Internet of Things (IoT) devices, desktop devices, and sensors. Third-party data producers may include, but are not limited to, Tryambak, Appneta, Oracle, Prognosis, ThousandEyes, Zabbix, ServiceNow, Density, Dyatrace, etc. Incident data 102 may include metadata indicating that the data belongs to a specific client or associated system.
[0055] Data pipeline system 100 may include a secondary collection point 110 to collect and preprocess incident data 102 from data source 101. The secondary collection point 110 may be utilized before data is transmitted to collection point 120. The secondary collection point 110 may be, for example, Apache Minifi software. In one instance, the secondary collection point 110 may run on a microprocessor for a third-party data producer. Each third-party data producer may have an instance of the secondary collection point 110 running on the microprocessor. The secondary collection point 110 may support data formats including, but not limited to, JSON, CSV, Avro, ORC, HTML, XML, and Parquet. The secondary collection point 110 may encrypt the incident data 102 collected from the third-party data producer. The secondary collection point 110 may encrypt the incident data, including but not limited to, via Mutual Authentication Transport Layer Security (mTLS), HTTP, SSH, PGP, IPsec, and SSL. The secondary collection point 110 may perform initial transformation or processing on the incident data 102. Secondary collection point 110 can be configured to collect data from various protocols, immediately generate data traceability, apply transformations and encryption to the data, and prioritize the data.
[0056] Data pipeline system 100 may include collection point 120. Collection point 120 may be a system configured to provide a security framework for routing, transforming, and delivering data from data source 101 to downstream processing devices (e.g., front-door processor 140). Collection point 120 may be, for example, software such as Apache NiFi. Collection point 120 may receive raw data and corresponding fields of the data, such as source name and ingestion time. Collection point 120 may run on a Linux virtual machine (VM) on a remote server. Collection point 120 may include one or more nodes. For example, collection point 120 may receive incident data 102 directly from data source 101. In another instance, collection point 120 may receive incident data 102 from a secondary collection point 110. The secondary collection point 110 may use, for example, a site-to-site protocol to transmit incident data 102 to collection point 120. Collection point 120 may include streaming algorithms. As described herein, streaming algorithms may connect different processors to transfer and modify data from one source to another. For each third-party data producer, collection point 120 may have a separate streaming algorithm. Each streaming algorithm may include a processing group. A processing group may include one or more processors. One or more processors may, for example, retrieve incident data 102 from a relational database. One or more processors may utilize the processing API of internal data 103 to make API calls to the relational database to retrieve incident data 102 from the incident table. One or more processors may also transmit incident data 102 to a destination system, such as front-door processor 140. Collection point 120 may encrypt data via HTTPS, Mutual Authentication Transport Layer Security (mTLS), SSH, PGP, IPsec, and / or SSL. Collection point 120 may support data formats including, but not limited to, JSON, CSV, Avro, ORC, HTML, XML, and Parquet. Collection point 120 may be configured to write messages to and communicate with the cluster of front-door processor 140.
[0057] Data pipeline system 100 may include a distributed event streaming platform, such as a front-door processor 140. Front-door processor 140 may connect to collection point 120 and be configured to receive data from collection point 120. Front-door processor 140 may be implemented in an Apache Kafka cluster software system. Front-door processor 140 may include one or more message brokers and corresponding nodes. Message brokers may be, for example, intermediate computer program modules that translate messages from the sender's formal messaging protocol to the receiver's formal messaging protocol. Message brokers may reside on a single node within front-door processor 140. The message brokers of front-door processor 140 may run on virtual machines (VMs) on a remote server. Collection point 120 may send incident data 102 to one or more message brokers within the message brokers of front-door processor 140. Each message broker may include a topic for storing incident data 102 of similar categories. Topics may be ordered logs of events. Each topic may include one or more subtopics. For example, one subtopic may store incident data 102 related to network problems, while another topic may store incident data 102 related to security breaches from third-party data producers. Each topic may further include one or more partitions. Partitioning can be a systematic approach to dividing a topic log file into numerous logs, each hosted on a separate server. Each partition can be configured to store up to one byte of incident data 102. Each topic can be evenly distributed across one or more message brokers for load balancing and scalability. The front-door processor 140 can be configured to classify received data into multiple client categories, thereby forming multiple datasets associated with the corresponding client categories. These datasets can be stored separately on storage devices, as described in more detail below. The front-door processor 140 can further pass data to storage devices and processors for further processing.
[0058] For example, the front door processor 140 can be configured to assign specific data to corresponding topics. Alarm sources can be assigned to alarm topics, and incident data can be assigned to incident topics. Change data can be assigned to change topics. Problem data can be assigned to problem topics.
[0059] Data pipeline system 100 may include a software framework for data storage device 150. Data storage device 150 may be configured for long-term storage and distributed processing. Data storage device 150 may be implemented using, for example, Apache Hadoop. Data storage device 150 may store incident data 102 transferred from front-door processor 140. Specifically, data storage device 150 may be used for distributed processing of incident data 102, and a Hadoop Distributed File System (HDFS) within the data storage device may be used to organize the communication and storage of incident data 102. For example, HDFS may replicate from any node of front-door processor 140. This replication may prevent hardware or software failures of front-door processor 140. The processing may be executed in parallel on multiple servers simultaneously.
[0060] Data storage device 150 may include HDFS configured to receive metadata (e.g., incident data). Data storage device 150 may also utilize the MapReduce algorithm to process the data. The MapReduce algorithm allows for parallel processing of large datasets. Data storage device 150 may utilize Yet Another Resource Negotiation (YARN) to further aggregate and store data. YARN can be used for cluster resource management and planning tasks related to the stored data. For example, a cluster computing framework such as processing platform 160 may be deployed to further utilize the HDFS of data storage device 150. For example, if data source 101 stops providing data, processing platform 160 may be configured to retrieve data from data storage device 150 directly or via front-door processor 140. Data storage device 150 may allow distributed processing of large datasets across a computer cluster using programming models. Data storage device 150 may include a master node and HDFS for distributed processing across multiple data nodes. The master node may store metadata, such as the number of blocks and their locations. The master node may maintain a file system namespace and regulate client access to said files. The master node can include files and directories and perform file system operations such as naming, closing, and opening files. Data storage device 150 can scale from a single server to thousands of machines, each providing local computing and storage. Data storage device 150 can be configured to store incident data in unstructured, semi-structured, or structured form. In a single instance, multiple datasets associated with corresponding client categories can be stored separately. The master node can store metadata, such as the location of individual datasets.
[0061] Data pipeline system 100 may include a real-time processing framework, such as processing platform 160. In one instance, processing platform 160 may be a distributed data stream engine without its own storage layer. For example, this could be the software platform Apache Flink. In another instance, the software platform Apache Spark may be utilized. Processing platform 160 may support both stream processing and batch processing. Stream processing may be a data processing type that performs continuous, real-time analysis on received data. Batch processing may involve receiving discrete datasets in batches. Processing platform 160 may include one or more nodes. Processing platform 160 may aggregate incident data 102 received from front-door processor 140 (e.g., incident data 102 that has already been processed by front-door processor 140). Processing platform 160 may include one or more operators to transform and process the received data. For example, a single operator may filter incident data 102 and then connect to another operator to perform further data transformation. Processing platform 160 may process incident data 102 in parallel. A single operator may reside on a single node within processing platform 160. Processing platform 160 can be configured to filter and send only specific processed data to a specific data destination layer. For example, depending on the data source of incident data 102 (e.g., whether the data is internal data 103 or third-party data 199), the data can be sent to a separate data destination layer (e.g., data destination layer 170 or data destination layer 171). Furthermore, additional data that is not needed at downstream modules (e.g., artificial intelligence module 180) can be filtered and excluded before being sent to a data destination layer.
[0062] Processing platform 160 can perform three general functions. First, processing platform 160 can perform data validation. The values, structure, and / or format of the data can be matched with the pattern of the destination (e.g., data sink layer 170). Second, processing platform 160 can perform data transformation. For example, source fields, target fields, functions, and parameters can be extracted from the data. Based on the extracted functions, specific transformations can be applied. The transformations can reformat the data for specific downstream uses. Users can be able to select a specific format for downstream uses. Third, processing platform 160 can perform data routing. For example, processing platform 160 can select the shortest and / or most reliable path to send data to the appropriate sink layer (e.g., sink layer 170 and / or sink layer 171).
[0063] In one instance, processing platform 160 can be configured to transmit a specific dataset to the data destination layer. For example, processing platform 160 can receive input variables from a specific artificial intelligence module 180. Processing platform 160 can then filter the data received from front-door processor 140 and transmit only the data relevant to the input variables of artificial intelligence module 180 to the data destination layer.
[0064] Data pipeline system 100 may include one or more data destination layers (e.g., data destination layer 170 and data destination layer 171). Incident data 102 processed from processing platform 160 may be transferred to and stored in data destination layer 170. In one instance, data destination layer 171 may be externally stored on a server for a specific client. Data destination layers 170 and 171 may be implemented using software such as, but not limited to, PostgreSQL, HIVE, Kafka, OpenSearch, and Neo4j. Data destination layer 170 may receive internal data 103 that has been processed and received from processing platform 160. Data destination layer 171 may receive third-party data 199 that has been processed and received from processing platform 160. Data destination layers may be configured to deliver incident data 102 to artificial intelligence module 180. Data destination layers may be data lakes, data warehouses, or cloud storage systems. Each data destination layer may be configured to store incident data 102 in either structured or unstructured formats. Data layer 170 can store incident data 102 in several different formats. For example, data layer 170 can support data formats such as JavaScript Object Notation (JSON), Comma-Separated Values (CSV), Avro, Optimized Row Columnar (ORC), Hypertext Markup Language (HTML), Extensible Markup Language (XML), or Parquet. Data layers (e.g., data layer 170 or data layer 171) can be accessed by one or more individual components. For example, data layers can be accessed by a Non-structured Query Language ("NoSQL") database management system (e.g., a Cassandra cluster), a graph database management system (e.g., a Neo4j cluster), further processing programs (e.g., a Kafka+Flink program), and a relational database management system (e.g., a Postgres cluster). Therefore, further processing can be performed before the artificial intelligence module 180 receives the processed data.
[0065] Data pipeline system 100 may include artificial intelligence module 180. Artificial intelligence module 180 may include machine learning components. Artificial intelligence module 180 may use the received data to train and / or use machine learning models. Machine learning models may be, for example, neural networks. Nevertheless, it should be noted that artificial intelligence module 180 may use other machine learning techniques and frameworks to perform the methods contemplated in this disclosure. For example, other types of supervised and unsupervised machine learning techniques may be used to implement the system and methods, such as regression problems, random forests, clustering algorithms, principal component analysis (PCA), reinforcement learning, or combinations thereof. Artificial intelligence module 180 may be configured to extract and receive data from data sink layer 170.
[0066] Figure 2 A flowchart of a method 200 for receiving and processing data using a data pipeline, according to one or more embodiments, is depicted. Flowchart 200 can depict the processing and transmission of data for... Figure 1 The exemplary methods used by the data pipeline system 100 described herein are described below. An exemplary process flow of the method 200 performed according to the data pipeline system 100 described above is described below.
[0067] It should be understood that the steps shown and described herein, and the order in which they are presented, are merely illustrative, such that various embodiments may include additional and / or fewer steps without departing from the scope of this disclosure.
[0068] At step 202, data may be received from a data source (e.g., data source 101). The received data may include, for example, incident data 102. Data may be received from a connected system or from a third-party data producer. The data may have been automatically generated by a monitoring system that generates alerts when warnings / critical events, outages, and / or failures occur in the IT environment. The received data may further include additional metadata. For example, an incident data alert may include metadata such as a reference code, a text field describing the incident, and a timestamp indicating when the incident occurred.
[0069] Once data is received at the data source, it can be collected by a secondary collection point (e.g., secondary collection point 110) if the data originates from one or more third-party data producers. In some embodiments, such as when the data originates from one or more internal systems (e.g., when the data is internal data 103), this step of collecting data by a secondary collection point may not be performed. A secondary collection point may collect incident data 102 from a single third-party data producer, or a single instance of a secondary collection point may collect incident data from one or more third-party data producers. The secondary collection point may encrypt the incident data 102 collected from the third-party producers. The secondary collection point may perform initial processing on the incident data 102. For example, the secondary collection point may apply transformations and encryption to the received data. The data may be further prioritized and transmitted for further processing.
[0070] At step 204, data can be transferred from the data source to a collection point (e.g., collection point 120). In one instance, the collection point may receive data directly from the data source (e.g., data source 101). In another instance, the collection point may receive data that has already been preprocessed by an additional collection point (e.g., secondary collection point 110). The collection point can be configured to manage and automate the data flow from the data source to the downstream processing system (e.g., front-door processor 140). For example, the collection point may receive raw data and corresponding fields of the raw data, such as source name and ingestion time. The collection point may utilize one or more processors to create streaming algorithms to transmit and modify received incident data before transmitting it for further processing. The collection point may create one or more streaming algorithms. For example, a separate streaming algorithm may be created for each secondary collection point with which the collection point interacts. The collection point may include processors configured to retrieve incident data 102 from secondary collection points using a site-to-site protocol. Furthermore, processors within the collection point may be interconnected to perform additional data processing or data transformation. For example, collection points can perform payload distribution of the received data and can be configured to provide data at a high transaction rate. Collection points can also buffer and queue data.
[0071] At step 206, data that may have already been organized / processed at the collection point is transmitted to the front-door processor (e.g., front-door processor 140). The front-door processor can perform additional processing on the received data. For example, the received data can be categorized into specific topics associated with a specific agent within the collection point. For example, alarm sources can be assigned an alarm topic, incident data can be assigned an incident topic, change data can be assigned a change topic, and problem data can be assigned a problem topic. Topics can further have corresponding partitions created in real time when the data is received. The created partitions can then be accessed by another processing device (e.g., processing platform 160). The created partitions can also be accessed by a storage device (e.g., data storage device 150).
[0072] At step 208, the processed data (e.g., incident data 102 already processed by the front-door processor 140) can be transferred from the front-door processor to a storage system (e.g., data storage device 150) and a processing platform (e.g., processing platform 160). The storage system can utilize nodes and HDFS to distribute the processing and incident data 102 across multiple nodes. In one embodiment, the storage system can be a long-term storage system.
[0073] At step 208, the front-door processor (e.g., front-door processor 140) can also send processed data (e.g., incident data 102 already processed by front-door processor 140) to a processing platform (e.g., processing platform 160). The processing platform can aggregate real-time incident data for real-time processing by operators. A single operator can filter incident data and then connect to another operator to perform further data transformation. The processing platform can provide continuous real-time processing while the front-door processor sends incident data as a continuous data stream. The processing platform can, for example, verify the received data, perform data transformations on the data, and route the data to a data destination layer (e.g., data destination layer 170 and / or data destination layer 171). The processing platform can transform and enrich the incident data and move it from one storage system to another in a continuous streaming mode. The processing platform can be further configured to output further processed data (e.g., incident data 102 already further processed by processing platform 160) to the data destination layer.
[0074] At step 210, the further processed data can be transferred from the processing platform to one or more data destination layers (e.g., data destination layer 170 and / or data destination layer 171). Each data destination layer can provide temporary storage for the further processed data. The data destination layer can store the received incident data in an optimized format for retrieval and further processing by a system (e.g., artificial intelligence module 180). For example, the optimized format of the processed data can be specific to a particular machine learning module (e.g., artificial intelligence module 180). The optimized format can depend on what type of machine learning system will retrieve the information. The data destination layer can be configured to store the received incident data in a structured or unstructured format on a cloud or local storage device. The data destination layer can further encrypt the received incident data. The data destination layer can be configured to transfer the received incident data to an artificial intelligence system (e.g., artificial intelligence module 180) upon request.
[0075] For example, data reception layer 170 can send further processed incident data 102 to artificial intelligence module 180. Artificial intelligence module 180 can be trained using supervised or unsupervised methods with the received incident data 102. After being trained on a specific system, artificial intelligence module 180 can further utilize the received incident data 102 in use cases. Artificial intelligence module 180 can be configured to analyze, aggregate, compare, and / or contrast received incident data 102 from one or more systems.
[0076] Figure 3 A flowchart is depicted for a method 300 for processing data via a data pipeline according to one or more embodiments.
[0077] At step 302, data from one or more data sources may be received by a collection point configured to perform at least one of extracting, transforming, or loading the data.
[0078] At step 304, data can be transmitted from the collection point to the front door processor, which is configured to process the data.
[0079] At step 306, the processed data can be transferred from the front-door processor to a data storage system configured to store the processed data.
[0080] At step 308, the processed data can be transferred from the front-door processor to the processing platform. The processed data transferred from the front-door processor to the processing platform includes data that has been classified by the front-door processor. The processing platform is configured to apply one or more real-time processing techniques, including filtering the processed data.
[0081] At step 310, the processed data can be transferred from the processing platform to one or more data destination layers, each of which is configured to provide short-term storage for the processed data.
[0082] In another aspect, one or more data sources include data from cloud-based environments and / or internal systems.
[0083] In another aspect, when data is received from a cloud-based environment, the data is transmitted to a secondary collection point that is configured to perform additional processing on the data before it is received at the collection point.
[0084] In another respect, data from one or more data sources includes at least one of the following: incident data, alarm data, or change data.
[0085] In another aspect, data from one or more data sources includes data in multiple formats.
[0086] In another aspect, data from one or more data sources undergoes format changes during data reception.
[0087] In another aspect, the front-door processor's data processing includes classifying the data into multiple client categories, thereby forming multiple datasets associated with the corresponding client categories, wherein the multiple datasets are stored separately in the data storage system.
[0088] In another aspect, transferring processed data from the processing platform to one or more data destination layers includes transferring multiple datasets to multiple data destination layers based on the associated corresponding client categories.
[0089] In another aspect, method 300 further includes: determining that the data is no longer received by the collection point; and, upon determining that the data is no longer received by the collection point, transferring the processed data from the data storage system to the processing platform.
[0090] In another aspect, the processed data transmitted from the front-door processor to the processing platform includes both stream processing data and batch processing data.
[0091] In another aspect, method 300 further includes: transferring the processed data from one or more data sinks to one or more machine learning systems.
[0092] Figure 4 Specific implementations of general-purpose computer systems capable of performing the techniques presented herein are illustrated.
[0093] Unless otherwise expressly stated, as will be apparent from the following discussion, it should be understood that throughout this specification, discussions using terms such as “processing,” “computing,” “calculating,” “determining,” “analyzing,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate data expressed as physical quantities (such as electronic quantities) and / or convert such data, expressed as physical quantities (such as electronic quantities), into other data similarly expressed as physical quantities.
[0094] In a similar manner, the term "processor" can refer to any device or part of a device that processes electronic data, for example, from registers and / or memory, to transform that electronic data into other electronic data, for example, that can be stored in registers and / or memory. "Computer," "computing machine," "computing platform," "computing device," or "server" can include one or more processors.
[0095] Figure 4 A specific implementation of computer system 400 is illustrated. Computer system 400 may include an instruction set that can be executed to cause computer system 400 to perform any or more of the methods or computer-based functions disclosed herein. Computer system 400 may operate as a stand-alone device or may be connected to other computer systems or peripheral devices, for example, via a network.
[0096] In a networked deployment, computer system 400 can operate as a server, or as a client computer in a server-client user network environment, or as a peer-to-peer (or distributed) computer system in a peer-to-peer (or distributed) network environment. Computer system 400 can also be implemented as or incorporated into a wide variety of devices, such as personal computers (PCs), tablet PCs, set-top boxes (STBs), personal digital assistants (PDAs), mobile devices, handheld computers, laptop computers, desktop computers, communication equipment, wireless telephones, landline telephones, control systems, cameras, scanners, fax machines, printers, pagers, personal trusted devices, network equipment, network routers, switches, or bridges, or any other machine capable of executing a set of instructions (sequentially or otherwise) specifying actions to be taken by that machine. In a particular implementation, computer system 400 can be implemented using electronic devices that provide voice, video, or data communication. Furthermore, although computer system 400 is exemplified as a single system, the term "system" should also be understood to include any collection of systems or subsystems that individually or jointly execute one or more sets of instructions to perform one or more computer functions.
[0097] like Figure 4 As illustrated, computer system 400 may include processor 402, such as a central processing unit (CPU), a graphics processing unit (GPU), or both. Processor 402 can be a component in various systems. For example, processor 402 may be part of a standard personal computer or workstation. Processor 402 may be one or more general-purpose processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), servers, networks, digital circuits, analog circuits, combinations thereof, or other devices now known or hereafter developed for analyzing and processing data. Processor 402 may implement software programs, such as manually generated (i.e., programmed) code.
[0098] Computer system 400 may include memory 404 that can communicate via bus 408. Memory 404 may be main memory, static memory, or dynamic memory. Memory 404 may include, but is not limited to, computer-readable storage media, such as various types of volatile and non-volatile storage media, including but not limited to random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media, etc. In one embodiment, memory 404 includes cache or random access memory for processor 402. In alternative embodiments, memory 404 is decoupled from processor 402, such as processor cache memory, system memory, or other memory. Memory 404 may be an external storage device or database for storing data. Examples include hard disk drives, optical discs (“Compact Disc CD”), digital video discs (“Digital Video Disc DVD”), memory cards, memory sticks, floppy disks, universal serial bus (“Universal Serial Bus USB”) storage devices, or any other device operable to store data. Memory 404 is operable to store instructions executable by processor 402. Functions, actions, or tasks illustrated in the figures or described herein can be performed by processor 402 executing the instructions stored in memory 404. Functions, actions, or tasks are independent of a particular type of instruction set, storage medium, processor, or processing strategy, and can be performed by software, hardware, integrated circuits, firmware, microcode, etc., operating individually or in combination. Similarly, processing strategies can include multiprocessing, multitasking, parallel processing, etc.
[0099] As shown, computer system 400 may further include display 410, such as a liquid crystal display (LCD), organic light-emitting diode (OLED), flat panel display, solid-state display, cathode ray tube (CRT), projector, printer, or other display device now known or hereafter developed for outputting determined information. Display 410 may serve as an interface for a user to view the functions of processor 402, or specifically as an interface with software stored in memory 404 or drive unit 406.
[0100] Further or alternatively, the computer system 400 may include an input device 412 configured to allow a user to interact with any component of the computer system 400. The input device 412 may be a numeric keypad, keyboard, or cursor control device such as a mouse, joystick, touchscreen display, remote control, or any other device operable to interact with the computer system 400.
[0101] Computer system 400 may also or alternatively include a drive unit 406, which is implemented as a disk or optical drive. Drive unit 406 may include a computer-readable medium 422 in which an instruction set 424 (e.g., software) may be embedded. Furthermore, the instructions 424 may embody one or more of the methods or logic described herein. During execution by computer system 400, the instructions 424 may reside wholly or partially within memory 404 and / or processor 402. Memory 404 and processor 402 may also include computer-readable media as discussed above.
[0102] In some systems, computer-readable medium 422 includes instructions 424, or receives and executes instructions 424 in response to a propagated signal, enabling devices connected to network 470 to transmit voice, video, audio, images, or any other data through network 470. Furthermore, instructions 424 can be sent or received via network 470 through a communication port or interface 420 and / or using bus 408. The communication port or interface 420 may be part of processor 402 or may be a separate component. The communication port or interface 420 may be created in software or may be a physical connection in hardware. The communication port or interface 420 may be configured to connect to network 470, external media, display 410, or any other component or combination thereof in computer system 400. Connection to network 470 may be a physical connection, such as a wired Ethernet connection, or wirelessly established as discussed below. Similarly, additional connections to other components of computer system 400 may be physical connections or may be wirelessly established. Network 470 may alternatively be directly connected to bus 408.
[0103] Although computer-readable medium 422 is shown as a single medium, the term "computer-readable medium" can include a single medium or multiple media, such as a centralized or distributed database, and / or associated caches and servers storing one or more sets of instructions. The term "computer-readable medium" can also include any medium capable of storing, encoding, or carrying a set of instructions for execution by a processor or for causing a computer system to perform any or more of the methods or operations disclosed herein. Computer-readable medium 422 can be non-transitory and can be tangible.
[0104] Computer-readable medium 422 may include solid-state memory, such as a memory card or other package containing one or more non-volatile read-only memories. Computer-readable medium 422 may be random access memory or other volatile rewritable memory. Further or alternatively, computer-readable medium 422 may include magneto-optical or optical media, such as disks or magnetic tapes or other storage devices, to capture carrier signals, such as signals transmitted via a transmission medium. Digital file attachments such as emails or other self-contained information archives or collections of archives can be considered as distribution media as tangible storage media. Therefore, this disclosure is considered to include any and more of computer-readable media or distribution media in which data or instructions are stored, as well as other equivalents and successor media.
[0105] In alternative embodiments, specialized hardware implementations, such as application-specific integrated circuits (ASICs), programmable logic arrays (PLA), and other hardware devices, can be configured to implement one or more of the methods described herein. Applications that can include a wide variety of implementations of the apparatus and systems can broadly encompass various electronic and computer systems. One or more embodiments described herein may use two or more specific interconnected hardware modules or devices to implement functionality, these modules or devices having associated control and data signals transmitted between and through modules, or as part of an ASIC. Therefore, the systems of the present invention encompass software, firmware, and hardware implementations.
[0106] Computer system 400 can be connected to network 470. Network 470 can define one or more networks, including wired or wireless networks. Wireless networks can be cellular telephone networks, 802.11, 802.16, 802.20, or WiMAX networks. Furthermore, such networks can include public networks (such as the Internet), private networks (such as intranets), or combinations thereof, and can utilize various networking protocols now available or developed in the future, including but not limited to TCP / IP-based networking protocols. Network 470 can include wide area networks (WANs), such as the Internet, local area networks (LANs), campus networks, metropolitan area networks, such as direct connections via universal serial bus (USB) ports, or any other network that allows data communication. Network 470 can be configured to couple one computing device to another to enable data communication between the devices. Network 470 can generally be able to use any form of machine-readable medium for transmitting information from one device to another. Network 470 can include communication methods through which information can be propagated between computing devices. Network 470 can be divided into subnets. A subnet may allow access to all other components connected to it, or a subnet may restrict access between components. Network 470 may be considered a public or private network connection and may include, for example, a virtual private network or encryption or other security mechanisms employed on the public Internet.
[0107] According to various embodiments of this disclosure, the methods described herein can be implemented by software programs executable by a computer system. Furthermore, in exemplary, non-limiting embodiments, the implementation may include distributed processing, component / object distributed processing, and parallel processing. Alternatively, virtual computer system processing may be configured to implement one or more of the methods or functionalities described herein.
[0108] Although this specification describes components and functions implemented in specific implementations with reference to particular standards and protocols, this disclosure is not limited to such standards and protocols. For example, standards used for transport on the Internet and other packet-switched networks (e.g., TCP / IP, UDP / IP, HTML, HTTP) represent examples of prior art. Such standards are periodically superseded by faster or more efficient equivalents with substantially the same functionality. Therefore, alternative standards and protocols with the same or similar functionality as those disclosed herein are considered their equivalents.
[0109] It should be understood that, in one embodiment, the steps of the method under discussion are performed by a suitable processor (or multiple processors) of a processing (i.e., computer) system that executes instructions (computer-readable code) stored in a storage device. It should also be understood that this disclosure is not limited to any particular specific implementation or programming technique, and that this disclosure may be implemented using any suitable technique for implementing the functionality described herein. This disclosure is not limited to any particular programming language or operating system.
[0110] It should be recognized that in the foregoing description of exemplary embodiments of this disclosure, the various features of this disclosure are sometimes combined in a single embodiment, drawing, or description thereof for the purpose of simplifying this disclosure and aiding in understanding one or more of the various aspects of the invention. However, this approach of the disclosure should not be construed as reflecting an intention that the claimed disclosure requires more features than expressly recited in each claim. Rather, as reflected in the appended claims, the inventive aspect lies in fewer than all the features of a single embodiment of the foregoing disclosure. Therefore, the claims following the detailed description are thus expressly incorporated into this detailed description, wherein each claim exists on its own as a separate embodiment of this disclosure.
[0111] Furthermore, while some embodiments described herein include features included in other embodiments but not all features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments, as understood by those skilled in the art. For example, any claimed embodiment may be used in any combination as described in the appended claims.
[0112] Furthermore, some embodiments described herein are described as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing such functions. Therefore, a processor having the necessary instructions for performing the elements of such methods or methods forms means for performing the elements of such methods or methods. Moreover, the elements of the apparatus embodiments described herein are examples of means for performing the functions performed by those elements to achieve the purposes of this disclosure.
[0113] Numerous specific details are set forth in the description provided herein. However, it should be understood that embodiments of this disclosure can be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0114] Similarly, it should be noted that when used in the claims, the term "coupled" should not be construed as limited to a direct connection. The terms "coupled" and "connected," and their derivatives, may be used. It should be understood that these terms are not intended to be synonymous with each other. Therefore, the scope of the statement "device A coupled to device B" should not be limited to devices or systems in which the output of device A is directly connected to the input of device B. This implies the existence of a path between the output of A and the input of B, which may include other devices or apparatuses. "Coupled" can refer to two or more elements in direct physical or electrical contact, or two or more elements that are not in direct contact with each other but still cooperate or interact with each other.
[0115] Therefore, while preferred embodiments considered to be those of this disclosure have been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of this disclosure, and it is intended that all such changes and modifications fall within the scope of this disclosure. For example, any formula given above is merely representative of procedures that can be used. Functions can be added to or removed from the block diagram, and operations can be interchanged between functional blocks. Within the scope of this disclosure, steps can be added to or removed from the described methods.
[0116] The subject matter discussed above should be considered illustrative rather than restrictive, and the appended claims are intended to cover all such modifications, improvements, and other embodiments that fall within the true spirit and scope of this disclosure. Therefore, to the fullest extent permitted by law, the scope of this disclosure will be determined by the broadest permissible interpretation of the appended claims and their equivalents, and should not be bound or limited by the foregoing detailed description. While a wide variety of embodiments of this disclosure have been described, it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of this disclosure. Therefore, this disclosure is not limited except as provided in the appended claims and their equivalents.
Claims
1. A method for processing data via a data pipeline, the method being executed by one or more processors and comprising: Data is received from one or more data sources by a collection point, which is configured to perform at least one of extracting, transforming, or loading the data. The data is transmitted from the collection point to the front door processor, which is configured to process the data; The processed data is transferred from the front-door processor to a data storage system configured to store the processed data; The processed data is transmitted from the front-door processor to the processing platform, the processed data transmitted from the front-door processor to the processing platform including data that has been classified by the front-door processor, and the processing platform is configured to apply one or more real-time processing technologies including filtering the processed data; as well as The processed data is transmitted from the processing platform to one or more data destination layers, each of which is configured to provide short-term storage of the processed data in an optimized format and output the processed data to the artificial intelligence module.
2. The method of claim 1, wherein the one or more data sources include data from a cloud-based environment and / or an internal system.
3. The method of claim 2, wherein when the data is received from the cloud-based environment, the data is transmitted to a secondary collection point configured to perform additional processing on the data before the data is received at the collection point.
4. The method of claim 1, wherein the data from the one or more data sources includes at least one of the following: incident data, alarm data, or change data.
5. The method of claim 1, wherein the data from the one or more data sources includes data in multiple formats.
6. The method of claim 5, wherein the data from the one or more data sources undergoes a format change during the receipt of the data.
7. The method of claim 1, wherein the processing of the data by the front-door processor includes: The data is classified into multiple client categories, thereby forming multiple datasets associated with the corresponding client categories, wherein the multiple datasets are stored separately in the data storage system.
8. The method of claim 7, wherein transmitting the processed data from the processing platform to one or more data destination layers comprises transmitting the plurality of datasets to the plurality of data destination layers based on the associated corresponding client categories.
9. The method according to claim 1, wherein the method further comprises: It is determined that the data will no longer be received by the collection point; as well as When it is determined that the data will no longer be received by the collection point, the processed data is transmitted from the data storage system to the processing platform.
10. The method of claim 1, wherein the processed data transmitted from the front-door processor to the processing platform includes stream processing data and batch processing data.
11. The method according to claim 1, wherein the method further comprises: The processed data is transmitted from the one or more data host layers to one or more machine learning systems.
12. A system for a data pipeline, the system comprising: A memory that stores processor-readable instructions; and At least one processor, the at least one processor being configured to access the memory and execute processor-readable instructions to perform operations including: Data is received from one or more data sources by a collection point, which is configured to perform at least one of extracting, transforming, or loading the data. The data is transmitted from the collection point to the front door processor, which is configured to process the data; The processed data is transferred from the front-door processor to a data storage system configured to store the processed data; The processed data is transmitted from the front-door processor to the processing platform, the processed data transmitted from the front-door processor to the processing platform including data that has been classified by the front-door processor, and the processing platform is configured to apply one or more real-time processing technologies including filtering the processed data; as well as The processed data is transmitted from the processing platform to one or more data destination layers, each of which is configured to provide short-term storage of the processed data in an optimized format and output the processed data to the artificial intelligence module.
13. The system of claim 12, wherein the one or more data sources include data from a cloud-based environment and / or an internal system.
14. The system of claim 13, wherein when the data is received from the cloud-based environment, the data is transmitted to a secondary collection point configured to perform additional processing on the data before the data is received at the collection point.
15. The system of claim 12, wherein the data from the one or more data sources includes at least one of the following: incident data, alarm data, or change data.
16. The system of claim 12, wherein the data from the one or more data sources includes data having multiple formats.
17. The system of claim 16, wherein the data from the one or more data sources undergoes a format change during the receipt of the data.
18. The system of claim 12, wherein the processing of the data by the front-door processor includes: The data is classified into multiple client categories, thereby forming multiple datasets associated with the corresponding client categories, wherein the multiple datasets are stored separately in the data storage system.
19. The system of claim 18, wherein transmitting the processed data from the processing platform to one or more data destination layers comprises transmitting the plurality of datasets to the plurality of data destination layers based on the associated corresponding client categories.
20. A non-transitory computer-readable medium storing processor-readable instructions, which, when executed by at least one processor, cause the at least one processor to perform operations including: Data is received from one or more data sources by a collection point, which is configured to perform at least one of extracting, transforming, or loading the data. The data is transmitted from the collection point to the front door processor, which is configured to process the data; The processed data is transferred from the front-door processor to a data storage system configured to store the processed data; The processed data is transmitted from the front-door processor to the processing platform, the processed data transmitted from the front-door processor to the processing platform including data that has been classified by the front-door processor, and the processing platform is configured to apply one or more real-time processing technologies including filtering the processed data; as well as The processed data is transmitted from the processing platform to one or more data destination layers, each of which is configured to provide short-term storage of the processed data in an optimized format and output the processed data to the artificial intelligence module.