Ensuring data quality through deployment automation in data streaming application

By introducing a framework of event ID counting and pausing data reading in the data streaming application, the reliability problem of data streaming services during the update period is solved, and a high reliability and availability data streaming services are achieved.

CN120106852APending Publication Date: 2025-06-06PAYPAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510158804.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-10
Filing Date
2020-04-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing data streaming services have difficulty providing high reliability during normal operation and deployment/update operations, which may lead to data delays and loss, affecting service availability.

Method used

By introducing a framework in a data streaming application, data pipeline event ID count is recorded and data reading is paused during updates to ensure the integrity of data processing. After restarting, collect data message counts to verify that the data is flowing as expected and detect and correct potential problems in a timely manner.

Benefits of technology

It realizes high reliability during normal operation and update operations of data streaming applications, ensures that data is not lost or delayed, and improves service availability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106852A_ABST
    Figure CN120106852A_ABST
Patent Text Reader

Abstract

A framework is described that allows a data streaming application to ensure high reliability both during update operations and during normal operations. A unique event ID count reflecting messages sent from a source to a streaming application may be recorded. Upon update and service restart, counts may be collected again to view whether data flows through the streaming application as expected. A unique database record count may be viewed (e.g., after reboot or during normal operation) to ensure that no records are accidentally discarded. Data content sampling may also be performed to view whether any data transformations are running normally. Corrective measures (after restart or during normal operation) may also be taken, including replaying discarded database messages or sending alarms.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese invention patent application filed on April 8, 2020, with application number 202080027662.1 and invention name “Ensuring data quality through deployment automation in data streaming applications”. Technical Field

[0002] According to various embodiments, the present disclosure relates to improvements in data streaming platforms, and more particularly to reliability and availability of data streaming during normal operations and deployment / update operations. Background Art

[0003] Data streaming services can provide a variety of information to users. In some cases, such as social media platforms, this may not be a problem if the platform does not make posts available to others for viewing until minutes or even hours later. For example, users may not notice (or care) if there is a service outage that causes some content to be displayed to them with a significant delay or not even displayed to them at all.

[0004] However, in other cases, users of the data streaming service may be particularly interested in viewing new streaming information displayed in a fast and reliable manner. When the data streaming service provides streaming content uninterruptedly (e.g., 24 hours a day, 7 days a week), it may be challenging to provide the highest level of reliability. Technical difficulties or other conditions encountered during normal operation may cause data delays and / or loss to the streaming application. Deployment of updates to the platform streaming service may also result in service interruptions. Applicants recognize that the data streaming service can be improved to provide better reliability during operation both during normal operation and during new code deployment / updates. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 A block diagram of a system including a user system, a front-end server, a back-end server, a data engine, and a database according to some embodiments is shown.

[0006] Figure 2 A diagram is shown relating to the flow of data through a stream processing application in accordance with some embodiments.

[0007] Figure 3 A diagram is shown relating to operations for examining count metrics of a stream processing application in accordance with some embodiments.

[0008] Figure 4 Diagrams related to operations for inspecting data content of a stream processing application are shown in accordance with some embodiments.

[0009] Figure 5 A flow chart related to a method of updating a stream processing application according to some embodiments is shown.

[0010] Figure 6 A flow chart relating to a method of operating a stream processing application and detecting possible problems is shown.

[0011] Figure 7 is a diagram of a computer-readable medium according to some embodiments.

[0012] Figure 8 is a block diagram of a system according to some embodiments. DETAILED DESCRIPTION

[0013] Techniques related to data streaming applications are described, and in particular, techniques that can be used to improve the reliability of data streaming (e.g., ensuring that data is sent to a consumer within a certain amount of time and that no data is dropped). This may be particularly desirable when reporting streaming data related to finance (e.g., records of completed electronic payment transactions). For example, a merchant who has just made a sale may want to see that sale reflected in a streaming data feed of the transaction within a short period of time (e.g., less than a minute). If transactions are not reported in a timely manner (or are never reported due to data loss or mishandling), this may have a negative impact on platform usage.

[0014] Therefore, data streaming applications may need to provide high reliability. A framework is described that allows data streaming applications to provide high reliability both during update operations and during normal (e.g., non-update) operations. For updates, a data pipeline event ID count reflecting the number of messages sent from the data source to the streaming application may be recorded. During updates, operations may be halted for a short period of time (even if in-progress database messages remain in the processing queue) while one or more executable files are updated. After the update, the streaming application may be restarted.

[0015] After a restart, data message counts can be collected to see if the data is flowing through the streaming application as expected. For each change to a record in the database, a unique key (e.g., a unique event ID) can be associated with the change. The unique event ID can be a database sequence number, for example, a unique sequence number assigned when a database record is written / modified for a particular database. Therefore, in various embodiments, for example, if a record is changed 3 times, 3 different unique event IDs will be generated. This is different from other types of database sequence numbers used to uniquely identify database records. Therefore, in various embodiments, the sequence numbers mentioned here can uniquely identify changes to records, because the streaming data can be a log stream of changes to records, where it may make sense to compare the counts of those logs.

[0016] Thus, the unique database record change counts can be checked - for example, after a restart or during normal operation - to ensure that no record changes are accidentally discarded. If the unique event ID counts for data entering the streaming application (e.g., at the beginning of the data pipeline, when loading data from the database) are the same as the unique event ID counts for data leaving the streaming application, this indicates that the streaming application is not discarding any record changes (or is unnecessarily duplicating record changes, such as performing 2 or more inserts when one insert would suffice). In various embodiments, event ID counting can be performed over a specific time window (e.g., one minute), with the expectation that all event ID counts within that window will later be seen to leave the streaming application.

[0017] Data content sampling can also be performed to see if any data processing is functioning properly - for example, to ensure that specific data within a database message is not being dropped, corrupted, or otherwise misprocessed. Typically, a stream processing application can act as a service between one or more source databases and an end user (e.g., a software application that displays data to an individual). Corrective actions can also be taken if post-restart monitoring or ongoing monitoring of normal operations indicates that there is a problem with the stream processing application. These corrective actions can include sending alert messages and replaying database messages (automatically and without human intervention) to ensure better quality of service.

[0018] ***

[0019] This specification includes references to "one embodiment," "some embodiments," or "an embodiment." The appearances of these phrases are not necessarily referring to the same embodiment. The particular features, structures, or characteristics may be combined in any suitable manner consistent with the present disclosure.

[0020] As used herein, the terms "first," "second," and the like are used as labels for the nouns that precede them and do not necessarily imply any type of ordering (eg, spatial, temporal, logical, cardinality, etc.).

[0021] Various components may be described or claimed as being "configured to" perform one or more tasks. In such contexts, "configured to" is used to imply structure by indicating that the component includes structure (e.g., stored logic) that performs the one or more tasks during operation. Thus, the component can be said to be configured to perform the task even if the component is not currently running (e.g., not started). Reference to a component being "configured to" perform one or more tasks expressly does not intend to invoke 35 U.S.C. §112(f) with respect to the component.

[0022] ***

[0023] Go to Figure 1, a block diagram of a system 100 according to various embodiments is shown. In the figure, the system 100 includes user systems 105A, 105B, and 105C. The system 100 also includes a front-end server 120, a back-end server 160, a database 165, a data engine 170, and a network 150. The techniques described herein can be used in the environment of the system 100 as well as in many other types of environments.

[0024] Note, imagine Figure 1 Many other arrangements of the present invention are shown (all figures are the same). Although certain connections (e.g., data link connections) between different components are shown, in various embodiments, there may be additional connections and / or components that are not depicted. As will be appreciated by those skilled in the art, various devices may be omitted from this figure for simplicity - thus, in various embodiments, routers, switches, load balancers, computing clusters, additional databases, servers, and firewalls, etc. may all be present. In this figure, components may be combined with each other and / or separated into one or more systems, as in other figures.

[0025] According to various embodiments, user systems 105A, 105B, and 105C ("user systems 105") may be any user computer system that is capable of potentially interacting with front-end server 120. Front-end server 120 may provide web pages that facilitate one or more services, such as account access and electronic payment transactions (such as those available from PayPal.com). TM The front-end server 120 may also facilitate access to various electronic resources, which may include accounts, data, and various software programs / functions, etc.

[0026] In some cases, the merchant may control user system 105A (as just one example) and may use that system to conduct sales transactions. An application on user system 105A may display a data feed to the merchant—for example, providing a record of electronic payment transactions that have been performed. Merchants may particularly want to see near real-time confirmation of their transactions (e.g., within a minute)—such as might be provided by a data stream listing all transactions made by an account. Of course, consumers may also receive a notification of their transactions (e.g., using PayPal). TM transactions made by an account).

[0027] The front-end server 120 can therefore be any computer system configured to provide access to electronic resources. In various embodiments, this can include providing web content and access to functions provided to network clients (or by other protocols, including but not limited to SSH, FTP, databases and / or API connections, etc.). The services provided can include providing web pages (e.g., in response to HTTP requests) and / or providing interfaces to the functions provided to the back-end server 160 and / or database 165. Database 165 can include various data, such as user account data, system data, and any other information. Of course, in various embodiments, multiple such databases can exist and can be distributed in one or more data centers, cloud computing services, etc. The front-end server 120 can include one or more computing devices, each of which has a processor and a memory. The network 150 can include all or part of the Internet.

[0028] In some embodiments, the front end server 120 may correspond to a server such as a server provided by PayPal. TM The electronic payment transaction service provided, but in other embodiments, the front-end server 120 may correspond to different services and functions. The front-end server 120 and / or the back-end server 160 may have a variety of associated user accounts, allowing users to make payments electronically and receive payments electronically. The user account may have a variety of associated funding mechanisms (e.g., linked bank accounts, credit cards, etc.), and may also maintain a monetary balance in the electronic payment account. Many possible different funding sources can be used to provide the source of funds (credit, check, balance, etc.). User devices (smartphones, laptops, desktops, embedded systems, wearable devices, etc.) can be used to access electronic payment accounts, such as PayPal. TM In various embodiments, amounts other than currency may be exchanged through the front-end server 120 and / or the back-end server 160, including, but not limited to, stocks, merchandise, gift cards, reward points (e.g., from an airline or hotel), etc. In some embodiments, the server system 120 may also correspond to a system that provides functionality such as API access, a file server, or other types of services with user accounts (and in various embodiments, such services may also be provided through the front-end server 120).

[0029] In the illustrated embodiment, database 165 may include a transaction database having records related to various transactions conducted by users of the transaction system. These records may include any number of detailed information, such as any information related to a transaction or an action taken by a user on a web page or an application installed on a computing device (e.g., a PayPal app on a smartphone). Many or all of the records in database 165 are transaction records, including details of a user sending or receiving money (or some other amount, such as credit card reward points, cryptocurrency, etc.). The database information may include the two or more parties involved in the electronic payment transaction, the date and time of the transaction, the monetary amount, whether the transaction is a recurring transaction, the source of funds / type of financing instrument, and any other details. Such information may be used for bookkeeping purposes and for risk assessment (e.g., fraud and risk determinations may be made using historical data; such determinations may be made using system and risk models, which are not described in detail for simplicity). Figure 1 This system and risk model are not depicted in the figure). It is understood that there may be more than one database in the system 100. The additional databases may include a variety of different data other than transaction data. Therefore, any description herein about database 165 may also apply to other (not shown) databases.

[0030] The back-end server 160 may be one or more computing devices, each of which has a memory and a processor that supports a variety of services. The back-end server 160 may be deployed in a variety of configurations. In some cases, all or part of the network service functions enabled by the back-end server 160 can only be accessed through the front-end server 120 (for example, some functions provided by the back-end server 160 may not be publicly accessible through the Internet unless the user accesses through the front-end server 120 or some other type of gateway system).

[0031] The data engine 170 can also be one or more computing devices, each of which has a memory and a processor. In various embodiments, the data engine 170 performs operations related to the stream processing application. In various embodiments, the data engine 170 can therefore send information to and / or receive information from multiple systems (including databases 165, front-end servers 120 and back-end servers 160, as well as other systems). The data engine 170 can therefore run one or more executable processes including the stream processing application 205 (discussed further below).

[0032] Steering Figure 2 , shows a diagram of a system 200 related to data flow through a stream processing application according to various embodiments. The concepts introduced with respect to this figure will be explained in more detail with respect to other figures further below.

[0033] exist Figure 2, streaming messages 203 are sent from database 165 to stream processing application 205 via input stream 202. One or more processing actions may be taken on these messages, which may then be sent as streaming messages 204 via output stream 204 to one or more output queues 225. Stream processing application 205 may thus process a constant stream of incoming data and send out a corresponding constant stream of outbound data.

[0034] The stream processing application may include one or more different executable files and associated data (e.g., configuration files, logs, images, text, etc.). In some cases, the executable file may include a JAR (Java Archive) file, but in other embodiments may be another type of executable file. The executable file may include various computer executable instructions, including but not limited to Java bytecodes and compiled native binary instructions, but may also include additional information (e.g., text, images, configuration information, and / or other media).

[0035] The correctness monitor 210 can interface with the stream processing application 205 to ensure that data is processed correctly. More specifically, the correctness monitor 210 can verify unique message counts from the stream processing application 205 (e.g., to ensure that no messages are dropped), and can also measure content validity (e.g., to ensure that the streaming message content is not corrupted or incorrectly transformed, which can be done through sampling techniques in some embodiments, as checking the content of all and every streaming message may be impractical).

[0036] The checkpoint repository 215 can contain checkpoint data for one or more checkpoints of the stream processing application 205. Checkpoints can be taken before an application update, for example, to allow the stream processing application 205 to automatically roll back to a previously stored version if a newer version of the stream processing application 205 does not appear to be running correctly. The deployer 220 can help deploy new code and perform related operations.

[0037] Steering Figure 3 , shows a diagram 300 related to operations for checking count metrics of a stream processing application 205 according to various embodiments. In various embodiments, the count metrics can be checked to ensure that the stream processing application 205 is functioning properly after an update, or more generally to ensure that the application is performing as expected during normal operation. In this figure, messages (e.g., data from data records in a database 165) can be passed between various components as part of data flows 302, 304, 306, and 308.

[0038] The database adapter 305 can be configured to communicate with a specific source database (e.g., Oracle TM Database, MySQL TM306). The router 315 may then perform one or more operations on the messages in the data stream 306, such as converting or otherwise processing the data messages. The messages are then output to the message queue 320, where further count metrics 307 are collected regarding the number of unique messages.

[0039] The aggregated metrics 309 may then be sent to a count aggregator / detector 330. The detector 330 may calculate whether all expected unique message count metrics match correctly (e.g., for a particular time window), or whether one or more messages may have been dropped (e.g., within the particular time window). If a message was dropped, the message may be replayed from the source database to ensure correct processing, but an alert may also be sent when a count mismatch is detected in operation 335. The alert may be sent to an analyst, for example, who may take steps to resolve the issue. Note that in some embodiments, a message count check is performed on all messages processed by the stream processing application 205 (i.e., the message count check may be performed not only after the application is updated and restarted, but may be performed continuously on all messages during normal operation).

[0040] Go to Figure 4 , shows a diagram 400 related to operations for inspecting data content of a stream processing application 205 according to some embodiments. In various embodiments, the data content at different stages (e.g., pre-processing and post-processing) can be inspected to ensure that the stream processing application 205 operates properly after an update, or more generally to ensure that the application performs as expected during normal operation. Figure 4 Some parts of Figure 3 and may function similarly or identically in various embodiments.

[0041] In operation 435, the data validator 430 performs a sample record query 435 to the database 165. The database 165 can then respond with sample data 440. The sample data 445 can also be received from the message queue 310 at the data validator 430 (before the message is processed by the router 315). The sample data 450 can also be received from the message queue 320 at the data validator 430 (after the message is processed by the router 315). The data validator 430 can then compare the content of various sampled data (for example, for a specific database record) to ensure that the content matches as expected. If an inconsistency is detected, this may indicate that the stream processing application 205 is faulty (for example, during normal operation, a newly updated executable file may cause data corruption, or some other problem may cause data corruption). When an inconsistency is detected, an alarm can be sent in operation 460. The alarm can be sent to an analyst, for example, who can take steps to resolve the problem. Note that in some embodiments, data content checks are performed only on some messages processed by the stream processing application 205. Sampling can be used after an application update (e.g., checking the content of one in a thousand messages, or some other frequency) and restart to ensure application correctness. Data content checking by sampling can also be performed continuously during normal operation.

[0042] Steering Figure 5 , shows a flow chart of one embodiment of a method 500 related to updating a stream processing application 205 according to various embodiments.

[0043] In various embodiments, the processing of the data may be performed by any suitable computer system and / or combination of computer systems (including the data engine 170). Figure 5 Describes the operation.

[0044] However, for convenience and ease of explanation, the operations described below will be discussed simply with respect to the data engine 170 rather than any other system. In addition, the various operations and elements of the operations discussed below may be modified, omitted, and / or used in a different manner or in a different order than indicated. Thus, in some embodiments, the data engine 170 may perform one or more operations while another system may perform one or more other operations.

[0045] In the description Figure 5Before operating on, note that the stream processing application 205 can process one or more data streams. When there are data transformation effects, the data stream can be defined by a tuple [source, destination] or [source, transformation, destination]. In various embodiments, a single data stream / stream processing pipeline does not consume data from more than one source database (for example, the stream processing pipeline can be a simple event processing pipeline rather than a complex event processing pipeline). In various embodiments, a simple event processing pipeline is logically easier to process because event IDs may not necessarily be unique between different databases, and combining data records from different databases for streaming purposes may therefore require additional consistency and validation mechanisms beyond those required for a data pipeline in which only a single source database is used.

[0046] An event ID for a database is more specifically a unique identifier for any changes made to that database. Consider the example of a particular database record that was modified five times at different times of the day. In this example, the database key (e.g., unique record ID) for that record would remain the same—it is the same record, just some of the data inside the record has changed.

[0047] Therefore, all changes to the database can be recorded with a unique event ID associated with each change. This can be thought of as a sequence number in the database "redo" log. As an example, the unique event IDs can be distributed based on the sort order of database commits. In some cases, the unique event IDs of the database pipeline may also correspond to data sources other than the database.

[0048] However, the event ID log associated with the record will be updated for each of the five different changes. Thus, each of the five record changes will have an associated unique event ID - five new event IDs will be generated as a result of the record changes. These unique event IDs can therefore be used to replay database messages, for example when the database message has been discarded somewhere within the pipeline of the data streaming application 205 (this aspect will be discussed in more detail below).

[0049] In operation 510, according to some embodiments, the data engine 170 stops reading data in one or more data pipelines of the stream processing application 205. Thus, reading from a source database used to provide data to the pipelines for processing may be suspended (at least temporarily) as update operations of the stream processing application 205 proceed. In some embodiments, operation 510 may be implemented using a maintenance mode (see further below).

[0050] In operation 520, according to some embodiments, the data engine 170 completes processing of the in-flight data of one or more data pipelines of the stream processing application 205. This may include sending the data through a transformation component such as a router 315 and / or transferring the in-flight data to an output queue (such as a message queue 320).

[0051] Completing processing of the in-flight data may include completely draining one or more message buffers for the stream processing application 205. For example, a particular data pipeline may have 100 unique database records (or a portion thereof) in the message queue 310. These records may all be processed and passed through the stream processing application 205, so that at the end of operation 520, the message queue 310 (and / or other data storage structure) does not have any data still waiting to be processed and transmitted to an end user (e.g., a consuming application and / or a person).

[0052] In various embodiments, suspending the stream processing application 205 may include operation 510 (stop reading data) and operation 520 (complete processing of data in transit). Suspending the stream processing application 205 may also include: stopping execution of one or more executable processes of the stream processing application. In some cases, stopping execution of a process may include killing the process, or may include pausing the process so that it does not execute any additional instructions (but may remain partially or completely in memory when suspended).

[0053] Maintenance mode can also be used in association with suspending the stream processing application 205. Maintenance mode can prevent restarting data pipelines that were suspended due to the maintenance mode being turned on. That is, maintenance mode can suspend multiple data pipelines (e.g., storage structures for data streams from one or more source databases and processed by the stream processing application 205). In normal operation, if a data pipeline is suspended (e.g., in response to the pipeline encountering some transient error), the self-healing feature of the stream processing application 205 can attempt to restart the data pipeline. However, according to various embodiments, this self-healing can cause problems during updates unless it is disabled (via maintenance mode). Similarly, maintenance mode can also prevent the stream processing application from starting new data pipelines.

[0054] In operation 530, according to some embodiments, the data engine 170 saves a backup checkpoint for the stream processing application 205. Saving the backup checkpoint may include storing various state and / or configuration information about the stream processing application 205, and storing one or more executable files of the stream processing application 205. For example, if one or more executable files will be replaced in an update to the stream processing application 205, a copy of one or more (older) executable files of the stream processing application may be saved (the one or more (older) executable files will be replaced by one or more newer executable files). In this way, the stream processing application 205 can be restored to a previous operating state in the event that the update is unsuccessful and / or appears to cause data streaming problems. Note that in various embodiments, data in transit (e.g., which may be stored in the message queue 310 or other location) will be processed and drained from any buffers before the backup checkpoint is saved.

[0055] In operation 540, according to various embodiments, after the suspension, the data engine 170 performs updates, including adding one or more new executable files to the stream processing application 205. The executable files of the stream processing application 205 can be used to perform various tasks, but in some cases, specific source data (e.g., from the database 165) can be obtained and one or more operations can be performed on the source data before sending it, for example, toward an output queue, where the processed data can be used by another application (e.g., a streaming data client that displays the streaming data to a user). Such a streaming data client can be a web application or a mobile phone application, for example, which displays the results of an electronic payment transaction to a user (e.g., PayPal). TM Users may see a record of all their successful and unsuccessful transactions in an app on their phone).

[0056] The operations performed on the source data by the executable file of the stream processing application 205 may include, for example, packaging data from one or more database records into a single data packet targeted to a specific end user (which may be an individual or a software application). The data may also be reformatted according to the needs of the end user. Some data from the source records may be discarded during such reformatting (for example, some data in the database source record may have privacy protection settings associated with it, so the data may be omitted when it is sent by the stream processing application 205).

[0057] The stream processing application 205 may also be configured to provide output information within a certain specified amount of time, for example, according to a service level agreement (SLA). For example, an SLA may specify that 100% (or some other percentage, such as 90%, 99.99%) of the data input to the stream processing application 205 must be output (e.g., to an end user and / or a message queue for an end user) within a certain amount of time (e.g., 15 seconds, 60 seconds, 2 minutes, or some other time period). Such an SLA may be particularly important for ensuring a good customer experience, for example, if a merchant conducts a payment transaction, she may want to see a record of the transaction in her payment application within a relatively short period of time (rather than wondering whether, for example, charging a customer $2,000 was completed correctly).

[0058] In operation 550, according to various embodiments, after the update, the data engine 170 restarts the stream processing application 205. Restarting the stream processing application 205 may include initiating execution of one or more executable processes of the stream processing application that were stopped (e.g., unpausing a paused process, and / or starting execution of a killed process). For example, an executable process that was killed because it was being updated with a new version may be fully restarted, while a paused process (e.g., not being updated) may be resumed in this operation.

[0059] In operation 560, according to various embodiments, the data engine 170 may perform validation on the stream processing application after the restart in operation 550. This validation process may be performed to ensure that the stream processing application operates properly after the update (without this validation process, the stream processing application may roll back to a previous version using a stored checkpoint).

[0060] The verification in operation 560 may include various steps, including a count check aspect to ensure that database records are not lost and a content check aspect to check the data content to verify that the data is being processed correctly (e.g., not corrupted or lost). Thus, the verification may include: for a first time window after the restart, collecting a first plurality of inbound unique event IDs, wherein each of the first plurality of inbound unique event IDs corresponds to a particular database record.

[0061] Collecting the first plurality of inbound unique event IDs may include collecting the event IDs at one or more locations associated with the stream processing application 205. Thus, for example, operation 560 may include collecting the event IDs at the database adapter 305 and / or the message queue 310. Each of these event IDs may be associated with a particular change within the database 165 (e.g., such as a record being added or modified).

[0062] In various embodiments, when data is sent from the database 165 to the stream processing application 205, a corresponding event ID is also included. Thus, each inbound unique event ID can correspond to a specific database record. However, note that since the event ID is associated with a data change, two or more of these event IDs may correspond to the same record. For example, in a set of 25 inbound unique event IDs, 22 of these unique event IDs may correspond to 22 different database records that were added to the database 165, while three of these unique event IDs may correspond to a single pre-existing database record that was modified three different times.

[0063] The first time window after the stream processing application 205 is restarted can be any length of time, but in some embodiments is 60 seconds. For example, immediately after the restart, all inbound unique event IDs for the data streaming application from 0.0 seconds to 60.0 seconds can be recorded, as seen at the database adapter 305. Subsequently, it is expected that all of these unique event IDs will eventually be processed by the stream processing application 205 and sent to the end user (e.g., a software application that consumes the generated streaming data).

[0064] Therefore, the verification in operation 560 also includes: after the end of the first time window, determining whether each of the first plurality of unique event IDs has been collected at a location outbound from the stream processing application and / or the message queue 310 and / or any other determined location. The unique event ID lists of two (or more) different locations are compared, and if an event ID is missing at a later location in the pipeline, it can be inferred that the event ID has been lost and needs to be replayed.

[0065] Determine whether all event IDs within a time window (e.g., 60 seconds after the stream processing application 205 is restarted) can include: waiting for a specific time period before determining whether a given unique event ID is lost. For example, an additional 30 seconds, 60 seconds, 5 minutes, 10 minutes, or some other time period can be used. In particular, after restarting, the stream processing application 205 may be reading a larger amount of data from the database 165 (because data reading is suspended during the update period). Therefore, some initial event IDs in the processing pipeline after restarting may have a longer delay than usual. However, if the inbound event ID seen after a certain amount of time and earlier in the pipeline (e.g., at the database adapter 305) is not seen in the later stages of the pipeline (e.g., the message queue 320), the stream processing application can determine that the data associated with the event ID has been lost and / or improperly processed.

[0066] In various embodiments, if the event ID is not seen in the later stages of the pipeline, the verification process will fail after restart. However, according to various embodiments, if all event IDs within a specific time window are seen in the later stages of the pipeline (e.g., all stages of the pipeline eventually reflect that the database message has experienced this event ID for processing), then the verification process succeeds.

[0067] During verification of the restart process, additional event IDs for additional time windows may also be verified by counting. Thus, a second time window (e.g., after the first time window) may cause inbound event IDs to be collected at one location and then collected again at a downstream location to determine if any event IDs are missing. This process may function similarly to that detailed above.

[0068] Pipeline count metrics can typically be based on a function of a time interval (window). The width of the time interval can be one minute (60 seconds) or some other number. Counts based on such time intervals can be published (e.g., pushed) to an orchestration or monitoring component. Note that typically in event processing, there are two different times that can be associated with an event: event time and processing time. Event time is typically related to the birth time of the event. As the event passes through different parts of the stream processing application, it may be processed. The time at which the event is processed at a component is called the processing time of event E at component C. Therefore, while an event has only a single event (birth time) time, in various embodiments it can have multiple processing times associated with it.

[0069] To track data within a pipeline (e.g., including count validation checks at different pipeline stages), counts of events can be captured based on their event time for a given one-minute wide window. This metric can be captured at each hop in the stream processing pipeline, and the metrics can be compared to determine if there is any data loss. In some embodiments, event time is a convenient tracking mechanism for data count validation because, while the processing time of an event may increase as it progresses down the pipeline, its event time (e.g., birth time) is constant. (Note that the term "data pipeline" or "pipeline" as used herein may include one or more intermediate storage structures for specific data originating from a specific database source and destined for one or more end users.)

[0070] The verification performed in operation 560 may also include comparing the pre-processed data content with the post-processed data content output by the stream processing application. The data content from database 165 (e.g., before the data streaming application 205 performs one or more operations on it) can be compared with the data that has been processed by the data streaming application to determine whether the data is being processed correctly. For example, if the transaction amount of a particular data record read from database 165 (or any other database) is $37.85, the accuracy of the amount can be checked again after processing to ensure that it matches. Figure 4 As shown, a data validator can perform such a check. Data operations performed by the data streaming application 205 can also be checked. For example, if the streaming application should remove a piece of data (e.g., some privacy or financially sensitive data should be deleted, such as a credit card expiration date), the output data from the data streaming application 205 can be checked to ensure that no part of this data is present after processing. Thus, a variety of different data can be checked to ensure correctness. Because data content checking can be an expensive operation, such checks can be performed on a sampled basis for data content (e.g., only once every 1,000 database records, or some other frequency).

[0071] Validation of the stream processing application 205 (e.g., after a restart after an update) can also include comparing count metrics to determine that an appropriate number of database messages (e.g., data from a database record) exist both before and after processing by the stream processing application. If the count metrics do not match, this may indicate that the stream processing application is discarding one or more messages, which may indicate an error. In other words, in various embodiments, the inbound message count and the outbound message count of the stream processing application 205 should match.

[0072] Thus, performing the validation in operation 560 may include comparing a count of unique database messages inbound to the stream processing application at a first point before these messages are processed with a count of unique database messages outbound from the stream processing application at a second point after the stream processing application processes these messages. Thus, for example, the message count may be checked at the database adapter 305 and / or the message queue 310 and compared to the message count at the (outbound) message queue 320 (note that in some cases, all three message counts may be compared).

[0073] Verifying the stream processing application 205 may also include verifying that the process state of the application does not include any errors. For example, upon restart, the data engine 170 may confirm that each of the one or more executable processes is up and running (eg, actively executed) and that there are no errors in these processes.

[0074] In operation 570, according to various embodiments, in response to the unsuccessful validation of operation 560, the data engine 170 rolls back the stream processing application 205 to the stored checkpoint. For example, a rollback may be performed if pipeline metrics do not meet expectations, if message data sampling shows that data content does not match expectations, and / or if message counts are incorrect.

[0075] When rolling back the stream processing application 205 to a stored checkpoint, the data engine 170 can remove the specific newly added executable file and replace it with the saved copy of the (older) executable file that has been replaced by the newly added executable file. More than one executable file may be replaced during a rollback. The configuration data can also be rolled back using the configuration data saved when the backup checkpoint was taken. Therefore, when a restart is deemed to have failed verification (e.g., a unique message count check indicates a problem and / or content sampling indicates a problem), one or more settings of the data streaming application 205 can be changed (restored) back to the previous settings.

[0076] Note that when the validation in operation 560 succeeds, the stream processing application will not actually roll back. Instead, according to various embodiments, operation 580 will be performed.

[0077] In operation 580, in response to the successful verification of operation 560, the data engine 170 determines not to roll back the stream processing application 205 to the stored checkpoint and allows the stream processing application to operate using the one or more added new executable files. When resuming normal operation of the stream processing application 205, the one or more data pipelines that were suspended can be resumed, and new pipelines can be opened again (new pipelines can be disabled during the update process and / or when the maintenance mode is turned on).

[0078] Steering Figure 6 , shows a flow chart of one embodiment of a method 600 according to various embodiments involving operating a stream processing application 205 and detecting possible problems (eg, for normal / ongoing operations other than update operations).

[0079] In various embodiments, the processing of the data may be performed by any suitable computer system and / or combination of computer systems (including the data engine 170). Figure 6 Describes the operation.

[0080] However, for convenience and ease of explanation, the operations described below will be discussed simply with respect to data engine 170 rather than any other system. In addition, various operational elements discussed below may be modified, omitted, and / or used in a different manner or in a different order than indicated. Thus, in some embodiments, data engine 170 may perform one or more operations while another system may perform one or more other operations. Note that any and all techniques and structures discussed above (e.g., with respect to method 500) are applicable to method 600, and vice versa, according to various embodiments.

[0081] In operation 610, according to some embodiments, the data engine 170 receives the content of a plurality of data records from one or more source databases. The content may be received at the stream processing application 205. In various embodiments, the data record may include any kind of data, but in some embodiments includes information related to an electronic payment transaction (e.g., identifying the buyer / source of funds, the seller / destination of funds, the transaction amount, the type of financing instrument, and / or other information). In operation 610, an entire database record (e.g., a row in a table) may be received, or only a portion of the record may be received. According to various embodiments, the received content will be processed by the stream processing application 205 as described below.

[0082] In operation 620, according to some embodiments, the data engine 170 collects a first plurality of inbound unique event IDs for a first time window, wherein each of the first plurality of inbound unique event IDs corresponds to a specific database record change of one or more source databases. By collecting these unique event IDs, the correct operation of the data processing application 205 can be verified, as further explained below.

[0083] The collection process can be compared with the above Figure 5 However, according to various embodiments, in Figure 6In the method, the stream processing application 205 is not updated and restarted, but these techniques can be applied to the "steady state" of the stream processing application. The counting process of unique event IDs can help verify that no database records (and / or record portions or other data) are discarded when data is migrated to and through the stream processing application 205. Using unique attributes such as event IDs to perform this counting may be particularly useful because, in some embodiments, duplicate records may be inadvertently introduced into the inbound data stream of the stream processing application 205. Therefore, what is important is not the total number of records (or portions thereof) inbound to the conversion component, but the total unique (net aggregate) number of records. As described above, in various embodiments, the event ID is unique within a specific data source attached to the data pipeline. In operation 620, each of the first plurality of inbound unique event IDs corresponding to a specific database record of one or more source databases may mean that each unique event ID corresponds to a change in a database record (e.g., adding a new record, modifying an existing record). Each event ID may correspond to a different record (but more than one event ID may correspond to the same record, for example, if the record is modified twice within a time window).

[0084] According to some embodiments, in operation 630, the data engine 170 processes the contents of the plurality of data records via a transformation component of a stream processing application. The transformation component may perform any number or type of operations on the data sent to it. In one embodiment, the transformation component is a router 315 (e.g., the transformation component may reside on a data path between a source database and one or more outbound message queues where the data may be retrieved by an end user (e.g., a software application).

[0085] Processing what is sent to the conversion component can include masking certain types of data. For example, a database record sent to the conversion engine may include different categories of information within the PCI (Payment Card Industry) compliance standard. The primary account number or expiration date associated with a payment card may be allowed to be stored on the data engine 170 and database 165, but not on a target system that will later retrieve the data from the outbound message queue. Such information may be masked (e.g., replaced with zeros or otherwise obfuscated / eliminated) before the conversion component sends out the data record (or portion thereof). Other policies besides PCI may also dictate what information the conversion component masks.

[0086] The conversion component can also change one or more types of data according to various rules. Floating point numbers can be truncated to a specific length (e.g., two decimal places), spaces can be removed from data fields, or delimiters can be replaced with something different. In general, in various embodiments, any arbitrary data conversion can be performed by the conversion component.

[0087] However, note that in some embodiments, no conversion component is required. Thus, router 315 (or another conversion component) may be omitted in various embodiments, and in such embodiments, data may be processed by stream processing application 205 without applying conversion.

[0088] In operation 640, according to some embodiments, the data engine 170 performs validation on the output content from the conversion component to a set of one or more outbound message queues. According to various embodiments, performing validation may include determining whether each of the first plurality of unique event IDs has been collected at a location outbound from the stream processing application. Therefore, a first number of unique data records (or portions thereof) may be counted and compared before and after processing by the conversion component. In various embodiments, it is expected that these counts are equal-no records should be discarded. However, if the post-processing count indicates a smaller number, this may indicate that the conversion component (or perhaps some other component) has discarded one or more records. According to various embodiments, if the pre-processing count of the database record is higher than the post-processing count, this indicates that one or more messages have been lost or improperly processed (thus, the validation process will fail). However, if the unique record counts match, the validation process is considered successful.

[0089] For verification purposes, unique event IDs may be compared within a single source database / data pipeline (e.g., a simple pipeline with only one source). Thus, operation 640 may include a verification process performed on each pipeline—for example, a pre-processing check and a post-processing check may be performed on all unique event IDs within a certain time window for data pipeline #1 / source database #1, then unique event IDs for data pipeline #2 / source database #2 may be checked, and so on. The stream processing application 205 may have many different pipelines. Verification may succeed for one pipeline but fail for another; thus, in some cases, corrective action may be taken for one data pipeline but not for another.

[0090] In operation 650, based on the validation failure, the data engine 170 may take corrective action regarding the operational state of the stream processing application. If one or more messages have been discarded, one corrective action may be to send an alert message regarding the discarded messages. For example, an administrator, software support specialist, or other person may receive an email, text SMS message, or some other type of communication containing detailed information about the discarded message(s).

[0091] Another corrective action is to replay one or more of the event IDs (e.g., source database messages) corresponding to the changes in the database 165. Consider a scenario where each database record is stamped with a unique submission time (e.g., a unique time when a database records the completion of an electronic payment transaction from a fund payer to a fund receiver). If one of the event IDs is discarded within the stream processing application 205, the event ID can be replayed until its associated data is successfully detected to have left the stream processing application 205 (in various embodiments, for example, when the event ID appears at the final outbound location associated with the application (e.g., message queue 320)).

[0092] Thus, in some embodiments, replay of lost event IDs continues indefinitely until 100% of unique inbound messages are reflected as being issued by a component, such as router 315. If one or more messages are dropped again, stream processing application 205 can continue to replay them until all messages are processed correctly, which can help ensure 100% reliability for end users, such as applications that can display completed transaction results to users of electronic payment services.

[0093] Note that in various embodiments, messages can be replayed from any different stage of the data pipeline. Thus, if a unique event ID is seen at database adapter 305 and message queue 310, but is not detected later at message queue 320, it can be determined that the message has been lost / mishandled by router 315 (using Figure 3 The architecture of is taken as an example). The lost messages can then be replayed from the message queue 310 without having to go back to the source database 165 (for example) to retrieve the lost database messages. This flexible architecture can reduce database reads for lost messages. A log of inbound data can be kept at each pipeline stage to facilitate replay from each pipeline stage.

[0094] A latency metric may also be associated with a given database message / event ID. For example, an SLA for a stream processing application 205 may specify that 99.9% (or some other amount) of messages must be delivered within 60 seconds (or some other time period). The SLA may also state that 100% of messages must be delivered within 120 seconds (or some other time period). If replaying a message (especially if the message is replayed more than once) results in a violation of this SLA time, a latency alert message may be sent to an administrator or other person to let them know that there may be a problem with the streaming application and that manual intervention is required. Message replay may also be multi-threaded - for example, a stream processing application may use different process threads so that new messages can continue to be processed even if old messages are being replayed, so that problems with a particular window message frame do not significantly affect future streaming operations.

[0095] In various embodiments, messages lost within the database pipeline will continue to be replayed until the message is successfully processed and output by the stream processing application 205. For example, in the event that a software bug causes a particular message to be repeatedly discarded, the latency metric for that message may continue to rise because the message failed to be processed correctly, and an alert may be sent for that message. If manual intervention subsequently fixes the underlying problem (e.g., through a bug fix, restarting certain pipeline components, etc.), the old message may be successfully processed. In various embodiments, because the stream processing application 205 is expected to process 100% of the messages, high latency for one or more messages may indicate a problem that needs to be addressed. (These latency metrics also apply after a restart and for updates - if one or more messages / event IDs have a high enough latency metric, then the update may be rolled back.)

[0096] In operation 660, based on the successful verification, the data engine 170 can determine that no corrective action is taken. Therefore, if the unique message count before data processing by the conversion component matches the post-processing unique message count, the data engine 170 can determine that the stream processing application 205 is operating correctly. If the application is operating correctly, corrective action (such as an alarm) may not be required, and therefore no additional steps need to be taken to respond to the verification process. Instead, the system can continue to operate as planned, and additional verification processes can be performed periodically. In various embodiments, the verification process can be repeated indefinitely so that the stream processing application 205 is continuously monitored. The process outlined in method 600 can be performed for a continuous time window, such as every 60 seconds (or some other time interval, larger or smaller), thereby ensuring and enhancing the reliability of the stream processing application 205.

[0097] Note that while many of the examples herein discuss streaming enhancements with respect to a source database having electronic transaction information, the techniques of this specification may be generalized to any number of other types of data streaming platforms.

[0098] In one embodiment, a system involving operating a stream processing application includes: a processor; a memory having instructions stored thereon, the instructions being executable to cause the system to perform operations, the operations including: receiving the contents of a plurality of data records from one or more source databases at the stream processing application; collecting a first plurality of inbound unique event IDs for a first time window, wherein each of the first plurality of inbound unique event IDs corresponds to a specific database record in one of the one or more source databases; processing the contents of the plurality of data records through a transformation component of the stream processing application; after the first time window ends, performing verification on output contents from the transformation component to a set of one or more outbound message queues, wherein performing verification includes determining whether each of the first plurality of unique event IDs has been collected at a location outbound from the stream processing application; based on a verification failure, taking corrective action regarding the operational state of the stream processing application; and based on a verification success, deciding not to take corrective action.

[0099] In various embodiments of the above system, the corrective action includes sending an alarm message about the operating status of the stream processing application; the corrective action includes replaying messages corresponding to the first time window from one or more source databases of the time window; each of the one or more source databases corresponds to a different data pipeline of the stream processing application; the conversion component is a router component residing between the source database and one or more outbound message queues; processing the content by the conversion component includes masking specific types of data; the masked specific types of data include detailed information related to financing instruments for electronic payment transactions; and / or processing the content by the conversion component includes changing one or more types of data.

[0100] In another embodiment, a method involving operating a stream processing application includes: receiving the contents of multiple data records from one or more source databases at the stream processing application; receiving the contents of multiple data records from one or more source databases at the stream processing application; collecting a first plurality of inbound unique event IDs for a first time window, wherein each of the first plurality of inbound unique event IDs corresponds to a specific database record of one of the one or more source databases; processing the contents of the multiple data records through a transformation component of the stream processing application; after the end of the first time window, performing verification on the output content from the transformation component to a set of one or more outbound message queues, wherein performing the verification includes determining whether each of the first plurality of unique event IDs has been collected at a location outbound from the stream processing application; and taking corrective action regarding the operational state of the stream processing application based on a verification failure.

[0101] In various embodiments of the above method, the corrective action includes sending an alarm message about the operating status of the stream processing application; each of the first plurality of unique event IDs includes a unique sequence number corresponding to data write or data modification for one of the one or more databases; a portion of the content of multiple data records from one or more source databases includes data from newly added records, while another portion includes data from modified records; the corrective action includes replaying messages from one or more source databases corresponding to the first time window in the time window; collecting the first plurality of inbound unique event IDs is performed at at least one of a database adapter or one or more inbound message queues, the database adapter or the one or more inbound message queues are configured to receive data and provide it to the transformation component; and / or the one or more source databases include a transaction database, which stores database records indicating multiple electronic transactions between multiple users of the electronic service.

[0102] In another embodiment, a non-transitory computer-readable medium stores instructions that, when executed by a computer system, cause the computer system to perform operations, including: receiving the contents of multiple data records from one or more source databases at a stream processing application; collecting a first plurality of inbound unique event IDs for a first time window, wherein each of the first plurality of inbound unique event IDs corresponds to a specific database record of one of the one or more source databases; processing the contents of the multiple data records through a transformation component of the stream processing application; after the first time window ends, performing verification on output contents from the transformation component to a set of one or more outbound message queues, wherein performing the verification includes determining whether each of the first plurality of unique event IDs has been collected at a location outbound from the stream processing application; based on a verification failure, taking corrective action regarding the operational state of the stream processing application, wherein the corrective action includes replaying messages from the one or more source databases; and based on a verification success, deciding not to take corrective action.

[0103] In various embodiments of the above-mentioned non-transitory computer-readable medium, the corrective action also includes sending an alarm message about the operational status of the stream processing application; collecting a first plurality of inbound unique event IDs for a first time window includes removing any duplicate lists of inbound unique event IDs; processing the content through a transformation component includes masking specific types of data; or one or more outbound message queues are configured to provide processed data to one or more software applications.

[0104] Computer readable medium

[0105] Steering Figure 7 , a block diagram of an embodiment of a computer readable medium 700 is shown. The computer readable medium may store Figure 5 and 6Thus, in one embodiment, instructions corresponding to the data engine 170 may be stored on a computer-readable medium 700.

[0106] Note that, more generally, the program instructions may be stored on a non-volatile medium such as a hard disk or flash drive, or may be stored in any other well-known volatile or non-volatile storage medium or device, such as ROM or RAM, or provided on any medium capable of launching program code, such as compact disc (CD) media, DVD media, holographic storage, network storage, etc. In addition, the program code or portions thereof may be transmitted and downloaded from a software source (e.g., via the Internet) or from another server, as is well known, or transmitted through any other well-known conventional network connection (e.g., extranet, VPN, LAN, etc.) using any well-known communication media and protocols (e.g., TCP / IP, HTTP, HTTPS, Ethernet, etc.). It will also be understood that the computer code used to implement aspects of the present invention may be implemented in any programming language that can be executed on a server or server system, such as, for example, in C, C+, HTML, Java, JavaScript, or any other scripting language, such as Perl. Note that, as used herein, the term "computer-readable medium" refers to a non-transitory computer-readable medium.

[0107] Computer Systems

[0108] exist Figure 8 , one embodiment of a computer system 800 is shown. Various embodiments of the system may be included in the front-end server 120, the back-end server 160, the data engine 170, or any other computer system.

[0109] In the illustrated embodiment, the system 800 includes at least one instance of an integrated circuit (processor) 810 coupled to an external memory 815. In one embodiment, the external memory 815 may form a main memory subsystem. The integrated circuit 810 is coupled to one or more peripheral devices 820 and the external memory 815. A power supply 805 is also provided that provides one or more supply voltages to the integrated circuit 810 and one or more supply voltages to the memory 815 and / or the peripheral devices 820. In some embodiments, more than one instance of the integrated circuit 810 may be included (and more than one external memory 815 may also be included).

[0110] The memory 815 may be any type of memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR6, etc.), SDRAM (including mobile versions of SDRAM, such as mDDR6, etc., and / or low power versions of SDRAM, such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. One or more memory devices may be coupled to a circuit board to form a memory module such as a single inline memory module (SIMM), a dual inline memory module (DIMM), etc. Alternatively, these devices may be mounted with the integrated circuit 810 in a chip-on-chip configuration, a package-on-package configuration, or a multi-chip module configuration.

[0111] Depending on the type of system 800, the peripheral device 820 may include any desired circuit. For example, in one embodiment, the system 800 may be a mobile device (e.g., a personal digital assistant (PDA), a smart phone, etc.) and the peripheral device 820 may include devices for various types of wireless communications (e.g., Wi-Fi, Bluetooth, cellular, global positioning system, etc.). The peripheral device 820 may include one or more network access cards. The peripheral device 820 may also include additional storage devices, including RAM storage devices, solid-state storage devices, or disk storage devices. The peripheral device 820 may include user interface devices, such as display screens (including touch display screens or multi-touch display screens), keyboards or other input devices, microphones, speakers, etc. In other embodiments, the system 800 may be any type of computing system (e.g., a desktop personal computer, a server, a laptop, a workstation, a desktop all-in-one, etc.). The peripheral device 820 may therefore include any networking or communication device. As a further explanation, in some embodiments, the system 800 may include multiple computers or computing nodes (e.g., computing clusters, server pools, cloud computing systems, etc.) configured to communicate together.

[0112] ***

[0113] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even if only a single embodiment is described with respect to a particular feature. Unless otherwise stated, the examples of features provided in the present disclosure are intended to be illustrative rather than restrictive. The above description is intended to cover such alternatives, modifications, and equivalents that are obvious to those skilled in the art having the benefit of the present disclosure.

[0114] The scope of the present disclosure includes any feature or combination of features disclosed herein (explicitly or implicitly), or any generalization thereof, whether or not it mitigates any or all of the problems solved by the various described embodiments. Accordingly, new claims may be made during the prosecution of the present application (or an application claiming priority thereto) for any such combination of features. In particular, with reference to the appended claims, features of dependent claims may be combined with features of the independent claims, and features of individual independent claims may be combined in any appropriate manner, and not merely in the specific combinations listed in the appended claims.

Claims

1. A system, include: processor; and A memory having instructions stored thereon, wherein the instructions can be executed to cause the system to perform operations, the operations comprising: Suspending the stream processing application, wherein suspending the stream processing application comprises: stopping reading data from a source database and draining one or more message buffers of the stream processing application; Saving a backup checkpoint for the stream processing application; After the suspending, performing an update of the stream processing application, the update comprising adding one or more new executable files to the stream processing application; After the updating, restarting the stream processing application; After the restart, performing verification on the stream processing application, including: for a first time window after the restart, collecting a first plurality of inbound unique event IDs, wherein each of the first plurality of inbound unique event IDs corresponds to a particular database record of the source database and is associated with a change to the particular database record; After the first time window ends, determining whether each of the first plurality of inbound unique event IDs has been collected at a location outbound from the stream processing application; and comparing the pre-processed data content with the post-processed data content output by the stream processing application; and In response to the verification being unsuccessful, rolling back the stream processing application to the backup checkpoint; or In response to the verification being successful, it is determined not to roll back the stream processing application to the backup checkpoint, and the stream processing application is allowed to operate using the one or more new executable files that are added.

2. The system of claim 1, wherein determining whether each of the first plurality of inbound unique event IDs has been collected at a location outbound from the stream processing application include: Waits a specific period of time before determining whether a given unique event ID is lost.

3. The system of claim 2, wherein the operation further include: The stream processing application is placed in a maintenance mode prior to the restarting, wherein the maintenance mode prevents restarting of a data pipeline that was suspended due to the maintenance mode being turned on. 4 . The system of claim 3 , wherein the maintenance mode prevents the stream processing application from starting new data pipelines.

5. The system of claim 1, wherein verification is performed on the stream processing application include: Verifying that the processing state of the stream processing application does not include any errors.

6. The system of claim 1, wherein the backup checkpoint is saved include: A copy of a first executable file of the stream processing application is saved, the first executable file being replaced by a specific new executable file of the one or more new executable files.

7. The system of claim 6, wherein the stream processing application is rolled back to the backup checkpoint include: The specific new executable file is removed and replaced with the saved copy of the first executable file.

8. The system of claim 1, wherein terminating the stream processing application include: The execution of one or more executable processes of the stream processing application is stopped.

9. The system of claim 8, wherein restarting the stream processing application include: Initiating execution of the stopped one or more executable processes of the stream processing application.

10. A method for updating a stream processing application, include: Suspending the stream processing application, wherein suspending the stream processing application comprises: stopping reading data from a source database and draining one or more message buffers of the stream processing application; Saving a backup checkpoint for the stream processing application; After the suspending, performing an update to the stream processing application, the update comprising adding one or more new executable files to the stream processing application; After the updating, restarting the stream processing application; After the restart, performing verification on the stream processing application, including: For a first time window after the restart, collecting a first plurality of inbound unique event IDs, wherein each of the first plurality of inbound unique event IDs corresponds to a particular database record of the source database and is associated with a change to the particular database record; and After the first time window ends, determining whether each of the first plurality of inbound unique event IDs has been collected at a location outbound from the stream processing application; comparing the pre-processed data content with the post-processed data content output by the stream processing application; and In response to the verification being unsuccessful, rolling back the stream processing application to the backup checkpoint; or In response to the verification being successful, it is determined not to roll back the stream processing application to the backup checkpoint, and the stream processing application is allowed to operate using the one or more new executable files that are added.

11. The method of claim 10, wherein determining whether each of the first plurality of inbound unique event IDs has been collected occurs at an outbound message queue to which the stream processing application has transmitted outbound data.

12. The method of claim 10, wherein performing the verification further include: For a second time window after the first time window following the restart, collecting a second plurality of inbound unique event IDs, wherein each of the second plurality of inbound unique event IDs corresponds to a second specific database record; and After the second time window ends, it is determined whether each of the second plurality of inbound unique event IDs has been collected at a location outbound from the stream processing application.

13. The method of claim 10, wherein the first plurality of inbound unique event IDs are collected via a database adapter coupled to the source database.

14. The method of claim 13, wherein the source database comprises a transaction database storing database records indicating a plurality of electronic transactions between a plurality of users of an electronic service.

15. The method of claim 10, wherein the location outbound from the stream processing application is a message queue configured to store outbound data for a plurality of pipelines.

16. The method of claim 10, wherein the stream processing application is configured to output streaming data within a specified amount of time after receiving information sent to the stream processing application.

17. A non-transitory computer readable medium storing instructions, which when executed by a computer system causes the computer system to perform operations, the operations include: Suspending the stream processing application, wherein suspending the stream processing application comprises: stopping reading data from a source database and draining one or more message buffers of the stream processing application; Saving a backup checkpoint for the stream processing application; After the suspending, performing an update of the stream processing application, including adding one or more new executable files to the stream processing application; After the updating, restarting the stream processing application; After the restart, performing verification on the stream processing application, including: for a first time window after the restart, collecting a first plurality of inbound unique event IDs, wherein each of the first plurality of inbound unique event IDs corresponds to a particular database record of the source database and is associated with a change to the particular database record; After the first time window ends, determining whether each of the first plurality of inbound unique event IDs has been collected at a location outbound from the stream processing application; and comparing the pre-processed data content with the post-processed data content output by the stream processing application; and In response to the verification being unsuccessful, rolling back the stream processing application to the backup checkpoint; or In response to the verification being successful, it is determined not to roll back the stream processing application to the backup checkpoint, and the stream processing application is allowed to operate using the one or more new executable files that are added.

18. The non-transitory computer readable medium of claim 17, wherein the backup checkpoint is saved include: A copy of one or more configuration settings is saved for the stream processing application, the one or more configuration settings being replaced by one or more new configuration settings.

19. The non-transitory computer readable medium of claim 17, wherein the one or more new executable files include content other than executable code.

20. The non-transitory computer readable medium of claim 17, wherein the backup checkpoint is saved include: A copy of a first executable file of the stream processing application is saved, the first executable file being replaced by a specific new executable file of the one or more new executable files.