Data stream processing method and device, program product and electronic equipment
By saving the source state and operator state of the data stream in the stream computing system and using the preset database and distributed state mechanism, the problem of low accuracy of calculation results after a failure of the stream computing system is solved, and the rapid recovery and consistency of data stream calculation results are achieved.
Patent Information
- Application Number
- CN202510882574.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
Smart Images

Figure CN120804050A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data monitoring, in particular to a data stream processing method and device, a program product and an electronic device. BACKGROUND
[0002] With the rapid development of digital technology, in the field of big data processing, the requirement for real-time computing of data stream is higher and higher, and the data stream engine needs to reasonably utilize cluster resources while meeting the real-time computing business requirements, so as to realize the monitoring of data stream to prevent faults in the data stream processing process.
[0003] Since the input data of stream computing is unbounded, there are uncertain factors such as message arrival delay, stream computing system delay, arrival order disorder, and unknown number / size of arrival data in the data stream computing system. In addition, like other types of distributed applications, stream computing systems are also affected by various unexpected factors and may fail, for example, traffic surge, network jitter, and cloud service resource allocation failure. After a failure occurs, the stream computing system needs to re-execute the computing task. However, under the premise of uncertain input order of stream data, it is difficult to design a fault-tolerant mechanism, so the accuracy of the computing result determined by re-executing the computing task of the data stream after a failure is low.
[0004] At present, there is no effective solution to the above problems. SUMMARY
[0005] The present application provides a data stream processing method, device, program product and electronic device to at least solve the technical problem of low accuracy of the computing result determined by re-executing the computing task of the data stream based on prior art after a failure in the processing of the data stream by the stream computing system.
[0006] According to one aspect of the present application, a data stream processing method is provided, comprising: obtaining a data tag of a data stream, wherein the data tag is used to represent the correlation between the computing result generated by the stream computing system after the data stream is input into the stream computing system and the input order of the data stream; in the case that the data tag is a target tag, storing first information corresponding to the data stream to a preset database, wherein the target tag is used to represent that there is an association relationship between the computing result and the order of the data stream input into the stream computing system, and the first information is used to represent the source state and the operator state corresponding to the data stream; in the case that a fault of the stream computing system is detected, determining a target result based on the first information, wherein the target result is the computing result generated after the data stream re-acquired based on the first information is input into the stream computing system.
[0007] Optionally, the step of storing the first information corresponding to the data stream to the preset database comprises: taking a consumption position of the data stream in a data source queue as a source state; taking an intermediate calculation result of the data stream generated based on a preset operator in the stream computing system as an operator state; packing the source state and the operator state to obtain the first information; and storing the first information to the preset database periodically based on a first parameter, a second parameter, a third parameter and a fourth parameter, wherein the first parameter is used to represent a time window for periodically storing the first information, the second parameter is used to represent an execution mode of storing the first information to the preset database, the third parameter is used to represent a maximum operation duration of a storage operation corresponding to the preset database at one time, and the fourth parameter is used to represent a maximum number of operation of executing the storage operation in parallel.
[0008] Optionally, the step of determining the target result based on the first information comprises: determining first data based on the source state in the first information, wherein the first data is source data re-read in the data source queue based on the consumption position corresponding to the source state; and determining the target result based on the first data and the operator state in the first information.
[0009] Optionally, after determining the target result based on the first information, the processing method of the data stream further comprises: writing the target result to a target system by a preset manner, wherein the target system is a downstream system connected with the stream computing system, and the preset manner comprises at least one of the following: an idempotent writing manner, used to write the target result to the target system based on a primary key-based update strategy; and a transactional writing manner, used to write the target result to the target system based on a two-phase commit manner through a preset interface.
[0010] Optionally, after obtaining the data tag of the data stream, the processing method of the data stream further comprises: in a case where the data tag is a target tag, obtaining a distribution state and second information of the data stream, wherein the distribution state is used to represent an elastic distributed state formed after the data stream is divided into at least two or more batches of sub-data and processed in batches, and the second information is used to represent a dependency relationship between the batches of sub-data; and storing the distribution state and the second information to a target log, wherein the target log is a prewrite log corresponding to the data stream.
[0011] Optionally, after storing the distribution state and the second information to the target log, the processing method of the data stream further comprises: analyzing the target log to obtain L logs, wherein L is a positive integer; de-duplicating the L logs to obtain M logs, wherein M is a positive integer less than or equal to L; and determining the target result based on the M logs obtained after de-duplication.
[0012] Optionally, before the data tag of the data stream is acquired, the processing method of the data stream further includes: acquiring third information, fourth information and fifth information, wherein the third information is used to represent a running state of a target program, the fourth information is used to represent a collection efficiency of data in a data source queue, and the fifth information is used to represent a data warehouse efficiency of a preset database corresponding to the data stream, and the target program is a program for monitoring the data stream; and the third information, the fourth information and the fifth information are used to determine alarm information, wherein the alarm information is used to prompt an operation and maintenance personnel that there is a fault in the processing process of the data stream.
[0013] According to another aspect of the present application, a processing device of a data stream is further provided, including: a first acquisition unit, configured to acquire a data tag of a data stream, wherein the data tag is used to represent a correlation between a calculation result generated by a stream computing system and an input order of the data stream after the data stream is input into the stream computing system; a first storage unit, configured to store first information corresponding to the data stream into a preset database in a case where the data tag is a target tag, wherein the target tag is used to represent that there is an association relationship between the calculation result and the order in which the data stream is input into the stream computing system, and the first information is used to represent a source state and an operator state corresponding to the data stream; and a first determination unit, configured to determine a target result based on the first information in a case where it is detected that there is a fault in the stream computing system, wherein the target result is a calculation result generated after a data stream reacquired based on the first information is input into the stream computing system.
[0014] According to another aspect of the present application, a computer program product is further provided, and the computer program product stores a computer program, wherein the computer program controls the computer program product to execute the processing method of the data stream of any one of the above aspects when the computer program runs.
[0015] According to another aspect of the present application, an electronic device is further provided, and the electronic device includes one or more processors and a memory, and the memory is used to store one or more programs, wherein the one or more programs enable the one or more processors to implement the processing method of the data stream of any one of the above aspects when the one or more programs are executed by the one or more processors.
[0016] In the present application, first, the data tag of the data stream is acquired, wherein the data tag is used to represent the correlation between the calculation result generated by the stream computing system after the data stream is input into the stream computing system and the input order of the data stream, then, in the case that the data tag is a target tag, the present application stores the first information corresponding to the data stream into a preset database, wherein the target tag is used to represent that there is an association relationship between the calculation result and the order of the data stream input into the stream computing system, and the first information is used to represent the source state and the operator state corresponding to the data stream, and then, in the case that it is detected that the stream computing system has a fault, the present application determines the target result based on the first information, wherein the target result is the calculation result generated after the data stream reacquired based on the first information is input into the stream computing system.
[0017] From the above, in the case that there is an association relationship between the calculation result of the data stream and the order of the data stream input into the stream computing system, the present application achieves the purpose of quickly recovering the calculation result of the data stream based on the first information after the stream computing system fails by saving the first information (including the source state and the operator state) corresponding to the data stream in the preset database, thereby enhancing the fault tolerance of the data stream computing system and the reliability of data processing, achieving the technical effect of quickly reproducing the calculation result of the data stream, and further solving the technical problem that the accuracy of the calculation result determined by re-executing the calculation task of the data stream based on the prior art after the failure is low in the case that the processing process of the stream computing system on the data stream has a fault. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and the explanation of the present application, and do not constitute improper limitations on the present application. In the drawings:
[0019] Figure 1 is a flow chart of an optional data stream processing method according to an embodiment of the present application;
[0020] Figure 2 is a framework diagram of an optional end-to-end data consistency implementation method according to an embodiment of the present application;
[0021] Figure 3 is a flow chart of an optional end-to-end data three-layer monitoring method according to an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of an optional stream data processing device according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of the present application.
[0025] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described accompanying drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device including a series of steps or units does not necessarily have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product, or device.
[0026] It should also be noted that the relevant information (including the login information of the target user and the relevant information of the system resources) and the data (including but not limited to the data for display and the analyzed data) involved in the present application are all information and data authorized by the user or authorized by all parties. For example, an interface is provided between the system and the relevant users or institutions. Before obtaining the relevant information, the interface needs to send a request to the aforementioned user or institution, and after receiving the consent information fed back by the aforementioned user or institution, the relevant information is obtained.
[0027] In addition, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant information and the relevant data involved in the present application all comply with the relevant laws, regulations, and standards of the relevant regions, and necessary security measures are taken, without violating public order and good customs. In addition, the present application provides a corresponding operation portal for the user to choose to authorize or refuse to authorize. If the user chooses to refuse to authorize, the corresponding expert decision-making process is entered.
[0028] According to the embodiments of the present application, an embodiment of a data flow processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0029] The application provides a processing system of a data stream (processing system for short) for executing a processing method of the data stream in the application, Figure 1 is a flow chart of an optional processing method of a data stream according to an embodiment of the application, as Figure 1 shown, the method comprises the following steps:
[0030] Step S101, obtaining a data tag of a data stream, wherein the data tag is used to represent the correlation between the calculation result generated by a stream computing system and the input order of the data stream after the data stream is input into the stream computing system.
[0031] Optionally, the data stream is a continuous real-time data sequence.
[0032] Optionally, the application judges the correlation between the calculation result generated by the stream computing system and the input order of the data stream, and in the case that the data tag is a target tag, that is, the calculation result is affected by the input order, the subsequent snapshot (that is, the first information) corresponding to the data stream is controlled by Spark (a general computing engine for large-scale data stream processing) and / or Flink (a distributed processing engine capable of performing state calculation on bounded and unbounded data streams) to perform persistent storage, thereby ensuring the purpose of quickly recovering the calculation result of the data stream with the target data tag in the case that the stream computing system fails.
[0033] Step S102, in the case that the data tag is a target tag, storing the first information corresponding to the data stream into a preset database, wherein the target tag is used to represent that there is an association relationship between the calculation result and the order of the data stream input into the stream computing system, and the first information is used to represent the source state and the operator state corresponding to the data stream.
[0034] Optionally, the source state of the data stream is used to represent the consumption position of the data stream in the data source queue, and the operator state of the data stream is used to represent the intermediate calculation result generated based on a preset operator in the stream computing system, and then the processing system packs the source state and the operator state corresponding to the data stream by using Flink to obtain the distributed consistent snapshot (that is, the first information) corresponding to the data stream.
[0035] Optionally, the preset database can be set as HDFS (Hadoop Distributed File System, a distributed file system), and the processing system can store data into the HDFS by using the FsStateBackend (a persistent storage method) mode, which is suitable for data processing tasks with long windows and key / value states.
[0036] Optionally, the processing system stores the first information into a preset database, thereby persistently storing the source state and the operator state of the data stream, so that when a failure occurs in the stream computing system, the processing system can determine the position of the data stream to be re-read and how to re-compute the data stream based on the first information, thereby maintaining the continuity and consistency of the data stream processing.
[0037] In step S103, when it is detected that the stream computing system has a failure, the target result is determined based on the first information, wherein the target result is a computing result generated after the data stream re-acquired based on the first information is input into the stream computing system.
[0038] Optionally, the processing system can monitor the running state of the stream computing system through the target program, and when it is found that the stream computing system runs abnormally (for example, node downtime, data loss, etc.), a failure report is timely made.
[0039] Optionally, when it is detected that the stream computing system has a failure, the processing system uses the previously stored first information to reconstruct the data stream processing flow, so as to ensure that the computing result after the failure recovery accurately reflects the input order of the data stream and the business logic, and through this mechanism, the stream computing system overcomes the problem of data processing interruption caused by the failure, and can continue to provide high-reliability data processing service, while maintaining the continuity of data processing and the consistency of the computing result.
[0040] From the above, in the case that there is a correlation between the computing result of the data stream and the order in which the data stream is input into the stream computing system, the present application achieves the purpose of quickly recovering the computing result of the data stream based on the first information after the failure of the stream computing system by storing the first information (including the source state and the operator state) corresponding to the data stream in the preset database, thereby enhancing the fault tolerance of the stream computing system and the reliability of data processing, achieving the technical effect of quickly reproducing the computing result of the data stream, and thereby solving the technical problem that the accuracy of the computing result determined by re-executing the computing task of the data stream based on the prior art after the failure of the stream computing system in the process of processing the data stream.
[0041] In an optional embodiment, the processing system first takes the consumption position of the data stream in the data source queue as the source state, then takes the intermediate calculation result of the data stream in the stream computing system based on the preset operator as the operator state, then packs the source state and the operator state to obtain first information, and finally stores the first information to the preset database periodically based on a first parameter, a second parameter, a third parameter and a fourth parameter, wherein the first parameter is used to represent a time window for periodically storing the first information, the second parameter is used to represent an execution mode for storing the first information to the preset database, the third parameter is used to represent a maximum operation duration of a storage operation corresponding to the preset database, and the fourth parameter is used to represent a maximum number of operations of the storage operation executed in parallel.
[0042] Optionally, the data source queue can be set as a Kafka (a distributed publish-subscribe message system) queue.
[0043] Optionally, the consumption position in the data source queue is determined by a specific partition, a topic and an offset parameter in the Kafka queue.
[0044] Optionally, the stream computing system can be set as Flink or Spark, which is used for real-time processing and calculation of the data stream.
[0045] Optionally, the preset operator refers to a function or operation in the stream computing system.
[0046] Optionally, the intermediate calculation result refers to a temporary result after the data stream is processed by the preset operator, and the temporary result is usually contained in the state of the stream computing engine.
[0047] Optionally, the processing system selects the consumption position as the source state, which can ensure that the stream computing system can determine which position in the data source queue to read the data stream from based on the first information after a failure occurs, thereby ensuring the accuracy of the calculation result determined by the stream computing system subsequently, and the processing system ensures the continuity and recoverability of the data stream processing through the operator state in the first information, so that the processing system can quickly locate the failure point through the operator state after the stream computing system fails, thereby quickly recovering the calculation process and avoiding recalculation from the beginning, and the fault tolerance and failure processing efficiency of the stream computing system are improved.
[0048] Optionally, the first parameter refers to the periodic execution time of checkpoint, i.e. the time interval between two storage operations, for example, the first parameter can be set to 2 minutes in Flink; the second parameter refers to the storage execution mode, for example, EXACTLY_ONCE (in a distributed system and a stream processing framework, data can be ensured to be processed once when processing, neither lost nor repeated), used to ensure the consistency of storage operation; the third parameter refers to the timeout time of checkpoint, i.e. the maximum allowed duration of a storage operation; the fourth parameter refers to the number of parallel checkpoint execution, i.e. the number of storage operations that the processing system decides to start at the same time, for example, setting the fourth parameter to 1 means that the stream computing system can only execute one checkpoint operation at a time.
[0049] In addition, the processing system can also set DELETE_ON_CANCELLATION (an action cancellation operation initiated through a user interface or a command line tool), i.e. when the job is cancelled, the processing system deletes the checkpoint, ensuring that the checkpoint can only be used when the job fails.
[0050] As can be seen from the above, by setting the first parameter, the second parameter, the third parameter and the fourth parameter, the processing system stores the first information into the preset database based on the above parameters. The periodic storage mechanism can ensure that the latest and consistent first information is used to recalculate the state when the stream computing system is restarted or fails to recover, thereby ensuring the continuity and accuracy of stream data processing.
[0051] In an optional embodiment, the processing system first determines the first data based on the source state in the first information, wherein the first data is the source data re-read in the data source queue based on the corresponding consumption position of the source state, and then determines the target result based on the first data and the operator state in the first information.
[0052] Optionally, when the stream computing system detects a failure and needs to recover data, the stream computing system accurately re-reads the data stream that needs to be re-read after the failure from the data source queue based on the source state in the first information, i.e. the first data, which is determined by the source state in the first information. The stream computing system can ensure that the data is not missed or re-read, thereby ensuring the continuity and integrity of the read data stream, providing an accurate data basis for subsequent recalculation.
[0053] Optionally, the processing system can reproduce the processing flow before the failure by using the operator state in the first information and the re-read first data, so as to generate the target result consistent with the result before the failure. This process ensures that the data processing state can be accurately restored by the stored first information even after the stream computing system encounters a failure, avoids data inconsistency, and enhances the fault tolerance of the system and the reliability of data processing.
[0054] In an optional embodiment, after determining the target result based on the first information, the processing system writes the target result to a target system in a preset manner, wherein the target system is a downstream system connected to the stream computing system, and the preset manner includes at least one of the following:
[0055] Idempotent writing, which is used to write the target result to the target system based on a primary key update strategy;
[0056] Transactional writing, which is used to write the target result to the target system based on a two-phase commit through a preset interface.
[0057] For example, in Spark, saveAsTextFile (a preset operator) is a typical idempotent writing method and is often used as a data output source. If the message source data contains a unique primary key, the stream processing system will not repeatedly write under the constraint of the primary key even if there are multiple repeated data in the message source, thereby realizing the Exactly-Once (data processing or data transmission should occur exactly once) semantics.
[0058] In Spark, reading Kafka data needs to satisfy the transactional writing of the output end, so a unique ID (which can be determined by the batch number, time, partition, and offset of the data stream) needs to be generated, and then the ID is combined with the calculation result to write to the target source in the same transaction. The commit and write operations are atomic, and the Exactly-Once semantics of the output end is realized.
[0059] Optionally, the target system is a downstream system of the stream computing system, i.e., the final storage or application destination of the calculation result, such as an HBase (a distributed columnar database) database or a Kafka system.
[0060] Optionally, through the idempotent write mode, no matter how many times the operation of writing data to the target system is repeated, only one execution effect is generated, that is, in the database with primary key constraint, the data update operation corresponding to the same primary key is only executed once, and through the update strategy of the processing system based on the primary key, it is ensured that even if the target result is written multiple times, the target system will de-duplicate according to the primary key and only keep the last update, thereby ensuring the consistency and correctness of the data.
[0061] Optionally, the transactional write refers to a method of two-phase commit (2pc, Two-Phase Commit), which ensures that a series of operations are eventually all successfully completed or all failed to roll back, thereby maintaining the integrity and consistency of the data.
[0062] Optionally, the transactional write of the target result through the preset interface can ensure the atomicity and consistency of the write operation, and even if a failure is encountered during the write process, the stream computing system can restore the data to a consistent state through the transaction rollback mechanism, thereby ensuring the integrity of the target result when writing to the target system, and avoiding the semi-written state or inconsistent state of the data.
[0063] For example, when the stream computing system is Flink, transactional write can be used, that is, the write operation in the target system is implemented through GenericWriteAheadSink (a template class for implementing Sink (an operator for writing the result of stream computing to an external system) operator) and TwoPhaseCommitSinkFunction (an interface for implementing Sink operator and supporting two-phase commit protocol) provided by DataStreamAPI (an interface for processing bounded and unbounded data streams), which guarantees the data consistency of the target result during the write process; when the stream computing system is Spark, the processing system uses the preset operator (such as saveAsTextFile) in the output source, combined with the unique primary key in the stream data, so that even if there are multiple repeated data in the source, the data repetition in the target system will not occur under the primary key constraint, thereby realizing the consistency of the data between the end-to-end.
[0064] In an optional embodiment, when the data tag is a target tag, the distribution state of the data stream and the second information are obtained, wherein the distribution state is used to represent the elastic distributed state formed after the data stream is divided into at least two or more batches of sub-data and processed in batches, and the second information is used to represent the dependency relationship between the multiple batches of sub-data, and then the processing system stores the distribution state and the second information to the target log, wherein the target log is a pre-write log corresponding to the data stream.
[0065] Optionally, when the stream computing system is Spark, the data stream is divided into multiple batches of sub-data for parallel processing by Spark, and the distributed state describes the local state after processing each sub-data, at this time, the second information refers to the lineage of RDD (Resilient Distributed Dataset).
[0066] Optionally, the target log is a log specially used for storing key states and dependency information in the process of stream data processing, that is, a WAL (Write Ahead Log) mechanism, which records the change of data state before the data modification operation, so as to ensure that the stream computing system can recover the data to the latest consistent state from the log in the event of failure.
[0067] For example, the checkpoint mechanism of Spark will restart a job after the current job (job, that is, a working unit of a computing task) is executed, so that the RDD that needs to be checkpointed is marked as MarkedForCheckpoint (a type of mark, representing the need to persist the corresponding RDD data to the disk or external storage system), and Spark re-executes the previous lineage of the RDD, saves the result to the checkpoint after completion, and deletes the old lineage of the RDD, thereby using the characteristics of the checkpoint to ensure that the RDD data result is not lost.
[0068] After that, the WAL mechanism is enabled, and if there is an abnormal scenario, the RDD result or job information will be lost, therefore, the processing system needs to persist the RDD result and job information to the WAL pre-write log, that is, by storing backup of metadata and intermediate data to prevent data loss, thereby providing a data basis for subsequent fault recovery.
[0069] Further, Spark also needs to manage the offset and save it to the checkpoint, and the Kafka partition and Spark RDD are one-to-one corresponding, and the data can be read in parallel, and the Executor (working process deployed on the cluster node) consumes data according to the offset range (that is, the start offset and end offset of the consumed message in the Kafka partition) and stores it locally, to ensure that the data is not lost.
[0070] Optionally, the distribution state and the second information obtained by the processing system can ensure that the processing flow of the data stream can be accurately reconstructed during fault recovery, i.e., the state of each batch of sub-data and the dependency relationship between the multiple batches of sub-data, and the pre-write log records the distribution state and the second information before the actual modification of the data, which means that in the case of a fault, the stream computing system can recover to the last consistent state of the data stream processing from the target log, and the mechanism improves the fault tolerance of the stream computing system, especially when processing large-scale and high-throughput data streams, the entire data stream processing can still be recovered and continued even in the case of partial node failure, while the integrity and consistency of the data are not damaged.
[0071] In an optional embodiment, after storing the distribution state and the second information into the target log, when the stream computing system fails and needs to recover data, the processing system first parses the target log to obtain L logs, where L is a positive integer, then the processing system de-duplicates the L logs to obtain M logs, where M is a positive integer less than or equal to L, and then the processing system determines the target result based on the M de-duplicated logs.
[0072] Optionally, the processing system parses the target log by using a regular expression, and during the processing, the problem data and the non-compliant data are stored in the HDFS for statistics, so as to ensure the consistency of the data.
[0073] Optionally, through the de-duplication processing, the processing system can eliminate the repeated log records caused by repeated data or data processing exceptions, so that the data processing flow can be based on accurate and unique data states after fault recovery, thereby avoiding the inconsistency of the target result caused by data duplication, and providing a guarantee for the accuracy of the subsequent recovery calculation result.
[0074] Optionally, based on the M de-duplicated logs, the stream computing system can accurately recover the state and reproduce the calculation flow, which means that the stream computing system can use these non-repeated log records to reconstruct the processing path of the stream data, including the original state of the data stream, the intermediate processing state and the output result, so as to ensure that the data processing flow after fault recovery is consistent with that before the fault, and the target result obtained is also completely consistent, thereby ensuring the Exactly-Once semantics of the stream data processing.
[0075] Optionally, the processing system can de-duplicate the target log by tracking the log watermarking mechanism.
[0076] Optionally, the tracking log has the following functions:
[0077] (1) User behavior audit: record the input and output of the user, and provide audit basis;
[0078] (2) User behavior analysis: analysis and statistics on user access data, providing decision basis;
[0079] (3) Fault analysis: post-fault analysis of business;
[0080] (4) High-level monitoring: for errors that application software cannot actively identify, through analysis of system tracking band logs, fault warning is performed based on preset rules.
[0081] Optionally, the tracking band log file format consists of three parts: control header, extension area, and content area, as follows:
[0082] (1) Control header: used to record time, recording software, and other basic information, the content of the control header needs to follow the following specifications: orderliness between log fields; if the value of a certain log field is missing, the delimiters before and after the field must be retained.
[0083] (2) Extension area: used to record the values of key business fields during system operation, structured information is used to establish query indexes, facilitating tracking band analysis and analysis, the extension area is a set of (Key, Value) pairs, which needs to follow the following specifications: in the form of "Key = Value (key-value pair)"; Key and Value appear in pairs; no order.
[0084] (3) Content area: used to record the data stream or key business processing information.
[0085] Optionally, if real-time data stream enters the Spark consumer end and there is duplicate data, the processing system can perform deduplication operation through the Spark program code combined with the fields of the tracking band, thereby realizing Exactly-Once consistency, the process is as follows:
[0086] (1) Use Hashset (a data structure that does not allow duplicate elements) structure to read the combination of tracking band timestamp, TransactionID, and GlobalID (TransactionID and GlobalID are fields in the target log), store in memory and perform deduplication judgment.
[0087] (2) By reading the fields of tracking band timestamp, TransactionID, and GlobalID, use groupByKey (an aggregation operator) to perform deduplication.
[0088] Optionally, the processing system can also track the time field in the log to handle time disorder and event disorder problems, set the timestamp type used in Flink to Event Time (a time semantic for processing event timestamps in Flink), assign time and watermarks by calling assignTimestampsAndWatermarks (a method for assigning timestamps and watermark markers to elements in a data stream), which can transmit AssignerWithPeriodicWatermar and AssignerWithPunctuatedWatermarks (two parameters), then generate watermarks directly at the data source based on the transmission parameters, set the time allowed for processing delayed data (allowedLateness) while collecting delayed data (sideOutputLateData).
[0089] In an optional embodiment, before obtaining the data label of the data stream, the processing system first obtains third information, fourth information, and fifth information, wherein the third information is used to represent the running state of the target program, the fourth information is used to represent the collection efficiency of the data in the data source queue, and the fifth information is used to represent the data warehouse efficiency of the preset database corresponding to the data, and the target program is a program for monitoring the data stream. Then, the processing system determines the alarm information based on the third information, the fourth information, and the fifth information, wherein the alarm information is used to prompt the operation and maintenance personnel that there is a fault in the processing process of the data stream.
[0090] Optionally, the processing system can obtain the third information by monitoring the heartbeat of the target program, obtain the fourth information by monitoring the offset change and data delay of Kafka, and obtain the fifth information based on the efficiency and state of data output from the stream computing system to the preset database, such as data writing speed and writing failure rate.
[0091] Optionally, through the above steps, the processing system realizes comprehensive monitoring of the stream data processing process, including three major links of data stream collection, processing, and storage. The monitoring mechanism can timely discover and report faults or performance bottlenecks in the data stream processing process, prompting the operation and maintenance personnel to timely intervene and repair the faults, thereby ensuring the stability of the data stream computing system and the consistency of the data stream processing.
[0092] From the above, in the case that there is a correlation between the calculation result of the data stream and the order in which the data stream is input to the stream computing system, the application achieves the purpose of quickly recovering the calculation result of the data stream based on the first information after the failure of the stream computing system by saving the first information (including the source state and the operator state) corresponding to the data stream in the preset database, thereby enhancing the fault tolerance of the data stream computing system and the reliability of data processing, achieving the technical effect of quickly reproducing the calculation result of the data stream, and further solving the technical problem of low accuracy of the calculation result determined by re-executing the calculation task of the data stream based on the prior art after the failure of the stream computing system.
[0093] According to another aspect of the embodiments of the application, a method for implementing end-to-end data consistency is also provided, Figure 2 is a framework diagram of an optional method for implementing end-to-end data consistency according to an embodiment of the application, as Figure 2 shown, the method is introduced as follows:
[0094] (1) Flink technical solution for end-to-end data consistency:
[0095] Based on the Flink project, the source state and the operator state corresponding to the data stream are packaged to obtain a distributed consistency snapshot, and then the snapshot data is persistently stored in the HDFS database. In addition, in order to ensure the consistency of data between the end-to-end, the processing system is optimized as follows:
[0096] Internal guarantee: use checkpoint to set checkpoints and failure recovery mechanism;
[0097] source end: ensure that the external data source can reset the data reading position;
[0098] sink end (destination of stream data processing, i.e. target system): when performing failure recovery, use idempotent and transactional writing methods to ensure that data is not repeatedly written to the external system;
[0099] Watermark mechanism: use a special tracking band log to add a timestamp watermark mechanism.
[0100] (2) Spark technical solution for end-to-end data consistency:
[0101] In the spark project, the entire data lineage is constructed by RDD. When failure recovery is needed, the entire RDD can be recalculated. In addition, in order to ensure the consistency of data between the end-to-end, the processing system is optimized as follows:
[0102] Internal guarantees based on the characteristics of RDD: checkpoint persistence mechanism, WAL mechanism;
[0103] Source end: through the built-in API interface, the Exactly-Once semantic is realized;
[0104] Sink end: idempotent writing mode, transactional writing mode;
[0105] Tracking tape log deduplication: according to the tracking tape log field, the content of the tracking tape log is deduplicated to prevent repeated consumption of the tracking tape log.
[0106] According to another aspect of the embodiments of the present application, an end-to-end data three-layer monitoring method is also provided, Figure 3 is a flow chart of an optional end-to-end data three-layer monitoring method according to an embodiment of the present application, as Figure 3 shown, the method includes three layers: the first layer is real-time project monitoring and statistics, the second layer is monitoring and statistics of the data source, and the third layer is monitoring and statistics of data warehousing.
[0107] Based on the above monitoring method, the consistency problem of data between the end-to-end can be investigated, and whether there is a data accumulation problem, a resource tilt problem, and whether the CPU and memory allocated to the project are reasonable, etc. can be quickly judged when the data volume increases, and then timely alarm can be given when a fault occurs in the processing process of the data stream, so as to prompt the operation and maintenance personnel to handle the fault according to the alarm information.
[0108] Optionally, when the stream computing system starts, the app-id (i.e. a unique identifier assigned to each application) is first acquired, and the running state is judged by calling the REST API (Representational State Transfer, a design architecture style for network application programs).
[0109] Optionally, the processing system can acquire the running state (such as running, succeeded, failed or unknown) of the project by periodic inquiry calling. Each start of the project will generate a new app-id, so the processing system needs to add a relationship table of app-id and project name.
[0110] For example, the process of acquiring the app-id and running state of TAM (a tenant) is as follows:
[0111] (1) Call API query:
[0112] http: / / ip:port / ws / v1 / cluster / apps?user=TAM&state=RUNNING&applicationTypes=SPARK;
[0113] (2) Partially intercept the query result:
[0114] <apps>
[0115] <app>
[0116] <id>application_163935898****_67741< / id>
[0117] <user>TAM< / user>
[0118] <name>TAMSAT Grapheye Online< / name>
[0119] <queue>root.TAM< / queue>
[0120] <state>RUNNING< / state>
[0121] ……
[0122] <app>
[0123] <apps>
[0124] (3) Call API to get the state of a certain project:
[0125] http: / / ip:port / ws / v1 / cluster / apps / application_163935898****_67741 / stat e;
[0126] Result:
[0127] <appstate>
[0128] <state>RUNNING< / state>
[0129] < / appstate>
[0130] Optionally, when monitoring the running index, the processing system can collect the batch processing time and queuing delay of the project to prevent the TASK task from being stuck.
[0131] For example, Spark can collect the batch processing time and queuing delay of the project through the following methods:
[0132] (1) The onStreamingStarted method realizes the record of the groupid (consumer group id) of the program. This method is executed once when the project starts and will not be executed thereafter. The groupid (consumer group id) is mainly stored to record the starting point related data of consumption.
[0133] (2) The onBatchStarted method realizes the record of the consumption of each batch, thereby laying the foundation for the subsequent statistics of the kafka data source.
[0134] (3) The onBatchSubmitted method realizes the record of the submission time of each batch and the record of the project heartbeat.
[0135] (4) The onBatchCompleted method realizes the record of the completion of each batch, the record of the project backlog, and the processing time of the program.
[0136] Optionally, in the process of monitoring the Kafka data source by the processing system, metadataDescription (a descriptive string used to provide metadata information about the current batch data processing) can be found for topic (topic) through debug (debug) and source code tracking, and the following log is printed:
[0137] StreamInputInfo(0, 248, Map(offsets -> List(
[0138] OffsetRange(topic: 'TEST_1', partition: 1, range: [47110** -> 47111**]),
[0139] OffsetRange(topic: 'TEST_1', partition: 3, range: [47111** -> 47112**]),
[0140] OffsetRange(topic: 'TEST_2', partition: 0, range: [33379** -> 33379**]),
[0141] OffsetRange(topic: 'TEST_1', partition: 2, range: [47109** -> 47110**]),
[0142] OffsetRange(topic: 'TEST_2', partition: 3, range: [33377** -> 33377**]),
[0143] OffsetRange(topic: 'TEST_2', partition: 2, range: [33379** -> 33379**]),
[0144] OffsetRange(topic: 'TEST_1', partition: 0, range: [47110** -> 47111**]), OffsetRange(topic: 'TEST_2', partition: 1, range: [33377** -> 33377**]),
[0145] topic: TEST_1 partition: 1 offsets: 47110** to 47111**
[0146] topic: TEST_1 partition: 3 offsets: 47111** to 47112**
[0147] topic: TEST_2 partition: 0 offsets: 33379** to 33379**
[0148] topic: TEST_1 partition: 2 offsets: 47109** to 47110**
[0149] topic: TEST 2 partition: 3 offsets: 33377 to 33377
[0150] topic: TEST 2 partition: 2 offsets: 33379 to 33379
[0151] topic: TEST 1 partition: 0 offsets: 47110 to 47111
[0152] topic: TEST 2 partition: 1 offsets: 33377 to 33377
[0153] Optionally, the processing system discovers Description by sending Kafka data, and Description is empty when no data is sent, and the processing system also needs to distinguish whether the project is one topic or multiple topics, and one topic directly obtains numRecods (the total number of records read or consumed from the data source in one processing batch), and multiple topics need to parse Description, and then group the statistical results, thereby reducing the time consumption of monitoring, and realizing high availability of the project, and here, the offset value of each partition of the topic of each batch of spark RDD is included, and after determining that consumption is failure-free, the offset is manually submitted, and if a batch is problematic in the processing process, the problem is recorded, and subsequent processing of the data is performed through intervention of an operation and maintenance personnel.
[0154] Optionally, when the database is monitored, the processing system can set a timing task, obtains the corresponding data volume in the database through scan rowkey (a data retrieval method), and performs early warning processing in the case of failure.
[0155] According to another aspect of the embodiments of the present application, a processing apparatus for stream data is further provided, Figure 4 is a schematic diagram of an optional processing apparatus for stream data according to an embodiment of the present application, as Figure 4 shown, the processing apparatus for stream data comprises a first obtaining unit 401, a first storage unit 402, and a first determining unit 403.
[0156] Optionally, the first obtaining unit 401 is configured to obtain a data tag of the data stream, where the data tag is used to represent a correlation between a calculation result generated by the stream computing system and an input order of the data stream after the data stream is input into the stream computing system; the first storage unit 402 is configured to store first information corresponding to the data stream into a preset database in a case where the data tag is a target tag, where the target tag is used to represent that there is an association between the calculation result and the order in which the data stream is input into the stream computing system, and the first information is used to represent a source state and an operator state corresponding to the data stream; and the first determining unit 403 is configured to determine a target result based on the first information in a case where it is detected that the stream computing system has a fault, where the target result is a calculation result generated after the data stream that is reacquired based on the first information is input into the stream computing system.
[0157] In an optional embodiment, the first storage unit comprises a first determining subunit, a second determining subunit, a packing unit, and a storage subunit.
[0158] Optionally, the first determining subunit is configured to take a consumption position of the data stream in the data source queue as the source state; the second determining subunit is configured to take an intermediate calculation result of the data stream generated based on a preset operator in the stream computing system as the operator state; the packing unit is configured to pack the source state and the operator state to obtain the first information; and the storage subunit is configured to store the first information into the preset database periodically based on a first parameter, a second parameter, a third parameter, and a fourth parameter, where the first parameter is used to represent a time window for periodically storing the first information, the second parameter is used to represent an execution mode for storing the first information into the preset database, the third parameter is used to represent a maximum operation duration of a storage operation corresponding to the preset database at one time, and the fourth parameter is used to represent a maximum number of operations of the storage operation executed in parallel.
[0159] In an optional embodiment, the first determining unit comprises a third determining subunit and a fourth determining subunit.
[0160] Optionally, the third determining subunit is configured to determine first data based on the source state in the first information, where the first data is source data re-read in the data source queue based on the consumption position corresponding to the source state; and the fourth determining subunit is configured to determine the target result based on the first data and the operator state in the first information.
[0161] In an optional embodiment, the processing apparatus of the data stream further comprises a writing unit.
[0162] Optionally, the writing unit is configured to write the target result into a target system by a preset mode, where the target system is a downstream system connected with the stream computing system, and the preset mode comprises at least one of the following:
[0163] An idempotent writing mode is used to write the target result into the target system based on a primary key-based update strategy;
[0164] A transactional writing mode is used to write the target result into the target system based on a two-phase commit through a preset interface.
[0165] In an optional embodiment, the data stream processing apparatus further comprises a second acquisition unit and a second storage unit.
[0166] Optionally, the second acquisition unit is configured to acquire a distribution state of the data stream and second information when the data tag is the target tag, the distribution state being used to represent an elastic distributed state formed after the data stream is divided into at least two batches of sub-data and the batches of sub-data are processed in batches, and the second information being used to represent a dependency relationship between the batches of sub-data; and the second storage unit is configured to store the distribution state and the second information into a target log, the target log being a pre-write log corresponding to the data stream.
[0167] In an optional embodiment, the data stream processing apparatus further comprises an analysis unit, a deduplication unit and a second determination unit.
[0168] Optionally, the analysis unit is configured to analyze the target log to obtain L logs, L being a positive integer; the deduplication unit is configured to deduplicate the L logs to obtain M logs, M being a positive integer less than or equal to L; and the second determination unit is configured to determine the target result based on the M logs obtained after the deduplication.
[0169] In an optional embodiment, the data stream processing apparatus further comprises a third acquisition unit and a third determination unit.
[0170] Optionally, the third acquisition unit is configured to acquire third information, fourth information and fifth information, the third information being used to represent a running state of a target program, the fourth information being used to represent a data collection efficiency of a data source queue, and the fifth information being used to represent a data warehouse efficiency of a preset database, the target program being a program for monitoring the data stream; and the third determination unit is configured to determine alarm information based on the third information, the fourth information and the fifth information, the alarm information being used to prompt an operation and maintenance personnel that there is a fault in the processing of the data stream.
[0171] From the above, in the case that there is a correlation between the calculation result of the data stream and the order in which the data stream is input into the stream computing system, the processing device achieves the purpose of quickly recovering the calculation result of the data stream based on the first information after the failure of the stream computing system by saving the first information (including the source state and the operator state) corresponding to the data stream in the preset database, thereby enhancing the fault tolerance of the data stream computing system and the reliability of data processing, achieving the technical effect of quickly reproducing the calculation result of the data stream, and further solving the technical problem that the accuracy of the calculation result determined by re-executing the calculation task of the data stream based on the prior art after the failure of the stream computing system in the process of processing the data stream by the stream computing system is low.
[0172] According to another aspect of the embodiments of the present application, a computer program product is also provided, which includes a stored computer program, wherein the computer program product controls the computer program to execute the processing method of the data stream according to any one of the above embodiments when the computer program runs.
[0173] According to another aspect of the embodiments of the present application, an electronic device is also provided, which includes a processor and a memory for storing executable instructions of the processor, wherein the processor is configured to execute the processing method of the data stream according to any one of the above embodiments by executing the executable instructions.
[0174] Optionally, Figure 5 is a schematic diagram of an optional electronic device according to the embodiments of the present application, as Figure 5 shown, the embodiments of the present application provide an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor, and the processor implements the processing method of the data stream according to any one of the above embodiments when executing the program.
[0175] The above embodiments or examples disclosed in the present application are not exhaustive, but only illustrate some embodiments or examples, and are not specific limitations on the protection scope of the present application. In the case of no contradiction, each step in a certain embodiment or example in the present application can be implemented as an independent example, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be exchanged arbitrarily, in addition, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; in addition, the embodiments or examples can be combined arbitrarily, for example, the steps of different embodiments or examples can be combined arbitrarily, a certain embodiment or example can be combined with the optional ways or optional examples of other embodiments or examples.
[0176] In the above-described embodiments of the present application, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be seen from the relevant description of other embodiments.
[0177] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks.
[0178] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks.
[0180] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. The memory can include non-persistent memory in the form of random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0181] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0182] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0183] Those skilled in the art will appreciate that embodiments of the present application can be provided as a method, system or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0184] The above merely provides embodiments of the present application and is not intended to limit the present application. Various modifications and changes can be made to the present application by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of the claims of the present application.< / apps> < / app> < / app> < / apps>
Claims
1. A method for processing a data stream, characterized in that: include: Obtaining a data tag of the data stream, wherein the data tag is used to represent the correlation between a calculation result generated by the stream computing system after the data stream is input into the stream computing system and the input order of the data stream; In a case where the data tag is a target tag, first information corresponding to the data flow is stored in a preset database, wherein the target tag is used to represent an association between the calculation result and the order in which the data flow is input into the stream computing system, and the first information is used to represent the source state and operator state corresponding to the data flow; When a fault is detected in the stream computing system, a target result is determined based on the first information, wherein the target result is a calculation result generated after the data stream re-acquired based on the first information is input into the stream computing system.
2. The data stream processing method according to claim 1, characterized in that: Storing the first information corresponding to the data stream in a preset database includes: The consumption position of the data stream in the data source queue is used as the source state; Using the intermediate calculation result of the data stream generated based on the preset operator in the stream computing system as the operator state; Packaging the source state and the operator state to obtain the first information; The first information is periodically stored in the preset database based on a first parameter, a second parameter, a third parameter, and a fourth parameter, wherein the first parameter is used to characterize a time window for periodically storing the first information, the second parameter is used to characterize an execution mode for storing the first information in the preset database, the third parameter is used to characterize a maximum operation duration of a storage operation corresponding to the preset database, and the fourth parameter is used to characterize a maximum number of operations for executing the storage operation in parallel.
3. The data stream processing method according to claim 1, characterized in that: Determining a target result based on the first information includes: Determining first data based on the source state in the first information, wherein the first data is source data re-read from a data source queue based on a consumption position corresponding to the source state; The target result is determined based on the first data and the operator state in the first information.
4. The data stream processing method according to claim 1, characterized in that: After determining the target result based on the first information, the data stream processing method further includes: Writing the target result to a target system in a preset manner, wherein the target system is a downstream system interconnected with the stream computing system, and the preset manner includes at least one of the following: An idempotent write mode, for writing the target result to the target system based on a primary key update strategy; The transactional writing mode is used to write the target result to the target system based on a two-phase commit mode through a preset interface.
5. The data stream processing method according to claim 1, characterized in that: After obtaining the data tag of the data stream, the data stream processing method further includes: When the data label is the target label, obtaining a distribution state of the data stream and second information, wherein the distribution state is used to represent an elastic distributed state formed after the data stream is divided into at least two or more batches of sub-data and processed in batches, and the second information is used to represent a dependency relationship between the sub-data in multiple batches; The distribution state and the second information are stored in a target log, wherein the target log is a write-ahead log corresponding to the data stream.
6. The data stream processing method according to claim 5, characterized in that: After storing the distribution state and the second information in the target log, the data stream processing method further includes: Parse the target log to obtain L logs, where L is a positive integer; Deduplication is performed on the L logs to obtain M logs, where M is a positive integer less than or equal to L; The target result is determined based on the M logs obtained after deduplication.
7. The data stream processing method according to claim 1, characterized in that: Before obtaining the data tag of the data stream, the data stream processing method further includes: Obtaining third information, fourth information, and fifth information, wherein the third information is used to characterize the running status of the target program, the fourth information is used to characterize the data collection efficiency in the data source queue, and the fifth information is used to characterize the data storage efficiency corresponding to the preset database, and the target program is a program that monitors the data flow; Alarm information is determined based on the third information, the fourth information, and the fifth information, wherein the alarm information is used to prompt operation and maintenance personnel that a fault exists during processing of the data stream.
8. A data stream processing device, characterized in that: include: A first acquiring unit is configured to acquire a data tag of a data stream, wherein the data tag is used to represent a correlation between a calculation result generated by the stream computing system after the data stream is input into the stream computing system and an input order of the data stream; a first storage unit, configured to store, when the data tag is a target tag, first information corresponding to the data flow in a preset database, wherein the target tag is used to indicate an association between the calculation result and the order in which the data flow is input into the stream computing system, and the first information is used to indicate a source state and an operator state corresponding to the data flow; The first determining unit is configured to determine a target result based on the first information when a fault is detected in the stream computing system, wherein the target result is a calculation result generated after inputting a data stream reacquired based on the first information into the stream computing system.
9. A computer program product, characterized in that The computer program product comprises a computer program, wherein when the computer program is run, the computer program product is controlled to execute the data stream processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The method comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the data stream processing method according to any one of claims 1 to 7.