A Distributed Processing Engine Integration Method, Storage Medium, and Computer Device
Through the method of custom development and decomposition of the Siddhi engine into Storm topology, the problem that Siddhi cannot be fully scripted is solved, and efficient big data processing and processing capabilities for complex business requirements are achieved.
Patent Information
- Application Number
- CN202510132168.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-06
AI Technical Summary
In the prior art, Siddhi is integrated into Storm and cannot be fully scripted, resulting in the failure of big data processing performance and poor processing capabilities of complex business requirements.
The Siddhi engine is customized to develop the running environment, syntax rules and processing rules to adapt to the Storm engine. Decompose the Siddhi engine into multiple operators, map these operators into a Storm topology, build a streaming channel between the Siddhi engine and the Storm engine, use the Storm component to process the operators, and generate the Storm topology.
It realizes the distributed integration of Siddhi based on Storm, supports full scripting, and improves development efficiency. Users can complete business development of complex event processing without writing Storm code.
Smart Images

Figure CN119556895B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method for integrating a distributed processing engine, a storage medium, and a computer device. Background Art
[0002] In the prior art method of integrating Siddhi into Storm, the CEP calculation process is still the single-threaded serial processing logic of the entire Siddhi process logic, and the performance still does not meet the standard. Additionally, Siddhi itself acts as a Spout / Bolt. After the upstream data is split through group and enters their respective Siddhi Spout / Bolt Tasks, these data actually lose their relevance within CEP. If the business logic is complex, one has to jump out of the current Siddhi Task, repartition, and then enter a new Siddhi script; this will result in CEP needing to be cut into many Siddhi scripts, losing the usability of the scripts, and thus having a very poor ability to handle complex business requirements. At the same time, Storm itself does not provide a high-level development entry to directly access the data interface of Siddhi. Therefore, for each business, one needs to manually write Topology code, and only the part involving CEP can reuse the code of a Bolt, and business development cannot be fully scripted. Summary of the Invention
[0003] This application mainly provides a method for integrating a distributed processing engine, a storage medium, and a computer device to solve the problem that integrating Siddhi into Storm cannot be fully scripted for big data processing.
[0004] To solve the above technical problems, one technical solution adopted by this application is: to provide a method for integrating a distributed processing engine based on Storm, including: customizing and developing the operating environment, syntax rules, and processing rules of the Siddhi engine to adapt to the Storm engine; decomposing the Siddhi engine into multiple operators according to a preset decomposition rule, and mapping all the operators into a Storm topology; the operators include Source, Sink, Query, and Window.
[0005] In some embodiments, mapping all the operators to a Storm topology includes: constructing a flow transmission channel between the Siddhi engine and the Storm engine to process messages between the Siddhi engine and the Storm engine; using components based on the Storm engine to process the operators corresponding to the respective components; and constructing the Storm topology based on the processed operators and the Storm engine.
[0006] In some embodiments, constructing the flow transmission channel between the Siddhi engine and the Storm engine includes: constructing a converter between the message unit of the Storm engine and the message unit of the Siddhi engine; and connecting the flow of the Storm engine to the message receiver and message processor of the Siddhi engine.
[0007] In some embodiments, using components based on the Storm engine to process the operators corresponding to the respective components includes: constructing Storm External to replace the Source operator or the Sink operator, or parsing the components generated by the corresponding syntax of the Siddhi engine to expand the plug-ins of the Siddhi engine; constructing Bolt to replace the Query operator and the calculation operations of the Siddhi engine; and replacing the Partition of the Siddhi engine with the Group capability of the Storm engine.
[0008] In some embodiments, constructing the Storm topology based on the processed operators and the Storm engine includes: parsing the Siddhi engine to generate a Siddhi Context context environment; generating corresponding Spout and Bolt based on the Storm engine according to all the processed operators in the SiddhiContext context environment, and combining the Bolt to generate the corresponding Storm topology; and generating a language environment based on the Siddhi Context context environment to process the flows and data corresponding to the Spout and the Bolt.
[0009] In some embodiments, customizing the operating environment of the Siddhi engine to adapt to the Storm engine includes: opening the operators, window functions, aggregation functions, and expressions saved in the operating environment to allow the Storm engine to access or control the operating environment.
[0010] In some embodiments, customizing and developing the syntax rules of the Siddhi engine to adapt to the Storm engine includes: adding distribution parameters, data response parameters, and resource configuration parameters to the domain-specific language of the Siddhi engine.
[0011] In some embodiments, customizing and developing the processing rules of the Siddhi engine to adapt to the Storm engine includes: adding an exception handling rule to pass the exception to the Storm engine and handle the exception through the exception reply handling logic of the Storm engine.
[0012] To solve the above technical problems, another technical solution adopted by this application is: providing a storage medium on which program data is stored, and when the program data is executed by a processor, the steps of the distributed processing engine integration method as described above are implemented.
[0013] This application also provides a computer device, including a processor and a memory connected to each other. The memory stores a computer program, and when the processor executes the computer program, the steps of the distributed processing engine integration method as described above are implemented.
[0014] The beneficial effects of this application are: Different from the prior art, this application discloses a distributed processing engine integration method, a storage medium, and a computer device. By customizing and developing the operating environment, syntax rules, and processing rules of the Siddhi engine to adapt to the Storm engine. The Siddhi engine is decomposed into multiple operators according to a preset decomposition rule, and all operators are mapped to a Storm topology, realizing the Siddhi distributed integration based on Storm. It realizes full scripting and improves development efficiency. Users no longer need to write Storm code, and only need to write an independent Siddhi to complete the business development of complex event processing, and the functional complexity will not be affected. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings, where:
[0016] Figure 1 is a flowchart of an embodiment of the distributed processing engine integration method provided by this application;
[0017] Figure 2 is as Figure 1Flow diagram of one embodiment of method step 200 shown;
[0018] Figure 3 is as Figure 2 Flow diagram of one embodiment of method step 210 shown;
[0019] Figure 4 is as Figure 2 Flow diagram of one embodiment of method step 220 shown;
[0020] Figure 5 is as Figure 2 Flow diagram of one embodiment of method step 230 shown;
[0021] Figure 6 Structural diagram of one embodiment of the storage medium provided by the present application;
[0022] Figure 7 Structural diagram of one embodiment of the computer device provided by the present application. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] The terms "first", "second", and "third" in the embodiments of the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0025] References herein to "embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0026] Refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of an embodiment of a distributed processing engine integration method provided by the present application. The Storm-based distributed processing engine integration method includes:
[0027] 100: Customize and develop the operating environment, grammar rules, and processing rules of the Siddhi engine to adapt to the Storm engine.
[0028] Perform some customization development on Siddhi to interface with the high performance and reliability of Storm. For example, open the API of the real-time operating environment for docking, modify the grammar rules of DSL, etc.
[0029] Among them, Storm is a distributed real-time computing engine for processing unbounded streaming data. It provides linear scalability, fast recovery, and data consistency guarantee, with a single node reaching up to one million per second per node; it is commonly used in scenarios such as real-time analysis, online machine learning, and real-time data cleaning. As a distributed high-performance and highly reliable computing framework, Storm can be horizontally scaled and is suitable for businesses dealing with massive data.
[0030] Siddhi is a complex event processing (CEP) engine that can perform real-time analysis, data processing, etc. using a DSL similar to SQL. It provides a high-level development interface through a customized DSL similar to SQL, provides functions commonly used in complex event processing such as PatternSequence, etc., which offers more development capabilities than SQL, and Siddhi itself provides many rich operators, helping to save development costs.
[0031] Optionally, customizing and developing the operating environment of the Siddhi engine to adapt to the Storm engine includes: opening the operators, window functions, aggregation functions, and expressions saved in the operating environment to allow the Storm engine to access or control the operating environment.
[0032] Split the API of the operating environment into multiple modules, each module responsible for different functions, such as data source management, data processing, data storage, etc. Develop a customized operating environment API.
[0033] Integrate Source and Sink with Storm's Spout and Bolt respectively to achieve the reading and writing of data streams. Use Storm's Topology Builder to build a real-time processing topology and connect the above components according to business requirements. Submit the Topology through Storm's client API and use Storm UI for monitoring and management.
[0034] Among them, Storm's API provides Spout and Bolt interfaces to process data streams and supports custom parallelism and grouping strategies.
[0035] Spouts are the source ends of the message flow in the topology, used to ingest external data and submit it to the topology. They can record some external states such as offset and other information, as well as respond to the replies (ack) from downstream bolts to ensure data reliability.
[0036] Bolts are the computing units in the topology that actually process data. They can filter, aggregate, correlate, etc. They can send the processed message flow to the downstream according to the topology's flow direction and can also reply (ack) to the message to tell the Spouts that the current message has been processed normally.
[0037] Optionally, customize and develop the grammar rules of the Siddhi engine to adapt to the Storm engine, including: adding distribution parameters, data response parameters, and resource configuration parameters to the domain-specific language of the Siddhi engine.
[0038] Add grammar in Siddhi's query language to support distribution parameters such as parallelism and partition fields. Identify the stream by adding specific query options or functions to reply to upstream messages. Set resource configuration parameters through Siddhi's configuration file or specific options in the query language.
[0039] Among them, the domain-specific language (DSL) of the Siddhi engine allows users to define streams, execute queries, process events, etc.
[0040] Optionally, customize and develop the processing rules of the Siddhi engine to adapt to the Storm engine, including: adding exception handling rules to pass exceptions to the Storm engine and handle the exceptions through the exception reply handling logic of the Storm engine.
[0041] The Storm engine provides a more powerful exception handling mechanism, including retry when a task fails, failover, etc. Storm's Topology can be configured to take specific actions for exception handling when a Spout or Bolt task fails.
[0042] By adding try-catch blocks to the query logic of Siddhi, exceptions that may occur during the query execution of Siddhi can be captured. The captured exceptions are passed to the Storm engine for processing through the interface between Siddhi and Storm.
[0043] Receive the exceptions passed from Siddhi in Storm's Bolt or Spout by listening to specific output streams or checking the exception information in the context. According to Storm's exception handling mechanism, configure the Topology to take appropriate actions when an exception is received.
[0044] Integrate the customized Siddhi engine with the Storm engine to ensure that they can communicate correctly and pass exceptions.
[0045] 200: Decompose the Siddhi engine into multiple operators according to the preset decomposition rules, and map all the operators to a Storm topology.
[0046] Analyze the Siddhi query, identify the data processing steps in each query, such as filtering, aggregation, window operations, etc. Define each processing step as an operator. These units should be independent, reusable, and capable of performing specific data processing tasks.
[0047] For each input stream, create a Storm Spout to read the data and pass it to the first Bolt in the topology. For each operator, create a Storm Bolt to perform the corresponding data processing task. Each Bolt can receive data from the upstream Bolt or Spout, process it, and pass the result to the downstream Bolt or output it to an external system, thus mapping all the operators to a Storm topology.
[0048] Among them, the Storm topology is composed of Spouts and Bolts linked in the form of a directed acyclic graph (DAG). Among them, the Spout is responsible for reading data from an external source, while the Bolt is responsible for processing the data and passing it to the next Bolt or outputting it to an external system.
[0049] In the Storm topology, the data flow is defined by the connections between Bolts.
[0050] See Figure 2 , 200 further includes:
[0051] Decompose the Siddhi engine into multiple operators according to the preset decomposition rules. The operators include Source, Sink, Query, and Window.
[0052] In Storm, the Source operator corresponds to the Spout component, which is used to implement a Storm Spout to read data streams from the Siddhi engine or other external sources. Ensure that the Spout can continuously receive data from the Siddhi engine and pass it as a Tuple (a data structure in Storm) to the downstream Bolt.
[0053] The Sink operator is not implemented as an independent component in Storm but as part of a Bolt or for final output processing. Implement a Bolt to process data from the upstream Bolt and send it to an external system (such as a database, file system, or another service).
[0054] The Query and Window operators correspond to the Bolt component in Storm. For each Query and Window operator, implement a Storm Bolt to process the input data, perform the corresponding query or window operation, and generate output data. Bolts can be connected through Storm's stream mechanism to transfer data.
[0055] 210: Build a stream transmission channel between the Siddhi engine and the Storm engine to process messages between the Siddhi engine and the Storm engine.
[0056] Further, refer to Figure 3 , 210 includes:
[0057] 211: Build a converter between the message unit of the Storm engine and the message unit of the Siddhi engine.
[0058] The message unit of the Storm engine (Storm Tuple) is the basic data unit in Storm, which is used to transfer data between different components in the topology.
[0059] The message unit of the Siddhi engine (Siddhi Event) is a data structure used to represent events, usually containing the timestamp of the event, data fields, etc.
[0060] Converting from the Storm engine to the Siddhi engine specifically includes mapping the fields in the Tuple to the corresponding fields in the Event, thereby extracting the data in the Storm Tuple and constructing a Siddhi Event object.
[0061] Converting from the Siddhi engine to the Storm engine specifically includes extracting the data in the Siddhi Event and constructing a Storm Tuple object. This also requires field mapping.
[0062] 212: Connect the stream of the Storm engine to the message receiver and message processor of the Siddhi engine.
[0063] Define a Spout in the Storm topology to read the data source and generate Storm Tuples.
[0064] Define one or more Bolts to process these Tuples and pass the processed data to the next Bolt or output it to an external system, which is the stream of the Storm engine.
[0065] The settings of the Siddhi message receiver and processor specifically include creating a Siddhi application and defining input and output streams. Use Siddhi's API to add queries that define how to extract data from the input stream and generate the output stream. Use Siddhi's message receiver (such as InputHandler) to receive external events and use the processor (such as QueryCallback) to process the query results.
[0066] In the Storm Bolt, use Siddhi's InputHandler to send Tuples to the input stream of Siddhi. In Siddhi's QueryCallback, convert the query results into Storm Tuples and use Storm's OutputCollector to send them to the next Bolt or output them to an external system to achieve the connection between the stream of the Storm engine and the message receiver and message processor of the Siddhi engine.
[0067] By constructing a converter between Storm Tuples and Siddhi Events and implementing the interconnection between the stream of Storm and the message receiver and processor of Siddhi, we can integrate Siddhi into Storm to utilize Siddhi's powerful complex event processing capabilities to enhance the functions of Storm.
[0068] 220: Use components based on the Storm engine to process the operators corresponding to each component.
[0069] During the process of parsing the Siddhi script into a Storm Topology, each operator in Siddhi (such as Source, Sink, Query, Window, etc.) needs to be mapped to the corresponding components in Storm (such as Spout, Bolt, etc.).
[0070] For each operator, select or customize appropriate Storm components according to its specific processing logic and functional requirements. For example, for the Window operator in Siddhi, a dedicated Bolt needs to be created to handle the window calculation logic.
[0071] In this way, the complex event processing logic in the Siddhi script is distributed to multiple nodes in the Storm cluster for parallel execution, thereby improving the processing performance and scalability of Storm.
[0072] See Figure 4 , 220 further includes:
[0073] 221: Build Storm External to replace the Source operator or Sink operator, or parse the components generated by the corresponding syntax of the Siddhi engine to expand the plug-ins of the Siddhi engine.
[0074] Use Storm External to replace the Source and Sink operators of Siddhi to utilize the high performance and data consistency design of Storm External, thereby enhancing the data processing ability and reliability of the entire system.
[0075] For Storm External such as Kafka, HBase, Hive, HDFS, etc. that have high-performance alternatives or do not exist in Siddhi, parse the corresponding Siddhi syntax to generate corresponding KafkaSpout / Bolt or HBaseBolt, HDFSBolt, etc. to expand the plug-ins of Siddhi itself.
[0076] Parse the corresponding syntax of the Siddhi engine and generate components compatible with the Storm engine according to these syntaxes to expand the plug-in library of the Siddhi engine. These newly generated components can be seamlessly integrated into the Storm topology, thereby further enhancing the flexibility and scalability of the system. Utilizing the data consistency design of Storm's External is conducive to ensuring the data reliability of the system.
[0077] 222: Build a Bolt to replace the Query operator and the calculation operations of the Siddhi engine.
[0078] A specialized Bolt is constructed to replace Siddhi's Query operators and computational operations. For example, the general QueryBolt processes general queries and patterns / sequences, etc., to address the performance bottlenecks faced by the Siddhi engine when dealing with complex queries. The Bolt is designed to efficiently handle various complex query logics, including pattern matching, sequence detection, multi-stream correlation, etc.
[0079] Specifically, the continuous aggregation Bolt corresponds to Siddhi's continuous aggregation calculation.
[0080] Optionally, two corresponding Bolts are created for Siddhi window calculation, one global window Bolt and one MapReduce-distributed grouped window, to enhance the window calculation ability for massive data.
[0081] By encapsulating these computational operations into the Bolt, it is beneficial to make full use of the features of the Storm engine such as asynchronous processing, caching mechanism, and backpressure control, thereby improving the overall performance and throughput of the system.
[0082] 223: Replace Siddhi engine's Partition with the Group ability of the Storm engine.
[0083] Use the Group ability of the Storm engine to replace Siddhi's Partition mechanism. The Group ability of the Storm engine allows the data stream to be grouped according to specific fields or rules, and the grouped data streams are distributed to different Bolts for processing.
[0084] In this way, finer-grained data control and more efficient resource utilization can be achieved. At the same time, the Storm engine also provides high-performance caching and concurrent processing mechanisms to ensure the consistency and reliability of data during grouping and processing. Make full use of the high-performance caching, concurrency, and backpressure processing methods provided by Storm itself to ensure performance.
[0085] 230: Build a Storm topology based on the processed operators and the Storm engine.
[0086] Define the structure of the Storm topology according to the operators defined in the Siddhi engine and their relationships. Determine which operators correspond to Spouts, which correspond to Bolts, and the data flow between them. Configure the appropriate parallelism for each Spout and Bolt according to the processing requirements. Set the appropriate grouping strategy for the streams in the Storm topology according to the operator relationships defined in the Siddhi engine. Submit the configured Storm topology to the Storm cluster for deployment.
[0087] Further, refer to Figure 5 , 230 includes:
[0088] 231: Parse the Siddhi engine to generate a Siddhi Context.
[0089] Read the Siddhi script file, parse elements such as Source, Sink, Query, Store, etc. in the script, as well as their dependency relationships and data flows. Generate a Siddhi Context that contains all the operators defined in the script and their configuration information.
[0090] Parse the configuration and rules of the Siddhi engine to generate a SiddhiContext that contains all the processed operators. The Siddhi Context is used to reflect the dependency relationships between operators, the direction of the data flow, and the specific processing logic of each operator.
[0091] 232: Generate corresponding Spouts and Bolts based on the Storm engine according to all the processed operators in the Siddhi Context, and combine the Bolts to generate the corresponding Storm topology.
[0092] Generate a Storm Spout according to the Source operator defined in the Siddhi Context. Generate Storm Bolts according to operators such as Query, Filter, Window, Aggregation, etc. defined in the SiddhiContext. Connect the Spout and Bolt in the Storm Topology according to the name and flow direction of the Stream to form a data flow. Configure the parallelism of each Spout and Bolt, as well as the grouping strategy of the data flow.
[0093] 233: Generate a language environment based on the Siddhi Context to process the streams and data corresponding to the Spout and Bolt.
[0094] Submit the generated Storm Topology to the Storm cluster for deployment. The Storm cluster will start the corresponding Worker processes and Tasks according to the configuration information of the Topology. When a Storm Task is initialized, a new Siddhi runtime environment will be generated for the current Task based on the previously generated Siddhi Context. Each Task will have an independent instance of the Siddhi runtime environment to process the data stream it is responsible for.
[0095] The Storm Topology starts running, and data flows in from the Source operator (i.e., Spout). The data flows in the StormTopology according to the defined flow direction and grouping strategy, and the corresponding calculation logic is executed in each Bolt. The processing results are finally output to an external system or storage through the Sink operator.
[0096] In fact, Siddhi completes the data flow between different components by directly sending and receiving messages between different Streams. In this application, by replacing or linking the Publisher and Receiver of the StreamJunction of Siddhi with the data sending and receiving interfaces inside the computing platform, correctly encapsulating the `Context`, and then starting the corresponding components in the `Launcher`, the splitting and deployment of Siddhi can be completed.
[0097] Specifically, the structure of a basic Siddhi Query in Runtime and its encapsulation as an independent Storm Bolt. According to the running structure of the Siddhi Query in the Siddhi App, using StreamJunction as an instance of the Siddhi Stream, when encapsulating it into the running structure of the Storm Bolt, it is necessary to connect the Storm interfaces execute and emit to the Siddhi InputHandler and Query Callback, and complete the basic connection after converting the tuple and Event.
[0098] Split its App into a DAG graph according to SiddhiQL, with Source as the starting point, Sink as the ending point, and each Query as a computing node; then turn the DAG into an actual deployment graph through the predefined Parallelism. Among them, the processing logics of multi-stream queries: Pattern, Sequence, and Join are the same as the above.
[0099] In addition to the basic Query, Siddhi also has the Window Query with load. There are two types of Window Queries in Siddhi. One is the chained Window Query, which is embedded in the ordinary Query and does not require special processing. The other is the explicitly defined Window Query, and any number of window query statements can be provided subsequently.
[0100] Inside the Siddhi Runtime, the Window Query is implemented jointly by multiple components. A StreamJunction is the StreamJunction of the explicitly defined Window Stream. A Query is the query of "insert into [window stream]". N Sub Queries are the query statements for window query and aggregation operations from the window stream; the Window Query and its SubQuery are merged into a Storm Bolt through operator merging. The data of the InputHandler is still a simple Event, but the output of the window's Query callback is not a simple Event, but a Complex Event, which is actually the data set of the entire window cache and cannot be sent as a message on the computing platform. Therefore, the downstream queries of the Window, that is, those Sub Queries, are merged into the operator where the Window Query is located to ensure that the data finally output by the operator is still a simple Event.
[0101] Optionally, the method for transforming the Window Query for distributed window computing includes:
[0102] In response to the existence of Group by in the Window sub query, the Window sub query containing this keyword can easily set its concurrency. Only by distributing according to fields upstream can the distributed transformation be completed, and only one WindowQueryBolt is required to complete it.
[0103] In response to the absence of Group by in the Window sub query, the Window sub query is a global window and needs to be split into two Storm Bolts, namely WindowQueryMapBolt and WindowQueryReduceBolt. The former consumes data according to the required concurrency and aggregates it into temporary results to be sent to the Reducer, while the latter is responsible for aggregating the intermediate results and outputting them to the downstream.
[0104] The new grammar rules added to Siddhi for adapting to Storm include:
[0105] Siddhi Source, which is used to generate a Spout with the corresponding concurrency in Storm and perform field grouping based on its Keys and concurrency.
[0106] Siddhi Query, which is used to consume the data stream distributed upstream according to keys by a Bolt with the corresponding concurrency.
[0107] Optionally, it also includes `delaySecs`, which is used to represent the maximum allowed message delay for the REDUCE Window. Messages that exceed the delay will not be aggregated but will be output as an independent data item.
[0108] Two new annotations, dcep_source and dcep_sink, are added to Siddhi and encapsulated as Spout / Bolt using Storm External to ensure query performance and data reliability.
[0109] Use Storm External to replace the Source and Sink of Siddhi, and use granular Siddhi components to complete the construction of the entire data stream pipeline, mapping a complete Siddhi App to a StormTopology.
[0110] Optionally, for simple Siddhi Queries that do not involve heavy intermediate states such as Join / Pattern / Sequence / chained Windows, and where the streams connected to the upstream or downstream and the grouping method are exactly the same, two Siddhi Bolts can be merged into one to reduce the consumption of Topology computing resource slots and unnecessary message communication I / O, etc. Of course, it is also possible to specify several Queries to be merged into one Bolt by adding parameters to the Siddhi script.
[0111] Based on the above embodiments of integrating Storm, those skilled in the art should understand that there are many other options for distributed real-time computing engines besides Storm, such as Heron, Spark Streaming, Flink, etc. They also provide flexible underlying APIs that can interface with Siddhi and different data reliability guarantees.
[0112] See Figure 6 , Figure 6 is a schematic structural diagram of an embodiment of the storage medium provided by the present application.
[0113] The storage medium 300 stores program data 310, and when the program data 310 is executed by a processor, it implements the steps of the distributed processing engine integration method as described in Figure 1 the distributed processing engine integration method as described in
[0114] The program data 310 is stored in a storage medium 300 and includes several instructions for causing a network device (which can be a router, a personal computer, a server, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application.
[0115] Optionally, the storage medium 300 can be various media that can store program data, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0116] See Figure 7 , Figure 7 is a schematic structural diagram of an embodiment of the computer device provided by the present application.
[0117] The computer device 400 includes a processor 420 and a memory 410 that are interconnected. The memory 410 stores a computer program, and when the processor 420 executes the computer program, it implements the steps of the distributed processing engine integration method as described above.
[0118] Different from the prior art, the present application splits it into different implementations such as Source, Sink, ordinary Query, multi-stream Query, merging Window operators (Window definition, Window stream, Sub Query stream), grouped Window, global Window, etc. according to script parsing; enables them to distinguish upstream and downstream and can be independently deployed to tasks of a distributed computing engine, realizing the Siddhi distributed integration based on Storm. Configuring Siddhi distributedly into Storm to achieve full scripting, only an independent Siddhi script needs to be written, and there is no need to write Storm Topology code to complete the business development of complex event processing, improving the development efficiency. Splitting the long serial processing logic of the Siddhi script into different Storm components and deploying them on different nodes can make full use of mechanisms such as asynchrony, caching, and backpressure of Storm itself to ensure processing performance. At the same time, the concurrency degree is set for each operator of Siddhi as needed. For example, statements with a large data intake and a high computing density can be set with a high concurrency, and vice versa with a low concurrency degree, so as to make full use of the resources of the Storm cluster. In addition, after the Siddhi script is split and incorporated into the Spout and Bolt of Storm, the emit and ack mechanisms of Storm itself can be used to perform data consistency verification and re-entry, which is beneficial to ensuring the data consistency of the system. During the process of encapsulating Siddhi as an operator of the Storm computing engine, by customizing the Siddhi syntax, opening context parameters, replacing External plugins, and customizing different Bolts according to different Queries, Siddhi can fully utilize the high performance and data reliability brought by the Storm computing engine.
[0119] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the storage medium embodiment and the computer device embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0120] The present application can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, network PCs, small computers, distributed computing environments including any of the above systems or devices, and so on.
[0121] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0122] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0124] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A distributed processing engine integration method based on Storm, characterized in that: include: Open the operators, window functions, aggregation functions and expressions saved in the Siddhi engine runtime environment to allow the Storm engine to access or control the runtime environment, add distribution parameters, data response parameters and resource configuration parameters to the domain-specific language of the Siddhi engine, and customize the processing rules of the Siddhi engine to adapt to the Storm engine; Decomposing the Siddhi engine into multiple operators according to a preset decomposition rule, and mapping all the operators into a Storm topology; the operators include Source, Sink, Query and Window; Constructing a converter between the message unit of the Storm engine and the message unit of the Siddhi engine, and connecting the stream of the Storm engine with the message receiver and message processor of the Siddhi engine, thereby constructing a stream transmission channel between the Siddhi engine and the Storm engine, so as to process messages between the Siddhi engine and the Storm engine; Using components based on the Storm engine to process the operators corresponding to each of the components; Parsing the Siddhi engine to generate a Siddhi Context context; Generate corresponding Spouts and Bolts based on the Storm engine according to all processed operators in the Siddhi Context context environment, and combine the Bolts to generate corresponding Storm topology; generate a language environment based on the Siddhi Context context environment to process the streams and data corresponding to the Spouts and Bolts; wherein the Storm topology is composed of the Spouts and Bolts linked in a directed acyclic graph (DAG) manner.
2. The distributed processing engine integration method according to claim 1, characterized in that: The using components based on the Storm engine to process the operators corresponding to the components includes: Building Storm External to replace the Source operator or the Sink operator, or parsing the components generated by the corresponding grammar of the Siddhi engine to expand the plug-in of the Siddhi engine; Building Bolt to replace the calculation operation of the Query operator and the Siddhi engine; The Partition of the Siddhi engine is replaced by the Group capability of the Storm engine.
3. The distributed processing engine integration method according to claim 1, characterized in that: The custom development of processing rules for the Siddhi engine to adapt to the Storm engine includes: New exception handling rules are added to pass exceptions to the Storm engine, and the exceptions are handled by the exception response processing logic of the Storm engine.
4. A storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the steps of the distributed processing engine integration method as described in any one of claims 1 to 3 are implemented.
5. A computer device, characterized in that: It comprises a processor and a memory connected to each other, the memory stores a computer program, and when the processor executes the computer program, the steps of the distributed processing engine integration method as described in any one of claims 1 to 3 are implemented.