Method, system, and computer program product for processing a tuple stream
By introducing the AASC technology in stream computing, the operation logic of stream applications is dynamically adjusted, real-time and scalability problems in stream computing are solved, and the accuracy and efficiency of data processing are improved.
Patent Information
- Application Number
- CN202111432697.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-17
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing stream calculation methods have problems such as insufficient real-time, inconsistent data format, unreasonable resource configuration, poor system scalability and inability to adjust dynamically when processing big data, resulting in low accuracy and efficiency of analysis results.
The aspect-aware flow computing (AASC) technology is used to insert a general program execution structure during the running period of the stream application and dynamically execute program code instructions to adjust and optimize some functions of the stream application without affecting the availability of the system.
It realizes dynamic adjustment and optimization of streaming applications, improves system flexibility and scalability, enhances data processing accuracy and efficiency, and reduces the risks of data loss and system failure.
Smart Images

Figure CN114647415B_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to stream computing and, more particularly, to influencing the operation of a stream application based on aspect program code.
[0002] Stream computing can be utilized to provide real-time analysis and processing of large amounts of data. Stream computing can be based on a pre-compiled set of fixed processing elements or stream operators. Summary of the Invention
[0003] According to an embodiment, disclosed is a method, system, and computer program product.
[0004] A tuple stream can be monitored. The tuple stream is processed by a plurality of processing elements operating on one or more computing nodes of a stream application, each processing element having one or more stream operators. A program request for executing a first set of program code instructions is received. A stream application target for the first set of program code instructions is identified based on the program request. A first portion of a set of one or more portions of the stream application is encapsulated using a general-purpose program execution structure. The encapsulation is performed during the operation of the stream application. The general-purpose program execution structure is configured to receive and execute program code instructions outside of pre-configured operations of the stream application. The first set of program code instructions is executed by the general-purpose program execution structure during a first time period of execution of the first portion of the stream application. The execution of the first set of program code instructions is in response to the program request, based on the stream application target, and during the operation of the stream application.
[0005] The above Summary of the Invention is not intended to describe every illustrated embodiment or every implementation of the present disclosure. Brief Description of the Drawings
[0006] The drawings included in this application are incorporated into and form a part of the specification. They illustrate embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. The drawings illustrate only certain embodiments and do not limit the present disclosure.
[0007] Figure 1 Depicts representative major components of an example computer system that can be used in accordance with some embodiments of the present disclosure;
[0008] Figure 2 Depicts an example stream computing application configured to operate on a tuple stream in accordance with some embodiments of the present disclosure;
[0009] Figure 3 Is an example stream application configured for aspect-aware stream computing in accordance with some embodiments of the present disclosure; and
[0010] Figure 4 Depicts a method of performing aspect-aware stream computing in accordance with some embodiments of the present disclosure.
[0011] While the present invention may have various modifications and alternative forms, details thereof have been shown by way of example in the drawings and will be described in detail. However, it should be understood that the intention is not to limit the present invention to the particular embodiments described. On the contrary, the present invention covers all modifications, equivalents, and alternatives falling within the spirit and scope of the present invention. Detailed Description
[0012] Aspects of the present disclosure relate to stream computing, and more particularly, aspects relate to influencing the operation of a stream application based on aspect program code. While the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure may be understood through discussion of various examples using this context.
[0013] One use of a computing system (alternatively, a computer system) is to collect available information, process the collected information, and make decisions based on the processed information. The computer system may operate on the information in the form of a database that allows a user to determine what has occurred and predict future results based on past events. These computer systems may receive information from various sources and then record the information in a persistent database. After the information has been recorded in the database, the computing system may run algorithms on the information—sometimes generating new information and then performing associated transformations on the new information and storing the new information—to make determinations and provide context to the user.
[0014] The ability of a computer system to analyze information and provide meaning to a user may be insufficient in some cases. The ability of large organizations such as corporations and governments to make decisions based on information analysis may be impaired by the limited scope of available information. Additionally, the analysis may have limited value because it relies on a stored structured database that may contain outdated information. This may result in decisions of limited value or, in some cases, inaccurate decisions. For example, a weather forecasting service may be unable to accurately predict precipitation in a given area, or a stock brokerage firm may make an incorrect decision regarding stock trading trends.
[0015] The analytical drawbacks of computer systems can be exacerbated by other factors. First, the world may be becoming increasingly IoT-enabled as previously non-intelligent devices are now becoming smart devices. Smart devices can include devices that historically did not provide analytical information but have now been equipped with sensors and can now do so (e.g., cars that can now provide diagnostic information to their owners or manufacturers, thermostats that now communicate to users via the web about daily temperature fluctuations in the home). Second, these drawbacks may also increase due to the increase in communication from information sources as previously isolated devices are now becoming interconnected (e.g., appliances within a home communicate with each other and with the power company to utilize electricity more efficiently). These new information sources can provide not only a volume of isolated data points but also the relationships between new smart devices.
[0016] A third compounding factor is that users of computing systems may prefer continuous analysis of information streams, while data acquisition methods may only provide event-based methods of analyzing pre-recorded information. For example, an analytics package may receive a limited amount of data and later apply analysis to that data. This approach may not work when dealing with continuous data streams. A fourth compounding factor is that computer systems may be deficient in handling large volumes of information and the unstructured nature of the information; for example, sensors, cameras, and other new data sources may not provide context or format but only raw information. The analytical methods of computing systems may need to modify and rearrange this data in order to provide any kind of context for the raw information. Modifying and rearranging may take time or resources that many computing systems may not be able to provide.
[0017] Another potential drawback is that computing systems may not be able to provide scalable solutions to new users. The emergence of smart and connected devices has provided new use cases for the analysis of continuous information streams. However, modern systems for large-scale data collection may require a great deal of user training and provide non-intuitive interfaces. For example, a farmer can equip every animal on the farm with sensors to monitor the health and location of the animals. Data from these sensors can enable the farmer to respond to the changing health of the animals, but only if the sensor data is collected and converted into a usable format to provide meaningful information to the farmer in real time. The farmer may not have the resources to provide to technical experts to build large-scale analytics packages, and the information obtained may go unused.
[0018] Figure 1Depicts representative major components of an example computer system 100 (alternatively, a computer) that can be used in accordance with some embodiments of the present disclosure. It should be understood that the various components can vary in complexity, quantity, type, and / or configuration. The specific examples disclosed are for illustrative purposes only and are not necessarily the only such variations. Computer system 100 can include a processor 110, a memory 120, an input / output interface (herein, I / O or I / O interface) 130, and a main bus 140. The main bus 140 can provide a communication path for the other components of computer system 100. In some embodiments, the main bus 140 can be connected to other components, such as a dedicated digital signal processor (not depicted).
[0019] The processor 110 of computer system 100 can include one or more cores 112A, 112B, 112C, 112D (collectively referred to as 112). The processor 110 can additionally include one or more memory buffers or caches (not shown) that provide temporary storage of instructions and data for the cores 112. The cores 112 can execute instructions on inputs provided from the cache or from the memory 120 and output the results to the cache or memory. The cores 112 can be composed of one or more circuits configured to execute one or more methods in accordance with embodiments of the present invention. In some embodiments, computer system 100 can include multiple processors 110. In some embodiments, computer system 100 can be a single processor 110 with a single core 112.
[0020] The memory 120 of computer system 100 can include a memory controller 122. In some embodiments, the memory 120 can include a random access semiconductor memory, storage device, or storage medium (volatile or non-volatile) for storing data and programs. In some embodiments, the memory can be in the form of a module (e.g., a dual in-line memory module). The memory controller 122 can communicate with the processor 110 to facilitate the storage and retrieval of information in the memory 120. The memory controller 122 can communicate with the I / O interface 130 to facilitate the storage and retrieval of inputs or outputs in the memory 120.
[0021] The I / O interface 130 may include an I / O bus 150, a terminal interface 152, a storage interface 154, an I / O device interface 156, and a network interface 158. The I / O interface 130 may connect the main bus 140 to the I / O bus 150. The I / O interface 130 may direct instructions and data from the processor 110 and the memory 120 to the various interfaces of the I / O bus 150. The I / O interface 130 may also direct instructions and data from the various interfaces of the I / O bus 150 to the processor 110 and the memory 120. The various interfaces may include the terminal interface 152, the storage interface 154, the I / O device interface 156, and the network interface 158. In some embodiments, the various interfaces may include a subset of the above interfaces (e.g., an embedded computer system in an industrial application may not include the terminal interface 152 and the storage interface 154).
[0022] Logic modules throughout the computer system 100—including but not limited to the memory 120, the processor 110, and the I / O interface 130—may communicate faults and changes in one or more components to a hypervisor or operating system (not depicted). The hypervisor or operating system may allocate the various resources available in the computer system 100 and track the location of data in the memory 120 and the location of processes assigned to each core 112. In embodiments where elements are combined or rearranged, aspects and capabilities of the logic modules may be combined or redistributed. Such variations are obvious to those skilled in the art.
[0023] I. Streaming Computation
[0024] Streaming computation may allow a user to process big data and continuously provide advanced metrics on the big data as it is generated by various sources. A streaming application may provide streaming computation by generating a configuration of one or more processing elements, each processing element containing one or more stream operators. For example, a streaming application may be compiled using fixed logic included in each processing element and / or stream operator. Each processing element and / or stream operator of a streaming application may process big data by generating and modifying information in the form of tuples. Each tuple may have one or more attributes (e.g., a tuple may be analogous to a row in a table and an attribute may be analogous to a column in a table).
[0025] A streaming application may deploy an instance of the configuration to a collection of hardware computing nodes. The streaming application may then manage the instance by adapting the hardware to execute the streaming application when it is configured, such as by load balancing processing elements onto, across, or a portion of a given computing node.
[0026] Figure 2Depicts an example stream computing application (stream application) 200 configured to operate on a stream of tuples, consistent with some embodiments of the present disclosure. The stream application 200 may be represented in the form of an operator graph 202. The operator graph 202 may visually represent to a user the data flow through the stream application 200. The operator graph 202 may define how tuples are routed through the various components of the stream application 200 (e.g., the execution path, Figure 2 the pre-compiled logical layout of processing and resource allocation represented by the curved arrow lines in). The stream application 200 may include one or more computing nodes 210-1, 210-2, 210-3, and 210-4 (collectively referred to as 210); a development system 220; a management system 230; one or more processing elements 240-1, 240-2, 240-3, 240-4, 240-5, and 240-6 (collectively referred to as 240); one or more stream operators 242-1, 242-2, 242-3, 242-4, 242-5, 242-6, 242-7 (collectively referred to as 242); and a network 250.
[0027] The stream application 200 may receive information from one or more sources 244. The stream application 200 may output information to one or more sinks 246. The input may come from outside the stream application 200, such as from multiple Internet of Things (IoT) devices. The stream network 250 may be a communication layer that handles connections, sends, and receives data between parts of the stream application 200. For example, the stream network 250 may be a transport layer for data packets within the stream application 200 and is configured to communicatively couple the processing elements 240.
[0028] The configuration of the stream application 200 depicted by the operator graph 202 is just an example stream application. The stream application may vary in the number of computing nodes, processing elements, or stream operators. The stream application may also change the roles and / or responsibilities performed by any component, or may include other components not depicted. For example, some or all of the functions of the development system 220 may be performed by the management system 230. In another example, the functions of the development system 220 and the management system 230 may be performed by a single management system (not depicted). The management system may be configured to perform these tasks without departing from the embodiments disclosed herein. In yet another example, the functions of the development system 220 and the management system 230 may be performed by multiple services (e.g., ten or more separate software programs, each configured to perform a specific function).
[0029] A compute node 210 can be a computer system and can each include the following components: a processor, a memory, and an input / output interface (herein, I / O). Each compute node 210 can also include an operating system or a hypervisor. In some embodiments, the compute node 210 can perform operations for the development system 220, the management system 230, the processing element 240, and / or the stream operator 242. The compute node 210 can be classified as a management host, an application host, or a hybrid host. A management host can perform operations for the development system 220 and / or the management system 230. An application host can perform operations for the processing element 240 and the stream operator 242. A hybrid host can perform operations for both the management host and the application host. Figure 1 Depicted is a computer system 100 that can be a compute node consistent with some embodiments.
[0030] A network (not shown) can couple each node 210 together (e.g., a local area network, the Internet, etc.). For example, node 210-1 can communicate with nodes 210-2, 210-3, and 210-4 via the network. The compute node 210 can communicate with the network via the I / O, and the network can include various physical communication channels or links. The link can be wired, wireless, optical, or any other suitable medium. The network can include various network hardware and software for performing routing, switching, and other functions, such as routers, switches, or bridges. The node 210 can communicate via various protocols (e.g., Internet Protocol, Transmission Control Protocol, File Transfer Protocol, Hypertext Transfer Protocol, etc.). In some embodiments, the node 210 can share the network with other hardware, software, or services (not shown).
[0031] The development system 220 can provide a user with the ability to create a stream application targeted at processing a specific data set. The development system 220 can operate on an instance of a computer system (not shown) such as the computer system 100. The development system 220 can operate on one or more of the compute nodes 210. The development system 220 can generate one or more configuration files that describe the stream computing application 200 (e.g., the processing element 240, the stream operator 242, the source 244, the sink 246, the allocation of the aforementioned compute node 210, etc.). The development system 220 can receive a request from the user to generate the stream application 200. The development system 220 can receive a request from the user to generate other stream applications (not depicted). The development system 220 can communicate with the management system 230 to transfer the configuration on any stream application that the development system 220 can create.
[0032] The development system 220 can generate a configuration by considering the performance characteristics of software components (e.g., processing elements 240, stream operators 242, etc.), hardware (e.g., compute nodes 210, network), and data (e.g., source 244, format of tuples, etc.). In a first example, the development system 220 can determine that running processing elements 240-1, 240-2, and 240-3 together on compute node 210-1 incurs an overhead that results in better performance than running them on separate compute nodes. In a second example, the development system 220 can determine that the memory footprint of placing stream operators 242-3, 242-4, 242-5, and 242-6 into a single processing unit 240-5 is larger than the cache of the first processor in compute node 210-2. To reserve memory space within the cache of the first processor, the development system 220 can decide to place only stream operators 242-4, 242-5, and 242-6 into a single processing unit 240-5, regardless of the inter-process communication latency with two processing units 240-4 and 240-5.
[0033] In a third example of considering performance characteristics, the development system 220 can identify a first operation (e.g., an operation executed on processing element 240-6 on compute node 210-3) that requires a larger amount of resources for the stream application 200. The development system 220 can allocate a larger amount of resources (e.g., operate processing element 240-6 on compute node 210-4 in addition to compute node 210-3) to assist in the execution of the first operation. The development system 220 can identify a second operation (e.g., an operation executed on processing element 240-1) that requires a smaller amount of resources within the stream application 200. The development system 220 can also determine that the stream application 200 can operate more efficiently with an increase in parallelization (e.g., more instances of processing element 240-1). The development system 220 can create multiple instances of processing element 240-1 (e.g., processing elements 240-2 and 240-3). The development system 220 can then assign processing elements 240-1, 240-2, and 240-3 to a single resource (e.g., compute node 210-1). Finally, the development system 220 can identify a third operation and a fourth operation (e.g., operations executed on processing elements 240-4 and 240-5), each of which requires low-level resources. The development system 220 can assign a smaller amount of resources to the two different operations (e.g., have them share the resources of compute node 210-2 instead of executing each operation on its own compute node).
[0034] The development system 220 may include a compiler (not shown) of compilation modules (e.g., processing element 240, stream operator 242, etc.). The modules may be source code or other program statements. The modules may be in the form of requests from a stream processing language (e.g., a computational language that includes declarative statements that allow a user to state a particular subset of information formatted in a particular way). The compiler may convert the modules into target code (e.g., machine code targeted at a particular instruction set architecture of the compute nodes 210). The compiler may convert the modules into an intermediate form (e.g., virtual machine code). The compiler may be a just-in-time compiler that executes as part of an interpreter. In some embodiments, the compiler may be an optimizing compiler. In some embodiments, the compiler may perform peephole optimizations, local optimizations, loop optimizations, interprocedural or whole-program optimizations, machine code optimizations, or any other optimizations that reduce the amount of time required to execute the object code, reduce the amount of memory required to execute the object code, or both.
[0035] The management system 230 may monitor and manage the stream application 200. The management system 230 may operate on an instance of a computer system (not shown), such as computer system 100. The management system 230 may operate on one or more compute nodes 210. The management system 230 may also provide an operator graph 202 of the stream application 200. The management system 230 may host services that make up the stream application 200 (e.g., services that monitor the health of the compute nodes 210, the performance of the processing element 240 and the stream operator 242, etc.). The management system 230 may receive requests from a user (e.g., requests for authenticating and authorizing users of the stream application 210, requests for viewing information generated by the stream application, requests for viewing the operator graph 202, etc.).
[0036] The management system 230 can provide a user with the ability to create multiple instances of the flow application 200 configured by the development system 220. For example, if a second instance of the flow application 200 is required to perform the same processing, the management system 230 can allocate a second set of computing nodes (not depicted) for executing the second instance of the flow application. The management system 230 can also reallocate the computing nodes 210 to relieve bottlenecks in the system. For example, as shown, the processing elements 240-4 and 240-5 are executed by the computing node 210-2, and the processing element 240-6 is executed by the computing nodes 210-3 and 210-4. In one case, the flow application 200 may experience performance issues because the processing elements 240-4 and 240-5 do not provide tuples to the processing element 240-6 before the processing element 240-6 enters the idle state. The management system 230 can detect these performance issues and can reallocate resources from the computing node 210-4 to execute a part or all of the processing element 240-4, thereby reducing the workload on the computing node 210-2. The management system 230 can also perform operations on the operating systems of the computing nodes 210, such as load balancing and resource allocation for the processing elements 240 and the flow operator 242. By performing operations on the operating system, the management system 230 can enable the flow application 200 to more effectively utilize the available hardware resources and improve performance (e.g., by reducing the overhead of the operating system and the multiprocessing hardware of the computing nodes 210).
[0037] The processing element 240 can execute the operations of the flow application 200. Each processing element 240 can operate on one or more computing nodes 210. In some embodiments, a given processing element 240 can operate on a subset of a given computing node 210, such as a processor or a single core of a processor of the computing node 210. In some embodiments, a given processing element 240 can correspond to an operating system process hosted by the computing node 210. In some embodiments, a given processing element 240 can operate on multiple computing nodes 210. The processing element 240 can be generated by the development system 220. Each processing element 240 can be in the form of a binary file and additional library files (e.g., an executable file and an associated library, a package file containing executable code and associated resources, etc.).
[0038] Each processing element in processing element 240 may include configuration information from development system 220 or management system 230 (e.g., resources and conventions required by the associated compute node 210 to which it is assigned, identities and credentials required to communicate with source 244 or sink 246, and identities and credentials required to communicate with other processing elements, etc.). Each processing element 240 may be configured by development system 220 to operate optimally on one of compute nodes 210. For example, processing elements 240-1, 240-2, and 240-3 may be compiled to take advantage of optimizations identified by the operating system running on compute node 210-1, and may also be optimized for the specific hardware of compute node 210-1 (e.g., instruction set architecture, configured resources such as memory and processors, etc.).
[0039] Each processing element 240 may include one or more stream operators 242 that perform the basic functions of stream application 200. As tuple streams flow through processing element 240 as directed by operator graph 202, they are passed from one stream operator to another (e.g., a first processing element may process a tuple and place the processed tuple in a queue assigned to a second processing element, a first stream operator may process a tuple and write the processed tuple to a memory region designated for a second stream operator, the processed tuple may not be moved, but metadata updates may be utilized to indicate that they are ready to be processed by a new processing element or stream operator, etc.). Multiple stream operators 242 within the same processing element 240 may benefit from architectural efficiencies (e.g., reduced cache misses, shared variables and logic, reduced memory swapping, etc.). Processing element 240 and stream operator 242 may utilize interprocess communication (e.g., network sockets, shared memory, message queues, message passing, semaphores, etc.). Processing element 240 and stream operator 242 may utilize different interprocess communication techniques depending on the configuration of stream application 200. For example: stream operator 242-1 may use a semaphore to communicate with stream operator 242-2; processing element 240-1 may use message QUE to communicate with processing element 240-3; and processing element 240-2 may use network sockets to communicate with processing element 240-4.
[0040] The stream operator 242 can perform the basic logic and operations of the stream application 200 (e.g., process tuples and pass the processed tuples to other components of the stream application). By separating the logic that can occur within a single larger program into basic operations performed by the stream operator 242, the stream application 200 can provide greater scalability. For example, dozens of compute nodes hosting hundreds of stream operators in a given stream application can enable the processing of millions of tuples per second. This logic can be created by the development system 220 prior to the runtime of the stream application 200. In some embodiments, the source 244 and the sink 246 can also be stream operators 242. In some embodiments, the source 244 and the sink 246 can link multiple stream applications together (e.g., the source 244 can be the sink for a second stream application, and the sink 246 can be the source for a third stream application). The stream operator 242 can be configured by the development system 220 to optimally execute the stream application 200 using the available compute nodes 210. The stream operator 242 can send and receive tuples from other stream operators. The stream operator 242 can receive tuples from the source 244 and can send the tuples to the sink 246.
[0041] The stream operator 242 can perform operations on the attributes of a tuple (e.g., conditional logic, iterative loop structures, type conversions, string formatting, filtering statements, etc.). In some embodiments, each stream operator 242 can perform only very simple operations and can pass the updated tuple to another stream operator in the stream application 200 - simple stream operators can be more scalable and easier to parallelize. For example, the stream operator 242-2 can receive a date value with a specific precision and can round the date value to a lower precision and pass the changed date value to the stream operator 242-4, which can change the changed date value from a 24-hour format to a 12-hour format. A given stream operator 242 can leave anything about the tuple unchanged. The stream operator 242 can perform operations on a tuple by adding new attributes or removing existing attributes.
[0042] The stream operator 242 can perform operations on a stream of tuples by routing some tuples to a first stream operator and other tuples to a second stream operator (e.g., stream operator 242-2 sends some tuples to stream operator 242-3 and other tuples to stream operator 242-4). The stream operator 242 can perform operations on a stream of tuples by filtering some tuples (e.g., discarding some tuples and passing a subset of the stream to another stream operator). The stream operator 242 can also perform operations on a stream of tuples by routing some in the stream to itself (e.g., stream operator 242-4 can perform simple arithmetic operations, and as part of its operations, it can perform a logical loop and direct a subset of the tuples to itself). In some embodiments, a particular tuple output by the stream operator 242 or the processing element 240 may not be considered the same tuple as the corresponding input tuple, even if the stream operator or processing element does not change the input tuple.
[0043] In some cases, a stream application can be primarily a static big data operation mechanism. Once configured, such a stream application may be immutable in the context it provides to users. For example, after developing the operator graph 202, the development system 220 can provide a generated form of the operator graph with the configuration of the stream application 200 to the management system 230 for execution by the compute nodes 210. Additionally, in some cases, such a stream application performs certain logic regarding how it processes tuples. Once configured, this logic may be non-updatable or changeable until a new stream application is compiled. Due to the real-time continuous nature of stream application and information stream application processing, attempting to provide an update to the processing element or stream operator of such a configured stream instance may be impractical. For example, any downtime, even in microseconds, can cause the stream application to not collect one or more tuples during the transition from the original configured processing element to the updated processing element. Losing a portion of the data may result in partial or complete failure of the stream application, and the stream application may be unable to provide the context to the big data source to the user.
[0044] Another problem can occur when too much or too little data flows through a streaming application. For example, the logic in a given streaming operator can provide a streaming application that processes only a subset, selection, or portion of the tuples. If too few tuples are processed based on the configuration of the streaming operator, it can lead to poor analytical values because the data set is too small to derive meaning. To compensate, a pre-compiled streaming application can be configured to ingest and process many tuples. If too many tuples are processed based on the configuration of the streaming operator, there can be bottlenecks in the system (e.g., processor or memory starvation). In other instances, one or more tuples can be discarded because a part of the streaming application is overwhelmed (e.g., when a large number of tuples are received, one or more computing nodes may not process some of the tuples). For example, when too much data (too many tuples) floods the streaming application, random tuples may be discarded, or the system becomes clogged and crashes.
[0045] The non-optimized configuration of a streaming application can be based not on a fault or bad intention, but rather on practical considerations. Specifically, a given streaming application can be generated based on the type or amount of test data created during the development phase of the streaming application. Since the analysis of big data is a relatively new phase of computing, it can be difficult to predict or anticipate the type, amount, and optimal route of tuple processing in a streaming application. For example, the streaming application 200 can be developed by the development system 220 based on test data or based on the source 244 under a different set of conditions. As a result, the development system 220 can allocate computing resources and data streams based on an incomplete or stale understanding of data processing (as represented by the curve in Figure 2 ). As the analysis of the situation evolves, the configuration of the operator graph 202 may not address the actual problems of production data, or update the real-world events or information that are changing, and thus affect the type and / or amount of data from the source 244.
[0046] Given some of these drawbacks, there may be a practical need to develop real-time analysis applications to address many different factors. The first factor could be the need for static analysis of good / bad data formats, values, etc. prior to developing the application. This need for static analysis may be required to properly develop and configure a data analysis system. The second factor could be planning for many failure scenarios, source changes, or incorrect assumptions about the data. For example, data processing can be performed based on assumptions about the data provided by the source and the best way to handle the data. The assumptions can help in designing and configuring prior to compiling, instantiating, and running various stream operators, processing elements, and the flow between them; the same assumptions can lead to later problems during the actual operation and runtime of the stream application (e.g., incorrect data tuple formats, data tuple value checks, or other misalignments of the configuration and the data). The third factor may be that, for various reasons, one or more parts of the data analysis or processing system may need to be restarted to enable certain functions, or to change the logic of a given system or the configuration of the auxiliary non-core logic. Example scenarios that can cause a stream application to restart can include changing the logging level, performing increased transaction processing, or alleviating resource bottlenecks by adjusting the conditions for processing tuples (e.g., passing, adding, or deleting tuples).
[0047] One attempt to mitigate these drawbacks could be to utilize a microservices architecture. A microservices architecture could attempt to utilize small applications that subscribe to data streams, perform certain functions, and then publish them. A microservices architecture can be an incomplete or partial solution. For example, a stream application analysis application needs to know and understand each and all potential subscription plans prior to deployment and runtime. A stream application may not be configurable for every scenario, and as a detriment, may have to start over to reconfigure or arrange the output of certain parts of the stream (e.g., the output of a processing element or a stream operator) to correct these deficiencies).
[0048] Another potential attempt to mitigate these drawbacks could be a job overlay architecture. A job overlay architecture can focus on determining the potential need to execute certain programming logic on a running data analysis application, but again, this particular architecture may have drawbacks for a stream application. For example, a job overlay architecture can focus on changing a part of the core logic ("job logic") of a part of a stream application, saving the new logic to the application package, and applying those changes to the running part. However, this solution does not address the usability issue because it requires restarting the running process to update the job logic. For example, once a part of a stream application is running, it may not be changed.
[0049] ii.d. and combination of ii. Aspect-aware stream computing
[0050] Aspect-Aware Streaming Computation (“AASC”) can overcome problems associated with other types of data analysis in the context of streaming applications. AASC can enable streaming applications to operate based on one or more aspect-oriented programming techniques. AASC can include exposing request handling logic (“hooks”) for receipt by aspects. The hooks can be in the form of a general program execution structure. The general program execution structure can be configured to receive or use any code, rules, or other logic provided by a requester. AASC can affect the operation of a streaming application by being executed by the general program execution structure. Specifically, a request can be sent that includes a set of one or more program code instructions. The set of program code instructions can operate on a tuple stream and outside of the logic built into the stream operators and / or processing elements of the streaming application.
[0051] AASC can be configured to affect the core or tertiary functions of a streaming application. In detail, a streaming application can have a first configuration developed prior to runtime that is formulated and operated during the runtime and operation of the streaming application. The general program execution structure can be executed during or outside of the operation of the streaming application. In a first example, the general program execution structure can execute a set of program code instructions immediately before and / or after a portion of the streaming application. Rules can change the functionality of the streaming application by rerouting, updating, deleting, adding, or otherwise modifying one or more tuples of the streaming application.
[0052] In some embodiments, AASC can operate by encapsulating a portion of a streaming application (e.g., a stream operator, a processing element). Encapsulating a portion of the streaming application can include executing the general program execution structure (which is configured to execute the received set of program code instructions) only outside of the logic of the constituent portion. For example, the encapsulation can be configured as the first programming step after or the last programming step before any code that is part of a given stream operator. Encapsulating a portion of the streaming application can include executing the execution of the general program execution structure as an initial or termination step of a portion of the streaming application. For example, the encapsulation can be configured as the first programming step and / or the last programming step of a given stream operator.
[0053] AASC can operate without affecting the availability and execution of streaming applications. Specifically, AASC does not cause downtime of running applications, and the core logic of each streaming operator and processing element executed based on pre-compiled logic is not affected or altered. For example, a first streaming operator can be configured to operate on tuples based on a filtering statement or other logic to affect only certain tuples of a tuple stream and ignore or pass on other tuples of the tuple stream. A general program execution structure can execute a set of program code instructions just before, after, external to, or separate from the original compiled logic of the streaming application. Thus, only the effects of the streaming application are changed (e.g., receipt, change, creation, deletion, and output of tuples), without changing the core logic.
[0054] AASC can offer advantages over other streaming computing methods. Specifically, AASC can facilitate the rapid or temporary injection and / or removal of aspects capable of performing various functions without the need to bounce back, restart, or otherwise temporarily take the streaming application offline. AASC can be used to continuously deliver aspects into an actively running and pre-compiled streaming application. This scalability of the streaming application through AASC may be a technical necessity in environments where the streaming application performs mission-critical or other infrastructure operations. For example, a power station can operate based on a streaming application that monitors multiple sensors that observe various energy particles, chemical mixtures, etc. AASC can also facilitate the tuning or fixing of streaming applications by developers and administrators of the streaming application (e.g., to handle changed source data and increase the accuracy of the calculations performed by the streaming application). Tuning or fixing can be performed during the execution of the streaming application without the risk of potential data loss or undergoing reconfiguration of the streaming application, as well as testing and restarting the streaming application. AASC can utilize dynamic proxy conventions, such as Java dynamic proxies.
[0055] The injection of program code instructions by AASC can benefit the debugging of applications through logging or other telemetry techniques. The injection of program code instructions can allow and / or deny certain components or connections for security. The injection of program code instructions can promote higher uptime and streaming availability by facilitating the rerouting of data to available services in the case of downstream network and / or computing node downtime or otherwise being offline. The injection of program code instructions can increase data accuracy by executing one or more data integrity routines to verify certain tuples. The injection of program code instructions can allow a production environment to demonstrate functionality by inserting demonstration data (e.g., tuples containing demonstration data) on a portion of the streaming application using a first set of program code instructions. For example, the first portion of the streaming application can be located just before a subset of processing elements and / or streaming operators.
[0056] The AASC can utilize a first general program execution structure that is just before the subset to execute a set of program code instructions configured to be inserted into the stream application downgraded tuples. A subset of the processing unit and / or the stream operator can operate on the downgraded tuples and generate, modify, or otherwise produce tuples based on the downgraded tuples. A second set of program code instructions can be executed by a second general program execution structure that is after the subset of the stream application. The second set of program code instructions can be configured to extract the tuples generated based on the downgraded tuples. In addition, the second set of program code instructions can be configured to identify any existing tuples modified by the subset of the stream application as a result of the downgraded tuples and undo the modification.
[0057] Figure 3 is an example stream application 300 configured as an AASC according to some embodiments of the present disclosure. The stream application 300 can be represented in the form of an operator graph 302. The operator graph 302 can visually represent to the user the data flow through the stream application 300. The operator graph 302 can have a configuration similar to that of Figure 2 the operator graph 202, except that the lines representing the tuple flow through the stream application 300 are not depicted for the sake of clarity. The operator subgraph 302 can define how to route tuples through various components of the stream application 300 (e.g., execution paths, pre-compiled logic layout of processing, and resource allocation). The stream application 300 can include one or more computing nodes 310-1, 310-2, 310-3, and 310-4 (collectively referred to as 310); a development system 320; a management system 330; one or more processing elements 340-1, 340-2, 340-3, 340-4, 340-5, and 340-6 (collectively referred to as 340); one or more stream operators 342-1, 342-2, 342-3, 342-4, 342-5, 342-6, 342-7 (collectively referred to as 342); and a network 350. The stream application 300 can also include multiple general program execution structures. Specifically, a first plurality of general program execution structures 362-1, 362-2 (collectively referred to as 362); a second plurality of general program execution structures 364-1, 364-2 (collectively referred to as 364); a third plurality of general program execution structures 366-1, 366-2, 366-3, and 366-4 (collectively referred to as 366); and a fourth plurality of general program execution structures 368-1 and 368-2 The stream application 300 can also include an aspect data repository 370.
[0058] The streaming application 300 can receive information from one or more sources 344. The streaming application 300 can output information to one or more sinks 346. The input can be from outside the streaming application 300, such as from multiple IoT devices. The streaming network 350 can be a communication layer that processes connections, sends, and receives data between parts of the streaming application 300. For example, the streaming network 350 can be a transport layer for data packets within the streaming application 300 and is configured to communicatively couple the processing elements 340.
[0059] Each of the general program execution structures 362, 364, 366, or 368 can be configured to receive, process, and execute requests. For example, the general program execution structure 364-2 can be configured to listen for a request processor or hook that executes a set of program code instructions. In another example, the general program execution structure 362-1 can start executing the received set of program code instructions in response to receiving a set of program code instructions. Each of the general program execution structures 362, 364, 366, or 368 can operate by using the streaming network 350. For example, the program execution structure 366-3 can operate through a network socket to accept a connection from the management system 330.
[0060] The general program execution structures 362, 364, 366, or 368 can be configured by the management system 330. Specifically, the management system 330 can default to operating by not invoking, not operating, or otherwise not providing processing cycles to any specific code outside the streaming application 300 (e.g., the processing element 340, the stream operator 342). The management system 330 can operate in a preemptive awareness mode. For example, during the operation of the streaming application 300, the management system can start programming code into one or more of the general program execution structures 362, 364, 366, 368 even before receiving a request to execute a set of program code instructions. The management system 330 can operate in a reactive awareness mode. For example, during the operation of the streaming application 300, the management system 330 can not encapsulate any part of the streaming application 300 until receiving a request to execute a set of program code instructions. The streaming application 300 can operate in a default awareness mode. For example, at the start of data processing, in addition to allocating the computing nodes 310 to the processing element 340 and the stream operator 342, the management system 330 can also allocate the general program code structures 362, 364, 366, 368 to parts of the streaming application 300.
[0061] The general program execution structures 362, 364, 366, and 368 can each encapsulate a part of the stream application 300. Specifically, the general program execution structure 362-1 can be configured to occur logically just before the processing elements 340-1, 340-2, and 340-3. Additionally, the general program execution structure 362-2 can be configured to occur logically just after the processing elements 340-1, 340-2, and 340-3. Thus, any set of program code instructions dispatched to the general program execution structure 362-1 can be executed respectively just before the processing elements 340-1, 340-2, and 340-3. As a result of the encapsulation, the set of program code instructions dispatched to the general program execution structure 362 can operate on the tuple stream either before or after the tuple stream is processed by the processing elements 340-1, 340-2, and 340-3. Thus, the stream operators 342-1 and 342-2 can also be affected by the general program execution structure 362.
[0062] A given part of the encapsulated stream application can be multiple processing elements. For example, the multiple processing elements 340-1, 340-2, and 340-3 can all be encapsulated respectively by only two general program execution structures 362-1 and 362-2. A given part of the encapsulated stream application can be a single processing element. For example, the general program execution structures 364-1 and 364-2 can encapsulate a single processing element 340-4. A given part of the encapsulated stream application can be multiple stream operators. For example, the general program execution structures 366-1 and 366-2 can encapsulate the stream operators 342-4 and 342-6. A given part of the encapsulated stream application can be a single stream operator. For example, the general program execution structures 366-3 and 366-4 can encapsulate the stream operator 342-5.
[0063] Computing resources can be allocated by the streaming application 300 to the general-purpose program execution structures 362, 364, 366, and 368. The allocated computing resources can be part of the computing resources for executing the streaming application 300. For example, one or more of the computing nodes 310 can be allocated to process the general-purpose program execution structures 362, 364, 366, and 368. The allocation and / or assignment of processing cycles and memory can be based on the operator graph 302 and based on the layout of portions of the streaming application 300. For example, the general-purpose program execution structures 362-1 and 362-2 can be allocated to the computing node 310-1 based on the structures 362-1 and 362-2 that encapsulate portions of the streaming application that are also allocated to the computing node 310-1 (e.g., the stream operator 342-1, the processing unit 340-2). In another example, based on the configuration of the stream operator 342-5 that is also assigned to the first execution thread, the general-purpose program execution structures 366-3 and 366-4 can be assigned to the first execution thread on the computing node 310-2. The assignment of similar threads or other multiprocessing structures as part of the encapsulation can result in performance efficiencies (e.g., reduced thread lock contention, elimination of process deadlocks, reduced inter-process communication). The general-purpose program execution structures 362, 364, 366, and 368 can be allocated to the computing resources of another computer (not shown) external to the streaming application 300. The assignment to another computer can improve the performance of the streaming application 300 by having additional computing resources process the tuple stream. Each of the general-purpose program execution structures 362, 364, 366, and 368 can be configured to receive and execute program code instructions outside of the preconfigured operations of the streaming application 300. For example, a given stream operator 342-1 or processing unit 340-2 can be configured and compiled to operate on a first subset of tuples. The corresponding general-purpose program execution structures 366-3 and / or 366-4 can be configured to operate on tuples outside of the first subset of tuples.
[0064] The management system 330 can be configured to respond to a request to execute program code instructions (e.g., request 380). The request 380 can be an example of a request provided to the AASC streaming application 300. The request 380 can be received from a user.
[0065] Request 380 may include a set of program code instructions. The set of program code instructions may be in the form of a package or other related software constructs, such as in the form of an aspect package. The set of program code instructions may include a compiled software library. The compiled software library may be configured in an executable format compatible with a particular part of the stream application. For example, the stream operator 342-7 may be configured in a first software language, and if a set of program code instructions is directed to the general program execution structure 368-1 or 368-2, then the set of program code instructions may need to be in the first software language for it to be executable. The set of program code instructions may be written only into a specific set of approved interfaces. For example, a single-functional implementation that will accept data tuples and return data tuples using a tuple pattern defined by an approved interface.
[0066] Request 380 may also include a set of program execution instructions. The set of program execution instructions may include a stream application target. The stream application target may be a specific processing element or stream operator of a given stream application. For example, request 380 may include a stream application target that includes a stream operator identifier specifically identifying stream operator 342-5, and request 380 may thus indicate that the included set of program code instructions is to be executed by the general program execution structures 366-3 and 366-4 (e.g., the general program execution structures encapsulating stream operator 342-5).
[0067] The management system 330 may instruct at least one given general program execution structure among the general program execution structures 362, 364, 366, or 368 to execute the received set of program code instructions (e.g., the program code instructions that are part of request 380). A specific part of a set of parts of the stream application 300 may be based on request 380. Specifically, the specific part may specify a part of the instructions that the network 350 may receive (e.g., an existing job ID, a given stream operator 342 based on location or port) and the set of program code instructions to be executed. The management system 330 may identify a specific stream application target based on the request. For example, request 380 may include a specific type of stream operation, such as all "SET" operations applicable to the included set of program code instructions. In response, the management system 330 may scan the operator graph 302 to identify a subset of one or more stream operators 342 or processing elements 340 that perform the "SET" operation. The management system 330 may instruct a given general program execution structure 362, 364, 366, or 368 to execute the received set of program code instructions under one or more conditions (e.g., the number of times to execute the program code instructions, the amount of time to execute the program code instructions up to a certain amount, execute the program code instructions before or after one or more conditions occur at a part of the stream application 300, or before or after input or output). The conditions may be received as part of the program execution instructions (e.g., part of request 380).
[0068] The management system 330 can also store program code instructions for later use. Specifically, once the management system 330 receives a request 380, the management system 330 can store the set of program code instructions in the aspect data repository 370. The aspect data repository 370 can be a database, a memory, a tertiary storage, or other data structure configured to store aspect bundles and other program code instructions. The storage in the aspect data repository 370 can achieve high availability of the AASC for the streaming application 300. Future requests may only need to identify a specific set of program code instructions (e.g., by file name or other unique identifier). In response to receiving a request with an identifier but without a specific set of program code instructions, the management system 330 can be configured to scan the aspect data repository 370 to retrieve the stored set of program code instructions. Other requests from different users can identify specific sets of program code instructions by the identifier to be run at different times.
[0069] The management system 330 can operate in a secure mode. Specifically, the management system 330 can be configured to handle the authentication and authorization of the submitted requests. Any submitted request can be authorized by the management system 330 before being allowed to pass to a specific general-purpose program execution structure 362, 364, 366, or 368. The management system 330 and any requester can communicate through a secure communication channel (e.g., through an encrypted network, using public / private certificates). The requester may not be able to directly execute the set of program code instructions, but can only submit and authenticate through the management system 330.
[0070] The AASC of the streaming application 300 can facilitate a set of extensible enhancements that can be invoked during streaming computing operations. Specifically, various sets of program code instructions can be executed by one or more of the general-purpose program execution structures 362, 364, 366, and / or 368 to perform various operations on the streaming application that result in specific use cases. For example, the first use case can be tuple generation and injection, which acts as a dynamic source to inject tuples into a stream operator for replaying a stream, debugging operator logic, and on-demand tuple processing. In another example, the second use case can be tuple validation, which checks whether a tuple is valid according to a schema or within a valid range; in addition, the set of program code instructions can also be configured to discard tuples that can be determined to be invalid. In yet another example, the third use case can be tuple transformation, which includes dynamic customization logic that can be added to manipulate certain tuples.
[0071] Other enhancements facilitated by the AASC of the streaming application 300 can include a set of program code for performing the following operations: logging, debugging, profiling, notification, caching, rerouting, discarding, and transaction processing. In a first example, logging can occur via debug or audit logging, which can be configured to include a given user, data, time, and whether the operation was successful. In a second example, profiling can occur by comparing timestamps before and after processing occurs to determine how long it takes the stream operator 342 to process a tuple. In a third example, a notification can be sent when certain criteria defined in a set of program code instructions are met. In a fourth example, caching can occur based on program code instructions configured to save certain tuples for offline processing. In a fifth example, rerouting can occur based on program code instructions that define rules for routing certain tuples to another stream operator, such as when the first stream operator is overloaded or experiencing certain performance issues. In a sixth example, discarding can occur based on program code instructions that define rules for discarding tuples, such as to relieve an overloaded stream operator, or based on updated filtering conditions. In a seventh example, transaction processing can occur based on program code instructions to split the data of a single tuple into multiple tuples. The logic in the seventh example can include program code instructions that gate the continuation of tuple processing to occur only if all tuples in a transaction are received by a given stream operator 342. In an eighth example, security can occur based on program code instructions configured to block data from certain connections or users deemed malicious.
[0072] Figure 4 Method 400 for performing aspect-aware stream computing in accordance with some embodiments of the present disclosure is depicted. Method 400 can generally be implemented in fixed function hardware, configurable logic, logic instructions, etc., or any combination thereof. For example, logic instructions can include assembly instructions, ISA instructions, machine-dependent instructions, microcode, status setting data, configuration data for an integrated circuit, and other structural components that localize to an electronic circuit and / or hardware (e.g., a host processor, a central processing unit / CPU, a microcontroller, etc.). Method 400 can be executed by a single computer such as computer system 100. Method 400 can be executed by a streaming application (e.g., streaming application 400).
[0073] From start 405, method 400 can begin by monitoring the activity of the streaming application. Monitoring of the streaming application can be performed by determining whether one or more processing elements and / or stream operators of the streaming application are operable and configured to actively process a tuple stream.
[0074] If the 420 stream application is active: Y, method 400 can continue at 430 via a listener request. The listener request can occur from a part of the stream application. For example, the management system of the stream application can operate via the listener request.
[0075] If a program request is received, then at 440: Y, method 400 can continue by identifying the stream application target of the stream application at 450. Specifically, the request can include a specific stream operator or processing element that is the target of the request (e.g., based on an identifier). The request can include a group of a specific type, category, or part of the stream application. For example, a stream application target specifying "any SET stream operator" can be included in the request. The request may not include a stream application target, but can indirectly identify the stream application target. For example, the request can include a specific identifier of a set of program code instructions to be executed. The request can include only the identifier of the program code instructions and not the target. Upon receiving the request, a program repository (e.g., the aspect data repository 370) can be scanned, searched, or otherwise accessed. The identification of the stream application target can be performed by retrieving the program code instructions from the program repository and scanning the program code instructions to determine the part of the stream application to be processed by the program code instructions.
[0076] At 460, a part of the stream application can be encapsulated. The stream application program can be encapsulated with a general program execution structure configured to receive and execute program code instructions. The stream application can be encapsulated during the runtime of the stream application. For example, a given stream operator or processing element can be configured with a first input and a first output. The first input can be configured to receive a plurality of tuples of a tuple stream, and the first output can be configured to output a second plurality of tuples of the tuple stream. The given stream operator or processing unit can also be configured with a second input; the second input can be configured to receive an aspect flag, such as a boolean value or a numerical value. By default, at the start of the execution of the stream application, the default configuration of the stream operator or processing element can operate without any general program execution structure being reachable or operable. In the absence of receiving any aspect flag, the stream application can operate without any general program execution structure. Upon receiving the flag, the stream application can start operating with the encapsulated general program execution structure.
[0077] At 470, a portion of the streaming application can execute the identified program code instructions. This portion of the streaming application can be a collection of one or more processing cycles outside of this portion of the streaming application. For example, a management system or other portions that are not part of the pre-compiled logic of the streaming application can transfer the program code instructions to a general-purpose program execution structure. The general-purpose program code execution structure can execute the program code instructions. After executing the program code instructions, method 400 can continue by monitoring the streaming application at 410. If at 410 the streaming application is not active: N, method 400 can end at 495.
[0078] The present invention can be a system, method, and / or computer program product at any possible technical detail integration level. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0079] The computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions used by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any appropriate combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0080] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0081] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, in order to carry out aspects of the present invention, an electronic circuit, including for example a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit.
[0082] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0083] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0084] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the blocks may not occur in the order noted in the figures. For example, two blocks shown in succession may actually be implemented as one step, executed simultaneously, substantially simultaneously, in partial or all-time overlap, or these blocks may sometimes be executed in the reverse order, depending on the functions involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0086] The description of the various embodiments of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A method for processing a tuple stream, the method comprising: Monitoring a tuple stream to be processed by a plurality of processing elements operating on one or more computing nodes of a stream application, each processing element having one or more stream operators; Receiving a program request to execute a first set of one or more program code instructions; Based on the program request, identifying a stream application target of the first set of program code instructions; During operation of the stream application, encapsulating a first portion of a set of one or more portions of the stream application using a general-purpose program execution structure configured to receive and execute program code instructions outside of pre-configured operations of the stream application; And During operation of the stream application, and in response to the program request and based on the stream application target and the general-purpose program execution structure, executing the first set of program code instructions during a first time period of execution of the first portion of the stream application.
2. The method according to claim 1, wherein The program request includes the first set of program code instructions.
3. The method according to claim 1, wherein The program request includes a program identifier associated with the first set of program code instructions.
4. The method according to claim 3, wherein, The first set of program code instructions is located in a program repository.
5. The method according to claim 4, wherein The program repository is part of the stream application.
6. The method according to claim 1, wherein The program request includes a first set of one or more program execution instructions associated with the first set of program code instructions.
7. The method according to claim 6, wherein, The first set of program execution instructions includes the stream application target.
8. The method according to claim 6, wherein, The first set of program execution instructions includes the number of times to execute the first set of program code instructions.
9. The method according to claim 6, wherein, The first set of program execution instructions includes a termination condition for stopping execution of the first set of program code instructions.
10. The method according to claim 1, wherein, The first portion of the set of portions of the stream application is a processing element of the stream application.
11. The method according to claim 1, wherein, The first portion of the set of portions of the stream application is a stream operator of the stream application.
12. The method according to claim 11, wherein, The first time period is before the stream application executes the stream operator and after the stream application executes any other portion of the set of portions of the stream application.
13. The method according to claim 12, wherein, The first time period is after the stream application executes the stream operator.
14. The method according to claim 11, wherein The first time period is after the stream application executes the stream operator and before the stream application executes any other portion of the set of portions of the stream application.
15. The method according to claim 1, further comprising: During operation of the stream application, encapsulating a second portion of the set of portions of the stream application using a second general-purpose program execution structure; And During operation of the stream application, and in response to the program request and based on the stream application target and through the second general-purpose program execution structure, executing the first set of program code instructions during a second time period of execution of the second portion of the stream application.
16. The method according to claim 1, wherein, The first set of program code instructions includes program code instructions for recording the operation of the stream application.
17. The method according to claim 1, wherein The first set of program code instructions includes program code instructions for generating one or more additional tuples.
18. The method according to claim 1, further comprising: During operation of the stream application, a second portion of the set of portions of the stream application is encapsulated using a second general program execution structure; Receive a second program request to execute a second set of one or more program code instructions; Based on the second program request, identify a second stream application target of the second set of program code instructions; And During operation of the stream application, and in response to the second program request and based on the second stream application target and through the second general program execution structure, execute the second set of program code instructions during a second time period of execution of the second portion of the stream application.
19. A system for processing a tuple stream, the system comprising: A processor; A memory containing one or more instructions that, when executed by the processor, perform the method according to any one of claims 1-18.
20. A computer program product, the computer program product comprising: Program instructions executable by a processor to cause the processor to perform the method according to any one of claims 1-18.
Citation Information
Patent Citations
System and method for embedding a sreaming media format header within a session description message
US20030236912A1
Operator to processing element assignment in an active stream processing job
US20200228587A1