Data processing method based on stream computing, server and program product
By setting a common software package that generalizes input parameters in the stream computing system, creating a stream computing task in response to a task startup request and passing parameters for processing, the problem of low flexibility in stream computing processing in the existing technology is solved, and the adaptation and flexibility improvement of different scenario requirements is achieved.
Patent Information
- Application Number
- CN202510199095.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
Existing stream computing systems need to modify code when modifying scenario requirements, resulting in low processing flexibility.
By setting up a common software package in the stream computing system, its input parameters are generalized into data source parameters and calculation parameters, creating a stream computing task in response to task startup requests, and passing data source parameters and calculation parameters for processing.
It realizes the adaptation of different stream computing scenario requirements, without modifying the package code, reduces the number of software package uploads, and improves the flexibility of stream computing processing.
Smart Images

Figure CN120216004A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of computing devices, and in particular, to a data processing method, a server, and a program product based on stream computing. Background Art
[0002] With the development of the Internet and information technology, the amount of data has shown an explosive growth. To meet the needs of real-time data processing and analysis, data stream computing has emerged.
[0003] In related technologies, for different scenario requirements of stream computing, such as data sources or calculation rules, developers need to write different software packages and upload them to the stream computing system. However, when modifying the scenario requirements in this way, the code needs to be modified, resulting in low flexibility in stream computing processing. Summary of the Invention
[0004] Embodiments of the present application provide a data processing method, a server, and a program product based on stream computing, which are used to improve the flexibility of stream computing processing.
[0005] In a first aspect, embodiments of the present application provide a data processing method based on stream computing, including:
[0006] Responding to a task start request, creating a stream computing task corresponding to a common software package; wherein, the task start request includes stream computing parameters; the stream computing parameters include data source parameters and calculation parameters;
[0007] Using the stream computing parameters as input parameters of the common software package of the stream computing task;
[0008] Starting the stream computing task, obtaining target data according to the data source parameters; and performing stream computing processing on the target data according to the calculation parameters to obtain a calculation result.
[0009] The beneficial effects of this embodiment: By setting a common software package with input parameters of data source parameters and calculation parameters in the stream computing system, the stream computing system can respond to a task start request carrying data source parameters and calculation parameters, create a stream computing task corresponding to the common software package, and transfer the data source parameters and calculation parameters to the input parameters of the common software package. After starting and executing the stream computing task, the stream computing system can obtain target data according to the data source parameters, and perform stream computing processing on the target data according to the calculation parameters to obtain a calculation result. In this way, by configuring different data source parameters and calculation parameters based on the common software package, the adaptation of different stream computing scenario requirements is realized, without modifying the code in the software package, and without uploading different software packages for different stream computing scenario requirements, reducing the number of software package uploads and improving the flexibility of stream computing processing.
[0010] In a possible implementation, creating a stream computing task corresponding to a public software package in response to a task start request includes:
[0011] Determine the identifier of the public software package according to the name of the public software package indicated by the task start request;
[0012] Call the task submission interface according to the identifier of the public software package to create the stream computing task.
[0013] Advantages of this embodiment: By specifying the name of the public software package in the task start request, different public software packages can be flexibly selected and used to execute stream computing tasks, which can meet the data processing requirements in more scenarios; the stream computing system can automatically create stream computing tasks and can respond to multiple task start requests to simultaneously create multiple stream computing tasks with different scenario requirements based on the public software package, improving the flexibility of stream computing processing.
[0014] In a possible implementation, the method further includes:
[0015] Obtain a data acquisition rule and a calculation rule;
[0016] Perform conversion processing on the data acquisition rule and the calculation rule to obtain the stream computing parameters;
[0017] Generate a task start request according to the stream computing parameters.
[0018] Advantages of this embodiment: Based on the data acquisition rule and the calculation rule, the stream computing system can adapt to various computing scenario requirements, improving the adaptability and flexibility of stream computing processing to multi-scenario requirements.
[0019] In a possible implementation, the data acquisition rule includes an identifier of a data source; performing conversion processing on the data acquisition rule and the calculation rule to obtain the stream computing parameters includes:
[0020] Perform first conversion processing on the data acquisition rule according to the identifier of the data source to obtain data source parameters;
[0021] Perform second conversion processing on the calculation rule according to the data source parameters to obtain calculation parameters.
[0022] Advantages of this embodiment: The stream computing system can automatically perform conversion processing on the data acquisition rule and the calculation rule to obtain data source parameters and calculation parameters applicable to the stream computing system, without manual coding processing, improving the flexibility of stream computing processing.
[0023] In a possible implementation manner, the first transformation process is performed on the data acquisition rule according to the identifier of the data source to obtain a data source parameter, including:
[0024] According to the identifier of the data source, obtain the transformation rule corresponding to the identifier of the data source;
[0025] Based on the transformation rule corresponding to the identifier of the data source, convert the data acquisition rule into an executable statement format corresponding to the data source to obtain the data source parameter.
[0026] Beneficial effects of this embodiment: By setting corresponding transformation rules based on different types of data sources, the accuracy of the transformation process is improved, and thus the accuracy of the data source parameter is improved.
[0027] In a possible implementation manner, the second transformation process is performed on the calculation rule according to the data source parameter to obtain a calculation parameter, including:
[0028] Convert the calculation rule into an executable statement format during stream computing processing and substitute the data source parameter to obtain the calculation parameter.
[0029] Beneficial effects of this embodiment: Different calculation rules reflect different scenario calculation requirements. Through the transformation process of the calculation rule, automatic coding processing of the scenario calculation requirements is realized, without manual code modification, improving the flexibility of the stream computing processing.
[0030] In a possible implementation manner, the obtaining of the data acquisition rule and the calculation rule includes:
[0031] In response to a trigger operation of the user based on the front-end interface, receive the data acquisition rule and the calculation rule indicated by the trigger operation; or
[0032] Receive the data acquisition rule and the calculation rule input by the user through the command-line interface.
[0033] Beneficial effects of this embodiment: Configure the data acquisition rule and the calculation rule in multiple ways to adapt to the requirements of multiple scenario calculations, without modifying the code in the software package.
[0034] In a possible implementation manner, the stream computing parameter further includes a save duration parameter; the method further includes:
[0035] Store the calculation result and configure the save duration attribute of the calculation result according to the save duration indicated by the save duration parameter.
[0036] Advantages of this embodiment: The preservation duration parameter included in the stream computing parameters can be flexibly adapted to the preservation duration of the calculation results required by different computing scenarios.
[0037] In a second aspect, an embodiment of the present application further provides a server, including a processor and a memory communicatively connected to the processor;
[0038] The memory is used to store computer execution instructions;
[0039] The processor is used to execute the computer execution instructions stored in the memory and is used to implement the data processing method based on stream computing as described in any item of the first aspect.
[0040] Advantages of this embodiment: The processor of the server can execute the computer execution instructions, so that the processor sets the input parameter as a common software package for the data source parameter and the calculation parameter. The stream computing system can respond to a task startup request carrying the data source parameter and the calculation parameter, create a stream computing task corresponding to the common software package, and transfer the data source parameter and the calculation parameter to the input parameter of the common software package. After starting and executing the stream computing task, the stream computing system can obtain target data according to the data source parameter, and perform stream computing processing on the target data according to the calculation parameter to obtain a calculation result. In this way, by configuring different data source parameters and calculation parameters based on the common software package, the adaptation of different stream computing scenario requirements is realized, without modifying the code in the software package, and there is no need to upload different software packages for different stream computing scenario requirements, reducing the number of software package uploads and improving the flexibility of stream computing processing.
[0041] In a third aspect, an embodiment of the present application provides a computer program product, including a computer program, which when executed by a processor implements the data processing method based on stream computing as described in the first aspect.
[0042] Advantages of this embodiment: The processor of the server can execute this computer program, enabling the processor to set the input parameters as a common software package for data source parameters and calculation parameters. The stream computing system can respond to a task startup request carrying data source parameters and calculation parameters, create a stream computing task corresponding to this common software package, and pass the data source parameters and calculation parameters into the input parameters of the common software package. After starting and executing this stream computing task, the stream computing system can obtain target data according to the data source parameters, and perform stream computing processing on the target data according to the calculation parameters to obtain a calculation result. In this way, by configuring different data source parameters and calculation parameters based on this common software package, the adaptation of different stream computing scenario requirements is achieved without modifying the code in the software package, and there is no need to upload different software packages for different stream computing scenario requirements, reducing the number of software package uploads and improving the flexibility of stream computing processing. Brief Description of the Drawings
[0043] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application;
[0044] Figure 2 A schematic diagram of another application scenario provided by an embodiment of the present application;
[0045] Figure 3 A flowchart of a data processing method based on stream computing provided by an embodiment of the present application;
[0046] Figure 4 A flowchart of a data processing method based on stream computing provided by an embodiment of the present application;
[0047] Figure 5 A schematic diagram of a front-end interface provided by an embodiment of the present application;
[0048] Figure 6 A schematic diagram of a front-end interface provided by an embodiment of the present application;
[0049] Figure 7 A schematic diagram of the structure of a data processing device based on stream computing provided by an embodiment of the present application;
[0050] Figure 8 A schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed Embodiments
[0051] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application under the inspiration of this embodiment fall within the scope of protection of this application.
[0052] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0053] First, the nouns involved in the embodiments of this application are explained:
[0054] JAR package (Java ARchive): It is an archive file of Java, used to combine many files into a compressed file. A JAR package can contain metadata for executing a Java application, such as information specifying the main class (i.e., the class containing the main method).
[0055] Stream computing: It is a computing model for real-time processing of data, which can process data immediately when it arrives. By dividing data into continuous and infinite data streams and processing each data one by one, real-time data analysis and processing are achieved.
[0056] Figure 1 It is a schematic diagram of an application scenario provided for the embodiments of this application. As Figure 1 shown, the stream computing system is communicatively connected to at least one data source. It should be noted that Figure 1 it is illustrated by taking 2 data sources as an example.
[0057] The stream computing system can be deployed in a server or a server cluster. When deployed in a server cluster, the deployment method of this server cluster in the embodiments of this application is not limited. For example, it can be all deployed in the cloud or distributedly deployed. This stream computing system can obtain data from the data source based on a stream computing task for stream computing processing.
[0058] A data source for storing data. The data source can be deployed on a server or a server cluster, and the embodiments of the present application do not limit the deployment method of the data source. It should be noted that the present application does not limit the type of the data source, and the data source can be a message queue, a database, a file system, a search server, etc.
[0059] The embodiments of the present application provide a data processing method based on stream computing. A common software package is set in the stream computing system, and the input parameters of the common software package are generalized into data source parameters and calculation parameters. By modifying the data source parameters and calculation parameters, the adaptation of different stream computing scenario requirements is realized without modifying the code of the common software package, which improves the flexibility of stream computing processing.
[0060] Figure 2 It is a schematic diagram of another application scenario provided by the embodiments of the present application. As Figure 2 shown, the stream computing system includes a stream computing service, a management service, and a configuration service. The execution subject of the embodiments of the present application can be the stream computing system.
[0061] The stream computing service is used to receive, process, and analyze data streams. It may include multiple stream processing engines, such as Apache Flink, Apache Storm, Spark Streaming, etc. These engines can perform operations such as transforming, aggregating, filtering, and pattern recognition on data streams in real time.
[0062] The management service is used to monitor the running state of the stream computing system, manage stream computing tasks, and manage software packages. The running state of the stream computing system can include, for example, the health state of data streams, the execution situation of stream computing tasks, and the utilization rate of system resources. The management of stream computing tasks can include, for example, operations such as creating, scheduling, stopping, and restarting stream computing tasks. The management of software packages can include, for example, uploading and deleting software packages. Among them, the software package is the basic unit for executing stream computing tasks, and the embodiments of the present application do not limit the type of the software package. For example, it can be a JAR package.
[0063] The management service can provide a Graphical User Interface (GUI) or an Application Programming Interface (API) to enable users to conveniently manage and control stream computing tasks.
[0064] A configuration service is used to configure data acquisition rules and calculation rules. Among them, the data acquisition rule represents the source of the data to be calculated, and the calculation rule represents the rule for stream computing processing. The configuration service can be a GUI interface or a command-line interface (CLI) for receiving the data acquisition rules and calculation rules configured by the user.
[0065] It should be noted that there are preset interfaces among the stream computing service, the management service, and the configuration service to achieve data transfer. Moreover, the stream computing service, the management service, and the configuration service can be distributed and deployed on different servers or integrated and deployed on the same server. The embodiments of the present application do not make any limitations here, which is specifically related to the settings of the stream computing system.
[0066] Next, the technical solutions of the embodiments of the present application will be described in detail through specific embodiments. It should be noted that these specific embodiments can be combined with each other, and the same or similar concepts or processes will not be repeated in some embodiments.
[0067] Figure 3 It is a schematic flowchart of a data processing method based on stream computing provided by the embodiments of the present application. As Figure 3 shown, the method includes:
[0068] S101. In response to a task start request, create a stream computing task corresponding to a common software package.
[0069] Exemplarily, a stream computing task represents a specific data processing job defined and executed on a stream computing system. The task start request represents creating a stream computing task based on a common software package, and the task start request may include stream computing parameters, and the stream computing parameters include data source parameters and calculation parameters.
[0070] The data source parameters can represent the data source, and the data source parameters may include the identifier of the data source, the identifier of the data table, fields, and data filtering conditions, etc. The calculation parameters can represent the calculation rule, and the calculation parameters may include the calculation window size (such as a time window or a data volume window) and calculation functions (such as summation, average value, maximum value), etc. It should be noted that the formats of the data source parameters and the calculation parameters are in the format of executable statements of the stream computing system.
[0071] The common software package includes metadata for executing an application, such as information specifying the main class, in other words, a class containing the main method. The input parameters of the common software package include data source parameters and calculation parameters. The embodiments of the present application do not limit the type of the common software package. For example, it can be a JAR package.
[0072] In some possible implementation manners, the common software package is pre-stored in the stream computing system, and the task start request has a request type, and the request type has a corresponding relationship with the common software package. For example, the request type may include a first type and a second type. Among them, the first type represents using the common software package, and the second type represents using other software packages. In this implementation manner, the stream computing system can, in response to the task start request triggered by the user, parse the request content, extract the request type and stream computing parameters carried in the task start request, and according to the request type, call the common software package, and then can call the task creation interface to create a stream computing task corresponding to the common software package.
[0073] In some possible implementation manners, the task start request may further include the name of the common software package. The stream computing system can, in response to the task start request triggered by the user, parse the request content, extract the name of the common software package and the stream computing parameters carried in the task start request, and according to the name of the common software package, call the common software package, and then can call the task creation interface to create a stream computing task corresponding to the common software package.
[0074] It should be noted that creating a stream computing task corresponding to the common software package can be regarded as creating a stream computing task object in the stream computing system and assigning a unique task ID to the stream computing task object, which is convenient for subsequent tracking and management of the stream computing task.
[0075] S102. Use the stream computing parameters as the input parameters of the common software package of the stream computing task.
[0076] Exemplarily, the stream computing system can, based on the stream computing task, call a preset transfer interface to transfer the stream computing parameters to the input parameters of the common software package. It should be noted that the transfer interface is related to the programming language of the common software package. For example, if the common software package is a JAR package, the transfer interface can be the reflection mechanism of Java.
[0077] It can be understood that when the main method of the common software package is developed, input parameters will be defined. For example, input parameter 1 is the data source parameter, and input parameter 2 is the calculation parameter. The stream computing system can call a preset transfer interface to transfer the stream computing parameters to the corresponding input parameters.
[0078] S103. Start the stream computing task, obtain target data according to the data source parameter; and perform stream computing processing on the target data according to the calculation parameter to obtain a calculation result.
[0079] Exemplarily, the target data represents the data to be calculated.
[0080] In some possible implementation manners, the stream computing system may take out the stream computing task from the task queue and start to execute it. When executing, the stream computing system will execute the main method of the common software package. The main method of the common software package may be, for example, an execution statement based on input parameters. Further, the stream computing system may obtain target data from the data source indicated by the data source parameter in the input parameters; and perform stream computing processing on the target data according to the computing parameters in the input parameters to obtain a computing result.
[0081] In this embodiment, by setting a common software package with data source parameters and computing parameters as input parameters in the stream computing system, the stream computing system may, in response to a task start request carrying data source parameters and computing parameters, create a stream computing task corresponding to the common software package, transfer the data source parameters and computing parameters to the input parameters of the common software package. After starting and executing the stream computing task, the stream computing system may obtain target data according to the data source parameters, and perform stream computing processing on the target data according to the computing parameters to obtain a computing result. By this means, different data source parameters and computing parameters are configured based on the common software package to adapt to different stream computing scenario requirements, without modifying the code in the software package, and without uploading different software packages for different stream computing scenario requirements, reducing the number of software package uploads and improving the flexibility of stream computing processing.
[0082] In some embodiments, the stream computing system may receive and store the common software package uploaded by the user based on the management service.
[0083] Figure 4 It is a schematic flowchart of a data processing method based on stream computing provided by an embodiment of the present application. On the basis of the embodiment shown in Figure 3 The embodiment provides a detailed description of the data processing method based on stream computing provided by the embodiment of the present application. As shown in Figure 4 shown, the method includes:
[0084] S201. Obtain a data acquisition rule and a computing rule.
[0085] Exemplarily, the data acquisition rule represents the data that needs to be subjected to stream computing processing, and may include, for example, the identifier of the data source, the identifier of the data table, fields, and data filtering conditions, etc. The computing rule represents how to process the data, and may include, for example, the type of computing operation (such as counting, averaging, summing), the computing window size (such as a time window or a data volume window), etc.
[0086] In some possible implementation manners, the stream computing system may, based on the configuration service, in response to a triggering operation of the user based on the front-end interface, receive the data acquisition rule and the computing rule indicated by the triggering operation. Refer toFigure 5 As shown Figure 5 This is a schematic diagram of a front - end interface provided by an embodiment of the present application. An editing area for data acquisition rules and an editing area for calculation rules can be displayed on this front - end interface, and users can perform editing based on the front - end interface.
[0087] Exemplarily, referring to Figure 5 As shown, users can input the data source name in the editing area of the data acquisition rules and select the data table of this data source in the form of a drop - down menu; if they need data of certain fields in the data table, they can trigger the control for adding fields and select the fields of this data table in the form of a drop - down menu. If multiple fields are needed, the logical relationship between multiple fields can be added; if they need to filter the values of the fields, logical operators and values can be added as the filtering conditions for the data. Similarly, multiple data acquisition rules can be added through the adding control for data acquisition rules.
[0088] Optionally, a display area for data acquisition rules can also be set on the front - end interface. This display area is configured with a display template, and based on the content configured in the editing area of the data acquisition rules and this display template, the configured data acquisition rules can be combined and displayed in this display area. Referring to Figure 6 As shown Figure 6 This is a schematic diagram of a front - end interface provided by an embodiment of the present application. After users edit the data acquisition rules, the configured data acquisition rules can be displayed on the front - end interface. For example, Figure 6 The data acquisition rules shown can represent: obtaining data with the value of severe or high - risk in the alarm level field and the data of the device IP field in the alarm table from data source 1, and obtaining the data of the device IP field in the asset table from data source 2.
[0089] Continuing to refer to Figure 5 As shown, users can select the type of configured calculation rules in the form of a drop - down menu in the editing area of the calculation rules. For example, it can include association, summation, filtering, etc.; input the added data source name and select the data table of this data source, as well as select fields, logical operators, calculation windows, etc. in the data table to configure the calculation rules between data.
[0090] Optionally, a display area for calculation rules can also be set on the front - end interface. This display area is configured with a display template, and based on the content configured in the editing area of the calculation rules and this display template, the configured calculation rules can be combined and displayed in this display area. Continuing to refer to Figure 6 As shown, after users edit the calculation rules, the configured calculation rules can be displayed on the front - end interface. For example, Figure 6The calculation rule shown can represent that the data in the device IP field in the alarm table of data source 1 is equal to the data in the device IP field in the asset table of data source 2.
[0091] After the data acquisition rule and the calculation rule are configured, the submission control is triggered, and based on the configuration service, the data acquisition rule and the calculation rule configured by the user can be obtained. In the above manner, the stream computing system can provide a visual interface for the user to configure the rules. Since the user does not need to modify the code, the development difficulty of stream computing processing is reduced.
[0092] In some possible implementation manners, the stream computing system can, based on the configuration service, receive the data acquisition rule and the calculation rule input by the user through the command line interface. For example, the stream computing system can pre-define templates for the data acquisition rule and the calculation rule, and the user can, based on the command line, input the data acquisition rule and the calculation rule according to the template.
[0093] S202. Perform conversion processing on the data acquisition rule and the calculation rule to obtain stream computing parameters.
[0094] Exemplarily, the data acquisition rule and the calculation rule obtained through the foregoing steps may be in the form of descriptive statement formats. For example, they may be in the form of key-value pairs. For example, the data table: alarm table may not be applicable to the stream computing system. Therefore, the stream computing system needs to perform conversion processing on the obtained data acquisition rule and calculation rule based on the configuration service to obtain stream computing parameters in an executable statement format that can be applied to the stream computing system. For example, the stream computing system can perform conversion processing on the data acquisition rule and the calculation rule according to the preset conversion rule to obtain stream computing parameters, where the conversion rule represents the corresponding relationship between each key in the data acquisition rule and the calculation rule and the executable statement.
[0095] Optionally, this step may include the following steps:
[0096] S2021. Perform first conversion processing on the data acquisition rule according to the identifier of the data source to obtain data source parameters.
[0097] Exemplarily, for different types of data sources, the formats of the executable statements that can be executed in the data source are different. For example, the topic in the Kafka message queue and the data table in the database are expressed differently in the data source. Therefore, it is necessary to determine the conversion method for the data acquisition rule according to the type of the data source in the data acquisition rule to obtain data source parameters in an executable statement format that can perform data query in the data source and can be executed in the stream computing system.
[0098] Specifically, the data acquisition rule includes the identifier of the data source. For example, the name of the data source mentioned above. The stream computing system can obtain the conversion rule corresponding to the identifier of the data source according to the identifier of the data source, and convert the data acquisition rule into the executable statement format corresponding to the data source based on the conversion rule corresponding to the identifier of the data source, so as to obtain the data source parameters. Among them, the conversion rule corresponding to the identifier of the data source can represent the corresponding relationship between each key in the data acquisition rule corresponding to the data source and the executable statement. Exemplarily, the conversion rule corresponding to the identifier of the data source can be a preset executable statement template, which includes the keys in the data acquisition rule. Substitute the value corresponding to the key in the data acquisition rule into the template to obtain the data source parameters.
[0099] For example, referring to Figure 6 the data acquisition rule shown: "Data source: Data source 2; Data table: Asset table; Field: Device IP", the data source parameters after conversion processing can be expressed as:
[0100] CREATE TABLE kafkatable(id BIGINT,name STRING,age STRING,startTimeTIMESTAMP(3),+"WATERMARK FOR startTime AS startTime-INTERVAL'5'SECOND)WITH('connector'='kafka',"+"'topic'='Asset table','properties.bootstrap.servers'='127.0.0.1:9092',"+"'properties.group.id'='Device IP','scan.startup.mode'='latest-offset','format'='json')
[0101] The data source parameters after the conversion processing define a table named kafkatable, determine that the data source is Kafka, the address of the Kafka bootstrap server (bootstrap.servers) is 127.0.0.1:9092, the topic is the asset table, and the ID (group.id) of the Kafka consumer group is the device IP.
[0102] It should be noted that if the data acquisition rule includes the acquisition rules of multiple data sources, corresponding conversion processing needs to be performed based on the conversion rules corresponding to each data source. Set corresponding conversion rules based on different types of data sources to improve the accuracy of the conversion processing.
[0103] S2022. Perform a second transformation process on the calculation rule according to the data source parameters to obtain calculation parameters.
[0104] Exemplarily, the stream computing system can perform a transformation process on the calculation rule according to a preset transformation rule to obtain calculation parameters in an executable statement format that can be executed in the stream computing system.
[0105] Specifically, the stream computing system can convert the calculation rule into an executable statement format during stream computing processing according to a preset transformation rule, and substitute the data source parameters to obtain calculation parameters. The preset transformation rule can represent the correspondence between each key in the calculation rule and the executable statement.
[0106] Exemplarily, as mentioned above, the calculation rule may include calculation operations. Therefore, the stream computing system can determine a calculation function template according to the calculation operation, and the calculation function template can indicate the executable statement format corresponding to the calculation function. For example, referring to Figure 6 the calculation rule shown: "Calculation operation: Count; Data source: Data source 1; Data table: Alarm table; Field: Device IP; Data source: Data source 2; Data table: Asset table; Field: Device IP; Calculation window: (1, 1)", the calculation parameters after the transformation process can be expressed as:
[0107] select count(*) as num, a.ip, a.startTime, b.endTime, + "HOP_START(a.startTime, INTERVAL '300' SECOND, INTERVAL '60' SECOND) as winStart, HOP_END(a.startTime, + "INTERVAL '300' SECOND, INTERVAL '60' SECOND) as winEnd from source_1a, source_2b where + "a.ip = b.ip group by HOP(a.startTime, INTERVAL '300' SECOND, INTERVAL '60' SECOND), a.ip.
[0108] The calculated parameters after the conversion process define a process of obtaining data from data source 1 and data source 2, grouping and counting. Among them, based on the value of the calculation operation in the calculation rule, the calculation function is determined as count(*) as num, and the count(*) function is used to calculate the number of records, naming the result as num. The HOP_START and HOP_END functions are used to define a hopping window (HOP window). The hopping window starts from a.startTime, with a window size of 300 seconds and a hopping interval of 60 seconds. winStart and winEnd are the start time and end time of the window respectively. Grouping is performed according to the condition that a.ip and b.ip are equal and the HOP function is grouped according to the hopping window.
[0109] S203. Generate a task start request according to the stream calculation parameters.
[0110] Exemplarily, the stream computing system can generate a task start request based on the configuration service, according to the preset data packet format, stream calculation parameters, and the name of the common software package, and call the REST interface to pass the task start request to the stream computing service of the stream computing system.
[0111] S204. In response to the task start request, determine the identifier of the common software package according to the name of the common software package indicated by the task start request.
[0112] Exemplarily, the task start request carries the name of the common software package. The stream computing system can determine the identifier of the common software package based on the stream computing service, according to the name of the common software package and the preset mapping table. Among them, the mapping table is used to indicate the corresponding relationship between the name and identifier of the software package, and the identifier of the common software package is used to uniquely represent the common software package, for example, it can be an ID assigned by the stream computing system.
[0113] By specifying the name of the common software package in the task start request, different common software packages can be flexibly selected and used to execute stream computing tasks, which can meet the data processing requirements in more scenarios and improve the flexibility and scalability of stream computing processing.
[0114] S205. According to the identifier of the common software package, call the task submission interface to create a stream computing task.
[0115] Exemplarily, the stream computing system can call the common software package based on the stream computing service according to the identifier of the common software package, and create a stream computing task by calling the task submission interface.
[0116] In this way, the stream computing system can automatically create stream computing tasks, and in response to multiple task startup requests, simultaneously create multiple stream computing tasks with different scenario requirements based on a common software package, improving the flexibility of stream computing processing.
[0117] S206. Use the stream computing parameters as input parameters of the common software package of the stream computing task.
[0118] It should be noted that this step is similar to the aforementioned step S102 and will not be elaborated here.
[0119] S207. Start the stream computing task, obtain target data according to the data source parameters; and perform stream computing processing on the target data according to the computing parameters to obtain the computing result.
[0120] Exemplarily, referring to the data source parameters exemplified in the aforementioned step S2021, the stream computing system can, based on the stream computing service, execute the data source parameters, obtain the data with the field name "device IP" from the topic "asset table" of the data source 2 (kafka), and store it in a table 1 (kafkatable). Correspondingly, obtain the data with the value of the alarm level field in the alarm table of the data source 1 being severe or high-risk, and the data of the device IP field in the alarm table, and store it in the table 2; and execute the above-mentioned computing parameters. Based on the data in the table 1 and the table 2, the stream computing system can count the number of data records with the same IP within a specific hopping window in the data of the table 1 (source_1) and the table 2 (source_2) to obtain the computing result.
[0121] S208. Store the computing result and configure the storage duration attribute of the computing result according to the storage duration indicated by the storage duration parameter.
[0122] Exemplarily, the stream computing parameters may further include a storage duration parameter for characterizing the storage duration of the computing result. The stream computing system can, based on the stream computing service, store the computing result in a preset storage location and configure the storage duration attribute of the computing result according to the storage duration indicated by the storage duration parameter. Through the storage duration parameter included in the stream computing parameters, the storage duration of the computing result can be flexibly adapted to different computing scenario requirements.
[0123] In this embodiment, by setting an input parameter in the stream computing system as a common software package for data source parameters and computing parameters, the stream computing system can receive data acquisition rules and computing rules configured by the user through a visual front-end interface, convert the data acquisition rules and computing rules into a format executable by the stream computing system to obtain stream computing parameters, and then generate a task start request. In response to the task start request carrying data source parameters and computing parameters, a stream computing task corresponding to the common software package is created, and the data source parameters and computing parameters are passed into the input parameters of the common software package. After starting and executing the stream computing task, the stream computing system can obtain target data according to the data source parameters and perform stream computing processing on the target data according to the computing parameters to obtain a computing result. In this way, using the common software package can meet the requirements of different stream computing scenarios. For different stream computing scenario requirements, configuration can be performed through a visual front-end interface without modifying the code in the software package, and there is no need to upload different software packages for different stream computing scenario requirements, reducing the number of software package uploads and improving the flexibility of stream computing processing.
[0124] Figure 7 FIG. is a schematic structural diagram of a data processing device based on stream computing provided by an embodiment of the present application. As Figure 7 shown, the data processing device 300 based on stream computing includes:
[0125] A creation module 301, configured to create a stream computing task corresponding to a common software package in response to a task start request; wherein, the task start request includes stream computing parameters; the stream computing parameters include data source parameters and computing parameters;
[0126] A transmission module 302, configured to use the stream computing parameters as input parameters of the common software package of the stream computing task;
[0127] A processing module 303, configured to start the stream computing task, obtain target data according to the data source parameters; and perform stream computing processing on the target data according to the computing parameters to obtain a computing result.
[0128] In a possible implementation manner, the creation module 301 is specifically configured to:
[0129] Determine an identifier of the common software package according to the name of the common software package indicated by the task start request;
[0130] Call a task submission interface according to the identifier of the common software package to create the stream computing task.
[0131] In a possible implementation manner, the device 300 further includes a generation module, configured to:
[0132] Obtain the data acquisition rule and the calculation rule;
[0133] Perform a conversion process on the data acquisition rule and the calculation rule to obtain the stream computing parameter;
[0134] Generate a task start request according to the stream computing parameter.
[0135] In a possible implementation manner, the generation module is specifically configured to:
[0136] According to the identifier of the data source, perform a first conversion process on the data acquisition rule to obtain a data source parameter;
[0137] According to the data source parameter, perform a second conversion process on the calculation rule to obtain a calculation parameter.
[0138] In a possible implementation manner, the generation module is specifically configured to:
[0139] According to the identifier of the data source, obtain the conversion rule corresponding to the identifier of the data source;
[0140] Based on the conversion rule corresponding to the identifier of the data source, convert the data acquisition rule into an executable statement format corresponding to the data source to obtain the data source parameter.
[0141] In a possible implementation manner, the generation module is specifically configured to:
[0142] Convert the calculation rule into an executable statement format during stream computing processing and substitute the data source parameter to obtain the calculation parameter.
[0143] In a possible implementation manner, the generation module is specifically configured to:
[0144] In response to a trigger operation of the user based on the front-end interface, receive the data acquisition rule and the calculation rule indicated by the trigger operation; or
[0145] Receive the data acquisition rule and the calculation rule input by the user through the command line interface.
[0146] In a possible implementation manner, the stream computing parameter further includes a save duration parameter; the apparatus 300 further includes a storage module for:
[0147] Store the calculation result and configure the save duration attribute of the calculation result according to the save duration indicated by the save duration parameter.
[0148] The data processing apparatus based on stream computing provided in this embodiment can execute the data processing method based on stream computing in the above method embodiment, and its beneficial effects are similar and will not be elaborated here.
[0149] Figure 8 This embodiment of the application provides a schematic structural diagram of a server. As Figure 8 shown, the server 400 includes a processor 401 and a memory 402 communicatively connected to the processor. Optionally, the device 400 further includes a communication component 403. Among them, the processor 401, the memory 402, and the communication component 403 are connected through a bus 404.
[0150] In a specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402, so that at least one processor 401 executes the above-mentioned method.
[0151] For the specific implementation process of the processor 401, reference may be made to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0152] In the above embodiment, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0153] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0154] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0155] The embodiments of the present application further provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the technical solutions provided in any of the foregoing method embodiments.
[0156] The embodiments of the present application further provide a computer program product, including a computer program, and when the computer program is executed by a processor, it is used to implement the technical solutions provided in the foregoing method embodiments.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, rather than to limit them; although the embodiments of the present application have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method based on stream computing, characterized in that: The method comprises: In response to a task start request, creating a stream computing task corresponding to the public software package; wherein the task start request includes stream computing parameters; the stream computing parameters include data source parameters and computing parameters; Using the stream computing parameters as input parameters of a common software package of the stream computing task; The stream computing task is started, and target data is acquired according to the data source parameters; and stream computing is performed on the target data according to the computing parameters to obtain computing results.
2. The method according to claim 1, characterized in that The step of creating a stream computing task corresponding to the public software package in response to the task initiation request includes: Determining an identifier of the public software package according to the name of the public software package indicated by the task initiation request; According to the identifier of the public software package, a task submission interface is called to create the stream computing task.
3. The method according to claim 1, characterized in that The method further comprises: Obtain data acquisition rules and calculation rules; Converting the data acquisition rule and the calculation rule to obtain the flow calculation parameter; A task start request is generated according to the stream computing parameters.
4. The method according to claim 3, characterized in that: The data acquisition rule includes an identifier of a data source; the converting process of the data acquisition rule and the computing rule to obtain the stream computing parameter includes: According to the identifier of the data source, the data acquisition rule is subjected to a first conversion process to obtain a data source parameter; According to the data source parameters, the calculation rules are subjected to a second conversion process to obtain calculation parameters.
5. The method according to claim 4, characterized in that The step of performing a first conversion process on the data acquisition rule according to the identifier of the data source to obtain a data source parameter includes: According to the identifier of the data source, obtaining a conversion rule corresponding to the identifier of the data source; Based on the conversion rule corresponding to the identifier of the data source, the data acquisition rule is converted into an executable statement format corresponding to the data source to obtain the data source parameter.
6. The method according to claim 4, characterized in that The step of performing a second conversion process on the calculation rule according to the data source parameter to obtain the calculation parameter includes: The calculation rule is converted into an executable statement format for stream computing processing, and the data source parameter is substituted into the format to obtain the calculation parameter.
7. The method according to claim 3, characterized in that The data acquisition rules and calculation rules include: In response to a trigger operation by a user based on a front-end interface, receiving a data acquisition rule and a calculation rule indicated by the trigger operation; or, Receive data acquisition rules and calculation rules input by the user through the command line interface.
8. The method according to any one of claims 1 to 7, characterized in that: The flow calculation parameters also include a storage duration parameter; the method also includes: The calculation result is stored, and a storage duration attribute of the calculation result is configured according to the storage duration indicated by the storage duration parameter.
9. A server, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory is used to store computer-executable instructions; The processor is used to execute the computer-executable instructions stored in the memory, so as to implement the data processing method based on stream computing as described in any one of claims 1 to 8.
10. A computer program product, characterized in that It includes a computer program, which, when executed by a server, implements the data processing method based on stream computing described in any one of claims 1 to 8.