Data processing method and apparatus, and cluster

By using read operators in the streaming computing device to identify and send data required for only computing logic, the storage overhead and memory consumption problems caused by the storage based on the state storage of the full data in streaming computing are solved, and more efficient computing performance and service processing are achieved.

WO2025102626A1PCT designated stage expired Publication Date: 2025-05-22HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/091496
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2024-05-07
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In streaming computing, state storage based on full data results in unnecessary storage overhead and memory consumption, affecting computing performance.

Method used

By introducing a read operator in the streaming computing device, the data required for only the calculation logic is identified and sent to the calculation operator, so that the calculation operator is stored in a state based on the data required for the calculation logic.

Benefits of technology

The amount of state storage data of the computing operator is reduced, memory consumption is saved, computing performance is improved, and the normality of business processing is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024091496_22052025_PF_FP_ABST
    Figure CN2024091496_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a data processing method and apparatus, and a cluster. The method comprises: a reading operator reads service data from a data source end of a service; the reading operator identifies, from the service data, data required by the calculation logic of the service, so as to obtain the data required for calculation, wherein the data required for calculation comprises first data required by first calculation logic; the reading operator sends the data required for calculation to a first calculation operator, so that the first calculation operator performs calculation on the first data on the basis of the first calculation logic to obtain a calculation result; and an output operator outputs a processing result of the service to a data target end of the service on the basis of the calculation result and the service data. The method can reduce the data volume of the state storage of the calculation operator while ensuring that the service is normally processed, thereby saving the memory of the calculation operator.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device and cluster

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 17, 2023, with application number 202311541722.3 and application name “A streaming computing method”, and the Chinese patent application filed with the State Intellectual Property Office of China on January 16, 2024, with application number 202410065781.6 and application name “A data processing method, device and cluster”, all of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a data processing method, device, and cluster. Background Art

[0003] Stream computing is a computing method that processes data in real time. Generally speaking, stream computing can be divided into stateful and stateless computing. In stateful computing, there is a relationship between previous and subsequent data, and operators need to calculate the current data based on previous data or previous calculation results. Therefore, operators performing stateful calculations store previous data and previous calculation results as state. When new data is received, the stored state (i.e., previous data and previous calculation results) is used to calculate the new data.

[0004] In related technologies, the full data set is sent to operators, causing them to store state based on the full data set. The data required for operator computation is typically only a small portion of the full data set. Therefore, storing state based on the full data set generates unnecessary storage overhead. Furthermore, because state storage occurs in memory, excessively large amounts of state storage lead to high operator memory consumption, impacting operator performance.

[0005] Summary of the Invention

[0006] The present application provides a data processing method, device, and cluster, which can reduce the amount of data stored in the state of computing operators while ensuring normal business processing, thereby saving the memory of computing operators.

[0007] In a first aspect, a data processing method is provided, which is applied to a streaming computing device, wherein the streaming computing device includes a reading operator, a first computing operator and an output operator, wherein the first computing operator is used to perform data calculation based on the first computing logic of the business; the method includes: the reading operator reads business data from the data source end of the business; the reading operator identifies the data required by the computing logic of the business in the business data, and obtains the data required for calculation; wherein the data required for calculation includes the first data required by the first computing logic; the reading operator sends the data required for calculation to the first computing operator, so that the first computing operator calculates the first data based on the first computing logic to obtain a calculation result; the output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data.

[0008] A compute operator stores state based on the data it receives. The data required by the operator's computational logic is typically only a small portion of the business data read by the reader. This method identifies the data required by the operator's computational logic within the business data and sends that data to the operator, allowing the operator to store state based on the data required by the computational logic. Compared to state storage based on business data, state storage based on the data required by the computational logic significantly reduces the amount of state stored, thereby reducing the operator's memory consumption. Furthermore, the data required by the computational logic meets the operator's computational requirements. The output operator obtains the business processing result based on the computation results and business data generated by the computation operator on the data required by the computational logic. In other words, if the reader operator only sends the data required by the computational logic to the compute operator, the output operator can obtain the business processing result without affecting business processing.

[0009] In short, the method provided in the embodiment of the present application can reduce the amount of data stored in the state of the computing operator while ensuring that the business can be processed normally, thereby saving the memory of the computing operator and ensuring the computing performance of the computing operator.

[0010] In one possible implementation, the streaming computing device corresponds to a storage space accessed by a reading operator and an output operator; the method further includes: the reading operator stores data not required for calculation in the storage space, where the data not required for calculation is data in the business data other than the data required for calculation; the output operator reads the data not required for calculation from the storage space; the output operator outputs the processing result of the business to the data target end based on the calculation result and the business data, including: the output operator outputs the processing result to the data target end based on the data not required for calculation and the calculation result.

[0011] This storage space is shared between the read operator and the output operator. The non-computationally required data and the calculation results of the calculation operator are used by the output operator to obtain the business processing results. In this implementation, the read operator stores the non-computationally required data in the shared storage space. Without the read operator having to pass the non-computationally required data through an operator (e.g., a calculation operator) between the read operator and the output operator, the output operator can obtain the non-computationally required data, thereby obtaining the business processing results based on the non-computationally required data and the calculation results of the calculation operator.

[0012] In one possible implementation, the streaming computing device corresponds to a storage space accessed by a reading operator and an output operator; the method further includes: the reading operator storing the business data in the storage space; and the output operator reading the business data from the storage space.

[0013] This storage space is shared between the read operator and the output operator. Business data and the calculation results of the calculation operator are used by the output operator to obtain the business processing results. In this implementation, the read operator stores business data in the shared storage space. Without the read operator having to pass business data through an operator (such as a calculation operator) between the read operator and the output operator, the output operator can obtain data not required for calculations, thereby obtaining the business processing results based on the non-computational data and the calculation results of the calculation operator.

[0014] In one possible implementation, the streaming computing device includes a second computing operator, which is used to perform data calculations according to the second computing logic of the business; in the data flow direction of the business, the second computing operator is located after the first computing operator; wherein, the data required for the calculation also includes the second data required by the second computing logic; the first computing operator is used to send the second data to the second computing operator, so that the second computing operator calculates the second data based on the second computing logic.

[0015] In this implementation, the computing operator can identify the data required by the computing logic of the downstream computing operator of the computing operator, and send the data required by the computing logic of the downstream computing operator to the downstream computing operator. While ensuring the computing needs of the downstream computing operator, it further reduces the amount of data stored in the state of the downstream computing operator, further saving the memory of the downstream computing operator.

[0016] In one possible implementation, the business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; the first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; a reading operator identifies data required by the calculation logic of the business in the business data to obtain the data required for the calculation, including: the reading operator identifies the first splicing key and the second splicing as the first data based on the first calculation logic; the output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator splices the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

[0017] In this implementation, when the computing operator performs data splicing services, the reading operator only needs to send the splicing key to the computing operator. While meeting the computing requirements of the computing operator, the computing operator can store the state based on the splicing key, reducing the amount of data stored in the computing operator's state.

[0018] In one possible implementation, the business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key meets the preset filtering conditions; the reading operator identifies the data required by the calculation logic of the business in the business data to obtain the data required for the calculation, including: the reading operator identifies the filter key as the first data based on the first calculation logic; the output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the calculation result indicates that the filter key meets the preset filtering conditions, the output operator uses the data to be filtered as the processing result.

[0019] In this implementation, when the calculation operator performs data filtering and output services, the reading operator only needs to send the filter key to the calculation operator. While meeting the calculation requirements of the calculation operator, the calculation operator performs state storage based on the filter key, reducing the amount of data stored in the state of the calculation operator.

[0020] In one possible implementation, business data is streaming data consisting of multiple data points, and each of the multiple data points includes at least one data; wherein, the reading operator reads different data points from the data source end at different times; the reading operator reads the business data from the data source end of the business, including: the reading operator reads the first data point from the multiple data points at a first moment; the reading operator identifies the data required by the calculation logic of the business in the business data, and obtains the data required for the calculation, including: the reading operator identifies the first data point as the first data based on the first calculation logic.

[0021] Whenever data is read, the read operator determines whether it is required for computation or not. If the data is required, the read operator sends it to the compute operator, allowing the compute operator to perform computations on it immediately, ensuring the real-time nature of stream computing.

[0022] In one possible implementation, a reading operator reads business data from a data source end of the business, including: the reading operator reads a second data point among multiple data points at a second moment; when the second data point is data other than the data required by the first calculation logic, the output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the second data, including: the output operator outputs the processing result to the data target end of the business based on the calculation result and the second data point.

[0023] In a possible implementation, the computing logic of the business is obtained based on a structured query language SQL statement of the business.

[0024] In this implementation, the business calculation logic can be obtained based on the business SQL statement, and the business calculation logic can be configured to the read operator. Thus, the read operator can identify the data required by the business calculation logic in the business data based on the business calculation logic.

[0025] In a possible implementation, the first calculation operator is obtained by arranging SQL statements based on the business.

[0026] In this implementation, the business's computational logic can be derived based on the business's SQL statements. Then, based on the business's computational logic, operators that carry the business's computational logic can be orchestrated. The orchestrated first computational operator is used to carry the business's first computational logic. That is, the first computational operator is used to perform data computations according to the first computational logic, thereby processing the business.

[0027] In a second aspect, a streaming computing device is provided, which includes a reading operator, a first computing operator and an output operator, wherein the first computing operator is used to perform data calculation based on the first computing logic of the business; wherein the reading operator is used to read business data from the data source end of the business; the reading operator is used to identify the data required by the computing logic of the business in the business data, and obtain the data required for calculation; wherein the data required for calculation includes the first data required by the first computing logic; the reading operator is used to send the data required for calculation to the first computing operator, so that the first computing operator calculates the first data based on the first computing logic to obtain a calculation result; the output operator is used to output the processing result of the business to the data target end of the business based on the calculation result and the business data.

[0028] In one possible implementation, the streaming computing device corresponds to a storage space accessed by a reading operator and an output operator; wherein the reading operator is also used to store data not required for calculation into the storage space, and the data not required for calculation is data in the business data other than the data required for calculation; the output operator is also used to read the data not required for calculation from the storage space; the output operator is also used to output the processing results to the data target end based on the data not required for calculation and the calculation results.

[0029] In one possible implementation, the streaming computing device corresponds to a storage space accessed by a reading operator and an output operator; wherein the reading operator is also used to store business data in the storage space; and the output operator is also used to read business data from the storage space.

[0030] In one possible implementation, the streaming computing device includes a second computing operator, which is used to perform data calculations according to the second computing logic of the business; in the data flow direction of the business, the second computing operator is located after the first computing operator; wherein, the data required for the calculation also includes the second data required by the second computing logic; the first computing operator is used to send the second data to the second computing operator, so that the second computing operator calculates the second data based on the second computing logic.

[0031] In one possible implementation, the business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; the first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; a reading operator is used to identify the first splicing key and the second splicing as the first data based on the first calculation logic; when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator is used to splice the first data to be spliced ​​and the second data to be spliced ​​to obtain a processing result.

[0032] In one possible implementation, the business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key meets the preset filtering conditions; a reading operator is used to identify the filter key as the first data based on the first calculation logic; when the calculation result indicates that the filter key meets the preset filtering conditions, the output operator is used to use the data to be filtered as the processing result.

[0033] In one possible implementation, business data is streaming data consisting of multiple data points, each of the multiple data points includes at least one data; wherein, the reading operator reads different data points from the data source end at different times; the reading operator is used to read the first data point among the multiple data points at a first time; the reading operator is used to identify the first data point as the first data based on the first computing logic.

[0034] In one possible implementation, the read operator is used to read a second data point among multiple data points at a second moment; when the second data point is data other than the data required by the first calculation logic, the output operator is used to output the processing result to the data target end of the business based on the calculation result and the second data point.

[0035] In a possible implementation, the computing logic of the business is obtained based on a structured query language SQL statement of the business.

[0036] In a possible implementation, the first calculation operator is obtained by arranging SQL statements based on the business.

[0037] In a third aspect, a computing device cluster is provided, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method provided in the first aspect.

[0038] In a fourth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method provided in the first aspect.

[0039] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computer device cluster, the computer device cluster executes the method provided in the first aspect.

[0040] The beneficial effects of the second to fifth aspects can be referred to the above introduction to the beneficial effects of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] FIG1 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0042] FIG2 is a schematic diagram of a computing node cluster provided in an embodiment of the present application;

[0043] FIG3 is a schematic diagram of a shared storage space provided in an embodiment of the present application;

[0044] FIG4 is a flow chart of a data processing method provided in an embodiment of the present application;

[0045] FIG5 is a schematic structural diagram of a stream computing device provided in an embodiment of the present application;

[0046] FIG6 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0047] FIG7 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0048] FIG8 is a schematic diagram of a structure of a computing device cluster connected via a network provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The following describes the solutions provided by the embodiments of the present application in conjunction with the accompanying drawings. In the embodiments of the present application, "plurality" refers to two or more, and "multiple" refers to two or more. Terms such as "first" and "second" are used only to distinguish similar objects and do not necessarily describe a specific order or quantity of objects.

[0050] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application are first introduced.

[0051] Streaming data, also known as time-series data, refers to data generated in chronological order, that is, data that is continuously generated over time. Streaming data consists of data points generated in chronological order, where each data point can include at least one piece of data. Different data points are generated at different times. In other words, a single point in time or time period can generate one or a segment of data, which is called a data point. Multiple data points generated at consecutive points in time or time periods constitute streaming data. Streaming data often features large data volumes, short intervals between data points, and continuous arrival of data points.

[0052] Stream computing: Also known as stream computing, stream computing is used to process data in real time. Stream computing can be a computing model triggered by input data, where each new input data acts as an event, triggering the computing model to perform calculations.

[0053] Stateless computing: A type of streaming computing. In stateless computing, data is processed independently, with no relationship between different data. For example, in a stateless computation of "input value + 5," if the current input value (i.e., the current data) is "10," then the input value is added to 5, resulting in a result of "15."

[0054] Stateful computation: Another type of streaming computation. Previous and subsequent input values ​​(i.e., input data) are related, requiring the current input value to be calculated based on the previous input value or calculation result. To facilitate computation, the previous input value or calculation result is stored as state. For example, in a stateful computation like accumulation, if the first input value is "5," the cumulative result is "5"; if the second input value is "100," the cumulative result is "5 + 100 = 105"; and if the third input value is "200," the cumulative result is "105 + 200 = 305." This shows that the current computation requires the previous result, so the previous result needs to be stored as state. Taking another stateful computation like associative computation as an example, data source A1 can be an employee information table, including the employee's identifier (ID), name, and department ID. Data source A2 can be department information, including the department ID and name. Specifically, an association calculation uses a department's ID to obtain information about employees within that department, such as the employee's ID, name, and department ID. The department ID here is a related field. The association calculation logic determines whether the related field (i.e., department ID) in the input value from data source A1 is identical to the related field (i.e., department ID) in the input value from data source A2. If they are identical (i.e., the association calculation results are identical), the employee information from data source A1 and the employee information from data source A2 are output. For example, if the input value from data source A1 is "1 (employee ID), Li Qiang (employee name), B10011 (employee's department ID)," and the input value from source data A2 is "B10011 (department ID), R&D Department (department name)," the output result will be "1, Li Qiang, R&D Department." Because employee information in the data source may change, such as if the employee's department changes, the previously output employee information must be updated. The previous input values ​​and output results need to be found, so they need to be stored as state. Furthermore, data may be duplicated, so to ensure the accuracy of the output results, duplicates need to be removed based on the previous input values. Therefore, the previous input values ​​also need to be stored as state.

[0055] An operator is a data processing unit that carries computational logic. An operator is the smallest executable unit in stream computing. Operators are used to perform calculations on related data based on the computational logic they carry. Common operators in stream computing devices include read operators, compute operators, and output operators. Typically, a stream computing device requires a read operator, at least one compute operator, and an output operator to execute a particular task.

[0056] Read operator: also known as source operator, is an operator in a stream computing device used to read business data from the data source of the business.

[0057] Computing operator: also known as task operator, is an operator in a stream computing device used to carry the computing logic of the business, that is, the computing operator is used to execute the computing logic of the business.

[0058] Output operator: also known as sink operator, is used to obtain the business processing results based on the calculation results of the calculation operator, and send the business processing results to the data target end of the business, such as the storage end for storing the business processing results.

[0059] Data required for calculation: refers to the data required by the calculation operator in the business data to execute the calculation logic, that is, the calculation object that the calculation operator calculates based on the calculation logic. Data required for calculation is also called data required for calculation logic. Among them, business data can be recorded in a data table, and the data required for calculation can be one or more fields in the data table. These one or more fields can be called required fields for calculation or calculated fields. For example, the calculation logic of an association calculation includes determining whether the associated fields of two tables to be associated are equal. In this case, the data required for this calculation logic is the associated fields of the two tables to be associated. For another example, the calculation logic of a filtering calculation includes determining whether the filtering field meets the filtering condition. In this case, the data required for this calculation logic is the filtering field.

[0060] Data not required for calculation: refers to business data other than data required for calculation.

[0061] State storage: This refers to the storage of the data received by an operator and its computational results as state. Typically, operators store state in local memory. This means that operator state storage consumes the operator's memory.

[0062] An embodiment of the present application provides a data processing method. In this method, a read operator identifies the data required by the business's calculation logic in the read business data based on the business's calculation logic, and obtains the data required for calculation. The data required for calculation is a portion of the business data, specifically the data required by the business's calculation logic in the business data. The read operator sends the data required for calculation rather than the business data to the calculation operator of the business, reducing the amount of data sent to the calculation operator. The calculation operator can perform state storage based on the data required for calculation, thereby reducing the amount of data stored in the state. The calculation operator calculates the data required for calculation based on the calculation logic carried by the calculation operator to obtain a calculation result. That is, the data required for calculation can meet the calculation requirements of the calculation operator. The output operator can obtain the processing result of the business based on the calculation result and the business data, and send the processing result to the data target end of the business. In this method, the read operator only sends the data required for calculation to the calculation operator, reducing the network overhead between the read operator and the calculation operator. In particular, the computing operator only receives the data required for calculation and performs state storage based on the data required for calculation, which reduces the amount of data stored in the state, thereby saving memory consumption for state storage and improving the computing performance of the computing operator.

[0063] Next, the data processing method provided in the embodiment of the present application is introduced.

[0064] Figure 1 shows a system architecture that can be used to implement the data processing method provided in an embodiment of the present application. As shown in Figure 1, the system architecture includes a user 110 and a computing node cluster 120. Among them, the computing node cluster 120 includes multiple computing nodes. The computing node can be a physical node, such as a server. The computing node can also be a virtual computing node such as a virtual machine (VM) or a container. The user can trigger the computing node cluster 120 to perform a business through a job request, such as a data association business, a data filtering business, etc. Among them, the data association business is also called the association computing business, which refers to associating or splicing data in two or more data tables based on association conditions. The data filtering business is also called data filtering computing, which refers to filtering or screening out data that meets the preset filtering conditions from the data.

[0065] The computing node cluster 120 is deployed with a master node 121. The master node 121 can be deployed on any one or more computing nodes in the computing node cluster 120. The master node 121 can receive a business job request from the user 110 and perform operator orchestration based on the job request to obtain a stream computing device 200, as described below.

[0066] A job request may include information about the business's computational logic. For example, a job request may include one or more Structured Query Language (SQL) statements. SQL statements reflect the business's computational logic. By parsing the SQL statements, the business's computational logic can be obtained. In one example, a job request may include the following SQL statements.

[0067] "Select table1.f1, table1.f2, table1.f3,.., table2.f2, table2.f3,...from table1 join table2 on table1.f1=table2.f2"

[0068] This SQL statement indicates that, with the f1 field in table 1 (table1) and the f2 field in table 2 (table2) as the associated fields, the f1 field, f2 field, f3 field, etc. in table 1 are concatenated with the f2 field, f3 field, etc. in table 2. The associated fields can be called concatenated fields or concatenated keys. The f2 field, f3 field, and other fields in table 1 and the f3 field, etc. in table 2 are data to be concatenated. The calculation logic included in this SQL statement is to determine whether the f1 field in table 1 and the f2 field in table 2 are the same. If they are the same, the f2 field, f3 field, and other fields in table 1 are concatenated with the f3 field, etc. in table 2.

[0069] In one example, the job request may also be configured to include the following SQL statement.

[0070] "Select table3.f1, table3.f2, table3.f3, ..from table3 where table3.f1=B10011"

[0071] This SQL statement indicates that, using field f1 in table 3 as the filter field and B10011 as the filter condition, it filters out fields f1, f2, f3, and so on in table 3. The calculation logic in this SQL statement determines whether field f1 in table 3 is identical to "B10011." If so, the corresponding fields are filtered out and output. The filter field can be called the filter key, and fields f2, f3, and so on can be called the data to be filtered.

[0072] In this way, the master node 121 can obtain the computing logic of the business by parsing the SQL statement in the job request.

[0073] The master node 121 can perform operator arrangement based on the obtained business computing logic to obtain one or more computing operators such as the computing operator 221, as well as the read operator 210 and the output operator 230. As described above, the business computing logic can be obtained by parsing the business SQL statement. Therefore, the computing operator 221 and one or more computing operators, as well as the read operator 210 and the output operator 230 are arranged based on the business SQL statement.

[0074] The job request also includes information about the data source. Master node 121 can configure a read operator based on this information, allowing the read operator to read business data from the data source. Master node 121 can also configure the obtained business calculation logic into the read operator, allowing the read operator to identify the data required by the business calculation logic within the business data and obtain the required data based on the business calculation logic. This will be described in detail in the method embodiments below and will not be repeated here.

[0075] The job request also includes information about the data destination. Based on this information, master node 121 configures an output operator, enabling it to output the business processing result to the data destination. Master node 121 also configures the output operator to generate the business processing result based on the computation results of the computation operator and the business data (or non-computational data within the business data). This will be described in detail in the following method embodiments and will not be repeated here.

[0076] 2 , the master node 121 can deploy one or more computing operators, such as a read operator 210, an output operator 230, and a compute operator 221, in the compute node cluster 120. Different operators among the read operator 210, the output operator 230, and the compute operator can be deployed on the same compute node in the compute node cluster 120 or on different compute nodes.

[0077] The stream computing device 200 comprises one or more computing operators, including the computing operator 221 deployed in the computing node cluster 120, along with the read operator 210 and the output operator 230. The stream computing device 200, also known as a streaming computing engine, is used to process business data in real time. Specifically, the read operator 210 reads business data from a data source, such as a data table or one or more rows within a data table. Based on the business's computational logic, the read operator 210 identifies the data required for the business's computational logic within the read business data, i.e., the data required for computation. The read operator 210 then inputs the required data into the compute operator 221. The compute operator 221 stores state based on the required data, thereby reducing the size of the state storage. The required data can meet the computational needs of the compute operator 221. The compute operator 221 can then perform computations on the required data based on the computational logic it carries, generating computation results. The output operator 230 can obtain the business processing results based on these computation results and the business data, and send the processing results to the data destination.

[0078] In some embodiments, as shown in FIG2 , a shared storage space 300 for the read operator 210 and the output operator 230 can be configured in the computing node cluster 120, that is, both the read operator 210 and the output operator 230 can access the storage space 300. Exemplarily, when the read operator 210 and the output operator 230 are on the same computing node, the shared storage space 300 can be located in the computing node. Exemplarily, when the read operator 210 and the output operator 230 are on different computing nodes, the shared storage space 300 can be located in the computing node where the read operator 210 is located or in the computing node where the output operator 230 is located. Exemplarily, the shared storage space 300 can be located in a computing node other than the computing node where the read operator 210 is located and the computing node where the output operator 230 is located.

[0079] In some embodiments, the read operator 210 may store the read business data in the shared storage space 300. The output operator 230 may obtain the business data from the shared storage space 300 and then, based on the business data and the calculation results of the calculation operator, obtain the business processing results. In some embodiments, the read operator 210 may treat the data in the business data other than the data required for calculation as non-computation required data and store the non-computation required data in the shared storage space 300. The output operator 230 may obtain the non-computation required data from the shared storage space 300 and then, based on the non-computation required data and the calculation results of the calculation operator, obtain the business processing results.

[0080] In some embodiments, as shown in FIG3 , shared storage space 300 can be implemented as a key (K)-value (V) storage system. That is, business data or non-computational data can be stored in shared storage space 300 in the form of KV pairs. In one example, shared storage space 300 can be implemented as an HBase database. In another example, shared storage space 300 can be implemented as a Redis database.

[0081] In some embodiments, a computing operator can perform state storage locally, i.e., the computing operator can store the data required for the computation and the computation results locally. In some embodiments, the computing operator can access the shared storage space 300. The computing operator can perform state storage in the shared storage space 300, i.e., the computing operator can store the data required for the computation and the computation results in the shared storage space 300, further reducing the local storage overhead of the operator.

[0082] The above examples introduce the system architecture and the stream computing device 200 provided in the embodiment of the present application. Next, the data processing method provided in the embodiment of the present application is described in conjunction with the system architecture and the stream computing device 200.

[0083] The method may be executed by the stream computing device 200, specifically by relevant operators in the stream computing device 200. As shown in FIG4 , the method includes the following steps.

[0084] In step 401 , the read operator 210 reads business data from a data source of the business.

[0085] In some embodiments, the data source may be a database, and the read operator 210 may read data from the database to obtain business data. In some embodiments, the data source may be a data acquisition terminal, i.e., the data source may acquire data from an acquisition object to obtain data. For example, the data source may be an environmental monitoring device that continuously monitors the environment and obtains data. The read operator 210 may read the most recently acquired data from the environmental monitoring device to obtain business data.

[0086] In some embodiments, the business data is streaming data, which is composed of multiple data points, wherein each data point includes at least one data. The reading operator 210 reads different data points from the data source at different reading times.

[0087] In some embodiments, there may be multiple data sources, that is, the read operator 210 may read data from multiple data sources simultaneously. For example, in a data association business scenario, the read operator 210 reads data from multiple data tables to concatenate data from different data tables.

[0088] Step 402 : The reading operator 210 identifies data required by the calculation logic of the business in the business data, and obtains the data required for calculation; wherein the data required for calculation includes the first data required by the first calculation logic.

[0089] As described above, the master node 121 can obtain the computing logic of the business by parsing the job request of the business. For example, the job request includes an SQL statement for executing the business, and the master node 121 parses the SQL statement to obtain the computing logic of the business. The master node 121 can configure the computing logic of the business into the read operator 210. Thus, the read operator 210 can identify the data required for the computing logic of the business based on the computing logic of the business in step 402. For example, for the computing logic of determining whether the f1 field in Table 1 and the f2 field in Table 2 are equal, the data required is the f1 field in Table 1 and the f2 field in Table 2. For another example, for the computing logic of determining whether the f1 field in Table 3 is equal to "B10011", the data required is the f1 field in Table 3.

[0090] In addition, the first calculation logic belongs to the calculation logic of the service, and specifically can be the calculation logic carried by calculation operator 221. Therefore, the data required for calculation obtained in step 402 includes the data required by the first calculation logic. For ease of description, the data required by the calculation logic carried by calculation operator 221 (first calculation logic) is referred to as first data.

[0091] In some embodiments, the business performed by the streaming computing device 200 includes splicing the data to be spliced ​​based on the splicing key of the data to be spliced. In this case, the computing logic carried by the computing operator 221 is to determine whether the splicing keys of the data to be spliced ​​are the same. The business data can be set to include the data to be spliced ​​C1, the splicing key C11 corresponding to the data to be spliced ​​C1, the data to be spliced ​​C2, and the splicing key C21 corresponding to the data to be spliced ​​C2. The computing logic carried by the computing operator 221 then includes determining whether the splicing key C11 and the splicing key C21 are the same. In step 402, based on the computing logic carried by the computing operator 221, it can be identified that the splicing key C11 and the splicing key C21 are the first data.

[0092] For example, the SQL statement of the business can be set as: "Select table1.f1, table1.f2, table1.f3, .., table2.f2, table2.f3, ... from table1 join table2 on table1.f1=table2.f2".

[0093] In this SQL statement, field f1 in table1 and field f2 in table2 serve as join keys and are the data required for the calculation. Fields f2 and f3 in table1, along with fields f3 in table2, are referred to as the data to be joined. Field f1 in table1 serves as join key C11, field f2 in table2 as join key C21, fields f2 and f3 in table1 as the data to be joined C1, and fields f3 in table2 as the data to be joined C2.

[0094] In some embodiments, the service performed by streaming computing device 200 includes outputting data that meets preset filtering conditions to a data destination. The service data includes the data to be filtered and a filter key corresponding to the data to be filtered. The computational logic implemented by computation operator 221 determines whether the filter key meets the preset filtering conditions. In step 402, the computational logic implemented by computation operator 221 identifies the filter key as the first data.

[0095] For example, the SQL statement of the business can be set as: "Select table3.f1, table3.f2, table3.f3, ..from table3 where table3.f1=B10011".

[0096] In this SQL statement, the f1 field in table3 is the filter key, B10011 is the preset filter condition, and the f2 field, f3 field, and other fields in table3 are the data to be filtered.

[0097] In some embodiments, as described above, the business data is streaming data, which is composed of multiple data points. The reading operator 210 reads different data points from the data source at different reading times. In step 402, each time the reading operator 210 reads a data point, the reading operator 210 can determine whether the data point is data required for calculation or data not required for calculation based on the calculation logic carried by the calculation operator 221. It can be set that at time T1, the reading operator 210 reads data point T11 from the data source. Based on the calculation logic carried by the calculation operator 221, the reading operator 210 identifies that data point T11 is data required for calculation, and then uses data point T11 as the first data. It can also be set that at time T2 after time T1, the reading operator 210 reads data point T21 from the data source. Based on the calculation logic carried by the calculation operator 221, the reading operator 210 identifies that data point T21 is not data required by the calculation logic carried by the calculation operator 221.

[0098] In step 403, the read operator 210 sends the data required for the calculation to the calculation operator 221, causing the calculation operator 221 to calculate the first data based on the first calculation logic to obtain a calculation result. The read operator 210 sends the data required for the calculation to the calculation operator 221, and also causes the calculation operator 221 to store the state based on the data required for the calculation.

[0099] The calculation operator 221 may also be referred to as the first calculation operator. The reading operator 210 may send the data required for calculation obtained in step 402 to the calculation operator 221. Upon receiving the data required for calculation, the calculation operator 221 may obtain the first data therein and perform calculations on the first data based on the calculation logic carried by the calculation operator 221 (i.e., the first calculation logic) to obtain a calculation result.

[0100] In some embodiments, as described above, the calculation logic carried by the calculation operator 221 is to determine whether the splicing keys of the data to be spliced ​​are the same, and the first data includes the splicing key C11 and the splicing key C21. In step 403, the calculation operator 221 can determine whether the splicing key C11 and the splicing key C21 are the same, and obtain a calculation result.

[0101] In some embodiments, as described above, the calculation logic carried by the calculation operator 221 is to determine whether the filter key meets the preset filter condition, and the first data includes the splicing key. The calculation operator 221 presets the filter condition. In step 403, the calculation operator 221 can determine whether the splicing key meets the filter condition to obtain a calculation result.

[0102] In some embodiments, as described above, the first data may be data point T11. In step 403, the calculation operator 221 may perform calculations on the data point T11 based on the first calculation logic to obtain a calculation result.

[0103] The calculation operator 221 stores the received data required for the calculation and the calculation result of the calculation operator 221 as a state. In one example, the calculation operator 221 can store the state in a local memory. In another example, the calculation operator 221 can store the state in the shared storage space 300.

[0104] The data received by the calculation operator 221 from the reading operator 210 is the data required for calculation, which is part of the business data read by the reading operator 210 in step 401. In other words, the data required for calculation is smaller than the business data read by the reading operator 210 in step 401. The calculation operator 221 stores state based on the data required for calculation, which can effectively reduce the storage overhead of state storage. In particular, state storage is typically stored in memory. Therefore, reducing the storage overhead of state storage can effectively save memory space and improve the computing performance of the calculation operator.

[0105] Step 404: The output operator 230 outputs the processing result of the business to the data target end of the business based on the calculation result and the business data.

[0106] After the calculation operator 221 completes the calculation of the first data, the obtained calculation result may be sent to the output operator 230 , so that the output operator 230 obtains the calculation result.

[0107] In some embodiments, the output operator 230 may obtain the business data obtained by the read operator 210 in step 401. In one example, the read operator 210 may store the business data obtained in step 401 in the shared storage space 300. In step 404, the output operator 230 may obtain the business data from the shared storage space 300.

[0108] In some embodiments, the output operator 230 can obtain data not required for computation from the business data. This data refers to data other than data required for computation from the business data obtained by the read operator 210 in step 401. In one example, the read operator 210 can store the data not required for computation in the shared storage space 300. In step 404, the output operator 230 can obtain the data not required for computation from the shared storage space 300.

[0109] The output operator 230 can obtain the processing result of the business based on the calculation result and data not required for calculation.

[0110] In some embodiments, as described above, the services executed by the streaming computing device 200 include a splicing key based on the data to be spliced. The computational logic carried by the computation operator 221 determines whether the splicing keys of the data to be spliced ​​are identical. In step 404, if the computation result indicates that the splicing keys C11 and C21 are identical, the output operator 230 splices the data to be spliced ​​C1 and C2 obtained from the service data or data not required for the computation. The output operator 230 then outputs the splicing result obtained by splicing the data to be spliced ​​C1 and C2. This splicing result serves as the processing result of the service.

[0111] In some embodiments, the service executed by stream computing device 200 includes outputting data that meets preset filtering conditions to a data destination. The computation logic carried by computation operator 221 determines whether a filter key meets the preset filtering conditions, and the second data includes the data to be filtered. In step 404, if the computation result indicates that the filter key meets the filtering conditions, output operator 230 retrieves the data to be filtered from the service data or data not required for the computation, and uses the data to be filtered as the result of the service processing.

[0112] In some embodiments, as described above, the calculation result is the result of calculation operator 221 performing a calculation on data point T11, and the second data is data point T21. Shared storage space 300 stores the business data obtained in step 401 or the non-computational data obtained in step 402. Both the business data and the non-computational data include data point T21. Therefore, output operator 230 can obtain data point T21 from shared storage space 300. In step 404, output operator 230 can obtain the business processing result based on the calculation result and data point T21.

[0113] When the output operator 230 obtains the processing result of the business, the output operator 230 can output the processing result to the data target end.

[0114] In some embodiments, referring to FIG2 , the stream computing device 200 further includes a computing operator 222 . The computing operator 222 may also be referred to as a second computing operator, and is used to carry the second computing logic of the service. That is, the computing operator 222 is used to perform data calculations according to the second computing logic. As shown in FIG2 , in the data flow direction of the service data, the computing operator 222 is located after the computing operator 221 , that is, the outflow data of the computing operator 221 is the inflow data of the computing operator 222 . The data required for the calculation obtained in step 402 includes the data required by the second computing logic. For the convenience of description, the data required by the second computing logic may be referred to as the second data.

[0115] In one example of this embodiment, after receiving the data required for calculation, operator 221 can identify second data from the data required for calculation and send the second data to operator 222, causing operator 222 to perform state storage based on the second data and causing operator 222 to perform calculations on the second data based on the second calculation logic to obtain a calculation result. The calculation result is used to output operator 230 to obtain the processing result of the business and send the processing result of the business to the data target end. In one example of this embodiment, operator 221 can identify the data required by the calculation logic of operator 222 and send the data required by the calculation logic of operator 222 to the current operator. In another example of this embodiment, operator 221 can remove the first data from the data required for calculation received by operator 221, obtaining data that includes the second data but does not include the first data. Then, operator 221 sends the data that includes the second data but does not include the first data to operator 222.

[0116] This example can further reduce the amount of data stored in the state of the computing operator.

[0117] In another example of this embodiment, after receiving the data required for calculation, calculation operator 221 may send the data required for calculation to calculation operator 222, causing calculation operator 222 to store the state based on the data required for calculation, and causing calculation operator 222 to calculate the second data based on the second calculation logic to obtain a calculation result. The calculation result is used to output the processing result of the business to operator 230, and the processing result of the business is sent to the data target end.

[0118] To sum up, in the data processing method provided in the embodiment of the present application, the reading operator can identify the data required for calculation in the business data, and then send the data required for calculation to the calculation operator, so that the calculation operator can perform state storage and calculation based on the data required for calculation, thereby meeting the computing requirements of the calculation operator while reducing the amount of data stored in the state of the calculation operator, saving the operator's memory, and improving the operator's computing performance.

[0119] Referring to FIG5 , an embodiment of the present application further provides a stream computing device 500. As shown in FIG5 , the stream computing device 500 includes a read operator 510, a first computing operator 520, and an output operator 530. The first computing operator 520 is used to perform data calculation based on the first computing logic of the business; wherein the read operator 510 is used to read business data from the data source end of the business; the read operator 510 is used to identify the data required by the computing logic of the business in the business data and obtain the data required for calculation; wherein the data required for calculation includes the first data required by the first computing logic; the read operator 510 is used to send the data required for calculation to the first computing operator 520, so that the first computing operator 520 calculates the first data based on the first computing logic to obtain a calculation result; the output operator 530 is used to output the processing result of the business to the data target end of the business based on the calculation result and the business data.

[0120] In some embodiments, the streaming computing device 500 corresponds to a storage space accessed by the read operator 510 and the output operator 520; wherein the read operator 510 is also used to store non-computationally required data in the storage space, and the non-computationally required data is data in the business data other than the data required for calculation; the output operator 530 is also used to read the non-computationally required data from the storage space; the output operator 530 is also used to output the processing result to the data target end based on the non-computationally required data and the calculation result.

[0121] In some embodiments, the streaming computing device corresponds to a storage space accessed by the read operator 510 and the output operator 530; wherein the read operator 510 is also used to store the business data in the storage space; the output operator 530 is also used to read the business data from the storage space.

[0122] In some embodiments, the streaming computing device 500 includes a second computing operator, which is used to perform data calculation according to the second computing logic of the business; in the data flow direction of the business, the second computing operator is located after the first computing operator 520; wherein, the data required for the calculation also includes the second data required by the second computing logic; the first computing operator 520 is used to send the second data to the second computing operator, so that the second computing operator calculates the second data based on the second computing logic.

[0123] In some embodiments, the business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; the first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; the read operator 510 is used to identify the first splicing key and the second splicing as the first data based on the first calculation logic; when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator 530 is used to splice the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

[0124] In some embodiments, the business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key meets the preset filtering conditions; the read operator 510 is used to identify the filter key as the first data based on the first calculation logic; when the calculation result indicates that the filter key meets the preset filtering conditions, the output operator 530 is used to use the data to be filtered as the processing result.

[0125] In some embodiments, the business data is streaming data consisting of multiple data points, each of the multiple data points includes at least one data; wherein, the reading operator 510 reads different data points of the multiple data points from the data source end at different times; the reading operator 510 is used to read the first data point of the multiple data points at a first time; the reading operator 510 is used to identify the first data point as the first data based on the first computing logic.

[0126] In an example of this embodiment, the read operator 510 is used to read a second data point among the multiple data points at a second moment; when the second data point is data other than the data required by the first calculation logic, the output operator 530 is used to output the processing result to the data target end of the business based on the calculation result and the second data point.

[0127] In some embodiments, the computing logic of the business is obtained based on a structured query language SQL statement of the business.

[0128] In some embodiments, the first computing operator 520 is arranged based on the SQL statement of the business.

[0129] Among them, the read operator 510, the first calculation operator 520, and the output operator 530 can all be implemented by software or by hardware. For example, the implementation of the read operator 510 is described below using the read operator 510 as an example. Similarly, the implementation of the first calculation operator 520 and the output operator 530 can refer to the implementation of the read operator 510.

[0130] As an example of a software functional unit, the module read operator 510 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the read operator 510 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone AZ or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0131] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0132] As an example of a hardware functional unit, the read operator 510 may include at least one computing device, such as a server. Alternatively, the read operator 510 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0133] The multiple computing devices included in the read operator 510 can be distributed in the same region or in different regions. The multiple computing devices included in the read operator 510 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the read operator 510 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0134] This application also provides a computing device 600. As shown in Figure 6, computing device 600 includes a bus 602, a processor 604, a memory 606, and a communication interface 608. Processor 604, memory 606, and communication interface 608 communicate with each other via bus 602. Computing device 600 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 600.

[0135] Bus 602 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG6 shows a single bus line, but this does not imply a single bus or type of bus. Bus 602 may include a path for transmitting information between various components of computing device 600 (e.g., memory 606, processor 604, and communication interface 608).

[0136] The processor 604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0137] The memory 606 may include volatile memory, such as random access memory (RAM). The memory 606 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0138] The memory 606 stores executable program code, and the processor 604 executes the executable program code to respectively implement the functions of the aforementioned read operator 510, the first calculation operator 520, and the output operator 530, thereby implementing the method shown in Figure 4. In other words, the memory 606 stores instructions for executing the method shown in Figure 4.

[0139] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or a communication network.

[0140] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0141] As shown in Figure 7 , the computing device cluster includes at least one computing device 600. The memory 606 in one or more computing devices 600 in the computing device cluster may store the same instructions for executing the method shown in Figure 4 .

[0142] In some possible implementations, the memory 606 of one or more computing devices 600 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 4. In other words, the combination of one or more computing devices 600 can jointly execute the instructions for executing the method shown in Figure 4.

[0143] It should be noted that the memory 606 in different computing devices 600 in the computing device cluster can store different instructions, each used to execute part of the functions of the streaming computing apparatus 500. In other words, the instructions stored in the memory 606 in different computing devices 600 can implement the functions of one or more modules among the reading operator 510, the first calculation operator 520, and the output operator 530.

[0144] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG8 illustrates a possible implementation. As shown in FIG8 , two computing devices 600A and 600B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 606 in the computing device 600A stores instructions for executing the function of the read operator 510. Simultaneously, the memory 606 in the computing device 600B stores instructions for executing the functions of the first computing operator 520 and the output operator 530.

[0145] It should be understood that the functionality of the computing device 600A shown in FIG8 may also be implemented by multiple computing devices 600. Similarly, the functionality of the computing device 600B may also be implemented by multiple computing devices 600.

[0146] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 7 and 8. However, the memory 606 of one or more computing devices 600 in this computing device cluster can store the same instructions for executing the method shown in Figure 4.

[0147] In some possible implementations, the memory 606 of one or more computing devices 600 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 4. In other words, the combination of one or more computing devices 600 can jointly execute the instructions for executing the method shown in Figure 4.

[0148] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method shown in FIG4 .

[0149] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a host migration device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method shown in FIG. 4 .

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method is applied to a stream computing device, the stream computing device includes a reading operator, a first computing operator and an output operator, wherein the first computing operator is used to perform data computing based on a first computing logic of a business; the method includes: The reading operator reads the business data from the data source of the business; The read operator identifies data required by the computing logic of the business in the business data, and obtains the data required for the computing; wherein the data required for the computing includes the first data required by the first computing logic; The read operator sends the data required for the calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data.

2. The method according to claim 1, characterized in that The stream computing device corresponds to a storage space accessible by the read operator and the output operator; The method further comprises: The read operator stores the data not required for calculation into the storage space, where the data not required for calculation is the data in the business data other than the data required for calculation; The output operator reads the data not required for calculation from the storage space; The output operator outputs the processing result of the business to the data target end based on the calculation result and the business data, including: the output operator outputs the processing result to the data target end based on the non-calculation required data and the calculation result.

3. The method according to claim 1, characterized in that The stream computing device corresponds to a storage space accessible by the read operator and the output operator; The method further comprises: The read operator stores the business data in the storage space; The output operator reads the business data from the storage space.

4. The method according to any one of claims 1 to 3, characterized in that The stream computing device includes a second computing operator, which is used to perform data computing according to a second computing logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator; wherein, The data required for calculation also includes second data required by the second calculation logic; the first calculation operator is used to send the second data to the second calculation operator, so that the second calculation operator calculates the second data based on the second calculation logic.

5. The method according to any one of claims 1 to 4, characterized in that The business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; The first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for calculation, including: the read operator identifies the first splicing key and the second splicing as the first data based on the first calculation logic; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: When the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator splices the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

6. The method according to any one of claims 1 to 4, characterized in that The business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key satisfies a preset filter condition; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for calculation, including: the read operator identifies the filter key as the first data based on the first calculation logic; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the calculation result indicates that the filter key meets the preset filtering condition, the output operator uses the data to be filtered as the processing result.

7. The method according to any one of claims 1 to 6, characterized in that The business data is streaming data composed of multiple data points, each of which includes at least one data; wherein the reading operator reads different data points from the data source end at different times; The reading operator reads the business data from the data source end of the business, including: the reading operator reads a first data point among the multiple data points at a first moment; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for the calculation, including: the read operator identifies the first data point as the first data based on the first calculation logic.

8. The method according to claim 7, characterized in that The reading operator reads the business data from the data source end of the business, including: the reading operator reads a second data point among the multiple data points at a second time; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the second data point is data other than the data required by the first calculation logic, the output operator outputs the processing result to the data target end of the business based on the calculation result and the second data point.

9. A streaming computing device, characterized in that: The device includes a reading operator, a first calculation operator and an output operator, wherein the first calculation operator is used to perform data calculation based on a first calculation logic of the business; wherein, The read operator is used to read business data from the data source end of the business; The read operator is used to identify data required by the computing logic of the business in the business data, and obtain the data required for the computing; wherein the data required for the computing includes the first data required by the first computing logic; The read operator is used to send the data required for the calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; The output operator is used to output the processing result of the business to the data target end of the business based on the calculation result and the business data.

10. The device according to claim 9, characterized in that The device corresponds to a storage space accessed by the read operator and the output operator; wherein, The read operator is further used to store data not required for calculation into the storage space, where the data not required for calculation is data in the business data other than the data required for calculation; The output operator is also used to read the non-computation required data from the storage space; The output operator is further used to output the processing result to the data target end based on the non-computation required data and the calculation result.

11. The device according to claim 9, characterized in that The device corresponds to a storage space accessed by the read operator and the output operator; wherein, The read operator is also used to store the business data in the storage space; The output operator is also used to read the business data from the storage space.

12. The device according to any one of claims 9 to 11, characterized in that The device includes a second computing operator, which is used to perform data calculation according to a second computing logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator; wherein, The data required for calculation also includes second data required by the second calculation logic; the first calculation operator is used to send the second data to the second calculation operator, so that the second calculation operator calculates the second data based on the second calculation logic.

13. The device according to any one of claims 9 to 12, characterized in that The business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; The first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; The read operator is used to identify the first splicing key and the second splicing as the first data based on the first calculation logic; When the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator is used to splice the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

14. The device according to any one of claims 8 to 12, characterized in that The business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key satisfies a preset filter condition; The read operator is used to identify the filter key as the first data based on the first calculation logic; When the calculation result indicates that the filter key satisfies the preset filter condition, the output operator is used to use the data to be filtered as the processing result.

15. The device according to any one of claims 8 to 14, characterized in that The business data is streaming data composed of multiple data points, each of which includes at least one data; wherein the reading operator reads different data points from the data source end at different times; The read operator is used to read a first data point among the multiple data points at a first moment; The read operator is used to identify the first data point as the first data based on the first calculation logic.

16. The device according to claim 15, characterized in that The read operator is used to read a second data point among the plurality of data points at a second time; When the second data point is data other than data required by the first calculation logic, the output operator is used to output the processing result to the data target end of the business based on the calculation result and the second data point.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.

19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device cluster, the computer device cluster is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method and device and cluster

    CN120020709A

  • Streaming computing engine running method and system for skew data

    CN110990059A

  • Neutral data application method, device and system

    CN111062057A

  • Streaming computing method and device based on DAG interaction

    CN111782371A

  • Stateless calculation data processing method, program product and electronic device

    CN115269038A