Data processing method and device and cluster

By introducing a read operator in the streaming computing device, identifying and sending only the data required for the calculation logic to the calculation operator, the storage overhead and memory consumption problems caused by the storage based on the state storage of the full data in streaming computing are solved, and more efficient computing performance and service processing are achieved.

CN120020709APending Publication Date: 2025-05-20HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410065781.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-01-16
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In streaming computing, state storage based on full data results in unnecessary storage overhead and memory consumption, affecting computing performance.

Method used

By introducing a read operator in the streaming computing device, the data required for only the calculation logic is identified and sent to the calculation operator, so that the calculation operator is stored in a state based on the data required for the calculation logic.

Benefits of technology

The amount of state storage data of the computing operator is reduced, memory consumption is saved, computing performance is improved, and business processing is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020709A_ABST
    Figure CN120020709A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device and a cluster. The method comprises the following steps that: a reading operator reads service data from a data source end of a service; reading data required by the operator for identifying calculation logic of the business in the business data to obtain data required by calculation; wherein the data required by calculation comprises first data required by the first calculation logic; the reading operator sends the data required for calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; and the output operator outputs a processing result of the service to a data target end of the service based on the calculation result and the service data. According to the method, the data volume stored in the state of the calculation operator can be reduced and the memory of the calculation operator can be saved while the normal processing of the service is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311541722.3 and the application title "A Streaming Computing Method" submitted to the China National Intellectual Property Administration on November 17, 2023. The entire content thereof is incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technologies, and in particular, to a data processing method, apparatus, and cluster. Background Art

[0003] Streaming computing is a computing method capable of real-time processing of data. Generally speaking, streaming computing can be divided into stateful computing and stateless computing. For stateful computing, there is a relationship between the front and back data. An operator needs to calculate the current data based on the previous data or the previous calculation results. Therefore, the operator performing stateful computing stores the previous data and the previous calculation results as a state to calculate the new data when receiving new data by using the stored state (i.e., the previous data and the previous calculation results).

[0004] In the related art, all data is sent to the operator, enabling the operator to perform state storage based on all data. Usually, only a small part of all data is required for the operator's calculation. Therefore, the state storage based on all data incurs unnecessary storage overhead. Moreover, since the state storage is performed in memory, the large amount of data for state storage causes a large memory consumption of the operator, which affects the computing performance of the operator. Summary of the Invention

[0005] This application provides a data processing method, apparatus, and cluster, which can reduce the amount of data stored in the state of the computing operator and save the memory of the computing operator while ensuring the normal processing of the service.

[0006] In a first aspect, a data processing method is provided. This method is applied to a streaming computing device, which includes a reading operator, a first computing operator, and an output operator. The first computing operator is used to calculate data based on the first computing logic of the service. The method includes: the reading operator reads service data from the data source end of the service; the reading operator identifies the data required by the computing logic of the service in the service data to obtain the data required for calculation; the data required for calculation includes the first data required by the first computing logic; the reading operator sends the data required for calculation to the first computing operator, enabling the first computing operator to calculate the first data based on the first computing logic to obtain a calculation result; the output operator outputs the processing result of the service to the data target end of the service based on the calculation result and the service data.

[0007] The computing operator stores its state based on the data received by the computing operator. Usually, only a small part of the service data read by the reading operator is the data required for the computing logic of the computing operator. This method identifies the data required for the computing logic of the computing operator in the service data and sends the data required for the computing logic of the computing operator to the computing operator, so that the computing operator stores its state based on the data required for the computing logic of the computing operator. Compared with the state storage based on the service data, the state storage based on the data required for the computing logic greatly reduces the amount of data for state storage, thereby reducing the memory consumption of the operator for state storage. Moreover, the data required for the computing logic meets the computing requirements of the computing operator. Based on the computing result of the computing operator on the data required for the computing logic and the service data, the output operator can obtain the processing result of the service. That is to say, in the case where the reading operator only sends the data required for the computing logic of the computing operator to the computing operator, the output operator can obtain the processing result of the service without affecting the processing of the service.

[0008] In short, the method provided by the embodiments of the present application can reduce the amount of data for state storage of the computing operator while ensuring the normal processing of the service, thereby saving the memory of the computing operator and ensuring the computing performance of the computing operator.

[0009] In a possible implementation, the streaming computing device corresponds to a storage space accessible by the reading operator and the output operator; the method further includes: the reading operator stores the non-computing required data in the storage space, where the non-computing required data is the data other than the computing required data in the service data; the output operator reads the non-computing required data from the storage space; the output operator outputs the processing result of the service to the data target end based on the computing result and the service data, including: the output operator outputs the processing result to the data target end based on the non-computing required data and the computing result.

[0010] This storage space is a shared storage space for the reading operator and the output operator. The non-computing required data and the computing result of the computing operator are used for the output operator to obtain the processing result of the service. In this implementation, the reading operator stores the non-computing required data in the shared storage space. In the case where the reading operator does not need to transfer the non-computing required data through an operator (such as a computing operator) between the reading operator and the output operator, the output operator can obtain the non-computing required data, so that the processing result of the service can be obtained based on the non-computing required data and the computing result of the computing operator.

[0011] In a possible implementation, the streaming computing device corresponds to a storage space accessible by the reading operator and the output operator; the method further includes: the reading operator stores the service data in the storage space; the output operator reads the service data from the storage space.

[0012] This storage space is a shared storage space for the read operator and the output operator. The service data and the calculation results of the calculation operator are used by the output operator to obtain the processing result of the service. In this implementation, the read operator stores the service data in the shared storage space. When the read operator does not need to transfer the service data through an operator (such as a calculation operator) between the read operator and the output operator, the output operator can obtain the non-calculation required data, so that the processing result of the service can be obtained based on the non-calculation required data and the calculation results of the calculation operator.

[0013] In a possible implementation, the streaming computing device includes a second calculation operator, and the second calculation operator is used to perform data calculation according to the second calculation logic of the service; in the data flow direction of the service, the second calculation operator is located after the first calculation operator; wherein, the calculation required data further includes the second data required by the second calculation logic; the first calculation operator is used to send the second data to the second calculation operator, so that the second calculation operator calculates the second data based on the second calculation logic.

[0014] In this implementation, the calculation operator can identify the data required by the calculation logic of the downstream calculation operator of the calculation operator, and send the data required for the calculation of the downstream calculation operator to the downstream calculation operator. While ensuring the calculation requirements of the downstream calculation operator, the data volume of the state storage of the downstream calculation operator is further reduced, and the memory of the downstream calculation operator is further saved.

[0015] In a possible implementation, the service data includes the first data to be spliced, the first splicing key corresponding to the first data to be spliced, the second data to be spliced, and the second splicing key corresponding to the second data to be spliced; the first calculation logic includes: judging whether the first splicing key and the second splicing key are the same; the read operator identifies the data required by the calculation logic of the service in the service data to obtain the calculation required data, including: the read operator identifies the first splicing key and the second splicing as the first data based on the first calculation logic; the output operator outputs the processing result of the service to the data target end of the service based on the calculation result and the service data, including: when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator splices the first data to be spliced and the second data to be spliced to obtain the processing result.

[0016] In this implementation, when the calculation operator executes the data splicing service, the read operator only needs to send the splicing key to the calculation operator. While meeting the calculation requirements of the calculation operator, the calculation operator performs state storage based on the splicing key, reducing the data volume of the state storage of the calculation operator.

[0017] In a possible implementation, the service data includes data to be filtered and a filtering key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filtering key meets a preset filtering condition; the reading operator identifies the data required by the calculation logic of the service from the service data to obtain the data required for calculation, including: the reading operator, based on the first calculation logic, identifies the filtering key as the first data; the output operator outputs the processing result of the service to the data target end of the service based on the calculation result and the service data, including: when the calculation result indicates that the filtering key meets the preset filtering condition, the output operator uses the data to be filtered as the processing result.

[0018] In this implementation, when the calculation operator executes data filtering and output for the service, the reading operator only needs to send the filtering key to the calculation operator. While meeting the calculation requirements of the calculation operator, it enables the calculation operator to perform status storage based on the filtering key, reducing the amount of data stored in the status of the calculation operator.

[0019] In a possible implementation, the service data is streaming data composed of multiple data points, and each data point in the multiple data points includes at least one piece of data; among them, the reading operator reads different data points at different times from the data source end of the service; the reading operator reads the service data from the data source end of the service, including: the reading operator reads the first data point among the multiple data points at the first moment; the reading operator identifies the data required by the calculation logic of the service from the service data to obtain the data required for calculation, including: the reading operator, based on the first calculation logic, identifies the first data point as the first data.

[0020] Whenever data is read, the reading operator can determine whether the data belongs to the data required for calculation or non-calculation required data. When the data belongs to the data required for calculation, the reading operator sends the data to the calculation operator, enabling the calculation operator to calculate the data in a timely manner and ensuring the real-time nature of the streaming calculation.

[0021] In a possible implementation, the reading operator reads the service data from the data source end of the service, including: the reading operator reads the second data point among the multiple data points at the second moment; when the second data point is data other than that required by the first calculation logic, the output operator outputs the processing result of the service to the data target end of the service based on the calculation result and the second data, including: the output operator outputs the processing result to the data target end of the service based on the calculation result and the second data point.

[0022] In a possible implementation, the calculation logic of the service is obtained based on the structured query language (SQL) statement of the service.

[0023] In this implementation manner, the calculation logic of the service can be obtained based on the SQL statement of the service, and the calculation logic of the service is configured into the reading operator. Thus, the reading operator can identify the data required by the calculation logic of the service from the service data according to the calculation logic of the service.

[0024] In a possible implementation manner, the first calculation operator is arranged based on the SQL statement of the service.

[0025] In this implementation manner, the calculation logic of the service can be obtained based on the SQL of the service. Then, based on the calculation logic of the service, the operator carrying the calculation logic of the service can be arranged. Among them, the arranged first calculation operator is used to carry the first calculation logic of the service, that is, the first calculation operator is used to perform data calculation according to the first calculation logic, so that the service is processed.

[0026] In a second aspect, a streaming computing device is provided. The streaming computing device includes a reading operator, a first calculation operator, and an output operator. The first calculation operator is used to perform data calculation based on the first calculation logic of the service. Among them, the reading operator is used to read service data from the data source end of the service; the reading operator is used to identify the data required by the calculation logic of the service from the service data to obtain the data required for calculation; among them, the data required for calculation includes the first data required by the first calculation logic; the reading operator is used to send the data required for calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; the output operator is used to output the processing result of the service to the data target end of the service based on the calculation result and the service data.

[0027] In a possible implementation manner, the streaming computing device corresponds to a storage space accessible by the reading operator and the output operator. Among them, the reading operator is further used to store the data that is not required for calculation into the storage space. The data that is not required for calculation is the data in the service data other than the data required for calculation; the output operator is further used to read the data that is not required for calculation from the storage space; the output operator is further used to output the processing result to the data target end based on the data that is not required for calculation and the calculation result.

[0028] In a possible implementation manner, the streaming computing device corresponds to a storage space accessible by the reading operator and the output operator. Among them, the reading operator is further used to store the service data into the storage space; the output operator is further used to read the service data from the storage space.

[0029] In a possible implementation, the streaming computing device includes a second computing operator, which is used to perform data calculation according to the second computing logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator; wherein, the data required for calculation further includes second data required by the second computing logic; the first computing operator is used to send the second data to the second computing operator, so that the second computing operator calculates the second data based on the second computing logic.

[0030] In a possible implementation, the service data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; the first computing logic includes: determining whether the first splicing key and the second splicing key are the same; the reading operator is used to identify the first splicing key and the second splicing key as first data based on the first computing logic; when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator is used to splice the first data to be spliced and the second data to be spliced to obtain a processing result.

[0031] In a possible implementation, the service data includes data to be filtered and a filtering key corresponding to the data to be filtered; the first computing logic includes: determining whether the filtering key meets a preset filtering condition; the reading operator is used to identify the filtering key as first data based on the first computing logic; when the calculation result indicates that the filtering key meets the preset filtering condition, the output operator is used to use the data to be filtered as the processing result.

[0032] In a possible implementation, the service data is streaming data composed of multiple data points, and each data point in the multiple data points includes at least one piece of data; wherein, the reading operator reads different data points in the multiple data points at different times; the reading operator is used to read the first data point in the multiple data points at the first moment; the reading operator is used to identify the first data point as first data based on the first computing logic.

[0033] In a possible implementation, the reading operator is used to read the second data point in the multiple data points at the second moment; when the second data point is data other than the data required by the first computing logic, the output operator is used to output a processing result to the data target end of the service based on the calculation result and the second data point.

[0034] In a possible implementation, the computing logic of the service is obtained based on the structured query language SQL statement of the service.

[0035] In a possible implementation, the first computing operator is arranged based on the SQL statement of the service.

[0036] In a third aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method provided in the first aspect.

[0037] In a fourth aspect, a computer-readable storage medium is provided, including computer program instructions, and when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method provided in the first aspect.

[0038] In a fifth aspect, a computer program product including instructions is provided, and when the instructions are run by a computer device cluster, the computer device cluster is caused to execute the method provided in the first aspect.

[0039] For the beneficial effects of the second to fifth aspects, reference may be made to the introduction of the beneficial effects of the first aspect above, and details are not repeated here. Description of the Drawings

[0040] Figure 1 A schematic diagram of a system architecture provided by an embodiment of the present application;

[0041] Figure 2 A schematic diagram of a computing node cluster provided by an embodiment of the present application;

[0042] Figure 3 A schematic diagram of a shared storage space provided by an embodiment of the present application;

[0043] Figure 4 A flowchart of a data processing method provided by an embodiment of the present application;

[0044] Figure 5 A schematic structural diagram of a streaming computing device provided by an embodiment of the present application;

[0045] Figure 6 A schematic structural diagram of a computing device provided by an embodiment of the present application;

[0046] Figure 7 A schematic structural diagram of a computing device cluster provided by an embodiment of the present application;

[0047] Figure 8 A schematic structural diagram of a computing device cluster connected through a network provided by an embodiment of the present application. Detailed Embodiments

[0048] The solutions provided in the embodiments of the present application will be described below in conjunction with the accompanying drawings. Among them, in the embodiments of the present application, "a plurality of" means two or more, and "a variety of" means two or more. "First", "second", etc. are only used to distinguish similar objects and do not necessarily need to describe a specific order or the number of objects.

[0049] To facilitate the understanding of the solutions provided in the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application will be introduced first.

[0050] Streaming data: Also known as time-series data, it refers to data generated in chronological order, that is, data continuously generated over time. Streaming data consists of data points generated in chronological order, where each data point may include at least one piece of data. The generation times of different data points are different. That is to say, at a certain time point or time period, one or a segment of data can be generated, and this one or a segment of data can be called a data point. Data points generated at multiple adjacent time points or time periods in sequence constitute streaming data. Streaming data often has characteristics such as a large amount of data, a short interval between data points, and continuous arrival of data points.

[0051] Streaming computing: Also known as stream computing, streaming computing is used for real-time processing of data. Among them, streaming computing can be a computing model triggered by the input data, and each newly input data can be used as an event to trigger the computing model to perform calculations.

[0052] Stateless computing: A type of streaming computing. In stateless computing, data is processed separately, and there is no relationship between different data. For example, for the stateless computing of "input value + 5", if the current input value (i.e., the current data) is "10", then the input is calculated by adding 5 to obtain the calculation result of "15".

[0053] Stateful computing: Another type in streaming computing. There is a relationship between the input values before and after (i.e., the data input before and after), and it is necessary to calculate the current input value based on the previous input value or the previous calculation result. For convenient calculation, the previous input value or the previous calculation result is stored as a state. Taking the cumulative calculation, a stateful calculation, as an example, it can be set that the first input value is "5", then the result of the cumulative calculation is "5"; the second input value is "100", then the result of the cumulative calculation is "5 + 100 = 105"; the third input value is "200", and the result of the cumulative calculation is "105 + 200 = 305". It can be seen that the current calculation needs to use the previous calculation result, so the previous calculation result needs to be stored as a state. Taking the association calculation, another stateful calculation, as an example, it can be set that data source A1 is an employee information table, including the identifier (ID) of the employee, the name of the employee, and the ID of the department to which the employee belongs, and data source A2 is department information, including the ID of the department and the name of the department. The association calculation specifically obtains the information of the employees in the department through the ID of the department, such as the ID of the employee, the name of the employee, the ID of the department to which the employee belongs, etc. Among them, the ID of the department here belongs to the association field. The calculation logic of the association calculation is to judge whether the association field in the input value from data source A1 (i.e., the ID of the department) and the association field in the input value from data source A2 (i.e., the ID of the department) are the same. If they are the same (i.e., the calculation result of the association calculation is the same), then output the information of the employees in the input value from data source A1 and the information of the employees in the input value from data source A1. It can be set that the input value from data source A1 is "1 (ID of the employee), Li Qiang (name of the employee), B10011 (ID of the department to which the employee belongs)", and the input value from data source A2 is "B10011 (ID of the department), R & D Department (name of the department)", then the output result is "1, Li Qiang, R & D Department". Since the information of the employees in the data source may change, for example, the department to which the employee belongs changes, this requires changing the previously output employee information. It is necessary to find the previous input value and the output result, so the previous input value and the input result need to be stored as a state. Also, there may be duplicate data. To ensure the accuracy of the output result, it is necessary to remove duplicates based on the previous input value. Therefore, the previous input value also needs to be stored as a state.

[0054] Operator: A data processing unit that carries the calculation logic. Among them, the operator is the smallest executable unit in streaming computing. The operator is used to calculate relevant data based on the calculation logic it carries. Common operators in a streaming computing device include read operators, calculation operators, and output operators. Usually, a streaming computing device needs a read operator, at least one calculation operator, and an output operator to execute a certain business.

[0055] Read operator: Also known as the source operator, it is an operator in a streaming computing device used to read business data from the data source end of the business.

[0056] Computation operator: Also known as the task operator, it is an operator in a streaming computing device used to carry the computing logic of the business, that is, the computation operator is used to execute the computing logic of the business.

[0057] Output operator: Also known as the sink operator, it is used to obtain the processing result of the business based on the calculation result of the computation operator, and send the processing result of the business to the data target end of the business, such as the storage end for storing the processing result of the business.

[0058] Data required for computation: It refers to the data in the business data required by the computation operator to execute the computing logic, that is, the computing object for which the computation operator performs calculations based on the computing logic. The data required for computation is also known as the data required by the computing logic. Among them, the business data can be recorded in a data table, and the data required for computation can be one or more fields in the data table, and the one or more fields can be called the fields required for computation or computing fields. For example, the computing logic of an association calculation includes determining whether the association fields of two tables to be associated are equal, then the data required by this computing logic is the association fields of the two tables to be associated. Another example is that the computing logic of a filtering calculation includes determining whether the filtering field meets the filtering condition, and the data required by this computing logic is the filtering field.

[0059] Data not required for computation: It refers to the data in the business data other than the data required for computation.

[0060] State storage: It means that the operator in the streaming computing stores the data received by the operator and the calculation result of the operator as the state. Usually, the operator stores the state in local memory. That is to say, the state storage of the operator will consume the memory of the operator.

[0061] An embodiment of the present application provides a data processing method. In this method, a reading operator, based on the calculation logic of a service, identifies the data required by the calculation logic of the service from the read service data to obtain the data required for calculation. Among them, the data required for calculation is a part of the service data, specifically the data required by the calculation logic of the service in the service data. The reading operator sends the data required for calculation instead of the service data to the calculation operator of the service, reducing the amount of data sent to the calculation operator. The calculation operator can perform state storage based on the data required for calculation, thereby reducing the amount of data for state storage. The calculation operator calculates the data required for calculation based on the calculation logic carried by the calculation operator to obtain a calculation result. That is, the data required for calculation can meet the calculation requirements of the calculation operator. The output operator can obtain the processing result of the service based on the calculation result and the service data, and send the processing result to the data target end of the service. In this method, the reading operator only sends the data required for calculation to the calculation operator, reducing the network overhead between the reading operator and the calculation operator. In particular, the calculation operator only receives the data required for calculation and performs state storage based on the data required for calculation, reducing the amount of data for state storage, thereby saving the memory consumption of state storage and improving the calculation performance of the calculation operator.

[0062] Next, the data processing method provided by the embodiment of the present application will be introduced.

[0063] Figure 1 Fig. shows a system architecture that can be used to implement the data processing method provided by the embodiment of the present application. As Figure 1 shown, the system architecture includes a user 110 and a computing node cluster 120. Among them, the computing node cluster 120 includes multiple computing nodes. The computing node can be a physical node, such as a server. The computing node can also be a virtual computing node such as a virtual machine (VM) or a container. The user can trigger the computing node cluster 120 to execute services, such as data association services and data filtering services, through a job request. Among them, the data association service is also called an association calculation service, which refers to associating or splicing the data in two or more data tables based on association conditions. The data filtering service is also called data filtering calculation, which refers to filtering or screening out the data that meets the preset filtering conditions from the data.

[0064] Among them, a master node 121 is deployed in the computing node cluster 120. Among them, the master node 121 can be deployed on any one or more computing nodes in the computing node cluster 120. The master node 121 can receive the job request of the service issued by the user 110 and perform operator orchestration based on the job request to obtain a streaming computing device 200, specifically as follows.

[0065] A job request may include the calculation logic information of a service. For example, a job request includes one or more Structured Query Languages (SQLs). SQL statements reflect the calculation logic of a service. The calculation logic of the service can be obtained by parsing the SQL statements. In one example, it can be set that the job request includes the following SQL statement.

[0066] “Select table1.f1, table1.f2, table1.f3,.., table2.f2, table2.f3,… from table1 join table2 on table1.f1=table2.f2”

[0067] This SQL statement means that taking the f1 field in table1 and the f2 field in table2 as the associated fields, concatenating the f1 field, f2 field, f3 field… in table1 and the f2 field, f3 field… in table2. Among them, the associated fields can be called concatenation fields or concatenation keys. The f2 field, f3 field, etc. in table1 and the f3 field, etc. in table2 belong to the data to be concatenated. The calculation logic included in this SQL is to judge whether the f1 field in table1 is the same as the f2 field in table2. If they are the same, then concatenate the f2 field, f3 field, etc. in table1 and the f3 field, etc. in table2.

[0068] In one example, it can also be set that the job request includes the following SQL statement.

[0069] “Select table3.f1, table3.f2, table3.f3,.. from table3 where table3.f1=B10011”

[0070] This SQL statement means that taking the f1 field in table3 as the filtering field and B10011 as the filtering condition, filtering out the f1 field, f2 field, f3 field… in table3. The calculation logic included in this SQL is to judge whether the f1 field in table3 is the same as “B10011”. If they are the same, then filter out and output the corresponding fields. Among them, the filtering field can be called the filtering key, and the f2 field, f3 field, etc. can be called the data to be filtered.

[0071] In this way, the master node 121 can obtain the calculation logic of the service by parsing the SQL statement in the job request.

[0072] The master node 121 can perform operator scheduling based on the obtained calculation logic of the service, to obtain one or more calculation operators such as the calculation operator 221, as well as the read operator 210 and the output operator 230. Among them, as described above, the calculation logic of the service can be obtained by parsing the SQL statement of the service. Therefore, one or more calculation operators such as the calculation operator 221, as well as the read operator 210 and the output operator 230 are scheduled based on the SQL statement of the service.

[0073] Among them, the job request further includes information about the data source end. The master node 121 can configure the read operator based on the information about the data source end, so that the read operator can read service data from the data source end. Among them, the master node 121 can also configure the obtained calculation logic of the service into the read operator, so that the read operator can identify the data required by the calculation logic of the service in the service data, to obtain the data required for calculation. Specific details will be introduced in the method embodiments below and will not be elaborated here.

[0074] The job request further includes information about the data target end. The master node 121 configures the output operator based on the information about the data target end, so that the output operator can output the processing result of the service to the data target end. The master node 121 also configures the ability of the output operator to obtain the service processing result based on the calculation result of the calculation operator and the service data (or the non-calculation-required data in the service data). Specific details will be introduced in the method embodiments below and will not be elaborated here.

[0075] Refer to Figure 2 , the master node 121 can deploy the read operator 210, the output operator 230, and one or more calculation operators such as the calculation operator 221 in the calculation node cluster 120. Among them, different operators among the read operator 210, the output operator 230, and the calculation operators can be deployed to the same calculation node in the calculation node cluster 120, or can be deployed to different calculation nodes.

[0076] Among them, one or more computing operators such as the computing operator 221 deployed in the computing node cluster 120, together with the reading operator 210 and the output operator 230, constitute the streaming computing device 200. The streaming computing device 200 is also called a streaming computing engine and is used for real-time processing of service data. Specifically, the reading operator 210 reads service data from the data source end, such as a data table or one or more rows in a data table. The reading operator 210 identifies the data required for the computing logic of the service in the read service data, that is, the data required for computing. Then, the reading operator 210 inputs the data required for computing into the computing operator 221. The computing operator 221 performs state storage based on the data required for computing, thereby reducing the scale of state storage. The data required for computing can meet the computing needs of the computing operator 221, and the computing operator 221 can perform computing on the data required for computing based on the computing logic it carries to obtain a computing result. The output operator 230 can obtain the processing result of the service based on the computing result and the service data, and send the processing result to the data target end.

[0077] In some embodiments, as Figure 2 shown, a shared storage space 300 for the reading operator 210 and the output operator 230 can be configured in the computing node cluster 120, that is, both the reading operator 210 and the output operator 230 can access the storage space 300. Exemplarily, when the reading operator 210 and the output operator 230 are in the same computing node, the shared storage space 300 can be located in this computing node. Exemplarily, when the reading operator 210 and the output operator 230 are in different computing nodes, the shared storage space 300 can be located in the computing node where the reading operator 210 is located or in the computing node where the output operator 230 is located. Exemplarily, the shared storage space 300 can be located in a computing node other than the computing node where the reading operator 210 is located and the computing node where the output operator 230 is located.

[0078] In some embodiments, the reading operator 210 can store the read service data into the shared storage space 300. The output operator 230 can obtain the service data from the shared storage space 300, and then, the output operator can obtain the processing result of the service based on the service data and the computing result of the computing operator. In some embodiments, the reading operator 210 can regard the data other than the data required for computing in the service data as non-computing required data and store the non-computing required data into the shared storage space 300. The output operator 230 can obtain the non-computing required data from the shared storage space 300, and then, the output operator can obtain the processing result of the service based on the non-computing required data and the computing result of the computing operator.

[0079] In some embodiments, as Figure 3As shown, the shared storage space 300 can be implemented as a key (K)-value (V) storage system. That is to say, business data or non-computationally required data can be stored in the shared storage space 300 in the form of KV pairs. In one example, the shared storage space 300 can be implemented as an HBase database. In another example, the shared storage space 300 can be implemented as a Redis database.

[0080] In some embodiments, the computing operator can perform state storage locally, that is, the computing operator can store the data required for computing and the computing results locally at the operator. In some embodiments, the computing operator can access the shared storage space 300. The computing operator can perform state storage in the shared storage space 300, that is, the computing operator can store the data required for computing and the computing results in the shared storage space 300 to further reduce the storage overhead at the local operator.

[0081] The above examples introduce the system architecture and the streaming computing device 200 provided by the embodiments of the present application. Next, the data processing method provided by the embodiments of the present application will be described in conjunction with this system architecture and the streaming computing device 200.

[0082] This method can be executed by the streaming computing device 200, specifically by the relevant operators in the streaming computing device 200. As Figure 4 shown, this method includes the following steps.

[0083] Step 401, the reading operator 210 reads business data from the data source end of the business.

[0084] In some embodiments, the data source end can be a database, and the reading operator 210 can read data from the database to obtain business data. In some embodiments, the data source end can be a data acquisition end, that is, the data source end can perform data acquisition on the acquisition object to obtain data. For example, the data source end can be an environmental monitoring device, and the environmental monitoring device continuously monitors the environment to obtain data. The reading operator 210 can read the data recently acquired by the environmental monitoring device to obtain business data.

[0085] In some embodiments, the business data is streaming data, which is composed of multiple data points, where each data point includes at least one piece of data. The reading operator 210 reads different data points from the data source end at different reading times.

[0086] In some embodiments, there can be multiple data source ends, that is, the reading operator 210 can read data from multiple data source ends simultaneously. For example, in a data association business scenario, the reading operator 210 reads data from multiple data tables to splice the data in different data tables.

[0087] Step 402: Read the data required by the operator 210 to identify the calculation logic of the service in the service data, and obtain the data required for calculation; wherein, the data required for calculation includes the first data required by the first calculation logic.

[0088] As described above, the master node 121 can obtain the calculation logic of the service by parsing the job request of the service. For example, the job request includes an SQL statement for executing the service, and the master node 121 parses the SQL statement to obtain the calculation logic of the service. The master node 121 can configure the calculation logic of the service into the read operator 210. Thus, the read operator 210 can, in step 402, identify the data required by the calculation logic of the service based on the calculation logic of the service. For example, for the calculation logic of determining whether the f1 field in table 1 is equal to the f2 field in table 2, the required data is the f1 field in table 1 and the f2 field in table 2. For another example, for the calculation logic of determining whether the f1 field in table 3 is equal to "B10011", the required data is the f1 field in table 3.

[0089] In addition, the first calculation logic belongs to the calculation logic of the service, and specifically may be the calculation logic carried by the calculation operator 221. Therefore, the data required for calculation obtained in step 402 includes the data required by the first calculation logic. For the convenience of description, the data required by the calculation logic (first calculation logic) carried by the calculation operator 221 is referred to as the first data.

[0090] In some embodiments, the service executed by the streaming computing device 200 includes splicing the data to be spliced based on the splicing key of the data to be spliced, and the calculation logic carried by the calculation operator 221 is to determine whether the splicing keys of the data to be spliced are the same. It can be set that the service data includes the data to be spliced C1, the splicing key C11 corresponding to the data to be spliced C1, the data to be spliced C2, and the splicing key C21 corresponding to the data to be spliced C2. Then the calculation logic carried by the calculation operator 221 includes: whether the splicing key C11 is the same as the splicing key C21. In step 402, based on the calculation logic carried by the calculation operator 221, it can be identified that the splicing key C11 and the splicing key C21 are the first data.

[0091] For example, it can be set that the SQL statement of the service is: "Select table1.f1, table1.f2, table1.f3,.., table2.f2, table2.f3,... from table1 join table2 on table1.f1 = table2.f2".

[0092] In this SQL statement, the f1 field in table1 and the f2 field in table2 are the concatenation keys and belong to the data required for calculation. The f2 field, f3 field, etc. in table1 and the f3 field, etc. in table2 can be referred to as the data to be concatenated. Among them, the f1 field in table1 can be used as the concatenation key C11, the f2 field in table2 can be used as the concatenation key C21, the f2 field, f3 field, etc. in table1 can be used as the data to be concatenated C1, and the f3 field, etc. in table2 can be used as the data to be concatenated C2.

[0093] In some embodiments, the operations performed by the streaming computing device 200 include outputting data that meets a preset filtering condition to a data target end. Among them, the service data includes data to be filtered and the corresponding filtering keys, and the computing logic carried by the computing operator 221 is to determine whether the filtering key meets the preset filtering condition. In step 402, based on the computing logic carried by the computing operator 221, it can be identified that the filtering key is the first data.

[0094] For example, the SQL statement of the service can be set as: "Select table3.f1, table3.f2, table3.f3,..from table3 where table3.f1 = B10011".

[0095] In this SQL statement, the f1 field in table3 is the filtering key, B10011 is the preset filtering condition, and the f2 field, f3 field, etc. in table3 are the data to be filtered.

[0096] In some embodiments, as described above, the service data is streaming data, which is composed of multiple data points. The reading operator 210 reads different data points from the data source end at different reading times. In step 402, whenever the reading operator 210 reads a data point, the reading operator 210 can determine whether the data point belongs to the data required for calculation or non-calculation required data based on the computing logic carried by the computing operator 221. It can be set that at time T1, the data point T11 read by the reading operator 210 from the data source end. The reading operator 210, based on the computing logic carried by the computing operator 221, identifies that the data point T11 belongs to the data required for calculation, and then uses the data point T11 as the first data. It can also be set that at time T2 after time T1, the data point T21 read by the reading operator 210 from the data source end. The reading operator 210, based on the computing logic carried by the computing operator 221, identifies that the data point T21 is not the data required by the computing logic carried by the computing operator 221.

[0097] Step 403: The reading operator 210 sends the data required for calculation to the calculation operator 221, enabling the calculation operator 221 to calculate the first data based on the first calculation logic to obtain a calculation result. Among them, the reading operator 210 sending the data required for calculation to the calculation operator 221 also enables state storage based on the data required for calculation by the calculation operator 221.

[0098] Among them, the calculation operator 221 can also be referred to as the first calculation operator. The reading operator 210 can send the data required for calculation obtained in step 402 to the calculation operator 221. When the calculation operator 221 receives the data required for calculation, it can obtain the first data therein and calculate the first data based on the calculation logic carried by the calculation operator 221 (i.e., the first calculation logic) to obtain a calculation result.

[0099] In some embodiments, as described above, the calculation logic carried by the calculation operator 221 is to determine whether the splicing keys of the data to be spliced are the same, and the first data includes the splicing key C11 and the splicing key C21. In step 403, the calculation operator 221 can determine whether the splicing key C11 and the splicing key C21 are the same to obtain a calculation result.

[0100] In some embodiments, as described above, the calculation logic carried by the calculation operator 221 is to determine whether the filtering key meets a preset filtering condition, and the first data includes the splicing key. Among them, the calculation operator 221 has preset the filtering condition. In step 403, the calculation operator 221 can determine whether the splicing key meets the filtering condition to obtain a calculation result.

[0101] In some embodiments, as described above, the first data can be the data point T11. In step 403, the calculation operator 221 can calculate the data point T11 based on the first calculation logic to obtain a calculation result.

[0102] The calculation operator 221 stores the received data required for calculation and the calculation result of the calculation operator 221 as a state for state storage. In one example, the calculation operator 221 can store the state in local memory. In one example, the calculation operator 221 can store the state in the shared storage space 300.

[0103] The data required for calculation received by the calculation operator 221 from the reading operator 210 is part of the service data read by the reading operator 210 in step 401. That is, compared with the service data read by the reading operator 210 in step 401, the amount of data required for calculation is small. The calculation operator 221 performing state storage based on the data required for calculation can effectively reduce the storage overhead of state storage. In particular, state storage is usually stored in memory. Therefore, reducing the storage overhead of state storage can effectively save memory space and improve the calculation performance of the calculation operator.

[0104] Step 404: The output operator 230 outputs the processing result of the service to the data target end of the service based on the calculation result and the service data.

[0105] After the calculation operator 221 completes the calculation of the first data, the obtained calculation result can be sent to the output operator 230, so that the output operator 230 obtains the calculation result.

[0106] In some embodiments, the output operator 230 can obtain the service data acquired by the reading operator 210 in step 401. In one example, the reading operator 210 can store the service data acquired in step 401 in the shared storage space 300. In step 404, the output operator 230 can obtain the service data from the shared storage space 300.

[0107] In some embodiments, the output operator 230 can obtain the data that is not required for calculation in the service data. The data that is not required for calculation refers to the data in the service data acquired by the reading operator 210 in step 401 other than the data required for calculation. In one example, the reading operator 210 can store the data that is not required for calculation in the shared storage space 300. In step 404, the output operator 230 can obtain the data that is not required for calculation from the shared storage space 300.

[0108] The output operator 230 can obtain the processing result of the service based on the calculation result and the data that is not required for calculation.

[0109] In some embodiments, as described above, the service executed by the streaming computing device 200 includes calculating based on the splicing key of the data to be spliced, and the calculation logic carried by the calculation operator 221 is to determine whether the splicing keys of the data to be spliced are the same. In step 404, if the calculation result indicates that the splicing key C11 and the splicing key C21 are the same, the output operator 230 splices the data to be spliced C1 and the data to be spliced C2 obtained from the service data or the data that is not required for calculation, and then, the output operator 230 obtains the splicing result of the data to be spliced C1 and the data to be spliced C2. This splicing result is used as the processing result of the service.

[0110] In some embodiments, the service executed by the streaming computing device 200 includes outputting the data that meets the preset filtering condition to the data target end, the calculation logic carried by the calculation operator 221 is to determine whether the filtering key meets the preset filtering condition, and the second data includes the data to be filtered. In step 404, if the calculation result indicates that the filtering key meets the filtering condition, the output operator 230 obtains the data to be filtered from the service data or the data that is not required for calculation, and uses the data to be filtered as the processing result of the service.

[0111] In some embodiments, as described above, the calculation result is the result of the calculation operator 221 calculating the data point T11, and the second data is the data point T21. The shared storage space 300 stores the service data obtained in step 401 or the non-calculation required data obtained in step 402. Both the service data and the non-calculation required data include the data point T21. Thus, the output operator 230 can obtain the data point T21 from the shared storage space 300. In step 404, the output operator 230 can obtain the processing result of the service based on the calculation result and the data point T21.

[0112] When the output operator 230 obtains the processing result of the service, the output operator 230 can output the processing result to the data target end.

[0113] In some embodiments, referring to Figure 2 , the streaming computing device 200 further includes a calculation operator 222. Among them, the calculation operator 222 can also be called the second calculation operator, which is used to carry the second calculation logic of the service. That is, the calculation operator 222 is used to perform data calculation according to the second calculation logic. As Figure 2 shown, in the data flow direction of the service data, the calculation operator 222 is located after the calculation operator 221, that is, the outflow data of the calculation operator 221 is the inflow data of the calculation operator 222. The calculation required data obtained in step 402 includes the data required by the second calculation logic. For the convenience of description, the data required by the second calculation logic can be called the second data.

[0114] In an example of this embodiment, after receiving the calculation required data, the calculation operator 221 can identify the second data in the calculation required data and send the second data to the calculation operator 222, so that the calculation operator 222 stores the state based on the second data, and so that the calculation operator 222 calculates the second data according to the second calculation logic to obtain a calculation result. Among them, the calculation result is used for the output operator 230 to obtain the processing result of the service and send the processing result of the service to the data target end. In an example of this example, the calculation operator 221 can identify the data required by the calculation logic of the calculation operator 222 and send the data required by the calculation logic of the calculation operator 222 to the current calculation operator. In another example of this example, the calculation operator 221 can remove the first data from the calculation required data received by the calculation operator 221 to obtain data including the second data but not including the first data. Then, the calculation operator 221 sends the data including the second data but not including the first data to the calculation operator 222.

[0115] This example can further reduce the amount of data stored in the state of the calculation operator.

[0116] In another example of this embodiment, after receiving the data required for calculation, the calculation operator 221 may send the data required for calculation to the calculation operator 222, so that the calculation operator 222 stores the state based on the data required for calculation, and so that the calculation operator 222 calculates the second data based on the second calculation logic to obtain a calculation result. The calculation result is used for the output operator 230 to obtain the processing result of the service and send the processing result of the service to the data target end.

[0117] In summary, in the data processing method provided in the embodiment of the present application, the reading operator can identify the data required for calculation in the service data, and then send the data required for calculation to the calculation operator, so that the calculation operator can store the state and perform calculations based on the data required for calculation, thereby meeting the calculation requirements of the calculation operator while reducing the amount of data stored in the state of the calculation operator, saving the memory of the operator, and improving the calculation performance of the operator.

[0118] Refer to Figure 5 , the embodiment of the present application also provides a streaming computing device 500. As Figure 5 shown, the streaming computing device 500 includes a reading operator 510, a first calculation operator 520, and an output operator 530. The first calculation operator 520 is used to perform data calculation based on the first calculation logic of the service. The reading operator 510 is used to read service data from the data source end of the service. The reading operator 510 is used to identify the data required by the calculation logic of the service in the service data to obtain the data required for calculation. The data required for calculation includes the first data required by the first calculation logic. The reading operator 510 is used to send the data required for calculation to the first calculation operator 520, so that the first calculation operator 520 calculates the first data based on the first calculation logic to obtain a calculation result. The output operator 530 is used to output the processing result of the service to the data target end of the service based on the calculation result and the service data.

[0119] In some embodiments, the streaming computing device 500 corresponds to a storage space accessible to the reading operator 510 and the output operator 520. The reading operator 510 is further used to store the data not required for calculation into the storage space. The data not required for calculation is the data in the service data other than the data required for calculation. The output operator 530 is further used to read the data not required for calculation from the storage space. The output operator 530 is further used to output the processing result to the data target end based on the data not required for calculation and the calculation result.

[0120] In some embodiments, the streaming computing device corresponds to a storage space accessible to the reading operator 510 and the output operator 530; wherein, the reading operator 510 is further configured to store the service data into the storage space; the output operator 530 is further configured to read the service data from the storage space.

[0121] In some embodiments, the streaming computing device 500 includes a second computing operator, which is configured to perform data calculation according to the second calculation logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator 520; wherein, the data required for the calculation further includes second data required by the second calculation logic; the first computing operator 520 is configured to send the second data to the second computing operator, so that the second computing operator calculates the second data based on the second calculation logic.

[0122] In some embodiments, the service data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; the first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; the reading operator 510 is configured to identify the first splicing key and the second splicing as the first data based on the first calculation logic; when the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator 530 is configured to splice the first data to be spliced and the second data to be spliced to obtain the processing result.

[0123] In some embodiments, the service data includes data to be filtered and a filtering key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filtering key meets a preset filtering condition; the reading operator 510 is configured to identify the filtering key as the first data based on the first calculation logic; when the calculation result indicates that the filtering key meets the preset filtering condition, the output operator 530 is configured to use the data to be filtered as the processing result.

[0124] In some embodiments, the service data is streaming data composed of multiple data points, and each data point in the multiple data points includes at least one piece of data; wherein, the reading operator 510 reads different data points in the multiple data points at different times; the reading operator 510 is configured to read the first data point in the multiple data points at a first time; the reading operator 510 is configured to identify the first data point as the first data based on the first calculation logic.

[0125] In an example of this embodiment, the reading operator 510 is used to read the second data point among the multiple data points at the second moment; when the second data point is data other than the data required by the first computing logic, the output operator 530 is used to output the processing result to the data target end of the service based on the computing result and the second data point.

[0126] In some embodiments, the computing logic of the service is obtained based on the structured query language SQL statement of the service.

[0127] In some embodiments, the first computing operator 520 is arranged based on the SQL statement of the service.

[0128] Among them, the reading operator 510, the first computing operator 520, and the output operator 530 can all be implemented by software or by hardware. Exemplarily, next, taking the reading operator 510 as an example, the implementation manner of the reading operator 510 will be introduced. Similarly, the implementation manners of the first computing operator 520 and the output operator 530 can refer to the implementation manner of the reading operator 510.

[0129] As an example of a software functional unit, the reading operator 510 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the reading operator 510 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone AZ or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Among them, generally, one region may include multiple AZs.

[0130] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same VPC or in multiple VPCs. Among them, generally, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0131] As an example of a hardware functional unit, the read operator 510 may include at least one computing device, such as a server. Alternatively, the read operator 510 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0132] The multiple computing devices included in the read operator 510 may be distributed in the same region or in different regions. The multiple computing devices included in the read operator 510 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the read operator 510 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0133] This application also provides a computing device 600. As Figure 6 shown, the computing device 600 includes: a bus 602, a processor 604, a memory 606, and a communication interface 608. The processor 604, the memory 606, and the communication interface 608 communicate with each other through the bus 602. The computing device 600 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 600.

[0134] The bus 602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 602 may include a path for transmitting information between various components of the computing device 600 (for example, the memory 606, the processor 604, and the communication interface 608).

[0135] The processor 604 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0136] The memory 606 may include volatile memory, such as random access memory (RAM). The memory 606 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0137] The memory 606 stores executable program code, and the processor 604 executes the executable program code to implement the functions of the foregoing reading operator 510, first calculation operator 520, and output operator 530 respectively, so as to implement Figure 4 the method shown. That is, the memory 606 stores instructions for executing Figure 4 the method shown.

[0138] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or communication networks.

[0139] The embodiments of the present application further provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0140] As Figure 7 shown, the computing device cluster includes at least one computing device 600. The memory 606 in one or more computing devices 600 in the computing device cluster may store the same instructions for executing Figure 4 the method shown.

[0141] In some possible implementation manners, the memory 606 in one or more computing devices 600 in the computing device cluster may also store instructions for executing Figure 4Portions of the instructions of the method shown. In other words, a combination of one or more computing devices 600 can jointly execute for performing Figure 4 the instructions of the method shown.

[0142] It should be noted that the memories 606 in different computing devices 600 in the computing device cluster can store different instructions, respectively for performing partial functions of the streaming computing device 500. That is to say, the instructions stored in the memories 606 of different computing devices 600 can implement the functions of one or more modules among the reading operator 510, the first computing operator 520, and the output operator 530.

[0143] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 8 Illustrates a possible implementation manner. As Figure 8 shown, two computing devices 600A and 600B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, the memory 606 in the computing device 600A stores instructions for performing the function of the reading operator 510. At the same time, the memory 606 in the computing device 600B stores instructions for performing the functions of the first computing operator 520 and the output operator 530.

[0144] It should be understood that Figure 8 the functions shown by the computing device 600A in

[0145] can also be completed by multiple computing devices 600. Similarly, the functions of the computing device 600B can also be completed by multiple computing devices 600. Figure 7 and Figure 8 The connection manner of the computing device cluster. The difference is that the memories 606 in one or more computing devices 600 in this computing device cluster can store the same instructions for performing Figure 4 the method shown.

[0146] In some possible implementation manners, the memories 606 in one or more computing devices 600 in this computing device cluster can also respectively store partial instructions for performing Figure 4 the method shown. In other words, a combination of one or more computing devices 600 can jointly execute for performing Figure 4 the instructions of the method shown.

[0147] The embodiments of the present application also provide a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute Figure 4 the method shown

[0148] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a host migration device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute Figure 4 the method shown

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method is applied to a stream computing device, the stream computing device includes a reading operator, a first computing operator and an output operator, wherein the first computing operator is used to perform data computing based on a first computing logic of a business; the method includes: The reading operator reads the business data from the data source of the business; The read operator identifies data required by the computing logic of the business in the business data, and obtains the data required for the computing; wherein the data required for the computing includes the first data required by the first computing logic; The read operator sends the data required for the calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data.

2. The method according to claim 1, characterized in that The stream computing device corresponds to a storage space accessible by the read operator and the output operator; The method further comprises: The read operator stores the data not required for calculation into the storage space, where the data not required for calculation is the data in the business data other than the data required for calculation; The output operator reads the data not required for calculation from the storage space; The output operator outputs the processing result of the business to the data target end based on the calculation result and the business data, including: the output operator outputs the processing result to the data target end based on the non-calculation required data and the calculation result.

3. The method according to claim 1, characterized in that The stream computing device corresponds to a storage space accessible by the read operator and the output operator; The method further comprises: The read operator stores the business data in the storage space; The output operator reads the business data from the storage space.

4. The method according to any one of claims 1 to 3, characterized in that The stream computing device includes a second computing operator, which is used to perform data computing according to a second computing logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator; wherein, The data required for calculation also includes second data required by the second calculation logic; the first calculation operator is used to send the second data to the second calculation operator, so that the second calculation operator calculates the second data based on the second calculation logic.

5. The method according to any one of claims 1 to 4, characterized in that The business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; The first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for calculation, including: the read operator identifies the first splicing key and the second splicing as the first data based on the first calculation logic; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: When the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator splices the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

6. The method according to any one of claims 1 to 4, characterized in that The business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key satisfies a preset filter condition; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for calculation, including: the read operator identifies the filter key as the first data based on the first calculation logic; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the calculation result indicates that the filter key meets the preset filtering condition, the output operator uses the data to be filtered as the processing result.

7. The method according to any one of claims 1 to 6, characterized in that The business data is streaming data composed of multiple data points, each of which includes at least one data; wherein the reading operator reads different data points from the data source end at different times; The reading operator reads the business data from the data source end of the business, including: the reading operator reads a first data point among the multiple data points at a first moment; The read operator identifies data required by the calculation logic of the business in the business data to obtain the data required for the calculation, including: the read operator identifies the first data point as the first data based on the first calculation logic.

8. The method according to claim 7, characterized in that The reading operator reads the business data from the data source end of the business, including: the reading operator reads a second data point among the multiple data points at a second time; The output operator outputs the processing result of the business to the data target end of the business based on the calculation result and the business data, including: when the second data point is data other than the data required by the first calculation logic, the output operator outputs the processing result to the data target end of the business based on the calculation result and the second data point.

9. A streaming computing device, characterized in that: The device includes a reading operator, a first calculation operator and an output operator, wherein the first calculation operator is used to perform data calculation based on a first calculation logic of the business; wherein, The read operator is used to read business data from the data source end of the business; The read operator is used to identify data required by the computing logic of the business in the business data, and obtain the data required for the computing; wherein the data required for the computing includes the first data required by the first computing logic; The read operator is used to send the data required for the calculation to the first calculation operator, so that the first calculation operator calculates the first data based on the first calculation logic to obtain a calculation result; The output operator is used to output the processing result of the business to the data target end of the business based on the calculation result and the business data.

10. The device according to claim 9, characterized in that The device corresponds to a storage space accessed by the read operator and the output operator; wherein, The read operator is further used to store data not required for calculation into the storage space, where the data not required for calculation is data in the business data other than the data required for calculation; The output operator is also used to read the non-computation required data from the storage space; The output operator is further used to output the processing result to the data target end based on the non-computation required data and the calculation result.

11. The device according to claim 9, characterized in that The device corresponds to a storage space accessed by the read operator and the output operator; wherein, The read operator is also used to store the business data in the storage space; The output operator is also used to read the business data from the storage space.

12. The device according to any one of claims 9 to 11, characterized in that The device includes a second computing operator, which is used to perform data calculation according to a second computing logic of the service; in the data flow direction of the service, the second computing operator is located after the first computing operator; wherein, The data required for calculation also includes second data required by the second calculation logic; the first calculation operator is used to send the second data to the second calculation operator, so that the second calculation operator calculates the second data based on the second calculation logic.

13. The device according to any one of claims 9 to 12, characterized in that The business data includes first data to be spliced, a first splicing key corresponding to the first data to be spliced, second data to be spliced, and a second splicing key corresponding to the second data to be spliced; The first calculation logic includes: determining whether the first splicing key and the second splicing key are the same; The read operator is used to identify the first splicing key and the second splicing as the first data based on the first calculation logic; When the calculation result indicates that the first splicing key and the second splicing key are the same, the output operator is used to splice the first data to be spliced ​​and the second data to be spliced ​​to obtain the processing result.

14. The device according to any one of claims 8 to 12, characterized in that The business data includes data to be filtered and a filter key corresponding to the data to be filtered; the first calculation logic includes: determining whether the filter key satisfies a preset filter condition; The read operator is used to identify the filter key as the first data based on the first calculation logic; When the calculation result indicates that the filter key satisfies the preset filter condition, the output operator is used to use the data to be filtered as the processing result.

15. The device according to any one of claims 8 to 14, characterized in that The business data is streaming data composed of multiple data points, each of which includes at least one data; wherein the reading operator reads different data points from the data source end at different times; The read operator is used to read a first data point among the multiple data points at a first moment; The read operator is used to identify the first data point as the first data based on the first calculation logic.

16. The device according to claim 15, characterized in that The read operator is used to read a second data point among the plurality of data points at a second time; When the second data point is data other than data required by the first calculation logic, the output operator is used to output the processing result to the data target end of the business based on the calculation result and the second data point.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.

19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device cluster, the computer device cluster is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Data processing method and apparatus, and cluster

    EP4797085A1

  • Data processing method and apparatus, and cluster

    WO2025102626A1