Data splicing method and apparatus, and cluster
By introducing storage space into the streaming computing device, the splicing operator can query and store data of the corresponding splicing keys, solving the problem of large amount of state storage data when splicing multiple data sources in the prior art, and achieving efficient data splicing.
Patent Information
- Application Number
- PCT/CN2024/091414
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-05-07
- Publication Date
- 2025-05-22
AI Technical Summary
In streaming calculation, when there are three or more data source terminals, the prior art requires two-dimensional splicing operations to be categorized by multiple splicing operators, resulting in large amounts of data storage in the state of splicing operators and low efficiency.
By introducing storage space in the streaming computing device, the splicing operator can query and store data of the corresponding splicing keys. When new data of the same splicing key is received, splicing and storing the results, avoiding the need for multiple splicing operators to cascade splicing.
The amount of data stored in the splicing operator is reduced, the efficiency of data splicing is improved, and the need for real-time data splicing is met.
Smart Images

Figure CN2024091414_22052025_PF_FP_ABST
Abstract
Description
Data splicing method, device and cluster
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 17, 2023, with application number 202311541760.9 and application name “A data splicing method”, and the Chinese patent application filed with the State Intellectual Property Office of China on January 16, 2024, with application number 202410063281.9 and application name “A data splicing method, device and cluster”, all of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a data splicing method, device, and cluster. Background Art
[0003] Real-time data collection and analysis has been applied in a variety of fields. For example, in logistics, real-time data collection and analysis allows companies to understand the location and transportation status of goods, helping to improve logistics efficiency. Because there may be multiple data sources, real-time data collection and analysis requires combining data from different sources.
[0004] Real-time data collection and analysis are carried out through streaming computing. The current streaming computing solution mainly uses a pairwise splicing method to splice data from different data sources. Among them, when there are three or more data sources, it is necessary to use a pairwise cascade splicing method for splicing. Specifically, the data from two of the data sources are spliced together by the first splicing operator to obtain a splicing result; then, the first splicing operator sends the splicing result to the second splicing operator, and the second splicing operator splices the splicing result with the data from the third data source, and so on. In this solution, when there are three or more data sources, splicing is required through multiple splicing operators, which requires multiple splicing operators to store the state of the received data, and the overall state storage data volume is large.
[0005] Summary of the Invention
[0006] The present application provides a data splicing method, device, and cluster, which can reduce the amount of data stored in the state of a splicing operator in a streaming computing device.
[0007] In a first aspect, a data splicing method is provided. The method is applied to a streaming computing device, the device comprising a splicing operator and a reading operator for reading data from multiple data sources. The method comprises: the reading operator reading first data from a first data source among the multiple data sources and sending it to the splicing operator, wherein the first data corresponds to a first splicing key; the splicing operator searching a storage space for second data corresponding to the first splicing key; when the second data is read by the reading operator from a data source other than the first data source among the multiple data sources, the splicing operator splicing the first data and the second data to obtain third data corresponding to the first splicing key; and the splicing operator storing the third data in the storage space. The third data is the splicing result of the first and second data.
[0008] In this method, the splicing operator can store the data corresponding to a certain splicing key in the storage space. When the splicing operator receives other data corresponding to the splicing key from the reading operator, the splicing operator can obtain the data corresponding to the splicing key from the storage space, and splice the data obtained from the storage space and the data received from the reading operator to obtain the splicing result. The splicing operator can store the splicing result in the storage space for splicing the subsequent received data. In this way, the data corresponding to the same splicing key in multiple (for example, three or more) data source ends can be spliced together by the same splicing operator. Accordingly, the same splicing operator performs state storage on the data in multiple data source ends. Compared with the two-by-two cascade splicing method, this can not only reduce the amount of data stored in the state of the splicing operator in the streaming computing device, but also improve the efficiency of data splicing, and meet the needs of related businesses for real-time data splicing.
[0009] In a possible implementation, the second data is a concatenation result of at least two data, each of the at least two data corresponds to a first concatenation key, and different data in the at least two data are read by a reading operator from different data sources.
[0010] In other words, the concatenation operator can store the concatenation result corresponding to a certain concatenation key in the storage space. When the concatenation operator receives data corresponding to the same concatenation key from the read operator again, the concatenation operator can obtain the concatenation result corresponding to the same concatenation key from the storage space and concatenate the concatenation result with the data to obtain a re-concatenated concatenation result. The concatenation operator can store the re-concatenated concatenation result in the storage space for concatenation with subsequent received data. In this way, the same concatenation operator can concatenate data corresponding to the same concatenation key from multiple data sources.
[0011] In one possible implementation, the device further includes an output operator connected to the data target end; the method further includes: the splicing operator sends the third data to the output operator, so that the output operator outputs the third data to the data target end.
[0012] The third data is the concatenation result of the second data and the first data. In this implementation, the concatenation operator can send the concatenation result to the data target end to meet the data target end's demand for streaming data.
[0013] In one possible implementation, the method further includes: when part or all of the first data and the second data are read by the read operator from the same data source, the splicing operator updates the second data based on the first data; and the splicing operator stores the updated second data in a storage space. The second data is the splicing result of the data corresponding to the first splicing key that was previously read by the read operator.
[0014] When the data corresponding to the first splicing key in one or more data sources is updated, the updated data can be updated to the splicing result corresponding to the first splicing key through this implementation.
[0015] In one possible implementation, the method further includes: the reading operator reads fourth data from the second data source end among the multiple data source ends and sends it to the splicing operator, wherein the fourth data corresponds to the second splicing key; when the splicing operator does not find data corresponding to the second splicing key in the storage space, the splicing operator stores the fourth data in the storage space.
[0016] In this implementation, if the reading operator has not read the data corresponding to the second splicing key in history, the splicing operator stores the currently received data corresponding to the second splicing key in the storage space, so that when the data corresponding to the second splicing key is subsequently received, the data corresponding to the second splicing key and the subsequently received data corresponding to the second splicing key are spliced.
[0017] In one possible implementation, the device includes multiple splicing operators, where different splicing operators correspond to different mapping ranges; the reading operator reads the first data from the first data source end among the multiple data source ends and sends it to the splicing operator, including: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, the reading operator sends the first data to the splicing operator.
[0018] The reading operator reads the mapping value corresponding to the splicing key of the data read from the data source, and sends the data to the splicing key of the corresponding mapping range based on the mapping value corresponding to the splicing key of the data, so that the splicing operator splices the data of the splicing key corresponding to the mapping value within the mapping range, so that the data of the splicing keys whose mapping values belong to different mapping ranges can be spliced by different splicing operators respectively, which reduces the load of the splicing operator, improves the concurrency of data splicing, and improves the efficiency of data splicing.
[0019] In a possible implementation, the data in the storage space has a life cycle; the method further includes: the splicing operator deleting the data whose life cycle has ended in the storage space.
[0020] Deleting data at the end of its lifecycle can free up the space occupied by this data, thereby ensuring that the storage space can be continuously used to store the splicing results of the splicing operator.
[0021] In a second aspect, a streaming computing device is provided, which includes a splicing operator and a reading operator for reading data from multiple data source terminals; wherein the reading operator is used to: read first data from a first data source terminal among the multiple data source terminals and send it to the splicing operator, wherein the first data corresponds to a first splicing key; the splicing operator is used to: query second data corresponding to the first splicing key in a storage space; the splicing operator is also used to: when the second data is read by the reading operator from a data source terminal other than the first data source terminal among the multiple data source terminals, splice the first data and the second data to obtain third data corresponding to the first splicing key; the splicing operator is also used to: store the third data in the storage space.
[0022] In a possible implementation, the second data is a concatenation result of at least two data, each of the at least two data corresponds to a first concatenation key, and different data in the at least two data are read by a reading operator from different data sources.
[0023] In a possible implementation, the device further includes an output operator connected to the data target end; wherein the splicing operator is further configured to: send third data to the output operator, so that the output operator outputs the third data to the data target end.
[0024] In one possible implementation, the splicing operator is also used to: when part or all of the first data and the second data are read by the reading operator from the same data source, update the second data based on the first data; and store the updated second data in the storage space.
[0025] In one possible implementation, the read operator is further used to: read fourth data from a second data source end among multiple data source ends and send it to the splicing operator, wherein the fourth data corresponds to the second splicing key; the splicing operator is further used to: when the splicing operator does not find data corresponding to the second splicing key in the storage space, store the fourth data in the storage space.
[0026] In one possible implementation, the device includes multiple splicing operators, where different splicing operators correspond to different mapping ranges; the reading operator is used to: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, send the first data to the splicing operator.
[0027] In a possible implementation, data in the storage space has a life cycle; the splicing operator is further used to delete data whose life cycle has ended in the storage space.
[0028] In a third aspect, a computing device cluster is provided, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method provided in the first aspect.
[0029] In a fourth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method provided in the first aspect.
[0030] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computer device cluster, the computer device cluster executes the method provided in the first aspect.
[0031] The beneficial effects of the second to fifth aspects can be referred to the above introduction to the beneficial effects of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a schematic diagram of a splicing key;
[0033] FIG2 is a schematic diagram of a data splicing solution;
[0034] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0035] FIG4 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0036] FIG5 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0037] FIG6 is a flow chart of a data splicing method provided in an embodiment of the present application;
[0038] FIG7 is a schematic structural diagram of a stream computing device provided in an embodiment of the present application;
[0039] FIG8 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0040] FIG9 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0041] FIG10 is a schematic diagram of a structure of a computing device cluster connected via a network provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following describes the solutions provided by the embodiments of the present application in conjunction with the accompanying drawings. In the embodiments of the present application, "plurality" refers to two or more. "First," "second," and the like are merely used to distinguish similar objects and are not necessarily used to describe a specific order or number of objects.
[0043] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application are first introduced.
[0044] Data table: It can be simply called a table and is used to record data. Data tables can be divided into horizontal tables and vertical tables.
[0045] A horizontal table is a type of data table. In a horizontal table, all data corresponding to the same identifier (ID) is recorded in the same row, or in other words, all values for the same key are recorded in the same row. In contrast to a horizontal table, a vertical table records only one piece of data per row. That is, one value for a key occupies one row, and multiple values occupy multiple rows.
[0046] Wide table: It is a horizontal table in the database field that records a large amount of data.
[0047] Splicing key: also known as association key, used to splice data from different data sources. For example, as shown in Figure 1, table t1 can be set to record the student's name, grade, and class, and table t2 can record the student's name, math score, physics score, etc. Table t1 and table t2 are different data sources. The student's student ID number can be used as the splicing key to splice the data in table t1 and table t2 into table t3. In table t3, the grade, class, math score, physics score, etc. corresponding to the same name are associated. For example, as shown in Figure 1, table t3 is a horizontal table, and the grade, class, math score, and physics score corresponding to the same name are recorded in the same row. In addition, name, grade, class, math score, and physics score are different fields, each corresponding to a column in the table. A student's name, grade, class, math score, and physics score are the field values of the corresponding fields.
[0048] A dataset is a collection of one or more data points. Data in the same dataset share the same join key. Data in the same dataset are linked together in a table. In horizontal tables (e.g., wide tables), data in the same dataset are recorded in the same row.
[0049] Splicing refers to combining data from different data sources that share the same splicing key into a single dataset. For example, as shown in Figure 1, in a horizontal table, splicing involves recording data from different data sources that share the same splicing key into the same row. A row in a table can be considered a dataset.
[0050] Stream computing: Also known as stream computing, stream computing is used to process data in real time. Stream computing can be a computing model triggered by input data, where each new input data acts as an event, triggering the computing model to perform calculations.
[0051] Operators are data processing units that carry computational logic. They are the smallest executable unit in stream computing. Operators are used to perform calculations on related data based on the computational logic they carry. Common operators in stream computing devices include read operators, compute operators, and output operators.
[0052] Read Operator: Also known as the source operator, it is an operator in a stream computing device that reads data from a data source. The read operator sends the read data to the splicing operator.
[0053] Splicing operator: It is a computing operator in stream computing, specifically an operator used to perform data splicing tasks.
[0054] Output operator: also known as sink operator, is used to send the calculation results of the calculation operator (such as the splicing results of the splicing operator) to the data target end.
[0055] State storage: This refers to the storage of the data received by the splicing operator or the splicing results of the splicing operator as state in stream computing. Typically, operators store state in local memory. This means that operator state storage consumes the operator's memory.
[0056] In one solution, StreamCompute uses join statements (e.g., full outer joins) in Structured Query Language (SQL) to join data from different data sources. Therefore, this solution is a pairwise join solution. When data from three or more data sources needs to be joined, a pairwise cascade join is required, which increases the amount of data stored in the overall state of the join operator.
[0057] Take the following statement as an example.
[0058] "Select C.ID,C.B1,C.B2,C.A1,C.A2,D.D1,D.D2
[0059] From
[0060] (Select A.ID,B.B1,B.B2,A.A1,A.A2.A.TIMESTAMP
[0061] FROM A
[0062] FULL OUTER JOIN B
[0063] ON A.ID=B.ID)C
[0064] FULL OUTER JOIN D
[0065] ON C.ID=D.ID”
[0066] In this statement, the data sources are Tables A, B, and D. This statement concatenates fields A1, A2, B1, and B2 from Table C with fields D1 and D2 from Table D, using ID as the concatenation key. Table C is created by concatenating fields A1 and A2 from Table A with fields B1 and B2 from Table B, using ID as the concatenation key. This statement requires two join operators, as shown below.
[0067] As shown in Figure 2, the two splicing operators that execute the join statement can be set as splicing operator 210 and splicing operator 220. When splicing operator 210 receives data from table A, splicing operator 210 records the data as a state in table A' in storage space 211. When splicing operator 210 receives data from table B, splicing operator 210 records the data as a state in table B' in storage space 211. That is, table A' records the data in table A received by splicing operator 210, and table B' records the data in table B received by splicing operator 210. When performing splicing, splicing operator 210 obtains table A' and table B' from storage space 211. Then, based on table A' and table B', data splicing is performed to obtain table C. Splicing operator 210 sends table C to splicing operator 220. When the concatenation operator 220 receives table C, it stores table C as a state in the storage space 221. When the concatenation operator 220 receives data in table D, it records the data as a state in table D' in the storage space 221.
[0068] As can be seen from the above description, the state of concatenation operator 210 stores the data in Table A and Table B. Table C is obtained by concatenating the data in Table A and Table B received by concatenation operator 210. The state of concatenation operator 220 stores the data in Table C and Table D. The total amount of data stored in the state is the sum of twice the amount of data in Table A, twice the amount of data in Table B, and the amount of data in Table D.
[0069] Moreover, in this solution, the more data sources there are, the more splicing operators are required, resulting in lower overall operating efficiency.
[0070] In addition, the data received by the splicing operator 210 from different data sources are stored in different tables. When splicing is performed, the splicing is performed based on different tables, which requires a large amount of computation and results in a large splicing delay.
[0071] In another solution, data arriving simultaneously is spliced in real time, while data arriving in different time periods is spliced periodically. This results in delayed splicing of data arriving in different time periods with the same splicing key, making it difficult to meet the requirements of real-time data analysis.
[0072] Embodiments of the present application provide a data splicing method applicable to a streaming computing device. This method can splice all data with the same splicing key using the same splicing operator. In other words, regardless of whether there are two, three, or more data sources, the same splicing operator can be used to splice data from these data sources. Therefore, when there are three or more data sources, there is no need to perform cascade splicing, thereby avoiding amplifying the amount of data stored in the state.
[0073] Next, the data splicing method provided in the embodiment of the present application is described.
[0074] Figure 3 shows a system architecture that can be used to implement the data processing method provided in the embodiment of the present application. As shown in Figure 3, the system architecture includes a stream computing device 300, multiple data source terminals and a data target terminal 500.
[0075] The multiple data sources may include data source 410, data source 420, data source 430, and data source 440. In some embodiments, the data source may be a database or a data table. In some embodiments, the data source may be a data acquisition terminal, i.e., the data source may acquire data from an object to be acquired. For example, the data source may be an environmental monitoring device that continuously monitors the environment and generates data.
[0076] The stream computing device 300 may also be referred to as a stream computing engine, and is used to implement stream computing. As shown in FIG3 , the stream computing device 300 includes a read operator 310 , a concatenation operator 320 , and an output operator 330 .
[0077] The read operator 310 is used to read data from multiple data sources and send the read data to the splicing operator 320. In some embodiments, as shown in Figure 3, the stream computing device 300 may include one read operator 310. In some embodiments, as shown in Figure 4, the stream computing device 300 may include multiple read operators 310. The multiple read operators 310 correspond one-to-one to the multiple data sources, and the read operators 310 are used to read data from the corresponding data sources.
[0078] The concatenation operator 320 is used to concatenate data read by the read operator 310 from different data sources based on a concatenation key. In some embodiments, the concatenation operator 320 concatenates data from multiple data sources based on a merge union statement in SQL. For example, assuming that data source 410, data source 420, data source 430, and data source 440 are tables 410, 420, 430, and 440, respectively, the concatenation operator 320 can concatenate the data from data sources 410, 420, 430, and 440 based on the following statement.
[0079] “Select table 410.colum1,…, table 420.colum1,…, table 430.colum1,…, table 440.colum1,…,
[0080] From Table 410
[0081] merge union table 420, table 430, table 440,
[0082] on table 410.key = table 420.key = table 430.key = table 440.key"
[0083] Among them, key represents splicing, colum represents the column to be spliced, and one column corresponds to one field, or colum represents the field.
[0084] The splicing operator 320 corresponds to a storage space 321, and the storage space 321 is used for state storage of the splicing operator 320. Specifically, the splicing operator 320 can store the received data or the splicing result of the splicing operator 320 as a state in the storage space 321. When the splicing operator 320 receives data from the reading operator 310, the splicing operator 320 can query the storage space 321 for a data set corresponding to the same splicing key as the data (wherein the data set may include data read by the reading operator from at least one data source end. And the data in the data set may refer to data that has been historically received by the splicing operator 320 but has not yet been spliced with other data, or may refer to the splicing result of the splicing operator historically splicing data from two or more data sources). If found, the received data and the queried data set are spliced to obtain a splicing result. Then, the splicing result is stored in the storage space 321. If not found, the received data is stored in the storage space 321.
[0085] In some embodiments, as shown in FIG. 3 or FIG. 4 , the streaming computing device 300 may include a splicing operator 320 .
[0086] In some embodiments, as shown in FIG5 , a stream computing device 300 may include multiple concatenation operators 320. The multiple concatenation operators 320 are arranged in parallel in the data transmission direction of the stream computing device 300. That is, in the data transmission direction, there is no order relationship between the different concatenation operators 320.
[0087] The concatenation key of the data read from the data source by the read operator 310 corresponds to a mapping value. Different concatenation operators 320 correspond to different mapping ranges. The mapping range corresponding to each concatenation operator 320 includes one or more mapping values. The concatenation operator 320 is configured to receive data whose mapping values fall within the mapping range corresponding to the concatenation operator 320 and concatenate the received data.
[0088] When reading data from the data source, the read operator 310 can calculate the mapping value of the data's splicing key and then identify the mapping range to which the mapping value belongs. After identifying the mapping range to which the mapping value belongs, the read operator 310 sends the data to the splicing operator 320 corresponding to the mapping range.
[0089] Among them, the read operator 310 can use a mapping algorithm to calculate the splicing key of the data to obtain the mapping value corresponding to the splicing key. Common mapping algorithms include hash range partitioning algorithm, key range partitioning algorithm, etc. Among them, the hash range partitioning algorithm includes consistent hashing algorithm (consistent hashing). The commonly used consistent hashing algorithm is: taking the remainder based on the total number of assignable mapping values. Specifically, the splicing key is hashed, and the resulting hash value is divided by the total number of assignable mapping values, and the remainder obtained is used as the mapping value of the splicing key.
[0090] As shown in FIG3 , FIG4 or FIG5 , the splicing operator 320 may send the splicing result to the output operator 330. The output operator 330 is used to output the splicing result to the data target end 500.
[0091] In some embodiments, as shown in FIG3 , FIG4 , or FIG5 , the stream computing device 300 can be deployed in a computing node cluster. The computing node cluster includes multiple computing nodes. The computing nodes can be physical nodes, such as servers. The computing nodes can also be virtual computing nodes, such as virtual machines (VMs) or containers.
[0092] Different operators in the stream computing device 300 can be deployed on the same computing node or on different computing nodes. Furthermore, the same operator in the stream computing device 300 can be deployed on the same computing node, meaning that the functionality of the operator is performed by that computing node. The same operator can also be deployed on multiple computing nodes, meaning that one of the multiple different computing nodes can perform a portion of the functionality of the operator.
[0093] Continuing with Figures 3, 4, or 5, the data target end 500 is the demand end for the splicing results and is used to receive the splicing results output by the streaming computing device 300. That is, the streaming computing device 300 can perform real-time splicing of the data required by the data target end 500. In some embodiments, the data target end 500 can be a database or data table for storing the data output by the streaming computing device 300. In some embodiments, the data target end 500 can be a data warehouse service (DWS) to splice the data to be stored in the DWS through the streaming computing device 300. In some embodiments, the data target end 500 is a data analysis end and can perform relevant analysis based on the splicing results of the streaming computing device 300.
[0094] The above examples introduce the system architecture and the stream computing device 300 provided in the embodiment of the present application. Next, the data processing method provided in the embodiment of the present application is described in conjunction with the system architecture and the stream computing device 300.
[0095] The method may be executed by the stream computing device 300, specifically by relevant operators in the stream computing device 300. As shown in FIG6 , the method includes the following steps.
[0096] In step 601, the read operator 310 sends data E1 read from a data source G1 among multiple data sources to the splicing operator 320. The data E1 corresponds to the splicing key F1.
[0097] The data source terminal G1 is one or more of the above-mentioned multiple data source terminals. For example, the data source terminal G1 can be any one or more of the data source terminal 410 , the data source terminal 420 , the data source terminal 430 , and the data source terminal 440 .
[0098] The read operator 310 reads data from the data source and sends the read data to the concatenation operator 320 whenever the read operator 310 reads data from the data source.
[0099] In some embodiments, as shown above, the streaming computing device 300 includes multiple splicing operators 320, wherein different splicing operators 320 correspond to different mapping ranges. The read operator 310 can calculate the mapping value corresponding to the splicing key corresponding to the read data, identify the mapping range to which the mapping value belongs, and then send the read data to the splicing operator corresponding to the mapping range to which the mapping value belongs. In other words, the read operator 310 can calculate the mapping value corresponding to the splicing key F1, identify the mapping range to which the mapping value corresponding to the splicing key F1 belongs, and then send the data E1 to the splicing operator 320 corresponding to the mapping range.
[0100] 6 , in step 602 , the concatenation operator 320 searches the storage space 321 of the concatenation operator 320 for data corresponding to the concatenation key F1 , ie, searches the storage space for data corresponding to the same concatenation key as the data E1 .
[0101] The storage space 321 stores data that has been historically read by the read operator 310 from multiple data sources. The history here refers to the period before the splicing operator 320 executes step 601. The splicing operator 320 can splice the data corresponding to the same splicing key that has been historically read by the read operator 310 from different data sources, and then store the splicing result in the storage space 321. In other words, when the read operator 310 reads data corresponding to the same splicing key from different data sources, the storage space 321 stores the splicing result of the data corresponding to the same splicing key that has been read by the read operator 310 from different data sources.
[0102] In some embodiments, the storage space 321 uses a data table to record data, so as to store the data in the storage space 321. Data corresponding to the same splicing key is recorded in the same row of the data table.
[0103] In step 603 , the splicing operator 320 may determine whether data corresponding to the splicing key F1 is found in the storage space 321 .
[0104] If the result of step 603 is yes, that is, data corresponding to the splicing key F1 is found in the storage space 321, the splicing operator 320 can splice the data E1 with the found data, or update the found data based on the data E1.
[0105] The data corresponding to the splicing key F1 found in the storage space 321 can be set as data E2. It can be determined whether data E1 and data E2 are data read by the read operator 310 from different data sources, that is, whether data E2 is read by the read operator 310 from a data source other than the data source G1. If data E2 is read by the read operator 310 from a data source other than the data source G1, the splicing operator 320 can splice data E1 and data E2 in step 604 to obtain a splicing result. This splicing result can be referred to as data E3 corresponding to the splicing key F1.
[0106] In some embodiments, data E2 may be composed of data read by the read operator 310 from at least two data sources. In other words, data E2 is derived from data read by the read operator 310 from at least two data sources. Specifically, data E2 is the concatenation result of data corresponding to the concatenation key F1 read by the read operator 310 from different data sources.
[0107] In one example, the data E1 may be set as shown in Table 1, and the data E2 may be set as shown in Table 2.
[0108] Table 1
[0109] It can be assumed that the data E1 is read by the read operator 310 from the data source end 410 , that is, the data in column A.1 and column A.2 are read by the read operator 310 from the data source end 410 .
[0110] Table 2
[0111] The data in columns C.1 and C.2 are read by the read operator 310 from the data source end 430 , and the data in columns D.1 and D.2 are read by the read operator 310 from the data source end 440 .
[0112] Then data E2 and data E1 are concatenated to obtain data E2' as shown in Table 3.
[0113] Table 3
[0114] In another example, data E1 may be data read from multiple data sources by the read operator 310, and the data read from different data sources may have the same concatenation key. In step 604, the data read from the multiple data sources may be concatenated with the data in data E2 to obtain data E3.
[0115] The data E1 may be set to include the data shown in Table 1 and the data shown in Table 4, wherein the data E2 is still as shown in Table 2.
[0116] Table 4
[0117] It can be assumed that the data in column B.1 and column B.2 are read by the read operator 310 from the data source end 420 .
[0118] The data in Table 1, Table 4 and Table 2 are combined to obtain data E3 as shown in Table 5.
[0119] Table 5
[0120] In some embodiments, data E2 may only include data read from a data source by the read operator 310. For example, data E1 may be set as shown in Table 1, and data E2 may be set as shown in Table 6.
[0121] Table 6
[0122] The data E3 obtained by concatenating the data E1 and the data E2 is shown in Table 7.
[0123] Table 7
[0124] Through the above method, the data in data E1 and data E2 can be spliced.
[0125] When the data in data E1 and data E2 are concatenated, the concatenation operator 320 may execute step 605a to store data E3 in the storage space 321. Data E3 may be stored in the storage space 321 as a state. For example, in step 605a, data E2 in the storage space 321 may be replaced with data E3. That is, while data E3 is being stored in the storage space 321, data E2 is deleted from the storage space 321, thereby reducing the amount of data stored in the storage space 321 and, in other words, reducing the amount of data stored in the state of the concatenation operator 320. Data E3 includes data E1 and data E2, and by storing the concatenated data E3, the storage of data E1 and data E2 is achieved.
[0126] When the splicing of the data in data E1 and data E2 is completed, the splicing operator 320 can execute step 605b to output data E3 to the output operator 330, so that the output operator 330 outputs data E3 to the data target end 500 to meet the data target end 500's demand for streaming data.
[0127] When part or all of data E1 and data E2 are read from the same data source by the read operator 310, that is, when data E2 corresponding to the splicing key F1 is found in the storage space 321, and part or all of data E1 and data E2 are read from the same data source by the read operator 310, the splicing operator 320 updates data E2 based on data E1. Specifically, data E2 can be set to include data E21. Data E21 can be part of data E2 or all of data E2. Data E21 and data E1 are read from the same data source by the read operator 310. In this case, the reading time of data E21 is earlier, and the reading time of data E1 is later, and data E1 is the updated data E21. Data E21 in data E2 can be replaced with data E1, thereby updating data E2.
[0128] For example, the data E2 may be set as shown in Table 3 described above, and the data E1 may be set as shown in Table 8.
[0129] Table 8
[0130] The data in columns A.1 and A.2 in Table 3 is data E21. The read timestamp for data E21 is "202305061230," while the read timestamp for data E1 is "202305061340." Data E21 is read earlier, while data E1 is read later. The data in columns A.1 and A.2 in Table 3 can be replaced with the data in columns A.1 and A.2 in Table 8 to obtain updated data E2. The updated data E2 is shown in Table 9.
[0131] Table 9
[0132] After completing the update of data E2, splicing operator 320 can send the updated data E2 to output operator 330, so that output operator 330 outputs the updated data E2 to data target 500, thereby meeting the streaming data requirements of data target 500. Splicing operator 320 can also store the updated data E2 in storage space 321 for subsequent use, such as splicing or updating.
[0133] Continuing with Figure 6 , if no data E2 corresponding to the same splicing key as data E1 is found in storage space 321, splicing operator 320 can execute step 606a to record data E1 in storage space 321 for subsequent splicing or updating. Specifically, data E1 can be set as shown in Table 1 described above, and the data in storage space 321 can be set as shown in Table 10.
[0134] Table 10
[0135] In the example, splicing keys F2, 3, and F4 are all different from splicing key F1. That is, the splicing keys of the data stored in storage space 321 are all different from the splicing keys of data E1. In this case, data E1 can be recorded in storage space 321. Specifically, data E1 can be recorded in Table 10 in storage space 321, resulting in Table 11.
[0136] Table 11
[0137] Continuing to refer to Figure 6, when the corresponding splicing key F1 data is not found in the storage space 321, the splicing operator 320 can also execute step 606b to send data E1 to the output operator 330, so that the output operator 330 sends data E1 to the data target end 500 to meet the data target end 500's demand for streaming data.
[0138] In some embodiments, read operator 310 sends data H1 read from data source G2 among multiple data sources to concatenation operator 320. Data H1 corresponds to concatenation key F5. Data source G2 can be one or more of the aforementioned multiple data sources. For example, data source G2 can be any one or more of data source 410, data source 420, data source 430, and data source 440.
[0139] The splicing operator 320 can query the storage space 321 for data corresponding to the splicing key F5 , that is, query the storage space 321 for data corresponding to the same splicing key as the data H1 .
[0140] If there is no data corresponding to the splicing key F5 in the storage space 321 , the splicing operator 320 stores the data H1 in the storage space 321 for subsequent splicing or updating.
[0141] The concatenation operator 320 sends the data H1 to the output operator 330 , so that the output operator 330 outputs the data H1 to the data target end 500 .
[0142] In some embodiments, the data in the storage space 321 can be aged based on a time to live (TTL) mechanism, and invalid data in the storage space 321 can be deleted in a timely manner to reduce the amount of data stored in the state. Specifically, a life cycle can be set for the data in the storage space 321. When the life cycle of the data ends, the data can be deleted from the storage space 321 to free up available space in the storage space 321. The starting point of the life cycle can be the storage moment when the data is stored in the storage space 321, and the length of the life cycle is related to the time it takes for all the data to be spliced with the same splicing key to reach the splicing operator 320.
[0143] Taking splicing key F1 as an example, the time at which data corresponding to splicing key F1 in different data sources arrives at splicing operator 320 may be different. Specifically, the time at which data corresponding to splicing key F1 in data source 410 arrives at splicing operator 320 first can be set. In other words, the time at which data corresponding to splicing key F1 in other data sources arrives at splicing operator 320 is later than the time at which data corresponding to splicing key F1 in data source 410 arrives at splicing operator 320. Alternatively, the time at which data corresponding to splicing key F1 in data source 420 arrives at splicing operator 320 last can be set. In other words, the time at which data corresponding to splicing key F1 in other data sources arrives at splicing operator 320 is earlier than the time at which data corresponding to splicing key F1 in data source 420 arrives at splicing operator 320. The time it takes for all the data to be spliced for the splicing key F1 to arrive at the splicing operator 320 is obtained by subtracting the time it takes for the data to arrive at the splicing operator 320 from the time it takes for the data to arrive at the splicing operator 320. For ease of description, the time it takes for all the data to arrive at the splicing operator 320 for the splicing key F1 is referred to as the time required for splicing for the splicing key.
[0144] Referring to the above method for calculating the time required for splicing the splicing key F1, the time required for splicing multiple splicing keys can be calculated. Then, based on the time required for splicing the multiple splicing keys, the length of the life cycle is set. For example, the time required for splicing the splicing key with the longest splicing time among the multiple splicing keys can be used as the length of the life cycle.
[0145] In summary, in the method provided by this application, the same concatenation operator can complete the concatenation of data from multiple (e.g., three or more) data sources. Thus, only one concatenation operator performs state storage on the data read from multiple data sources. The amount of data stored in the state = ∑Ti, where Ti is the data read from the i-th data source. Compared to multiple concatenation operators performing pairwise cascade concatenation on data from multiple data sources, the method provided by the embodiments of this application reduces the amount of data stored in the state and saves storage overhead.
[0146] Furthermore, in the method provided herein, upon receiving data, the splicing operator retrieves data corresponding to the data in the same splicing pattern from the storage space and splices the data retrieved from the storage space with the received data to obtain a splicing result. The splicing operator can store the splicing result in the storage space. Compared to splicing based on data tables in the storage space, the method provided herein requires less computation, saves computing resources, and improves splicing efficiency, thereby completing the splicing faster and enhancing the real-time nature of the splicing.
[0147] In addition, the method provided in the embodiment of the present application can achieve splicing and output regardless of which data stream arrives first. There is no dependency between different data streams, so that data can be output to the data target end in a timely manner, ensuring the data target end's demand for data.
[0148] In addition, the method provided in this application sets a life cycle for the data in the storage space. When the data life cycle ends, the data is deleted, so that unnecessary data can be deleted from the storage space in a timely manner, further reducing the amount of data stored in the state.
[0149] Referring to FIG7 , the embodiment of the present application further provides a stream computing device 700. As shown in FIG7 , the stream computing device 700 includes a read operator 710 and a splice operator 720. The splice operator 710 is used to read data from multiple data sources;
[0150] The read operator 710 is configured to: read first data from a first data source among the multiple data sources and send the first data to the splicing operator, wherein the first data corresponds to a first splicing key;
[0151] The splicing operator 720 is used to: query the storage space for the second data corresponding to the first splicing key;
[0152] The concatenation operator 720 is further configured to: when the second data is read by the read operator from a data source end other than the first data source end among the multiple data source ends, concatenate the first data and the second data to obtain third data corresponding to the first concatenation key;
[0153] The concatenation operator 720 is further configured to store the third data in the storage space.
[0154] In some embodiments, the second data is a concatenation result of at least two data, each of the at least two data corresponds to the first concatenation key, and different data in the at least two data are read by the read operator 710 from different data sources.
[0155] In some embodiments, the streaming computing device 700 also includes an output operator 730 connected to the data target end; wherein, the splicing operator 720 is also used to: send the third data to the output operator 730, so that the output operator 730 outputs the third data to the data target end.
[0156] In some embodiments, the splicing operator 720 is also used to: when part or all of the first data and the second data are read by the reading operator from the same data source, update the second data based on the first data; and store the updated second data in the storage space.
[0157] In some embodiments, the read operator 710 is further used to: read fourth data from a second data source end among multiple data source ends and send it to the splicing operator, wherein the fourth data corresponds to a second splicing key; the splicing operator 720 is further used to: when the splicing operator 720 does not find data corresponding to the second splicing key in the storage space, store the fourth data in the storage space.
[0158] In some embodiments, the streaming computing device 700 includes multiple splicing operators, wherein different splicing operators correspond to different mapping ranges; the reading operator 710 is used to: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, send the first data to the splicing operator 720.
[0159] In some embodiments, the data in the storage space has a life cycle; the splicing operator 720 is further configured to delete data whose life cycle has ended in the storage space.
[0160] Among them, the read operator 710, the splice operator 720, and the output operator 730 can all be implemented by software or hardware. For example, the implementation of the read operator 710 is described below using the read operator 710 as an example. Similarly, the implementation of the splice operator 720 and the output operator 730 can refer to the implementation of the read operator 710.
[0161] As an example of a software functional unit, the module read operator 710 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the read operator 710 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone AZ or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Generally, a region may include multiple AZs.
[0162] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0163] As an example of a hardware functional unit, the read operator 710 may include at least one computing device, such as a server. Alternatively, the read operator 710 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0164] The multiple computing devices included in the read operator 710 can be distributed in the same region or in different regions. The multiple computing devices included in the read operator 710 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the read operator 710 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0165] This application also provides a computing device 800. As shown in Figure 8, computing device 800 includes a bus 802, a processor 804, a memory 806, and a communication interface 808. Processor 804, memory 806, and communication interface 808 communicate with each other via bus 802. Computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 800.
[0166] Bus 802 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG8 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 802 may include a path for transmitting information between various components of computing device 800 (e.g., memory 806, processor 804, and communication interface 808).
[0167] The processor 804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0168] The memory 806 may include volatile memory, such as random access memory (RAM). The memory 806 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0169] The memory 806 stores executable program code, and the processor 804 executes the executable program code to respectively implement the functions of the aforementioned read operator 710, splice operator 720, and output operator 730, thereby implementing the method shown in Figure 6. In other words, the memory 806 stores instructions for executing the method shown in Figure 6.
[0170] The communication interface 808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.
[0171] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0172] As shown in Figure 9 , the computing device cluster includes at least one computing device 800. The memory 806 in one or more computing devices 800 in the computing device cluster may store the same instructions for executing the method shown in Figure 6 .
[0173] In some possible implementations, the memory 806 of one or more computing devices 800 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 6. In other words, the combination of one or more computing devices 800 can jointly execute the instructions for executing the method shown in Figure 6.
[0174] It should be noted that the memory 806 in different computing devices 800 in the computing device cluster can store different instructions, each used to execute part of the functions of the streaming computing apparatus 500. In other words, the instructions stored in the memory 806 in different computing devices 800 can implement the functions of one or more modules among the reading operator 710, the splicing operator 720, and the output operator 730.
[0175] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG10 shows a possible implementation. As shown in FIG10 , two computing devices 800A and 800B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 806 in the computing device 800A stores instructions for executing the function of the read operator 710. At the same time, the memory 806 in the computing device 800B stores instructions for executing the functions of the splicing operator 720 and the output operator 730.
[0176] It should be understood that the functionality of the computing device 800A shown in FIG10 may also be implemented by multiple computing devices 800. Similarly, the functionality of the computing device 800B may also be implemented by multiple computing devices 800.
[0177] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 9 and 10. However, the memory 806 in one or more computing devices 800 in this computing device cluster can store the same instructions for executing the method shown in Figure 6.
[0178] In some possible implementations, the memory 806 of one or more computing devices 800 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 6. In other words, the combination of one or more computing devices 800 can jointly execute the instructions for executing the method shown in Figure 6.
[0179] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method shown in FIG6 .
[0180] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a host migration device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method shown in FIG6 .
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data splicing method, characterized in that: The method is applied to a stream computing device, the device comprising a splicing operator and a reading operator for reading data from multiple data source ends; the method comprises: The read operator reads first data from a first data source end among the multiple data source ends and sends the first data to the splicing operator, wherein the first data corresponds to a first splicing key; The concatenation operator searches for second data corresponding to the first concatenation key in the storage space; When the second data is read by the read operator from a data source end other than the first data source end among the multiple data source ends, the concatenation operator concatenates the first data and the second data to obtain third data corresponding to the first concatenation key; The concatenation operator stores the third data in the storage space.
2. The method according to claim 1, characterized in that: The second data is a concatenation result of at least two data, each of the at least two data corresponds to the first concatenation key, and different data in the at least two data are read by the read operator from different data sources.
3. The method according to claim 1 or 2, characterized in that: The device also includes an output operator connected to the data target end; the method also includes: the splicing operator sends the third data to the output operator, so that the output operator outputs the third data to the data target end.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: When part or all of the first data and the second data are read by the read operator from the same data source, the concatenation operator updates the second data based on the first data; The concatenation operator stores the updated second data in the storage space.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The read operator reads fourth data from a second data source end among the multiple data source ends and sends the fourth data to the splicing operator, wherein the fourth data corresponds to the second splicing key; When the concatenation operator does not find data corresponding to the second concatenation key in the storage space, the concatenation operator stores the fourth data in the storage space.
6. The method according to any one of claims 1 to 5, characterized in that The device includes a plurality of splicing operators, wherein different splicing operators correspond to different mapping ranges; The reading operator reads first data from a first data source end among multiple data source ends and sends the first data to the splicing operator, including: When the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, the read operator sends the first data to the splicing operator.
7. The method according to any one of claims 1 to 6, characterized in that The data in the storage space has a life cycle; the method further includes: the concatenation operator deleting data whose life cycle has ended in the storage space.
8. A streaming computing device, characterized in that: The device includes a splicing operator and a reading operator for reading data from multiple data source ends; wherein, The read operator is used to: read first data from a first data source end among the multiple data source ends and send it to the splicing operator, wherein the first data corresponds to a first splicing key; The concatenation operator is used to: query the storage space for the second data corresponding to the first concatenation key; The concatenation operator is further used for: when the second data is read by the read operator from a data source end other than the first data source end among the multiple data source ends, concatenating the first data and the second data to obtain third data corresponding to the first concatenation key; The concatenation operator is further used to: store the third data into the storage space.
9. The device according to claim 8, characterized in that The second data is a concatenation result of at least two data, each of the at least two data corresponds to the first concatenation key, and different data in the at least two data are read by the read operator from different data sources.
10. The device according to claim 8 or 9, characterized in that The device also includes an output operator connected to the data target end; wherein the splicing operator is further used to: send the third data to the output operator, so that the output operator outputs the third data to the data target end.
11. The device according to any one of claims 8 to 10, characterized in that The concatenation operator is also used to: When part or all of the first data and the second data are read by the read operator from the same data source, Based on the first data, updating the second data; The updated second data is stored in the storage space.
12. The device according to any one of claims 8 to 11, characterized in that The read operator is further used to: read fourth data from a second data source end among the multiple data source ends and send it to the splicing operator, wherein the fourth data corresponds to the second splicing key; The concatenation operator is further used for: when the concatenation operator fails to find data corresponding to the second concatenation key in the storage space, storing the fourth data in the storage space.
13. The device according to any one of claims 8 to 12, characterized in that The device includes a plurality of splicing operators, wherein different splicing operators correspond to different mapping ranges; The read operator is used to send the first data to the splicing operator when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator.
14. The device according to any one of claims 8 to 13, characterized in that The data in the storage space has a life cycle; the splicing operator is further used to delete data whose life cycle has ended in the storage space.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for realizing data splicing and system for realizing data splicing
CN112632053A
Data stream processing method and device, electronic equipment and storage medium
CN115481153A
Cross-data-source query method, system and equipment for database and storage medium
CN115982230A
Distributed Indexing and Aggregation
US20200192947A1