Data splicing method and device and cluster
By introducing storage space into the streaming computing device, the splicing operator can query and store data of the corresponding splicing keys. When new data is received, the splicing operator obtains corresponding data from the storage space for splicing, solving the problem of large amount of state storage data and low efficiency caused by cascade splicing of multiple splicing operators in the prior art, and realizing more efficient data splicing.
Patent Information
- Application Number
- CN202410063281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-01-16
- Publication Date
- 2025-05-20
AI Technical Summary
In streaming calculation, when there are three or more data source terminals, the prior art requires cascade splicing through multiple splicing operators, resulting in large amounts of data storage in the state of splicing operators and low efficiency.
By introducing storage space into the streaming computing device, the splicing operator can query and store data of the corresponding splicing keys. When new data is received, the splicing operator acquires corresponding data from the storage space for splicing, reducing the use of multiple splicing operators and state storage.
The amount of data stored in the splicing operator in the streaming computing device is reduced, the efficiency of data splicing is improved, and the need for real-time data splicing is met.
Smart Images

Figure CN120020756A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202311541760.9 and the application title "A Data Splicing Method" submitted to the China National Intellectual Property Administration on November 17, 2023. The entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular, to a data splicing method, apparatus, and cluster. Background Art
[0003] The real-time collection and analysis of data have been applied in various fields. For example, in the logistics field, by collecting and analyzing logistics information in real time, enterprises can master the location and transportation status of goods, which helps to improve logistics efficiency. Since there may be multiple data source ends, when collecting and analyzing data in real time, it is necessary to splice data from different data source ends.
[0004] The real-time collection and analysis of data are carried out through stream computing. The current stream computing solutions mainly use a pairwise splicing method to splice data from different data source ends. Among them, when there are three or more data source ends, a two-level cascaded splicing method needs to be used for splicing. Specifically, the first splicing operator splices the data of two of the data source ends to obtain a splicing result; then, the first splicing operator sends the splicing result to the second splicing operator, and the second splicing operator splices the splicing result and the data of the third data source end, and so on. In this solution, when there are three or more data source ends, multiple splicing operators are required for splicing, which requires multiple splicing operators to store the status of the received data, and the overall amount of data stored in the status is large. Summary of the Invention
[0005] This application provides a data splicing method, apparatus, and cluster, which can reduce the amount of data stored in the status of the splicing operator in the stream computing device.
[0006] In a first aspect, a data splicing method is provided. This method is applied to a stream computing device, which includes a splicing operator and a reading operator for reading data from multiple data source ends. The method includes: the reading operator sends the first data read from the first data source end among the multiple data source ends to the splicing operator, where the first data corresponds to the first splicing key; the splicing operator queries the second data corresponding to the first splicing key in the storage space; when the second data is read by the reading operator from a data source end other than the first data source end among the multiple data source ends, the splicing operator splices the first data and the second data to obtain the third data corresponding to the first splicing key; the splicing operator stores the third data in the storage space. Among them, the third data is the splicing result of the first data and the second data.
[0007] In this method, the splicing operator can store the data corresponding to a certain splicing key into the storage space. When the splicing operator receives other data corresponding to the splicing key from the reading operator, the splicing operator can obtain the data corresponding to the splicing key from the storage space, and splice the data obtained from the storage space and the data received from the reading operator to obtain a splicing result. The splicing operator can store the splicing result into the storage space for splicing the subsequently received data. In this way, the data corresponding to the same splicing key in multiple (such as three or more) data source ends can be spliced together by the same splicing operator. Correspondingly, the same splicing operator stores the states of the data in multiple data source ends. Compared with the two-stage cascaded splicing method, this can not only reduce the amount of data stored in the state of the splicing operator in the streaming computing device, but also improve the efficiency of data splicing, meeting the needs of related services for real-time data splicing.
[0008] In a possible implementation manner, the second data is a splicing result of at least two pieces of data. Each piece of data in the at least two pieces of data corresponds to a first splicing key, and different data in the at least two pieces of data are read by the reading operator from different data source ends.
[0009] That is to say, the splicing operator can store the splicing result corresponding to a certain splicing key into the storage space. When the splicing operator receives the data corresponding to the splicing key from the reading operator again, the splicing operator can obtain the splicing result corresponding to the splicing key from the storage space, and splice the splicing result and the data to obtain a spliced result of the second splicing. The splicing operator can store the spliced result of the second splicing into the storage space for splicing the subsequently received data. In this way, the data corresponding to the same splicing key in multiple data source ends can be spliced together by the same splicing operator.
[0010] In a possible implementation manner, the device further includes an output operator connected to the data target end; the method further includes: the splicing operator sends third data to the output operator, so that the output operator outputs the third data to the data target end.
[0011] The third data is a splicing result of the second data and the first data. In this implementation manner, the splicing operator can send the splicing result to the data target end to meet the requirements of the data target end for streaming data.
[0012] In a possible implementation manner, the method further includes: when part or all of the first data and the second data are read by the reading operator from the same data source end, the splicing operator updates the second data based on the first data; the splicing operator stores the updated second data into the storage space. Wherein, the second data is a splicing result of the data corresponding to the first splicing key read by the reading operator historically.
[0013] When the data corresponding to the first splicing key in a certain or certain data source ends is updated, through this implementation method, the updated data can be updated to the splicing result corresponding to the first splicing key.
[0014] In a possible implementation method, the method further includes: the reading operator sends the fourth data read from the second data source end among multiple data source ends to the splicing operator, where the fourth data corresponds to the second splicing key; when the splicing operator does not query the data corresponding to the second splicing key in the storage space, the splicing operator stores the fourth data in the storage space.
[0015] In this implementation method, if the reading operator has not read the data corresponding to the second splicing key in history, the splicing operator stores the currently received data corresponding to the second splicing in the storage space, so that when the data corresponding to the second splicing key is subsequently received, the data corresponding to the second splicing key and the subsequently received data corresponding to the second splicing key are spliced.
[0016] In a possible implementation method, the device includes multiple splicing operators, where different splicing operators correspond to different mapping ranges; the reading operator sends the first data read from the first data source end among multiple data source ends to the splicing operator, including: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, the reading operator sends the first data to the splicing operator.
[0017] The mapping value corresponding to the splicing key of the data read by the reading operator from the data source end is sent to the splicing key of the corresponding mapping range pair based on the mapping value corresponding to the splicing key of the data, so that the splicing operator splices the data corresponding to the splicing key of the mapping value belonging to this mapping range, so that the data corresponding to the splicing keys with mapping values belonging to different mapping ranges can be spliced by different splicing operators respectively, reducing the load of the splicing operator, increasing the concurrency of data splicing, and improving the efficiency of data splicing.
[0018] In a possible implementation method, the data in the storage space has a life cycle; the method further includes: the splicing operator deletes the data whose life cycle has ended in the storage space.
[0019] Deleting the data whose life cycle has ended can release the space occupied by these data, so as to ensure that the storage space can be continuously used to store the splicing results of the splicing operator.
[0020] In a second aspect, a streaming computing device is provided. The device includes a splicing operator and a reading operator for reading data from multiple data source ends. Among them, the reading operator is used to: send the first data read from the first data source end among the multiple data source ends to the splicing operator, where the first data corresponds to a first splicing key; the splicing operator is used to: query the second data corresponding to the first splicing key in the storage space; the splicing operator is further used to: when the second data is read by the reading operator from a data source end other than the first data source end among the multiple data source ends, splice the first data and the second data to obtain a third data corresponding to the first splicing key; the splicing operator is further used to: store the third data in the storage space.
[0021] In a possible implementation, the second data is the splicing result of at least two data. Each of the at least two data corresponds to the first splicing key, and different data among the at least two data are read by the reading operator from different data source ends.
[0022] In a possible implementation, the device further includes an output operator connected to the data target end. Among them, the splicing operator is further used to: send the third data to the output operator, so that the output operator outputs the third data to the data target end.
[0023] In a possible implementation, the splicing operator is further used to: when part or all of the first data and the second data are read by the reading operator from the same data source end, update the second data based on the first data; and store the updated second data in the storage space.
[0024] In a possible implementation, the reading operator is further used to: send the fourth data read from the second data source end among the multiple data source ends to the splicing operator, where the fourth data corresponds to a second splicing key; the splicing operator is further used to: when the splicing operator does not query the data corresponding to the second splicing key in the storage space, store the fourth data in the storage space.
[0025] In a possible implementation, the device includes multiple splicing operators. Among them, different splicing operators correspond to different mapping ranges. The reading operator is used to: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, send the first data to the splicing operator.
[0026] In a possible implementation, the data in the storage space has a life cycle. The splicing operator is further used to: delete the data whose life cycle has ended in the storage space.
[0027] In a third aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method provided in the first aspect.
[0028] In a fourth aspect, a computer-readable storage medium is provided, including computer program instructions, and when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method provided in the first aspect.
[0029] In a fifth aspect, a computer program product including instructions is provided, and when the instructions are run by a computer device cluster, the computer device cluster is caused to execute the method provided in the first aspect.
[0030] For the beneficial effects of the second aspect to the fifth aspect, reference may be made to the introduction of the beneficial effects of the first aspect above, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of a splicing key;
[0032] Figure 2 Schematic diagram of a data splicing solution;
[0033] Figure 3 Schematic diagram of a system architecture provided by an embodiment of the present application;
[0034] Figure 4 Schematic diagram of a system architecture provided by an embodiment of the present application;
[0035] Figure 5 Schematic diagram of a system architecture provided by an embodiment of the present application;
[0036] Figure 6 Flowchart of a data splicing method provided by an embodiment of the present application;
[0037] Figure 7 Schematic diagram of the structure of a streaming computing device provided by an embodiment of the present application;
[0038] Figure 8 Schematic diagram of the structure of a computing device provided by an embodiment of the present application;
[0039] Figure 9 Schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application;
[0040] Figure 10 Schematic diagram of the structure of a computing device cluster connected through a network provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The solutions provided in the embodiments of this application will be described below in conjunction with the accompanying drawings. Among them, in the embodiments of this application, "a plurality of" means two or more. "First", "second", etc. are only used to distinguish similar objects and do not necessarily need to describe a specific order or the number of objects.
[0042] To facilitate understanding of the solutions provided in the embodiments of this application, the technical terms that may be involved in the embodiments of this application will be introduced first.
[0043] Data table: It can be abbreviated as table, which is used to record data. Data tables can be divided into horizontal tables and vertical tables.
[0044] Horizontal table: A type of data table. In a horizontal table, all data records corresponding to the same identifier (ID) are recorded in the same row, or in other words, all values (value) of the same key are recorded in the same row. Opposite to the horizontal table is the vertical table. Among them, in a vertical table, each row only records one piece of data, that is, one value of a key occupies one row, and multiple values occupy multiple rows.
[0045] Wide table: It is a horizontal table in the database field that records a large amount of data.
[0046] Joining key: Also known as an association key, it is used to join data from different data source ends. For example, as Figure 1 shown, it can be set that table t1 records the names, grades, classes, etc. of students, and table t2 records the names, math scores, physics scores, etc. of students. Table t1 and table t2 are different data source ends. The student ID number can be used as the joining key to join the data in table t1 and table t2 into table t3. In table t3, the grades, classes, math scores, physics scores, etc. corresponding to the same name are associated. For example, as Figure 1 shown, table t3 is a horizontal table, and the grades, classes, math scores, physics scores corresponding to the same name are recorded in the same row. In addition, name, grade, class, math score, and physics score are different fields, corresponding to one column in the table respectively. The name, grade, class, math score, and physics score of a certain student are the field values of the corresponding fields respectively.
[0047] Data set: It refers to a set containing one or more data. Among them, the data in the same data set correspond to the same joining key. Among them, the data in the same data set are associated and recorded in the table. In the case where the table is a horizontal table (such as a wide table), the data in the same data set are recorded in the same row in the table.
[0048] Joining: It means including the data from different data source ends that correspond to the same joining key into the same data set. Exemplarily, as Figure 1As shown, in the case of a horizontal table, splicing refers to recording data records corresponding to the same splicing key from different data source ends in the same row of the table. Here, a row in the table can be considered a data set.
[0049] Streaming computing: Also known as stream computing, streaming computing is used for real-time data processing. Among them, streaming computing can be a computing model triggered by input data, and each newly input data can be regarded as an event to trigger the computing model to perform calculations.
[0050] Operator: It is a data processing unit that carries computing logic. Among them, the operator is the smallest executable unit in streaming computing. The operator is used to perform calculations on relevant data based on the computing logic it carries. Common operators in a streaming computing device include read operators, compute operators, and output operators.
[0051] Read operator: Also known as the source operator, it is an operator in a streaming computing device used to read data from the data source end. Among them, the read operator sends the read data to the splicing operator.
[0052] Splicing operator: It is a type of compute operator in streaming computing, specifically an operator used to perform data splicing tasks.
[0053] Output operator: Also known as the sink operator, it is used to send the calculation results of the compute operator (such as the splicing results of the splicing operator) to the data target end.
[0054] State storage: It means that the splicing operator in streaming computing stores the data received by the splicing operator or the splicing results of the splicing operator as the state. Usually, the operator stores the state in local memory. That is to say, the state storage of the operator consumes the memory of the operator.
[0055] In one solution, the stream computing uses a join statement (such as a full outer join statement) in the structured query language (SQL) to splice data from different data source ends. Therefore, this solution is a pairwise splicing solution. When data from three or more data source ends needs to be spliced, pairwise cascaded splicing is required, which amplifies the amount of data stored in the overall state of the splicing operator.
[0056] Take the following statement as an example.
[0057] “Select C.ID,C.B1,C.B2,C.A1,C.A2,D.D1,D.D2
[0058] From
[0059] (Select A.ID,B.B1,B.B2,A.A1,A.A2.A.TIMESTAMP
[0060] FROM A
[0061] FULL OUTER JOIN B
[0062] ON A.ID = B.ID)C
[0063] FULL OUTER JOIN D
[0064] ON C.ID = D.ID”
[0065] In this statement, the data source side is Table A, Table B, and Table D. This statement means that using ID as the concatenation key, the fields A1, A2, B1, and B2 in Table C are concatenated with the fields D1 and D2 in Table D. Among them, Table C is obtained by concatenating the fields A1 and A2 in Table A with the fields B1 and B2 in Table B using ID as the concatenation key. The execution of this statement requires two concatenation operators for executing join statements, specifically as follows.
[0066] As Figure 2 shown, the two concatenation operators for executing join statements can be set as concatenation operator 210 and concatenation operator 220. When concatenation operator 210 receives the data in Table A, concatenation operator 210 takes this data as the state and records it in Table A' in storage space 211. When concatenation operator 210 receives the data in Table B, concatenation operator 210 takes this data as the state and records it in Table B'
[0067] in it. That is to say, Table A' records the data in Table A received by concatenation operator 210, and Table B' records the data in Table B received by concatenation operator 210. When performing concatenation, concatenation operator 210 obtains Table A' and Table B' from storage space 211. Then, based on Table A' and Table B', data concatenation is performed to obtain Table C. Concatenation operator 210 sends Table C to concatenation operator 220. When concatenation operator 220 receives Table C, concatenation operator 220 takes Table C as the state and stores it in storage space 221. When concatenation operator 220 receives the data in Table D, concatenation operator 220 takes this data as the state and records it in Table D' in storage space 221.
[0068] As described above, the state of the concatenation operator 210 stores the data in Table A and the data in Table B. Table C is obtained by concatenating the data in Table A and the data in Table B received by the concatenation operator 210, and the state storage of the concatenation operator 220 stores the data in Table C and the data in Table D. The total amount of data stored in the state is the sum of twice the data volume of Table A, twice the data volume of Table B, and the data volume of Table D.
[0069] Moreover, in this solution, the more data source ends there are, the more concatenation operators are required, resulting in a relatively low overall operating efficiency.
[0070] In addition, the data received by the concatenation operator 210 from different data source ends are stored in different tables respectively. When performing concatenation, the concatenation is based on different tables, resulting in a large amount of concatenation calculation and a large concatenation delay.
[0071] In another solution, real-time concatenation is performed on the data that arrives simultaneously, and the data that arrives at different times is concatenated periodically. This results in a long delay in concatenating the data with the same concatenation key that arrives at different times, making it difficult to meet the requirements of real-time data analysis.
[0072] The embodiment of the present application provides a data concatenation method, which can be applied to a streaming computing device. This method can complete the concatenation of all data with the same concatenation key through the same concatenation operator. That is to say, regardless of whether there are two, three or more data source ends, the data from these data source ends can be concatenated by the same concatenation operator. Thus, when there are three or more data source ends, there is no need to perform two-level cascaded concatenation, thereby avoiding amplifying the amount of data stored in the state.
[0073] Next, the data concatenation method provided by the embodiment of the present application will be described.
[0074] Figure 3 Fig. shows a system architecture that can be used to implement the data processing method provided by the embodiment of the present application. As Figure 3 shown, the system architecture includes a streaming computing device 300, multiple data source ends, and a data target end 500.
[0075] Among them, the multiple data source ends may include a data source end 410, a data source end 420, a data source end 430, a data source end 440, etc. In some embodiments, the data source end may be a database or a data table. In some embodiments, the data source end may be a data acquisition end, that is, the data source end can collect data from the acquisition object to obtain data. For example, the data source end may be an environmental monitoring device, and the environmental monitoring device continuously monitors the environment to generate data.
[0076] The streaming computing device 300 may also be referred to as a streaming computing engine and is used to implement streaming computing. As Figure 3As shown, the streaming computing device 300 includes a reading operator 310, a splicing operator 320, and an output operator 330.
[0077] The reading operator 310 is used to read data from multiple data source ends and send the read data to the splicing operator 320. In some embodiments, as Figure 3 shown, the streaming computing device 300 may include one reading operator 310. In some embodiments, as Figure 4 shown, the streaming computing device 300 may include multiple reading operators 310. Among them, the multiple reading operators 310 correspond to the multiple data source ends one by one, and the reading operator 310 is used to read data from the corresponding data source end.
[0078] The splicing operator 320 is used to splice the data read by the reading operator 310 from different data source ends based on a splicing key. In some embodiments, the splicing operator 320 splices the data of multiple data source ends based on the merge union statement in SQL. For example, assuming that the data source ends 410, 420, 430, and 440 are the tables 410, 420, 430, and 440 respectively, the splicing operator 320 may splice the data in the data source ends 410, 420, 430, and 440 based on the following statement.
[0079] "Select table 410.colum1,…, table 420.colum1,…, table 430.colum1,…, table 440.colum1,…,
[0080] From table 410
[0081] merge union table 420, table 430, table 440,
[0082] on table 410.key = table 420.key = table 430.key = table 440.key"
[0083] Where key represents splicing, and colum represents the columns to be spliced. Among them, one column corresponds to one field, or it can be said that colum represents the field.
[0084] The splicing operator 320 corresponds to a storage space 321, which is used for the splicing operator 320 to store its state. Specifically, the splicing operator 320 can store the received data or the splicing result of the splicing operator 320 as the state in the storage space 321. When the splicing operator 320 receives data from the reading operator 310, the splicing operator 320 can query the data set with the same splicing key corresponding to the data in the storage space 321 (where the data set can include the data read by the reading operator from at least one data source end. And the data in the data set can refer to the data received by the splicing operator 320 historically but not yet spliced with other data, or the splicing result of the splicing operator historically splicing the data from two or more data source ends). If found, the received data and the found data set are spliced to obtain a splicing result. Then, the splicing result is stored in the storage space 321. If not found, the received data is stored in the storage space 321.
[0085] In some embodiments, as Figure 3 or Figure 4 shown, the streaming computing device 300 may include a splicing operator 320.
[0086] In some embodiments, as Figure 5 shown, the streaming computing device 300 may include multiple splicing operators 320. Among them, in the data transmission direction of the streaming computing device 300, the multiple splicing operators 320 are arranged in parallel. That is to say, in the data transmission direction, there is no sequence relationship between different splicing operators 320.
[0087] The mapping value corresponding to the splicing key of the data read by the reading operator 310 from the data source end. Different splicing operators 320 correspond to different mapping ranges. Among them, each mapping range corresponding to a splicing operator 320 includes one or more mapping values. The splicing operator 320 is used to receive the data whose mapping value is within the mapping range corresponding to the splicing operator 320 and splice the received data.
[0088] When the reading operator 310 reads data from the data source end, it can calculate the mapping value of the splicing key of the data, and then identify the mapping range to which the mapping value belongs. After identifying the mapping range to which the mapping value belongs, the reading operator 310 sends the data to the splicing operator 320 corresponding to the mapping range.
[0089] Among them, the reading operator 310 can calculate the splicing key of the data using a mapping algorithm to obtain the mapping value corresponding to the splicing key. Common mapping algorithms include the hash range partitioning algorithm, the key range partitioning algorithm, etc. Among them, the hash range partitioning algorithm includes the consistent hashing algorithm. A commonly used consistent hashing algorithm is: taking the remainder based on the total number of assignable mapping values. Specifically, the splicing key is hashed, and the obtained hash value is divided by the total number of assignable mapping values, and the obtained remainder is used as the mapping value of the splicing key.
[0090] As Figure 3 , Figure 4 or Figure 5 shown, the splicing operator 320 can send the splicing result to the output operator 330. The output operator 330 is used to output the splicing result to the data target end 500.
[0091] In some embodiments, as Figure 3 , Figure 4 or Figure 5 shown, the streaming computing device 300 can be deployed in a computing node cluster. Among them, the computing node cluster includes multiple computing nodes. The computing node can be a physical node, such as a server. The computing node can also be a virtual computing node such as a virtual machine (VM) or a container.
[0092] Different operators in the streaming computing device 300 can be deployed in the same computing node or in different computing nodes. In addition, the same operator in the streaming computing device 300 can be deployed in the same computing node, that is, the function of the operator is executed by the computing node. The same operator can also be deployed in multiple computing nodes, that is, one of the multiple different computing nodes can execute a part of the function of the operator.
[0093] Continuing to refer to Figure 3 , Figure 4 or Figure 5 , the data target end 500 is the demand side of the splicing result and is used to receive the splicing result output by the streaming computing device 300. That is to say, the data required by the data target end 500 can be spliced in real time through the streaming computing device 300. In some embodiments, the data target end 500 can be a database or a data table for storing the data output by the streaming computing device 300. In some embodiments, the data target end 500 can be a data warehouse service (DWS) to splice the data to be stored in the DWS through the streaming computing device 300. In some embodiments, the data target end 500 is a data analysis end and can perform relevant analysis based on the splicing result of the streaming computing device 300.
[0094] The above examples introduced the system architecture provided by the embodiments of the present application and the streaming computing device 300. Next, the data processing method provided by the embodiments of the present application will be described in combination with the system architecture and the streaming computing device 300.
[0095] This method can be executed by the streaming computing device 300, specifically by relevant operators in the streaming computing device 300. As Figure 6 shown, this method includes the following steps.
[0096] Step 601, the reading operator 310 sends the data E1 read from the data source end G1 among multiple data source ends to the splicing operator 320. Among them, the data E1 corresponds to the splicing key F1.
[0097] The data source end G1 is one or more of the above-mentioned multiple data source ends. For example, the data source end G1 can be any one or more of the data source ends 410, 420, 430, and 440.
[0098] The reading operator 310 reads data from the data source end. Whenever the reading operator 310 reads data from the data source end, it can send the read data to the splicing operator 320.
[0099] In some embodiments, as shown above, the streaming computing device 300 includes multiple splicing operators 320, where different splicing operators 320 correspond to different mapping ranges. The reading operator 310 can calculate the mapping value corresponding to the splicing key of the read data, identify the mapping range to which the mapping value belongs, and then send the read data to the splicing operator corresponding to the mapping range to which the mapping value belongs. That is, the reading operator 310 can calculate the mapping value corresponding to the splicing key F1, identify the mapping range to which the mapping value corresponding to the splicing key F1 belongs, and then send the data E1 to the splicing operator 320 corresponding to this mapping range.
[0100] Refer to Figure 6 In step 602, the splicing operator 320 queries the data corresponding to the splicing key F1 in the storage space 321 of the splicing operator 320, that is, queries the data with the same splicing key as the data E1 in the storage space.
[0101] Among them, the storage space 321 stores the data read by the reading operator 310 from multiple data source ends historically. Here, "historically" means before the splicing operator 320 executes step 601 this time. Among them, the splicing operator 320 can splice the data read by the reading operator 310 from different data source ends corresponding to the same splicing key, and then store the splicing result in the storage space 321. That is to say, when the data read by the reading operator 310 from different data source ends corresponds to the same splicing key, what is stored in the storage space 321 is the splicing result of the data read by the reading operator 310 from different data source ends corresponding to the same splicing key.
[0102] In some embodiments, the storage space 321 uses a data table to record data to store the data in the storage space 321. Among them, the data corresponding to the same splicing key is recorded in the same row of the data table.
[0103] Among them, in step 603, the splicing operator 320 can determine whether the data corresponding to the splicing key F1 is queried in the storage space 321.
[0104] If the judgment result of step 603 is yes, that is, the data corresponding to the splicing key F1 is queried in the storage space 321. At this time, the splicing operator 320 can splice the data E1 and the queried data, or update the queried data based on the data E1. Specifically as follows.
[0105] It can be set that the data corresponding to the splicing key F1 queried in the storage space 321 is data E2. It can be determined whether the data E1 and the data E2 are the data read by the reading operator 310 from different data source ends, that is, to determine whether the data E2 is the data read by the reading operator 310 from a data source end other than the data source end G1. If the data E2 is the data read by the reading operator 310 from a data source end other than the data source end G1, then the splicing operator 320 can, in step 604, splice the data E1 and the data E2 to obtain a splicing result. Among them, this splicing result can be called the data E3 corresponding to the splicing key F1.
[0106] In some embodiments, the data E2 can be composed of the data read by the reading operator 310 from at least two data source ends, or rather, the data E2 is obtained by splicing the data read by the reading operator 310 from at least two data source ends. Specifically, the data E2 is the splicing result of the data read by the reading operator 310 from different data source ends corresponding to the splicing key F1.
[0107] In an example, it can be set that the data E1 is as shown in Table 1, and the data E2 is as shown in Table 2.
[0108] Table 1
[0109] Splicing key A.1 A.2 Timestamp F1 11 21 202305061230
[0110] Among them, it can be set that the data E1 is read by the reading operator 310 from the data source end 410, that is, the data in column A.1 and column A.2 is read by the reading operator 310 from the data source end 410.
[0111] Table 2
[0112] Splicing key C.1 C.2 D.1 D.2 Timestamp F1 1111 2111 11111 21111 202305061230
[0113] Among them, the data in column C.1 and column C.2 is read by the reading operator 310 from the data source end 430, and the data in column D.1 and column D.2 is read by the reading operator 310 from the data source end 440.
[0114] Then the data E2 and the data E1 are spliced, and the obtained data E2' is shown in Table 3.
[0115] Table 3
[0116] Splicing key A.1 A.2 C.1 C.2 D.1 D.2 Timestamp F1 11 21 1111 2111 11111 21111 202305061230
[0117] In another example, it can be set that the data E1 is the data read by the reading operator 310 from multiple data source ends, and the splicing keys of the data read from different data source ends are the same. In step 604, the data read from multiple data source ends and the data in the data E2 can be spliced to obtain the data E3.
[0118] It can be set that the data E1 includes the data shown in Table 1 and the data shown in Table 4, where the data E2 is still shown in Table 2.
[0119] Table 4
[0120] Splicing key B.1 B.2 Timestamp F1 111 211 202305061230
[0121] Among them, it can be set that the data in column B.1 and column B.2 is read by the reading operator 310 from the data source end 420.
[0122] The data in Table 1, Table 4 and Table 2 are spliced, and the obtained data E3 is shown in Table 5.
[0123] Table 5
[0124] Splicing key A.1 A.2 B.1 B.2 C.1 C.2 D.1 D.2 Timestamp F1 11 21 111 211 1111 2111 11111 21111 202305061230
[0125] In some embodiments, the data E2 can only include the data read by the reading operator 310 from one data source end. For example, it can be set that the data E1 is still shown in Table 1 and the data E2 is shown in Table 6.
[0126] Table 6
[0127] Splicing key C.1 C.2 Timestamp F1 1111 2111 202305061230
[0128] Then, the data E3 obtained by concatenating data E1 and data E2 is shown in Table 7 as follows.
[0129] Table 7
[0130] Splicing key A.1 A.2 C.1 C.2 Timestamp F1 11 21 1111 2111 202305061230
[0131] In the above manner, the concatenation of the data in data E1 and data E2 can be completed.
[0132] When the concatenation of the data in data E1 and data E2 is completed, the concatenation operator 320 can execute step 605a to store data E3 in the storage space 321. Among them, data E3 can be stored in the storage space 321 as a state. Exemplarily, in step 605a, data E2 in the storage space 321 can be replaced with data E3. That is to say, when storing data E3 in the storage space 321, data E2 is deleted from the storage space 321, thereby reducing the amount of data stored in the storage space 321, that is, reducing the amount of data stored in the state of the concatenation operator 320. Among them, data E3 includes data E1 and data E2. By storing the concatenated data E3, the storage of data E1 and data E2 is realized.
[0133] When the concatenation of the data in data E1 and data E2 is completed, the concatenation operator 320 can execute step 605b to output data E3 to the output operator 330, so that the output operator 330 outputs data E3 to the data target end 500 to meet the requirements of the data target end 500 for streaming data.
[0134] When some or all of the data in data E1 and data E2 is read by the read operator 310 from the same data source end, that is to say, when data E2 corresponding to the concatenation key F1 is queried in the storage space 321, and some or all of the data in data E1 and data E2 is the data read by the read operator 310 from the same data source end, the concatenation operator 320 updates data E2 based on data E1. Specifically, it can be set that data E2 includes data E21. Data E21 can be a part of the data in data E2 or all of the data in data E2. Data E21 and data E1 are read by the read operator 310 from the same data source end. Among them, the reading time of data E21 is earlier, and the reading time of data E1 is later. Data E1 is the updated data E21. The data E21 in data E2 can be replaced with data E1, thereby updating data E2.
[0135] For example, it can be set that data E2 is as shown in Table 3 described above, and data E1 is as shown in Table 8.
[0136] Table 8
[0137] Splicing key A.1 A.2 Timestamp F1 11-1 21-1 202305061340
[0138] Among them, the data in column A.1 and column A.2 in Table 3 are data E21. Among them, the read timestamp of data E21 is "202305061230", and the read timestamp of data E1 is "202305061340". The read time of data E21 is earlier, and the read time of data E1 is later. The data in column A.1 and column A.2 in Table 3 can be replaced with the data in column A.1 and column A.2 in Table 8 to obtain the updated data E2. The updated data E2 is shown in Table 9.
[0139] Table 9
[0140] Splicing key A.1 A.2 C.1 C.2 D.1 D.2 Timestamp F1 11-1 21-1 1111 2111 11111 21111 202305061340
[0141] After the splicing operator 320 completes the update of data E2, it can send the updated data E2 to the output operator 330, so that the output operator 330 outputs the updated data E2 to the data target end 500 to meet the requirements of the data target end 500 for streaming data. The splicing operator 320 can also store the updated data E2 in the storage space 321 for subsequent use, such as splicing or updating.
[0142] Continue to refer to Figure 6 , when no data E2 corresponding to the same splicing key as data E1 is found in the storage space 321, the splicing operator 320 can execute step 606a to record data E1 in the storage space 321 for subsequent splicing or updating. Specifically, data E1 can be set as shown in Table 1 described above, and the data in the storage space 321 is shown in Table 10.
[0143] Table 10
[0144] Splicing key A.1 A.2 B.1 B.2 Timestamp F2 12 22 112 212 202305061230 F3 13 23 Empty Empty 202305061250 F4 Empty Empty 113 213 202305061340
[0145] Among them, the splicing key F2, the splicing key 3, and the splicing key F4 are all different from the splicing key F1, that is, the splicing keys of the data stored in the storage space 321 are all different from the splicing key of data E1. In this case, data E1 can be recorded in the storage space 321. Among them, data E1 can be recorded in Table 10 in the storage space 321 to obtain Table 11.
[0146] Table 11
[0147] Splicing key A.1 A.2 B.1 B.2 Timestamp F1 11 21 Empty Empty 202305061230 F2 12 22 112 212 202305061230 F3 13 23 Empty Empty 202305061250 F4 Empty Empty 113 213 202305061340
[0148] Continue to refer to Figure 6, when the data corresponding to the splicing key F1 is not found in the storage space 321, the splicing operator 320 can also execute step 606b to send the data E1 to the output operator 330, so that the output operator 330 sends the data E1 to the data target end 500 to meet the requirements of the data target end 500 for the streaming data.
[0149] In some embodiments, the reading operator 310 sends the data H1 read from the data source end G2 among the multiple data source ends to the splicing operator 320. Among them, the data H1 corresponds to the splicing key F5. The data source end G2 is one or more of the above-mentioned multiple data source ends. For example, the data source end G2 can be any one or more of the data source ends 410, 420, 430, and 440.
[0150] The splicing operator 320 can query the data corresponding to the splicing key F5 in the storage space 321, that is, query the data with the same splicing key corresponding to the data H1 in the storage space 321.
[0151] If there is no data corresponding to the splicing key F5 in the storage space 321, the splicing operator 320 stores the data H1 in the storage space 321 for subsequent splicing or updating.
[0152] The splicing operator 320 sends the data H1 to the output operator 330, so that the output operator 330 outputs the data H1 to the data target end 500.
[0153] In some embodiments, based on the time to live (TTL) mechanism, the data in the storage space 321 can be aged to timely delete the invalid data in the storage space 321 and reduce the amount of data stored in the state. Specifically, a life cycle can be set for the data in the storage space 321. When the life cycle of the data ends, the data can be deleted from the storage space 321 to release the available space in the storage space 321. Among them, the starting point of the life cycle can be the storage moment when the data is stored in the storage space 321, and the length of the life cycle is related to the duration when all the data to be spliced with the same splicing key reaches the splicing operator 320.
[0154] Taking the splicing key F1 as an example, the time when the data corresponding to the splicing key F1 in different data source ends reaches the splicing operator 320 may be different. Among them, it can be set that the data corresponding to the splicing key F1 in the data source end 410 reaches the splicing operator 320 first. That is to say, the time when the data corresponding to the splicing key F1 in other data source ends reaches the splicing operator 320 is later than the time when the data corresponding to the splicing key F1 in the data source end 410 reaches the splicing operator 320. It can also be set that the data corresponding to the splicing key F1 in the data source end 420 reaches the splicing operator 320 latest. That is to say, the time when the data corresponding to the splicing key F1 in other data source ends reaches the splicing operator 320 is earlier than the time when the data corresponding to the splicing key F1 in the data source end 420 reaches the splicing operator 320. Subtract the time when the data corresponding to the splicing key F1 in the data source end 420 reaches the splicing operator 320 from the time when the data corresponding to the splicing key F1 in the data source end 410 reaches the splicing operator 320, and the duration for all the data to be spliced corresponding to the splicing key F1 to reach the splicing operator 320 is obtained. Among them, for the convenience of description, the duration for all the data to be spliced corresponding to the splicing key to reach the splicing operator 320 can be called the splicing required duration of this splicing key.
[0155] Referring to the above method for calculating the splicing required duration of the splicing key F1, the splicing required durations of multiple splicing keys can be calculated. Then, based on the splicing required durations of multiple splicing keys, the length of the life cycle is set. For example, the splicing required duration of the splicing key with the longest splicing required duration among the multiple splicing keys can be used as the length of the life cycle.
[0156] To sum up, in the method provided in this application, the same splicing operator can complete the splicing of data in multiple (such as three or more) data source ends. In this way, only one splicing operator performs state storage on the data read from multiple data source ends, and the amount of data for state storage = ∑Ti, where Ti is the data read from the i-th data source end. Compared with multiple splicing operators performing two-level cascaded splicing on the data in multiple data source ends, the method provided in the embodiments of this application reduces the amount of data for state storage and saves storage overhead.
[0157] Moreover, in the method provided in this application, when the splicing operator receives data, it obtains the data corresponding to the same splicing as the received data from the storage space, and splices the data obtained from the storage space and the received data to obtain a splicing result. The splicing operator can store the splicing result in the storage space. Compared with splicing based on the data table in the storage space, the method provided in this application has less computational complexity, can save computing resources and improve splicing efficiency, so that splicing can be completed as soon as possible, and the real-time performance of splicing is improved.
[0158] In addition, no matter which data stream arrives first, the method provided in the embodiments of the present application can achieve splicing and output. There is no dependency relationship between different data streams, so that data can be output to the data target end in a timely manner, ensuring the data requirements of the data target end.
[0159] In addition, the method provided in the present application sets a life cycle for the data in the storage space. When the life cycle of the data ends, the data is deleted, so that unnecessary data can be deleted from the storage space in a timely manner, further reducing the amount of data stored in the state.
[0160] Refer to Figure 7 , the embodiments of the present application also provide a streaming computing device 700. As Figure 7 shown, the streaming computing device 700 includes a reading operator 710 and a splicing operator 720. The splicing operator 710 is used to read data from multiple data source ends; among them,
[0161] The reading operator 710 is used to: send the first data read from the first data source end among the multiple data source ends to the splicing operator, where the first data corresponds to a first splicing key;
[0162] The splicing operator 720 is used to: query the second data corresponding to the first splicing key in the storage space;
[0163] The splicing operator 720 is further used to: when the second data is read by the reading operator from a data source end other than the first data source end among the multiple data source ends, splice the first data and the second data to obtain the third data corresponding to the first splicing key;
[0164] The splicing operator 720 is further used to: store the third data in the storage space.
[0165] In some embodiments, the second data is the splicing result of at least two data. Each of the at least two data corresponds to the first splicing key, and different data among the at least two data are read by the reading operator 710 from different data source ends.
[0166] In some embodiments, the streaming computing device 700 further includes an output operator 730 connected to the data target end; among them, the splicing operator 720 is further used to: send the third data to the output operator 730, so that the output operator 730 outputs the third data to the data target end.
[0167] In some embodiments, the splicing operator 720 is further configured to: when some or all of the first data and the second data are read by the reading operator from the same data source end, update the second data based on the first data; and store the updated second data in the storage space.
[0168] In some embodiments, the reading operator 710 is further configured to: send the fourth data read from the second data source end among multiple data source ends to the splicing operator, where the fourth data corresponds to a second splicing key; the splicing operator 720 is further configured to: when the splicing operator 720 does not query the data corresponding to the second splicing key in the storage space, store the fourth data in the storage space.
[0169] In some embodiments, the streaming computing device 700 includes a plurality of splicing operators, where different splicing operators correspond to different mapping ranges; the reading operator 710 is configured to: when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, send the first data to the splicing operator 720.
[0170] In some embodiments, the data in the storage space has a life cycle; the splicing operator 720 is further configured to: delete the data with the ended life cycle in the storage space.
[0171] Among them, the reading operator 710, the splicing operator 720, and the output operator 730 can all be implemented by software or by hardware. Exemplarily, next, taking the reading operator 710 as an example, the implementation manner of the reading operator 710 will be introduced. Similarly, the implementation manners of the splicing operator 720 and the output operator 730 can refer to the implementation manner of the reading operator 710.
[0172] As an example of a software functional unit, the reading operator 710 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the reading operator 710 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same available zone AZ or in different AZs, and each AZ includes a data center or multiple geographically proximate data centers. Among them, generally one region may include multiple AZs.
[0173] Similarly, multiple hosts / virtual machines / containers for running the code can be distributed within the same VPC or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, communication gateways need to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateways.
[0174] As an example of a hardware functional unit, the reading operator 710 may include at least one computing device, such as a server. Alternatively, the reading operator 710 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0175] The multiple computing devices included in the reading operator 710 can be distributed in the same region or in different regions. The multiple computing devices included in the reading operator 710 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the reading operator 710 can be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0176] This application also provides a computing device 800. As Figure 8 shown, the computing device 800 includes: a bus 802, a processor 804, a memory 806, and a communication interface 808. The processor 804, the memory 806, and the communication interface 808 communicate with each other through the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 800.
[0177] The bus 802 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only one line is used in Figure 8 , but it does not mean that there is only one bus or one type of bus. The bus 802 can include a path for transmitting information between various components of the computing device 800 (for example, the memory 806, the processor 804, and the communication interface 808).
[0178] The processor 804 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0179] The memory 806 can include volatile memory, such as random access memory (RAM). The memory 806 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0180] The executable program code is stored in the memory 806, and the processor 804 executes the executable program code to respectively implement the functions of the aforementioned reading operator 710, stitching operator 720, and output operator 730, thereby implementing Figure 6 the method shown. That is, the instructions for executing the method shown are stored on the memory 806. Figure 6 shown in the method.
[0181] The communication interface 808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.
[0182] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0183] As Figure 9 shown, the computing device cluster includes at least one computing device 800. Instructions for executing the Figure 6 shown method can be stored in the memory 806 of one or more of the computing devices 800 in the computing device cluster.
[0184] In some possible implementation manners, instructions for executing the Figure 6 shown method can also be stored separately in the memory 806 of one or more of the computing devices 800 in the computing device cluster. In other words, a combination of one or more computing devices 800 can jointly execute the instructions for executing the Figure 6 shown method.
[0185] It should be noted that the memories 806 in different computing devices 800 in the computing device cluster can store different instructions, respectively for executing partial functions of the streaming computing device 500. That is, the instructions stored in the memories 806 of different computing devices 800 can implement the functions of one or more modules among the read operator 710, the splicing operator 720, and the output operator 730.
[0186] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 10 shows a possible implementation manner. As Figure 10 shown, two computing devices 800A and 800B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the function of the read operator 710 are stored in the memory 806 of the computing device 800A. At the same time, instructions for executing the functions of the splicing operator 720 and the output operator 730 are stored in the memory 806 of the computing device 800B.
[0187] It should be understood that Figure 10 the function of the computing device 800A shown in
[0188] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to Figure 9 and Figure 10 the connection mode of the computing device cluster. The difference is that the same instructions for executing Figure 6 the method shown may be stored in the memory 806 of one or more computing devices 800 in the computing device cluster.
[0189] In some possible implementation manners, the memory 806 of one or more computing devices 800 in the computing device cluster may also store partial instructions for executing Figure 6 the method shown respectively. In other words, the combination of one or more computing devices 800 can jointly execute the instructions for executing Figure 6 the method shown.
[0190] The embodiments of the present application also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute Figure 6 the method shown.
[0191] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a host migration device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions, and the instructions instruct the computing device to execute Figure 6 the method shown.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data splicing method, characterized in that: The method is applied to a stream computing device, the device comprising a splicing operator and a reading operator for reading data from multiple data source ends; the method comprises: The read operator reads first data from a first data source end among the multiple data source ends and sends the first data to the splicing operator, wherein the first data corresponds to a first splicing key; The concatenation operator searches for second data corresponding to the first concatenation key in the storage space; When the second data is read by the read operator from a data source end other than the first data source end among the multiple data source ends, the concatenation operator concatenates the first data and the second data to obtain third data corresponding to the first concatenation key; The concatenation operator stores the third data in the storage space.
2. The method according to claim 1, characterized in that The second data is a concatenation result of at least two data, each of the at least two data corresponds to the first concatenation key, and different data in the at least two data are read by the read operator from different data sources.
3. The method according to claim 1 or 2, characterized in that: The device also includes an output operator connected to the data target end; the method also includes: the splicing operator sends the third data to the output operator, so that the output operator outputs the third data to the data target end.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: When part or all of the first data and the second data are read by the read operator from the same data source, the concatenation operator updates the second data based on the first data; The concatenation operator stores the updated second data in the storage space.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The read operator reads fourth data from a second data source end among the multiple data source ends and sends the fourth data to the splicing operator, wherein the fourth data corresponds to the second splicing key; When the concatenation operator does not find data corresponding to the second concatenation key in the storage space, the concatenation operator stores the fourth data in the storage space.
6. The method according to any one of claims 1 to 5, characterized in that The device includes a plurality of splicing operators, wherein different splicing operators correspond to different mapping ranges; The reading operator reads first data from a first data source end among multiple data source ends and sends the first data to the splicing operator, including: When the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator, the read operator sends the first data to the splicing operator.
7. The method according to any one of claims 1 to 6, characterized in that The data in the storage space has a life cycle; the method further includes: the concatenation operator deleting data whose life cycle has ended in the storage space.
8. A streaming computing device, characterized in that: The device includes a splicing operator and a reading operator for reading data from multiple data source ends; wherein, The read operator is used to: read first data from a first data source end among the multiple data source ends and send it to the splicing operator, wherein the first data corresponds to a first splicing key; The concatenation operator is used to: query the storage space for the second data corresponding to the first concatenation key; The concatenation operator is further used for: when the second data is read by the read operator from a data source end other than the first data source end among the multiple data source ends, concatenating the first data and the second data to obtain third data corresponding to the first concatenation key; The concatenation operator is further used to: store the third data into the storage space.
9. The device according to claim 8, characterized in that The second data is a concatenation result of at least two data, each of the at least two data corresponds to the first concatenation key, and different data in the at least two data are read by the read operator from different data sources.
10. The device according to claim 8 or 9, characterized in that The device also includes an output operator connected to the data target end; wherein the splicing operator is further used to: send the third data to the output operator, so that the output operator outputs the third data to the data target end.
11. The device according to any one of claims 8 to 10, characterized in that The concatenation operator is also used to: When part or all of the first data and the second data are read by the read operator from the same data source, the second data is updated based on the first data; The updated second data is stored in the storage space.
12. The device according to any one of claims 8 to 11, characterized in that The read operator is further used to: read fourth data from a second data source end among the multiple data source ends and send it to the splicing operator, wherein the fourth data corresponds to the second splicing key; The concatenation operator is further used for: when the concatenation operator fails to find data corresponding to the second concatenation key in the storage space, storing the fourth data in the storage space.
13. The device according to any one of claims 8 to 12, characterized in that The device includes a plurality of splicing operators, wherein different splicing operators correspond to different mapping ranges; The read operator is used to send the first data to the splicing operator when the mapping value corresponding to the first splicing key belongs to the mapping range corresponding to the splicing operator.
14. The device according to any one of claims 8 to 13, characterized in that The data in the storage space has a life cycle; the splicing operator is further used to delete data whose life cycle has ended in the storage space.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 7.