Data stream writing control method, system, device, equipment and readable storage medium

By adjusting the operator routing table of the data stream processing engine and optimizing the write path of the data stream, the problem of poor data stream writing performance in traditional methods is solved, and more efficient data stream writing is achieved.

CN119576250BActive Publication Date: 2025-06-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510140576.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-03
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

Traditional data stream writing methods cause data streams to be stored in multiple small files in the file system, resulting in low read efficiency, high write cost and poor performance.

Method used

By receiving the write performance characteristics of the data stream processing engine, adjusting the operator routing table to optimize the write path of the data stream, dynamically adjusting the number of data write operators used to write the data stream during each cycle, avoiding the generation of unnecessary small files.

Benefits of technology

While ensuring data read and write efficiency, it reduces the cost and delay of data stream writing and improves data stream writing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576250B_ABST
    Figure CN119576250B_ABST
Patent Text Reader

Abstract

The present application relates to a data stream writing control method, system, device, equipment, readable storage medium and program product. It includes: receiving the writing performance characteristics of the data stream processing engine in the first cycle; adjusting the first operator routing table of the first cycle according to the writing performance characteristics to obtain the second operator routing table of the second cycle, where the second operator routing table contains routing records within the second cycle, and the routing records indicate a data stream to be written in the data stream processing engine and a corresponding data writing operator; sending the second operator routing table to the data stream processing engine, and the second operator routing table is used to instruct the data stream processing engine in the second cycle to use the data writing operators indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written. Using this method can effectively improve the data stream writing performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data stream writing control method, system, device, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] A data stream refers to a sequence of data with the same data type and originating from the same data source. For example, the log data of a video application each time it plays a video, the log data of a shared device each time it is used, and the log data of a terminal device each time it is started.

[0003] In traditional technologies, most often, all data writing operators in a data writing engine are first used to write a data stream into an intermediate file system. When each data writing operator writes the data stream into the intermediate file system, a corresponding small file is generated, so that the same data stream is stored in the intermediate file system in the form of multiple small files. Then, the small files of the same data stream in the intermediate file system are integrated, and the large file corresponding to the integrated data stream is stored in the final file system.

[0004] However, although this method solves the problem that the same data stream is stored in the file system in the form of multiple small files, resulting in poor data reading efficiency. However, setting up the intermediate file system and integrating the small files of the same data stream in the intermediate file system not only makes the writing cost of the data stream relatively high, but also makes the writing efficiency of the data stream relatively low, thereby resulting in poor data stream writing performance. Summary of the Invention

[0005] Based on this, it is necessary to provide a data stream writing control method, system, device, computer device, computer-readable storage medium, and computer program product for the above technical problems. While ensuring the data reading and writing efficiency, it can effectively improve the data stream writing performance.

[0006] In a first aspect, this application provides a data stream writing control method, including:

[0007] Receiving the writing performance characteristics of a data stream processing engine in a first cycle, where the writing performance characteristics characterize the writing performance of the data stream processing engine for writing each data stream in the first cycle;

[0008] According to the writing performance characteristics, adjusting a first operator routing table in the first cycle to obtain a second operator routing table in a second cycle, where the first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle, and each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator;

[0009] Send the second operator routing table to the data stream processing engine. The second operator routing table is used to instruct the data stream processing engine to use the data writing operators indicated by each routing record in the second operator routing table to write the data into the corresponding data stream to be written during the second cycle, and write the data into the file corresponding to the data writing operator and the data stream to be written.

[0010] In a second aspect, the present application further provides a data stream writing control method, including:

[0011] Obtain the first operator routing table of the first cycle. The first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator;

[0012] Use the data writing operators indicated by each routing record to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written;

[0013] For each data stream to be written, obtain the writing performance characteristics of writing through at least one data writing operator during the first cycle;

[0014] Send the writing performance characteristics. The sent writing performance characteristics are used to indicate that the routing records in the first operator routing table are adjusted according to the writing performance characteristics to obtain the second operator routing table of the second cycle. The writing performance characteristics characterize the writing performance of writing each data stream during the first cycle.

[0015] In a third aspect, the present application further provides a data stream writing control system, including a control node, a data stream processing engine, a file system, and a cache. The data stream processing engine includes a data writing operator and a data source operator;

[0016] The data source operator is used to receive the first operator routing table of the first cycle from the control node. The first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator; read the data segments of each data stream from the cache, and allocate the data segments to each data writing operator according to the routing records in the operator routing table; wherein, the writing performance characteristics are obtained and sent by the data writing operator;

[0017] The data writing operator is used to write the data stream to be written corresponding to the data writing operator indicated by the routing record into the file corresponding to the data writing operator and the data stream to be written in the file system; obtain the writing performance characteristics of writing through the data writing operator during the first cycle; send the writing performance characteristics to the control node;

[0018] The control node is configured to receive the write performance characteristics of the data writing operator in the first cycle. The write performance characteristics characterize the write performance of the data stream processing engine for writing each data stream in the first cycle. According to the write performance characteristics, the first operator routing table in the first cycle is adjusted to obtain a second operator routing table in the second cycle. The first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator. The second operator routing table is sent to the data source operator. The operator routing table is used to indicate that in the second cycle, the data stream processing engine uses the data writing operator indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data writing operator into the file system, specifically into the file corresponding to the data writing operator and the data stream to be written.

[0019] In a fourth aspect, the present application further provides a data stream writing control device, including:

[0020] A receiving module, configured to receive the write performance characteristics of the data stream processing engine in the first cycle. The write performance characteristics characterize the write performance of the data stream processing engine for writing each data stream in the first cycle.

[0021] A first execution module, configured to adjust the first operator routing table in the first cycle according to the write performance characteristics to obtain a second operator routing table in the second cycle. The first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator.

[0022] A sending module, configured to send the second operator routing table to the data stream processing engine. The second operator routing table is used to indicate that in the second cycle, the data stream processing engine uses the data writing operator indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written.

[0023] In one embodiment, the first execution module is specifically configured to determine the write performance metric values of the respective data streams written in the first cycle according to the write performance characteristics; for each data stream written in the first cycle, if the write performance metric value is greater than the first threshold, perform routing expansion based on the existing routing record indicating the data stream in the first operator routing table of the first cycle to generate an expanded routing record indicating the data stream; and generate the second operator routing table of the second cycle according to the existing routing record and the expanded routing record.

[0024] In one embodiment, the first execution module is further configured to obtain the number of operators of the data write operator on the data stream processing engine, and obtain the concurrent routing upper limit value of each data write operator on the data stream processing engine; determine the concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value; the first execution module is specifically configured to, for each data stream written in the first cycle, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value and the write performance metric value is greater than the first threshold, generate an expanded routing record indicating the data stream based on the existing routing record indicating the data stream in the first operator routing table of the first cycle.

[0025] In one embodiment, the first execution module is specifically configured to, for each data stream written in the first cycle, if the write performance metric value is greater than the first threshold, determine the number of expanded routing records for the data stream according to a preset expansion strategy; determine the candidate data write operators that have no routing records established with the data stream according to the existing routing record indicating the data stream in the first operator routing table of the first cycle; for each candidate data write operator, count the number of routing records corresponding to the candidate data write operator according to the routing records in the first operator routing table; select the number of new data write operators that is the same as the number of expanded routing records from the candidate data write operators in ascending order of the priority of the number of routing records; and respectively establish routing records indicating each new data write operator and the data stream to obtain an expanded routing record indicating the data stream.

[0026] In one embodiment, the first execution module is specifically configured to, for each data stream written in the first cycle, if the write performance metric value is greater than the first threshold and an expanded routing record indicating the data stream has been generated for the first cycle, determine the timing sequence of this routing expansion in the continuous routing expansion operations; determine the number of expanded routing records for the data stream according to the timing sequence, where the number of expanded routing records is positively correlated with the timing sequence; and generate an expanded routing record indicating the data stream based on the existing routing record indicating the data stream in the first operator routing table of the first cycle and according to the number of expanded routing records.

[0027] In one embodiment, the first execution module is specifically configured to determine, according to the write performance characteristics, the write performance metric values of each data stream written in the first cycle; for each data stream written in the first cycle, if the write performance metric value is less than or equal to a second threshold, perform routing reduction based on the existing routing record indicating the data stream in the first operator routing table of the first cycle, and determine a reduced routing record indicating the data stream; generate a second operator routing table for the second cycle according to the existing routing record and the reduced routing record.

[0028] In one embodiment, the first execution module is specifically configured to, for each data stream written in the first cycle, if the write performance metric value is less than or equal to a second threshold, determine the number of reduced routing records for the data stream according to a preset reduction strategy; determine, according to the existing routing record indicating the data stream in the first operator routing table of the first cycle, candidate data write operators that have established routing records with the data stream; for each of the candidate data write operators, count the number of routing records corresponding to the candidate data write operator according to the routing record in the first operator routing table; select, from the candidate data write operators, reduced data write operators with the same number as the number of reduced routing records in descending order of the priority of the number of routing records, and obtain a reduced routing record indicating the data stream and the reduced data write operators.

[0029] In one embodiment, the first execution module is specifically configured to determine a multi-cycle period to be statistically analyzed, where the multi-cycle period includes the first cycle and at least one historical cycle before the first cycle; obtain the write performance characteristics of each of the at least one historical cycle; perform statistical analysis of the multi-cycle period according to the write performance characteristics of each of the multi-cycle periods, and obtain the write performance metric values of each data stream written in the first cycle.

[0030] In one embodiment, each data stream includes multiple data segments. The write performance characteristics include a publish timestamp and a consumption timestamp for each data segment. The publish timestamp represents the time when the data segment is written to the cache, and the consumption timestamp represents the time when the data segment read from the cache is written to the file. The write performance metric values of each data stream written in the first cycle period include the data segment latency statistical values for multiple cycle periods. The first execution module is specifically configured to determine the latency value of each data segment according to the publish timestamp and the consumption timestamp of each data segment written in each cycle period within the multiple cycle periods; for each data stream, determine the total latency value of the data stream according to the latency values of the data segments in the data stream written in the multiple cycle periods; for each data stream, count the total number of data segments in the data stream written in the multiple cycle periods; for each data stream, determine the data segment latency statistical value of the data stream in the multiple cycle periods according to the total latency value of the data stream and the total number of data segments in the data stream.

[0031] In one embodiment, the write performance characteristics include the amount of data written to each data stream in each cycle period. The write performance metric values of each data stream written in the first cycle period include the data stream load statistical values for multiple cycle periods. The first execution module is specifically configured to, for each data stream written in the first cycle period, count the total amount of data written to the data stream in the multiple cycle periods according to the amount of data written to the data stream in each cycle period; for each data stream written in the first cycle period, count the number of routing records indicating the data stream from the first operator routing table; for each data stream written in the first cycle period, determine the data stream load statistical value of the data stream in the multiple cycle periods according to the total amount of data, the number of routing records, and the number of cycle periods in the multiple cycle periods.

[0032] In one embodiment, the first execution module is further configured to scan the data streams newly added by the data stream processing engine in the first cycle period; determine the default number of routing records of the newly added data streams according to the default operator allocation policy; count the number of routing records of each data write operator on the data stream processing engine according to the routing records in the first operator routing table; select target data write operators with the same number as the default number of routing records in ascending order of the number of routing records from the data write operators on the data stream processing engine; add target routing records indicating the target data write operators and the data streams to the second operator routing table in the second cycle period.

[0033] In one embodiment, the first execution module is specifically configured to traverse the data streams indicated by the routing records in the first operator routing table; in response to the traversed data stream being in a locked state, continue to traverse the next data stream; in response to the traversed data stream being in an unlocked state, adjust the routing records for the traversed data stream according to the write performance characteristics; until the traversal is complete, obtain the second operator routing table for the second cycle.

[0034] In a fifth aspect, the present application further provides a data stream writing control device, including:

[0035] A first acquisition module, configured to acquire a first operator routing table for a first cycle, where the first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator;

[0036] A second execution module, configured to use the data writing operator indicated by each routing record to write the data stream to be written corresponding to the data writing operator into a file corresponding to the data writing operator and the data stream to be written;

[0037] A second acquisition module, configured to, for each data stream to be written, acquire write performance characteristics for writing in the first cycle through at least one data writing operator;

[0038] A sending module, configured to send the write performance characteristics, and the sent write performance characteristics are used to indicate that the routing records in the first operator routing table are adjusted according to the write performance characteristics to obtain a second operator routing table for a second cycle, and the write performance characteristics characterize the write performance of writing each data stream in the first cycle.

[0039] In one embodiment, the first acquisition module is specifically configured to use a data source operator to receive the first operator routing table for the first cycle from a control node; the first acquisition module is further configured to use the data source operator to read data segments of each data stream from a cache, and allocate the data segments to each data writing operator according to the routing records in the first operator routing table; wherein, the write performance characteristics are acquired and sent by the data writing operator.

[0040] In a sixth aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method in any one of the embodiments in the first aspect are implemented.

[0041] In a seventh aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the embodiments in the first aspect are implemented.

[0042] In an eighth aspect, the present application further provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method in any one of the above first aspects.

[0043] For the above data stream writing control method, system, device, computer equipment, computer-readable storage medium and computer program product, first receive the writing performance characteristics of the data stream processing engine in the first cycle, where the writing performance characteristics characterize the writing performance of the data stream processing engine for writing each data stream in the first cycle; then, according to the writing performance characteristics, adjust the first operator routing table in the first cycle to obtain the second operator routing table in the second cycle, where the first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle, and each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator; then, send the second operator routing table to the data stream processing engine, and the second operator routing table is used to instruct the data stream processing engine in the second cycle to use the data writing operator indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written. The data stream writing control method provided by the present application dynamically adjusts the data writing operator used to write the data stream in each cycle according to the writing performance characteristics of the data stream. Since not all data writing operators are used to write the same data stream, but an appropriate number of data writing operators are used to write the data stream, it makes the same data stream correspond to as few small files as possible, without setting up an additional intermediate file system and performing small file integration processing, while ensuring the data reading and writing efficiency, reducing the cost of data stream writing, improving the writing efficiency of the data stream, and thus effectively improving the writing performance of the data stream. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is an application environment diagram of the data stream writing control method in an embodiment;

[0046] Figure 2 It is a flowchart of the data stream writing control method in an embodiment;

[0047] Figure 3Schematic diagram of the first operator routing table in an embodiment;

[0048] Figure 4 Schematic diagram of the second operator routing table in an embodiment;

[0049] Figure 5 Flowchart of routing capacity reduction and routing capacity expansion in an embodiment;

[0050] Figure 6 Schematic flowchart of the data stream writing control method in another embodiment;

[0051] Figure 7 Schematic flowchart of the data stream writing control method in another embodiment;

[0052] Figure 8 Schematic diagram of the data stream processing engine in an embodiment;

[0053] Figure 9 Schematic diagram of the data stream processing engine in another embodiment;

[0054] Figure 10 Schematic diagram of the data stream processing engine in another embodiment;

[0055] Figure 11 Schematic diagram of the intermediate file system in an embodiment;

[0056] Figure 12 Schematic diagram of the data stream writing control system in an embodiment;

[0057] Figure 13 Structural block diagram of the data stream writing control device in an embodiment;

[0058] Figure 14 Structural block diagram of the data stream writing control device in another embodiment;

[0059] Figure 15 Internal structure diagram of a computer device in an embodiment;

[0060] Figure 16 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] In order to describe the technical solution of the present application clearly and facilitate the understanding of the technical solution of the present application, the key concepts involved in the present application will be explained below.

[0063] Data stream refers to a data sequence with the same data type and originating from the same data source.

[0064] Data stream processing engine refers to an engine that has the function of processing data streams and can process data streams to write them into the file system.

[0065] Cycle period refers to a time period with a fixed duration set when the data stream processing engine performs the write operation for the data stream.

[0066] The first cycle period refers to any one of the multiple cycle periods when the data stream processing engine performs the write operation for the data stream.

[0067] The second cycle period refers to a cycle period after the first cycle period.

[0068] Write performance characteristics characterize the write performance of the data stream processing engine for writing each data stream in the first cycle period.

[0069] The operator routing table contains routing records, and each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data write operator.

[0070] Data write operator refers to an operator set in the data stream processing engine for writing data streams.

[0071] The file corresponding to the data write operator and the data stream to be written refers to the file created when the data stream to be written is first written into the file system by the data write operator.

[0072] Write performance metric value refers to a specific value used to quantify the performance of the data stream processing engine when writing each data stream in the first cycle period.

[0073] Routing expansion refers to the operation of adding new routing records on the basis of the original routing records of the data stream.

[0074] Concurrent routing upper limit value refers to the maximum value of the routes that the data write operator can concurrently perform.

[0075] Concurrent routing constraint value refers to the maximum value of the routes that the data stream processing engine can concurrently perform.

[0076] Timing order refers to the cumulative number of times that the current routing expansion occurs in consecutive routing expansion operations.

[0077] The data stream write control method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the control node in the computer device can be communicatively connected to the data stream processing engine. The control node and the data stream processing engine can be set in the same computer device or in different computer devices.

[0078] Exemplarily, the computer device can receive the write performance characteristics of each data stream written by the data stream processing engine in the first cycle, and adjust the first operator routing table in the first cycle according to the write performance characteristics to obtain the second operator routing table in the second cycle. The first operator routing table contains routing records in the first cycle, and the second operator routing table contains routing records in the second cycle. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data write operator.

[0079] Furthermore, the computer device can send the second operator routing table to the data stream processing engine to instruct the data stream processing engine to use the data write operators indicated by each routing record in the second operator routing table in the second cycle to write the data stream to be written corresponding to the data write operator into the file corresponding to the data write operator and the data stream to be written.

[0080] Among them, the computer device can be a terminal or a server. The terminal can be, but is not limited to, various personal computers, laptop computers, and tablet computers. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0081] In an exemplary embodiment, as Figure 2 shown, a data stream writing control method is provided. The method includes the following steps:

[0082] Step 201: Receive the write performance characteristics of the data stream processing engine in the first cycle.

[0083] A data stream refers to a data sequence with the same data type and originating from the same data source. Exemplarily, the data stream can be a log file type data stream, a user interaction type data stream, a sensor network type data stream, etc. Specifically, log files exist in various systems and application programs. For example, server log files can record user access to websites, and database log files can record the operation history of databases. These log files all exist in the form of data streams.

[0084] An engine refers to a key component or program module with specific functions that can process inputs and generate outputs to support the operation of specific tasks. A data stream processing engine refers to an engine with data stream processing capabilities that can process data streams to write the data streams into a file system.

[0085] Exemplarily, the data stream processing engine can be Apache Flink. Apache Flink is an open-source stream processing framework developed by the Apache Software Foundation. Its core is a distributed stream data engine written in Java and Scala. Apache Flink executes any stream data program in a data-parallel and pipelined manner. The pipelined runtime system of Apache Flink can execute batch processing and stream processing programs, and the runtime itself of Apache Flink also supports the execution of iterative algorithms.

[0086] The cycle period refers to a time period with a fixed duration set when the data stream processing engine executes the write operation for the data stream. The duration of each cycle period can be pre-set by technicians according to actual needs. The duration of each cycle period can be 1 minute, 10 minutes, 30 minutes, 1 hour, 5 hours, 10 hours, 24 hours, etc.

[0087] Exemplarily, assuming the duration of each cycle period is 1 minute, then starting from when the data stream processing engine begins to execute the write operation for the data stream, it is the start of the first cycle period. After 1 minute, it is the end of the first cycle period and the start of the second cycle period.

[0088] The first cycle period refers to any one of the multiple cycle periods when the data stream processing engine executes the write operation for the data stream.

[0089] The write performance characteristics of the first cycle period characterize the write performance of the data stream processing engine for writing each data stream in this first cycle period. The write performance characteristics can specifically include the data stream identifier (dataFlowId), the data write operator identifier (Sink taskIndex), the cycle period duration, the consumption time, the publishing time, the total data volume, and the total number of data items.

[0090] Among them, the data in the data stream is first stored in the cache, and then the data stream processing engine fetches it from the cache for writing. The consumption time refers to the time when the data in the data stream is written, and the publishing time refers to the time when the data in the data stream is cached. The total data volume refers to the total data volume of each data stream in the first cycle period, and the total number of data items refers to the total number of data items of each data stream in the first cycle period.

[0091] Further, if the data in a data stream is written by a data writing operator in a data stream processing engine, the total data volume of the data stream can be characterized by the total size. If the data in a data stream is written by multiple data writing operators in a data stream processing engine, the total data volume of the data stream can be characterized by the total number of fields.

[0092] In some exemplary embodiments, after the end of the first cycle period, the data stream processing engine can actively report the writing performance characteristics of the first cycle period to the computer device.

[0093] Further, after the data stream processing engine actively reports the writing performance characteristics of the first cycle period to the computer device, the computer device can receive the writing performance characteristics of the data processing engine in the first cycle period.

[0094] In some other exemplary embodiments, the computer device can also first send a writing performance characteristic acquisition request to the data stream processing engine. After the data stream processing engine receives the writing performance characteristic acquisition request sent by the computer device, it will send the writing performance characteristics of the first cycle period to the computer device after the end of the first cycle period.

[0095] Further, after the data stream processing engine sends the writing performance characteristics of the first cycle period to the computer device, the computer device can receive the writing performance characteristics of the data processing engine in the first cycle period.

[0096] Step 202: According to the writing performance characteristics, adjust the first operator routing table of the first cycle period to obtain the second operator routing table of the second cycle period.

[0097] The second cycle period refers to a cycle period after the first cycle period. For example, if the first cycle period is the first cycle period, the second cycle period is the second cycle period. If the first cycle period is the fifth cycle period, the second cycle period is the sixth cycle period.

[0098] The operator routing table contains routing records, and each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator. That is, the operator routing table is used to indicate the mapping relationship between the data stream and the data writing operator for writing the data stream. The operator routing table can specifically be the Sink routing table in the data stream processing engine.

[0099] The first operator routing table can be as Figure 3 shown. The first operator routing table contains routing records within the first cycle period, and the routing records within the first cycle period indicate a data stream to be written in the data stream processing engine during the first cycle period and a corresponding data writing operator.

[0100] The second operator routing table can be as follows Figure 4 As shown, the second operator routing table contains routing records within a second cycle. The routing records within the second cycle indicate a data stream to be written and a corresponding data writing operator in the data stream processing engine during the second cycle.

[0101] The data writing operator refers to an operator set in the data stream processing engine for writing data streams. The data writing operator can be a Sink. A Sink is a computing unit in the data stream processing engine for writing data streams. Multiple Sinks can be set in the data stream processing engine, and each Sink has an identification number.

[0102] In some exemplary embodiments, after receiving the writing performance characteristics of the data stream processing engine in the first cycle, the computer device can adjust the first operator routing table in the first cycle according to the writing performance characteristics to obtain the second operator routing table in the second cycle.

[0103] Specifically, since the writing performance characteristics in the first cycle can characterize the writing performance of the data stream processing engine for writing each data stream, and the data stream processing engine writes the data stream using the data writing operator in the first cycle. Therefore, when the writing performance characteristics in the first cycle indicate that the writing performance of a certain data stream is poor, the data writing operator corresponding to the data stream can be adjusted to improve the writing performance of the data stream.

[0104] For example, after receiving the writing performance characteristics of the data stream processing engine in the first cycle, if the computer device determines that the writing performance of data stream a is poor in the first cycle according to the writing performance characteristics, the data writing operator corresponding to data stream a can be adjusted. Specifically, assuming that the data volume of data stream a is large and it is written by data writing operator x in the first cycle, adjusting the data writing operator corresponding to data stream a can be to make data writing operator x and data writing operator y jointly write data stream a in the second cycle, that is, adding a routing record of data writing operator y and data stream a in the first operator routing table to obtain the second operator routing table. The above process is only the adjustment process for a certain data stream. In practical applications, the above process can be executed for each data stream to obtain the second operator routing table.

[0105] Further, in addition to the case of adding routing records, there is also the case of deleting routing records. For example, the amount of data written in the N cycle periods before the first cycle period of data stream a is large, so that in the N+1 operator routing table in the N+1 cycle period, there are multiple routing records corresponding to this data stream a, such as data stream a - data write operator x, data stream a - data write operator y, data stream a - data write operator z.

[0106] After the computer device receives the write performance characteristics of the data stream processing engine in the first cycle period, it determines according to the write performance characteristics that the amount of data of data stream a in the first cycle period has decreased significantly, that is, there is no need to use more data write operators to write this data stream a, so the data write operators corresponding to this data stream a can be adjusted.

[0107] Specifically, in the first operator routing table, the routing records of data stream a can be deleted, such as deleting the routing records of "data stream a - data write operator x, data stream a - data write operator y", to obtain the second operator routing table. The above process is only the adjustment process for a certain data stream. In actual applications, the above process can be executed for each data stream to obtain the second operator routing table.

[0108] Step 203: Send the second operator routing table to the data stream processing engine.

[0109] The second operator routing table is used to instruct the data stream processing engine in the second cycle period to use the data write operators indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data write operator into the file corresponding to the data write operator and the data stream to be written.

[0110] The file corresponding to the data write operator and the data stream to be written refers to the file created when the data stream to be written is first written into the file system by the data write operator.

[0111] Exemplarily, the file system can be a distributed file system (HDFS, Hadoop Distributed FileSystem).

[0112] In some exemplary embodiments, after the computer device adjusts the first operator routing table in the first cycle period according to the write performance characteristics to obtain the second operator routing table in the second cycle period, it can send the second operator routing table to the data stream processing engine.

[0113] Further, after the data stream processing engine receives the second cycle table sent by the computer device, it can use the data writing operator indicated by each routing record in the second operator routing table to write the data into the data stream to be written corresponding to the data writing operator, and write it into the file corresponding to the data writing operator and the data stream to be written.

[0114] For the above data stream writing control method, first, receive the writing performance characteristics of the data stream processing engine in the first cycle, where the writing performance characteristics characterize the writing performance of the data stream processing engine for writing each data stream in the first cycle; then, according to the writing performance characteristics, adjust the first operator routing table in the first cycle to obtain the second operator routing table in the second cycle. The first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator; then, send the second operator routing table to the data stream processing engine. The second operator routing table is used to instruct the data stream processing engine in the second cycle to use the data writing operator indicated by each routing record in the second operator routing table to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written. The data stream writing control method provided in this application dynamically adjusts the data writing operator used to write the data stream in each cycle according to the writing performance characteristics of the data stream. Since not all data writing operators are used to write the same data stream, but an appropriate number of data writing operators are used to write the data stream, it makes the same data stream correspond to as few small files as possible, without setting an additional intermediate file system and performing small file integration processing. While ensuring the data reading and writing efficiency, it reduces the cost of data stream writing, improves the writing efficiency of the data stream, and thus effectively improves the writing performance of the data stream.

[0115] In an exemplary embodiment, adjusting the first operator routing table in the first cycle according to the writing performance characteristics to obtain the second operator routing table in the second cycle includes: determining the respective writing performance metric values of each data stream written in the first cycle according to the writing performance characteristics; for each data stream written in the first cycle, if the writing performance metric value is greater than the first threshold, perform routing expansion based on the existing routing record indicating the data stream in the first operator routing table in the first cycle to generate an expanded routing record indicating the data stream; and generate the second operator routing table in the second cycle according to the existing routing record and the expanded routing record.

[0116] The write performance metric value refers to a specific value used to quantify the performance of a data stream processing engine when writing to each data stream during the first cycle. This write performance metric value can be a write speed metric value, a throughput metric value, a latency metric value, or a resource utilization metric value.

[0117] The first threshold can be pre-set by a technician according to actual needs and is used to compare with the write performance metric value of the data stream to determine whether to perform a routing expansion operation on the routing of the data stream.

[0118] Routing expansion refers to the operation of adding new routing records based on the existing routing records of the data stream.

[0119] In some exemplary embodiments, after receiving the write performance characteristics of the data stream processing engine in the first cycle, the computer device can determine the write performance metric value of each data stream written in the first cycle according to the write performance characteristics. Specifically, the computer device can use a preset algorithm and the write performance characteristics to determine the write performance metric value of each data stream written in the first cycle.

[0120] Furthermore, after the computer device determines the write performance metric value of each data stream written in the first cycle, for each data stream written in the first cycle, if the computer device determines that the write performance metric value of the data stream is greater than the first threshold, it performs routing expansion based on the existing routing record indicating the data stream in the first operator routing table of the first cycle to generate an expanded routing record indicating the data stream.

[0121] Specifically, for the data stream a written in the first cycle, if the computer device determines that the write performance metric value of the data stream a is greater than the first threshold, it first determines the routing record of the data stream a in the first operator routing table. Assuming that the routing record indicating the data stream in the first operator routing table is "data stream a - data write operator x", then it can perform routing expansion for the data stream a based on other data write operators in the data stream processing engine except the data write operator x. Assuming that the selected other data write operator is the data write operator y, the generated expanded routing record can be "data stream a - data write operator y".

[0122] If the computer device determines that the write performance metric value of the data stream is less than or equal to the first threshold, it does not perform routing expansion based on the existing routing record indicating the data stream in the first operator routing table of the first cycle.

[0123] Furthermore, for each data stream written in the first cycle period, the computer device can perform the above operations. After determining the expansion routing records of each data stream, the computer device can generate the second operator routing table for the second cycle period based on the existing routing records and the expansion routing records. Specifically, the expansion routing records can be added to the first operator routing table to generate the second operator routing table for the second cycle period.

[0124] The above method determines the write performance metric values of each data stream written in the first cycle period according to the write performance characteristics; for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, then route expansion is performed based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period to generate an expansion routing record indicating the data stream; and the second operator routing table for the second cycle period is generated according to the existing routing record and the expansion routing record. In the case where the write performance metric value of the data stream is greater than the first threshold, route expansion is performed on the data stream to improve the write efficiency of the data stream, effectively improving the write performance of the data stream.

[0125] In an exemplary embodiment, the method further includes: obtaining the number of operators of the data write operator on the data stream processing engine, and obtaining the concurrent routing upper limit value of each data write operator on the data stream processing engine; determining the concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value; for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, then route expansion is performed based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period to generate an expansion routing record indicating the data stream, including: for each data stream written in the first cycle period, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value and the write performance metric value is greater than the first threshold, then an expansion routing record indicating the data stream is generated based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period.

[0126] The concurrent routing upper limit value refers to the maximum number of routes that the data write operator can concurrently route. Exemplarily, if a data write operator can concurrently route 10 routes, then the concurrent routing upper limit value of the data write operator can be determined to be 10.

[0127] The concurrent routing constraint value refers to the maximum number of routes that the data stream processing engine can concurrently route.

[0128] Exemplarily, if there are 10 data write operators set in the data stream processing engine and the concurrent routing upper limit value of each data write operator is 10, then the concurrent routing constraint value of the data stream processing engine can be 10×10.

[0129] In some exemplary embodiments, a computer device may obtain the number of operators of a data writing operator on a data stream processing engine, and obtain the respective concurrent routing upper limit values of each data writing operator on the data stream processing engine.

[0130] Further, the computer device may determine a concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value. For each data stream written in the first cycle period, when the computer device determines that the written performance metric value is greater than a first threshold, it also needs to determine whether the number of routing records in the first operator routing table is less than the concurrent routing constraint value.

[0131] If the number of routing records in the first operator routing table is less than the concurrent routing constraint value, it may be determined that the data stream processing engine can add new routing records, and an expanded routing record indicating the data stream is generated based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period.

[0132] If the number of routing records in the first operator routing table is greater than or equal to the concurrent routing constraint value, it may be determined that the data stream processing engine cannot add new routing records, and the first operator routing table of the first cycle period is determined as the second operator routing table of the second cycle period.

[0133] For each data stream written in the first cycle period, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value, and the written performance metric value is greater than the first threshold, the method of generating an expanded routing record indicating the data stream based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period first determines whether the current number of routing records is less than the concurrent routing constraint value of the data stream processing engine before generating the expanded routing record, avoiding the abnormal problem caused by the number of routing records being greater than the concurrent routing constraint value and resulting in the data stream processing engine being full, thereby effectively improving the writing performance of the data stream.

[0134] In an exemplary embodiment, for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, route expansion is performed based on the existing route record indicating the data stream in the first operator routing table of the first cycle period to generate an expanded route record indicating the data stream, including: for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, determine the number of expanded route records for the data stream according to a preset expansion strategy; according to the existing route record indicating the data stream in the first operator routing table of the first cycle period, determine the candidate data write operators that have not established route records with the data stream; for each candidate data write operator, count the number of route records corresponding to the candidate data write operator according to the route records in the first operator routing table; select the number of new data write operators that is the same as the number of expanded route records from the candidate data write operators in ascending order of the priority of the number of route records; respectively establish route records indicating each new data write operator and the data stream to obtain an expanded route record indicating the data stream.

[0135] The preset expansion strategy refers to the strategy that technicians preset according to actual needs and use to determine the number of expanded route records when performing route expansion.

[0136] Exemplarily, the preset expansion strategy can indicate determining the number of expanded route records based on the degree of exceeding the performance metric. Specifically, the number of expanded route records can be determined according to the amplitude by which the write performance metric value exceeds the first threshold. For example, if the write speed metric value is 20% higher than the first threshold, then 1 expanded route record is added; if it is 50% higher, then 3 expanded route records are added.

[0137] The preset expansion strategy can also indicate expansion according to a fixed ratio. Specifically, the number of expanded route records can be determined according to a fixed ratio. For example, regardless of how much the write performance metric value exceeds the first threshold, expansion is performed according to a certain ratio of the existing number of route records. Assuming the existing number of route records is 5 and the preset expansion strategy stipulates expansion by 20%, then 1 expanded route record needs to be added.

[0138] In some exemplary embodiments, for each data stream written in the first cycle period, if the computer device determines that the write performance metric value is greater than the first threshold, it determines the number of expanded route records for the data stream according to the preset expansion strategy.

[0139] Further, after determining the number of expanded routing records, the computer device can determine the specific expansion objects, that is, the data writing operators for expanding the data stream. Specifically, the computer device can determine the candidate data writing operators that have not established routing records with the data stream according to the existing routing records indicating the data stream in the first operator routing table of the first cycle period.

[0140] For example, if the data stream processing engine is provided with 5 data writing operators, namely data writing operator 0 - data writing operator 4, and the existing routing record indicating data stream a in the first operator routing table is "data stream a - data writing operator 0", then the determined candidate data writing operators that have not established routing records with the data stream are data writing operator 1, data writing operator 2, data writing operator 3, and data writing operator 4.

[0141] Further, for each such candidate data writing operator, the computer device can count the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table.

[0142] Specifically, the computer device can index in the first operator routing table for each candidate data writing operator to obtain the number of routing records corresponding to each candidate data writing operator. For example, if there are 3 routing records of data writing operator 1 in the first operator routing table, it can be determined that the number of routing records corresponding to data writing operator 1 is 3.

[0143] After the computer device determines the number of routing records corresponding to each candidate data writing operator according to the routing records in the first operator routing table, it can sort the candidate data writing operators in ascending order of the number of routing records. For example, if the number of routing records corresponding to data writing operator 1 is 3, the number of routing records corresponding to data writing operator 2 is 4, the number of routing records corresponding to data writing operator 3 is 5, and the number of routing records corresponding to data writing operator 4 is 6, then after sorting in ascending order of the number of routing records, it is "data writing operator 1, data writing operator 2, data writing operator 3, and data writing operator 4".

[0144] Further, after sorting the candidate data writing operators in ascending order of the number of routing records, the computer device can select, from the candidate data writing operators, the number of new data writing operators that is the same as the number of expanded routing records in the order of ascending priority of the number of routing records, and respectively establish routing records indicating each new data writing operator and the data stream to obtain the expanded routing records indicating the data stream.

[0145] Specifically, assume that the number of expanded routing records is 1. Then, select the data writing operator ranked first after sorting as the new data writing operator, and establish the routing record of this new data writing operator and this data stream to obtain the expanded routing record indicating this data stream. Assume that the number of expanded routing records is 2. Then, select the data writing operators ranked first and second after sorting as the new data writing operators, and establish the routing records of these new data writing operators and this data stream respectively to obtain the expanded routing record indicating this data stream.

[0146] For each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, then according to the preset expansion strategy, determine the number of expanded routing records for this data stream; based on the existing routing records indicating this data stream in the first operator routing table of the first cycle period, determine the candidate data writing operators that have not established routing records with this data stream; for each such candidate data writing operator, according to the routing records in the first operator routing table, count the number of routing records corresponding to this candidate data writing operator; from these candidate data writing operators, select new data writing operators with a number consistent with the number of expanded routing records in the ascending priority order of the number of routing records; establish the routing records indicating each such new data writing operator and this data stream respectively to obtain the expanded routing record indicating this data stream. The method can first determine the candidate data writing operators that have not established routing records with this data stream during routing expansion, and then determine the new data writing operators from the candidate data writing operators according to the number of routing records of the candidate data writing operators and the number of expanded routing records. Since the number of routing records can reflect the idle degree of the data writing operator to some extent, the relatively idle data writing operator can be selected as the new data writing operator through the number of routing records, realizing the load balancing of the data writing operator and improving the stability of the data stream processing engine.

[0147] In an exemplary embodiment, for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold, then generate the expanded routing record indicating this data stream based on the existing routing records indicating this data stream in the first operator routing table of the first cycle period, including: for each data stream written in the first cycle period, if the write performance metric value is greater than the first threshold and the expanded routing record indicating this data stream has been generated for the first cycle period, then determine the timing order of this routing expansion in the continuous routing expansion operations; according to this timing order, determine the number of expanded routing records for this data stream; based on the existing routing records indicating this data stream in the first operator routing table of the first cycle period and according to the number of expanded routing records, generate the expanded routing record indicating this data stream.

[0148] The timing order refers to the cumulative number of times this expansion routing occurs in consecutive routing expansion operations. For example, if two expansion routings have been continuously performed in the two cycle periods before this expansion routing, then the timing order of this routing expansion is 3. If one expansion routing has been performed in one cycle period before this expansion routing, then the timing order of this routing expansion is 2.

[0149] The number of expansion routing records is positively correlated with this timing order. That is, as the timing order increases, the number of expansion routing records also increases. For example, when the timing order is 1, the number of expansion routings can be 1; when the timing order is 2, the number of expansion routings can be 2; when the timing order is 3, the number of expansion routings can be 4; when the timing order is 4, the number of expansion routings can be 8.

[0150] It should be noted that after the consecutive routing expansion operation is interrupted and the routing expansion is performed again, the number of expansion routing records will return to the initial number of expansion routing records. For example, the initial number of expansion routing records is 1. After the routing expansion operation is performed in three consecutive cycle periods, the number of expansion routing records in the third cycle is 4. If the routing expansion operation is not performed in the fourth cycle, then if the routing expansion operation is performed in the fifth cycle, the number of expansion routing records is 1.

[0151] In some exemplary embodiments, for each data stream written in the first cycle period, if the computer device determines that the write performance metric value of the data stream is greater than the first threshold and an expansion routing record indicating the data stream has been generated for the first cycle period, it can determine the timing order of this routing expansion in the consecutive routing expansion operation.

[0152] Specifically, for each data stream written in the first cycle period, if the computer device determines that the write performance metric value of the data stream is greater than the first threshold, it can first determine whether an expansion routing record indicating the data stream has been generated in the first cycle period. If it has been generated, it can determine that it is a consecutive routing expansion operation, and then determine the timing order of this routing expansion in the consecutive routing expansion operation. If it has not been generated, it can determine that the timing order of this routing expansion is 1.

[0153] Furthermore, after the computer device determines the timing order, it can, according to this timing order, determine the number of expansion routing records for the data stream, and based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period, generate an expansion routing record indicating the data stream according to this number of expansion routing records.

[0154] For each data stream written in the first cycle, if the write performance metric value is greater than the first threshold and an expansion routing record indicating the data stream has been generated for the first cycle, determine the timing order of the current routing expansion in the continuous routing expansion operations; according to the timing order, determine the number of expansion routing records for the data stream; based on the existing routing records indicating the data stream in the first operator routing table of the first cycle and according to the number of expansion routing records, generate the method for the expansion routing records indicating the data stream. According to the timing order of the current routing expansion in the continuous routing expansion operations, determining the number of expansion routing records for the data stream can accelerate the expansion rate during continuous expansion operations and effectively improve the write performance of the data stream.

[0155] In an exemplary embodiment, adjusting the first operator routing table of the first cycle according to the write performance characteristics to obtain the second operator routing table of the second cycle includes: determining the write performance metric value of each data stream written in the first cycle according to the write performance characteristics; for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, perform routing contraction based on the existing routing records indicating the data stream in the first operator routing table of the first cycle to determine the contraction routing records indicating the data stream; generate the second operator routing table of the second cycle according to the existing routing records and the contraction routing records.

[0156] The second threshold can be preset by those skilled in the art according to actual needs and is used to compare with the write performance metric value of the data stream to determine whether to perform routing contraction operations on the routing of the data stream.

[0157] Routing contraction refers to the operation of deleting the original routing records of the data stream.

[0158] In some exemplary embodiments, after receiving the write performance characteristics of the data stream processing engine in the first cycle, the computer device can determine the write performance metric value of each data stream written in the first cycle according to the write performance characteristics. Specifically, the computer device can use a preset algorithm and the write performance characteristics to determine the write performance metric value of each data stream written in the first cycle.

[0159] Further, after the computer device determines the write performance metric value of each data stream written in the first cycle, for each data stream written in the first cycle, if the computer device determines that the write performance metric value of the data stream is less than or equal to the second threshold, it performs routing contraction based on the existing routing records indicating the data stream in the first operator routing table of the first cycle to determine the contraction routing records indicating the data stream.

[0160] Specifically, for the data stream a written in the first cycle, if the computer device determines that the write performance metric value of the data stream a is less than or equal to the second threshold, it first determines the routing record of the data stream a in the first operator routing table. Assuming that the routing record indicating the data stream in the first operator routing table is "data stream a - data write operator x, data stream a - data write operator y", then it can perform routing capacity reduction for the data stream a based on the original data write operator of the data stream. Assuming that the selected data write operator is data write operator y, it determines that the reduced-capacity routing record indicating the data stream is "data stream a - data write operator y".

[0161] If the computer device determines that the write performance metric value of the data stream is greater than the second threshold, it does not perform routing capacity reduction based on the existing routing record indicating the data stream in the first operator routing table of the first cycle.

[0162] Furthermore, for each data stream written in the first cycle, the computer device can perform the above operations. After determining the reduced-capacity routing record for each data stream, it can generate the second operator routing table for the second cycle based on the existing routing record and the reduced-capacity routing record. Specifically, it can delete the reduced-capacity routing record in the first operator routing table to generate the second operator routing table for the second cycle. For example, if the routing record indicating the data stream in the first operator routing table is "data stream a - data write operator x, data stream a - data write operator y", and the reduced-capacity routing record of the data stream is "data stream a - data write operator y", then the routing record of the data stream in the second operator routing table is "data stream a - data write operator x".

[0163] The above method of determining the write performance metric value of each data stream written in the first cycle according to the write performance characteristics; for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, performing routing capacity reduction based on the existing routing record indicating the data stream in the first operator routing table of the first cycle to determine the reduced-capacity routing record indicating the data stream; and generating the second operator routing table for the second cycle according to the existing routing record and the reduced-capacity routing record. In the case where the write performance metric value of the data stream is less than or equal to the second threshold, it performs routing capacity reduction on the data stream to reduce the small files corresponding to the data stream, thereby improving the efficiency of the later data acquisition process.

[0164] In an exemplary embodiment, for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, route reduction is performed based on the existing route record indicating the data stream in the first operator routing table of the first cycle, and a reduced route record indicating the data stream is determined, including: for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, determine the number of reduced route records for the data stream according to a preset reduction strategy; according to the existing route record indicating the data stream in the first operator routing table of the first cycle, determine the candidate data write operators that have established route records with the data stream; for each candidate data write operator, count the number of route records corresponding to the candidate data write operator according to the route records in the first operator routing table; select the reduced data write operators with the same number as the number of reduced route records from the candidate data write operators in the descending priority order of the number of route records, and obtain the reduced route record indicating the data stream and the reduced data write operators.

[0165] The preset reduction strategy refers to a strategy that is pre-set by a technician according to actual needs and is used to determine the number of reduced route records when performing route reduction. Exemplarily, the preset reduction strategy may indicate that the number of reduced route records for each reduction is 1.

[0166] In some exemplary embodiments, for each data stream written in the first cycle, if the computer device determines that the write performance metric value is less than or equal to the second threshold, it determines the number of expanded route records for the data stream according to the preset reduction strategy.

[0167] Further, after the computer device determines the number of reduced routes, it can determine the specific reduction object, that is, the data write operator for reducing the data stream. Specifically, the computer device can determine the candidate data write operators that have established route records with the data stream according to the existing route record indicating the data stream in the first operator routing table of the first cycle.

[0168] For example, the data stream processing engine is provided with 5 data write operators, namely data write operator 0 - data write operator 3, and the existing route records indicating data stream a in the first operator routing table are "data stream a - data write operator 0, data stream a - data write operator 1, data stream a - data write operator 2", then the determined candidate data write operators that have established route records with the data stream are data write operator 0, data write operator 1, and data write operator 2.

[0169] Further, for each candidate data write operator, the computer device can count the number of route records corresponding to the candidate data write operator according to the route records in the first operator routing table.

[0170] Specifically, the computer device can index each candidate data writing operator in the first operator routing table to obtain the number of routing records corresponding to each candidate data writing operator. For example, if there are 3 routing records for data writing operator 0 in the first operator routing table, it can be determined that the number of routing records corresponding to data writing operator 0 is 3.

[0171] After the computer device determines the number of routing records corresponding to each candidate data writing operator according to the routing records in the first operator routing table, it can sort the candidate data writing operators in descending order of the number of routing records. For example, if the number of routing records corresponding to data writing operator 0 is 5, the number of routing records corresponding to data writing operator 1 is 4, and the number of routing records corresponding to data writing operator 2 is 3, then after sorting in descending order of the number of routing records, it is "data writing operator 0, data writing operator 1, data writing operator 2".

[0172] Further, after the computer device sorts the candidate data writing operators in descending order of the number of routing records, it can select, from the candidate data writing operators, reduction data writing operators with a number consistent with the number of the scaling-down routing records in the priority order of descending number of routing records, to obtain the scaling-down routing records indicating the data stream and the reduction data writing operators.

[0173] Specifically, assuming that the number of the scaling-down routing records is 1, the data writing operator ranked first after sorting is selected as the reduction data writing operator, and the scaling-down routing record indicating the data stream and the reduction data writing operator is obtained. Assuming that the number of the scaling-up routing records is 2, the data writing operators ranked first and second after sorting are selected as the reduction data writing operators, and the scaling-down routing records indicating the data stream and the reduction data writing operators are obtained respectively.

[0174] For each data stream written in the first cycle period, if the write performance metric value is less than or equal to the second threshold, determine the number of scaled-down routing records for the data stream according to a preset scaling strategy; based on the existing routing records indicating the data stream in the first operator routing table of the first cycle period, determine the candidate data write operators that have established routing records with the data stream; for each such candidate data write operator, according to the routing records in the first operator routing table, count the number of routing records corresponding to the candidate data write operator; from the candidate data write operators, select the reduced data write operators with the same number as the scaled-down routing record number in the order of descending priority of the routing record number, to obtain the method of the scaled-down routing record indicating the data stream and the reduced data write operator. When performing routing scaling, the candidate data write operators that have established routing records with the data stream can be determined first, and then the reduced data write operators can be determined from the candidate data write operators according to the routing record number of the candidate data write operator and the scaled-down routing record number. Since the routing record number can reflect the idle degree of the data write operator indirectly, the relatively busy data write operators can be selected as the reduced data write operators through the routing record number, realizing the load balancing of the data write operators and improving the stability of the data stream processing engine.

[0175] In an exemplary embodiment, determining the write performance metric value of each data stream written in the first cycle period according to the write performance characteristic includes: determining the multi-cycle periods to be statistically analyzed; obtaining the write performance characteristics of each of the at least one historical cycle period; and performing statistical analysis of the multi-cycle periods according to the write performance characteristics of the multi-cycle periods to obtain the write performance metric value of each data stream written in the first cycle period.

[0176] The multi-cycle periods include the first cycle period and at least one historical cycle period before the first cycle period.

[0177] In some exemplary embodiments, the computer device can first determine the multi-cycle periods to be statistically analyzed, and then obtain the write performance characteristics of each of the at least one historical cycle period.

[0178] Further, after the computer device obtains the write performance characteristics of the first cycle period and each of the at least one historical cycle period, it can perform statistical analysis of the multi-cycle periods according to the write performance characteristics of the multi-cycle periods to obtain the write performance metric value of each data stream written in the first cycle period.

[0179] The method for determining the multi-cycle period to be counted; obtaining the write performance characteristics of each of the at least one historical cycle period; and performing statistical analysis on the multi-cycle period according to the write performance characteristics of each of the multi-cycle periods to obtain the write performance metric values of each data stream written in the first cycle period. By determining the write performance metric values of each data stream written in the first cycle period according to the write performance characteristics of each of the multi-cycle periods, the reliability and accuracy of the write performance metric values can be effectively improved.

[0180] In an exemplary embodiment, each data stream includes a plurality of data segments, and the write performance characteristics include the publish timestamp and the consume timestamp of each data segment. The publish timestamp represents the time when the data segment is written into the cache, and the consume timestamp represents the time when the data segment read from the cache is written into the file; the write performance metric values of each data stream written in the first cycle period include the data segment latency statistical value of the multi-cycle period; performing statistical analysis on the multi-cycle period according to the write performance characteristics of each of the multi-cycle periods to obtain the write performance metric values of each data stream written in the first cycle period includes: determining the latency value of each data segment according to the publish timestamp and the consume timestamp of each data segment written in each cycle period within the multi-cycle period; for each data stream, determining the total latency value of the data stream according to the latency values of the data segments in the data stream written in the multi-cycle period; for each data stream, counting the total number of data segments of the data stream written in the multi-cycle period; for each data stream, determining the data segment latency statistical value of the data stream in the multi-cycle period according to the total latency value of the data stream and the total number of data segments of the data stream.

[0181] Exemplarily, a data stream can be understood as an ordered arrangement of a series of data packets in the time dimension. When data is transmitted in the form of a data stream, the data will be split into multiple data packets for transmission. Each data packet is a part of the data stream and is transmitted in sequence one after another, and finally forms a complete data stream. A data segment is the data obtained by parsing a data packet. That is, a data stream is composed of a series of data packets, and a data packet is composed of a series of data segments.

[0182] The cache can be a storage device for temporarily storing data segments of a data stream. The cache can be Pulsar, TubeMQ. It should be noted that in the case where the cache is TubeMQ, there is no publish timestamp, so the above process is not involved. Only in the case where the cache is Pulsar, there is a publish timestamp and the above process is involved.

[0183] The latency value of each data segment refers to the duration that the data segment is stored in the buffer waiting to be written. The total latency value of the data stream refers to the total duration that the data segments included in the data stream are stored in the buffer waiting to be written during multiple loop cycles. The statistical value of the data segment latency in multiple loop cycles can be the maximum value, the median value, or the average value. In the embodiments of the present application, the statistical value of the data segment latency in multiple loop cycles refers to the average duration that the data segments included in the data stream are stored in the buffer waiting to be written.

[0184] In some exemplary embodiments, the computer device can first obtain the publication timestamp and consumption timestamp of each data segment written in each loop cycle within multiple loop cycles, and then determine the latency value of each data segment according to the publication timestamp and consumption timestamp of each data segment written in each loop cycle within the multiple loop cycles. Specifically, the latency value is the absolute value of the difference between the consumption timestamp and the publication timestamp, that is, latency value = |consumption timestamp - publication timestamp|.

[0185] Furthermore, for each data stream, the computer device determines the total latency value of the data stream according to the latency values of the data segments in the data stream written in the multiple loop cycles. Specifically, the total latency value is the sum of the latency values of the data segments, that is, total latency value = latency value of data segment a + latency value of data segment b + latency value of data segment c + ......

[0186] For each data stream, the computer device counts the total number of data segments in the data stream written in the multiple loop cycles. For each data stream, the computer device determines the statistical value of the data segment latency in the multiple loop cycles of the data stream according to the total latency value of the data stream and the total number of data segments of the data stream. Specifically, the statistical value of the data segment latency = total latency value / total number of data segments.

[0187] Furthermore, in the case where the write performance metric value is the statistical value of the data segment latency in multiple loop cycles, the first threshold can be 10000 ms, and the second threshold can be 3000 ms.

[0188] Determine the latency value of each data segment according to the publish timestamp and consumption timestamp of each data segment written in each cycle within the multi-cycle period; for each data stream, determine the total latency value of the data stream according to the latency values of the data segments in the data stream written in the multi-cycle period; for each data stream, count the total number of data segments of the data stream written in the multi-cycle period; for each data stream, determine the data segment latency statistical value of the data stream in the multi-cycle period according to the total latency value of the data stream and the total number of data segments of the data stream. The method can determine the data segment latency statistical value of the data stream in the multi-cycle period according to the latency value and total amount of data segments in each data stream, and use this value as the write performance metric value, so that the routing expansion and routing contraction based on the write performance metric value can be more accurate and reliable.

[0189] In an exemplary embodiment, the write performance characteristics include the amount of data written to each data stream in each cycle; the write performance metric values of each data stream written in the first cycle, including the data stream load statistical value of the multi-cycle period; according to the write performance characteristics of each multi-cycle period, perform statistical analysis on the multi-cycle period to obtain the write performance metric values of each data stream written in the first cycle, including: for each data stream written in the first cycle, according to the amount of data written to the data stream in each cycle, count the total amount of data written to the data stream in the multi-cycle period; for each data stream written in the first cycle, count the number of routing records indicating the routing records of the data stream from the first operator routing table; for each data stream written in the first cycle, determine the data stream load statistical value of the data stream in the multi-cycle period according to the total amount of data, the number of routing records, and the number of cycles of the multi-cycle period.

[0190] In some exemplary embodiments, for each data stream written in the first cycle, the computer device can obtain the amount of data written to the data stream in each cycle, and count the total amount of data written to the data stream in the multi-cycle period according to the amount of data written to the data stream in each cycle. The total amount of data can be the total size, that is, the total number of bytes, or the total number of fields, that is, the product of the total number of data records and the number of fields of each data record.

[0191] Further, for each data stream written in the first cycle period, the computer device counts the number of routing records indicating the routing records of the data stream from the first operator routing table. For each data stream written in the first cycle period, the computer device determines the data stream load statistical value of the data stream in the multi-cycle period according to the total amount of data, the number of routing records, and the number of cycle periods of the multi-cycle period. Specifically, the data stream load statistical value of the data stream in the multi-cycle period = total amount of data / number of routing records / number of cycle periods of the multi-cycle period. Optionally, the data stream load statistical value of the multi-cycle period may also be the median, the maximum value, etc.

[0192] Further, when the write performance metric value is the data stream load statistical value of the data stream in the multi-cycle period, the first threshold may be 500MB / 5000000, and the second threshold may be 50MB / 5000000.

[0193] In an alternative embodiment, the following will be combined with Figure 5 , to illustrate the complete process of routing capacity reduction and routing capacity expansion. Specifically, after the computer device receives the write performance characteristics of the data stream processing engine in the first cycle period, for each data stream, it may first determine the data stream load statistical value and the data segment delay statistical value according to the write performance characteristics of the data stream.

[0194] Further, after determining the data stream load statistical value and the data segment delay statistical value according to the write performance characteristics of the data stream, it may first determine whether the data segment delay statistical value is greater than the first delay threshold. If it is greater, then according to the preset capacity expansion policy, determine the number of expanded routing records for the data stream, and determine the new data writing operator from the data writing operator in the data stream engine according to the routing records in the first operator routing table, and generate expanded routing records to obtain the second operator routing table.

[0195] If it is less, then determine whether the data stream load statistical value is greater than the first load threshold. If it is greater, then according to the preset capacity expansion policy, determine the number of expanded routing records for the data stream, and determine the new data writing operator from the data writing operator in the data stream engine according to the routing records in the first operator routing table, and generate expanded routing records to obtain the second operator routing table.

[0196] If it is less, then determine whether the data segment delay statistical value is less than or equal to the second delay threshold. If it is less, then according to the preset capacity reduction policy, determine the number of reduced routing records for the data stream, and generate reduced routing records according to the routing records in the first operator routing table to obtain the second operator routing table.

[0197] If it is greater, determine whether the data flow load statistical value is less than or equal to the second load threshold. If it is less, determine the number of scaling routing records for the data flow according to the preset scaling strategy, and generate scaling routing records based on the routing records in the first operator routing table to obtain the second operator routing table. If it is greater, use the first operator routing table as the second operator routing table.

[0198] For each data flow written in the first cycle period, according to the amount of data written to the data flow in each cycle period, count the total amount of data written to the data flow in the multi-cycle period; for each data flow written in the first cycle period, from the first operator routing table, count the number of routing records indicating the data flow; for each data flow written in the first cycle period, according to the total amount of data, the number of routing records, and the number of cycle periods of the multi-cycle period, determine the data flow load statistical value of the data flow in the multi-cycle period. Using the data flow load statistical value of the data flow in the multi-cycle period as the write performance metric value makes the routing expansion and routing scaling more accurate and reliable based on the write performance metric value.

[0199] In an exemplary embodiment, the method further includes: scanning the data flows newly added by the data flow processing engine in the first cycle period; determining the default number of routing records for the newly added data flow according to the default operator allocation strategy; according to the routing records in the first operator routing table, counting the number of routing records of each data write operator on the data flow processing engine; selecting, from the data write operators on the data flow processing engine, target data write operators whose number is the same as the default number of routing records in ascending order of the number of routing records; adding target routing records indicating the target data write operators and the data flow to the second operator routing table in the second cycle period.

[0200] The default operator allocation strategy refers to an operator allocation strategy preset by technicians according to actual needs.

[0201] Exemplarily, the default operator allocation strategy may indicate that when the newly added data flow is not a preset data flow, allocate one data write operator to the newly added data flow, that is, the default number of routing records for the newly added data flow is 1; when the newly added data flow is a preset data flow, allocate a preset number of data write operators to the newly added data flow, that is, the default number of routing records for the newly added data flow is the preset number. The preset data flow is a data flow determined by technicians according to prior knowledge that requires multiple data write operators to be set.

[0202] The target data write operator refers to the data write operator used to write the newly added data flow.

[0203] In some exemplary embodiments, the computer device may also scan the data stream newly added by the data stream processing engine in the first cycle, and determine the default routing record quantity of the newly added data stream according to the default operator allocation policy.

[0204] Further, after determining the default routing record quantity of the newly added data stream, the computer device may, according to the routing records in the first operator routing table, count the respective routing record quantities of the data writing operators on the data stream processing engine, and select, from the data writing operators on the data stream processing engine, target data writing operators whose quantities are consistent with the default routing record quantity in the ascending priority order of the routing record quantities.

[0205] Further, after determining the target writing operator, the computer device may add, to the second operator routing table in the second cycle, target routing records indicating the target data writing operator and the data stream.

[0206] The method of scanning the data stream newly added by the data stream processing engine in the first cycle; determining the default routing record quantity of the newly added data stream according to the default operator allocation policy; counting the respective routing record quantities of the data writing operators on the data stream processing engine according to the routing records in the first operator routing table; selecting, from the data writing operators on the data stream processing engine, target data writing operators whose quantities are consistent with the default routing record quantity in the ascending priority order of the routing record quantities; and adding, to the second operator routing table in the second cycle, target routing records indicating the target data writing operator and the data stream can, when it is necessary to write data to the data stream, first allocate as few data writing operators as possible according to the default operator allocation policy to perform the writing operation, avoid generating a large number of small files, and improve the efficiency of the subsequent data acquisition process.

[0207] In an exemplary embodiment, adjusting the first operator routing table in the first cycle according to the writing performance characteristics to obtain the second operator routing table in the second cycle includes: traversing the data streams indicated by the routing records in the first operator routing table; in response to the traversed data stream being in a locked state, continuing to traverse the next data stream; in response to the traversed data stream being in an unlocked state, adjusting the routing record for the traversed data stream according to the writing performance characteristics; and until the traversal is completed, obtaining the second operator routing table in the second cycle.

[0208] A data stream in a locked state is a data stream whose routing record cannot be adjusted, that is, a data stream that cannot be expanded or contracted in routing.

[0209] A data stream in an unlocked state is a data stream whose routing record can be adjusted, that is, a data stream that can be expanded or contracted in routing.

[0210] In some exemplary embodiments, for some data streams with a relatively small amount of data but time-consuming data writing, the status of these data streams can be set to the locked state. During the process of the computer device traversing the data streams indicated by the routing records in the first operator routing table, if it is determined that the traversed data stream is in the locked state, the computer device continues to traverse the next data stream. If it is determined that the traversed data stream is in the unlocked state, according to the writing performance characteristics, the routing record of the traversed data stream is adjusted until the traversal is completed, and the second operator routing table of the second cycle is obtained.

[0211] The method of traversing the data streams indicated by the routing records in the first operator routing table; in response to the traversed data stream being in the locked state, continuing to traverse the next data stream; in response to the traversed data stream being in the unlocked state, adjusting the routing record of the traversed data stream according to the writing performance characteristics; until the traversal is completed, obtaining the second operator routing table of the second cycle. For special data streams, their status can be set to the locked state, so as not to perform routing expansion and routing contraction on them, avoiding low writing efficiency of special data streams.

[0212] In an exemplary embodiment, as Figure 6 shown, a data stream writing control method is provided, and the method includes the following steps:

[0213] Step 601, obtain the first operator routing table of the first cycle.

[0214] The cycle refers to a time period with a fixed duration set when the data stream processing engine executes the writing operation for the data stream. The duration of each cycle can be pre-set by those skilled in the art according to actual needs. The duration of each cycle can be 1 minute, 10 minutes, 30 minutes, 1 hour, 5 hours, 10 hours, 24 hours, etc.

[0215] Exemplarily, assuming that the duration of each cycle is 1 minute, then starting from when the data stream processing engine starts to execute the writing operation for the data stream, it is the start of the first cycle. After 1 minute, it is the end of the first cycle and the start of the second cycle.

[0216] The first cycle refers to any cycle in the multiple cycles when the data stream processing engine executes the writing operation for the data stream.

[0217] The operator routing table contains routing records, and each routing record indicates a data stream to be written and a corresponding data writing operator. That is, the operator routing table is used to indicate the mapping relationship between the data stream and the data writing operator for writing the data stream. The operator routing table can specifically be the Sink routing table in the data stream processing engine.

[0218] The first operator routing table contains at least one routing record. The first operator routing table contains routing records within the first cycle period, and the routing records within the first cycle period indicate that within the first cycle period, a data stream to be written corresponds to a corresponding data writing operator.

[0219] In some exemplary embodiments, before the start of each cycle period, the data stream processing engine can obtain the operator routing table corresponding to this cycle period. Specifically, before the start of the first cycle period, the data stream processing engine can obtain the first operator routing table corresponding to the first cycle period.

[0220] Step 602: Use the data writing operator indicated by each such routing record to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written.

[0221] The data writing operator refers to an operator set in the data stream processing engine for writing data streams. The data writing operator can be a Sink. A Sink is a computing unit in the data stream processing engine for writing data streams. There can be multiple Sinks set in the data stream processing engine, and each Sink has an identification serial number.

[0222] The data stream refers to a data sequence with the same data type and originating from the same data source. Exemplarily, the data stream can be a data stream of the log file type, a data stream of the user interaction type, a data stream of the sensor network type, etc. Specifically, there are log files in various systems and application programs. For example, server log files can record the access situations of users to websites, and database log files can record the operation histories of databases. These log files all exist in the form of data streams.

[0223] The file corresponding to the data writing operator and the data stream to be written refers to the file created when the data stream to be written is initially written into the file system by the data writing operator. Exemplarily, the file system can be a distributed file system (HDFS, Hadoop Distributed File System).

[0224] In some exemplary embodiments, after the data stream processing engine obtains the first operator routing table of the first cycle period, it can use the data writing operator indicated by each such routing record in the first operator routing table to write the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written.

[0225] Step 603: For each data stream to be written, obtain the writing performance characteristics of the data written through at least one data writing operator within the first cycle.

[0226] The writing performance characteristics table shows the writing performance of each data stream during the first cycle. The writing performance characteristics may specifically include the data stream identifier (dataFlowId), the data writing operator identifier (Sink taskIndex), the cycle duration, the consumption time, the publishing time, the total data volume, and the total number of data items.

[0227] Among them, the data in the data stream is first stored in the cache, and then retrieved by the data stream processing engine from the cache for writing. The consumption time refers to the time when the data in the data stream is written, and the publishing time refers to the time when the data in the data stream is cached. The total data volume refers to the total data volume of each data stream within the first cycle, and the total number of data items refers to the total number of data items of each data stream within the first cycle.

[0228] Furthermore, if the data in a data stream is written by one data writing operator, the total data volume of the data stream can be characterized by the total size. If the data in a data stream is written by multiple data writing operators, the total data volume of the data stream can be characterized by the total number of fields.

[0229] In some exemplary embodiments, for each data stream to be written, the data stream processing engine can obtain the writing performance characteristics of the data written through at least one data writing operator within the first cycle through the data writing operator.

[0230] Step 604: Send the writing performance characteristics.

[0231] Among them, the sent writing performance characteristics are used to indicate that the routing records in the first operator routing table are adjusted according to the writing performance characteristics to obtain the second operator routing table for the second cycle, and the writing performance characteristics represent the writing performance of each data stream written during the first cycle.

[0232] In some exemplary embodiments, after determining the writing performance characteristics, the data stream processing engine can send the writing performance characteristics to the computer device. Specifically, it can be the control node in the computer device.

[0233] Furthermore, after receiving the writing performance characteristics, the control node in the computer device can adjust the routing records in the first operator routing table according to the writing performance characteristics to obtain the second operator routing table for the second cycle.

[0234] The above method for obtaining the first operator routing table of the first cycle, where the first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator; using the data writing operator indicated by each such routing record, writing the data stream to be written corresponding to the data writing operator into the file corresponding to the data writing operator and the data stream to be written; for each data stream to be written, obtaining the writing performance characteristics written through at least one data writing operator within the first cycle; the method of sending the writing performance characteristics, by sending the writing performance characteristics of each cycle, enabling the computer device to determine the operator routing table of the next cycle according to the writing performance characteristics, making the obtained operator routing table more accurate, and thus effectively improving the data stream writing efficiency.

[0235] In an exemplary embodiment, the method for obtaining the first operator routing table of the first cycle includes: using a data source operator to receive the first operator routing table of the first cycle from a control node; the method further includes: using the data source operator to read data segments of each data stream from a cache, and allocating the data segments to each data writing operator according to the routing records in the first operator routing table.

[0236] Wherein, the writing performance characteristics are obtained and sent by the data writing operator.

[0237] The data source operator refers to an operator set in a data stream processing engine, which is used to obtain data segments from a cache and allocate the data segments to data writing operators so that the data writing operators can write the data segments. The data source operator can be a Source, and Source is a computing unit in the data stream processing engine responsible for reading source data. There can be multiple Sources set in the data stream processing engine, and each Source has an identification serial number.

[0238] In some exemplary embodiments, the data stream processing engine can first use the data source operator to receive the first operator routing table of the first cycle from a control node. Specifically, the first operator routing table is stored in a routing table repository. The routing table repository can be Zookeeper.

[0239] Further, after receiving the first operator routing table, the data stream processing engine can use the data source operator to read data segments of each data stream from a cache, and allocate the data segments to each data writing operator according to the routing records in the first operator routing table.

[0240] The method of using the data source operator to read data segments of each data stream from the cache and allocate the data segments to each data writing operator according to the routing records in the first operator routing table can allocate the data segments to each data writing operator according to the routing records in the first operator routing table, so that the data writing operator can write the data segments, which can effectively improve the writing performance of the data stream.

[0241] In an optional embodiment of the present application, the method further includes: during the process of writing the data stream, using a checkpoint mechanism to periodically generate a status snapshot.

[0242] The checkpoint mechanism refers to the Checkpoint mechanism. Exemplarily, the checkpoint mechanism can be used to ensure fault tolerance and persistence during the data stream writing process. By periodically generating a status snapshot during the data stream processing, it can be used for recovery when the system fails.

[0243] In an optional embodiment of the present application, the data writing operator can write the data segments in the data stream into a distributed file system. There are two types of nodes in the distributed file system, namely NameNode and DataNode. The NameNode is responsible for the storage-related scheduling in the distributed file system, and the DataNode is responsible for the specific storage tasks. Specifically, the NameNode and DataNode will be described below in combination with the writing process of the distributed file system. The data writing operator can call the create() method through the distributed file system API to create a new file or append to an existing file. The data writing operator first communicates with the NameNode, requests to create a file, and obtains a group of DataNodes for storing data blocks. The NameNode returns a list of DataNodes according to the number of replicas of the data blocks and the load conditions of the DataNodes, and these DataNodes will store the data blocks in sequence. The data writing operator will establish a connection with the first DataNode and request it to be ready to receive data. After the first DataNode confirms, the data writing operator starts to send data to it. After the first DataNode receives the data, it forwards it to the next DataNode, forming a data stream pipeline. During the data segment writing process, the data segment is written into the DataNodes in the pipeline in the form of data blocks. After each DataNode receives the data block, it will send a confirmation message to the data writing operator. When the data writing operator receives the confirmation messages from all DataNodes, it can determine that the data segment is successfully written. After the data segment writing is completed, the data writing operator can call the close() method to close the file.

[0244] In an exemplary embodiment, such as Figure 7As shown, a data stream writing control method is provided, and the method includes the following steps:

[0245] Step 701: The data stream processing engine uses a data source operator to receive a first operator routing table of a first cycle from a control node, and uses the data source operator to read data segments of each data stream from a cache, and according to the routing records in the first operator routing table, uses the data writing operator indicated by each routing record to write the data into the data stream to be written corresponding to the data writing operator, and write it into the file corresponding to the data writing operator and the data stream to be written; for each data stream to be written, obtain the writing performance characteristics of writing through at least one data writing operator within the first cycle; send the writing performance characteristics.

[0246] Wherein, the first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator.

[0247] In some exemplary embodiments, the data stream processing engine may first use a data source operator to receive a first operator routing table of a first cycle from a control node.

[0248] Further, after receiving the first operator routing table, the data stream processing engine may use the data source operator to read data segments of each data stream from the cache, and according to the routing records in the first operator routing table, use the data writing operator indicated by each routing record to write the data into the data stream to be written corresponding to the data writing operator, and write it into the file corresponding to the data writing operator and the data stream to be written.

[0249] Further, for each data stream to be written, the data stream processing engine may obtain the writing performance characteristics of writing through at least one data writing operator within the first cycle, and send the writing performance characteristics to the control node in the computer device.

[0250] Step 702: The computer device receives the writing performance characteristics of the data stream processing engine in the first cycle; traverse the data streams indicated by the routing records in the first operator routing table; in response to the traversed data stream being in a locked state, continue to traverse the next data stream.

[0251] The writing performance characteristics characterize the writing performance of the data stream processing engine for writing each data stream in the first cycle.

[0252] In some exemplary embodiments, for some data streams with a relatively small amount of data but time-consuming data writing, the states of these data streams can be set to a locked state. During the process of the computer device traversing the data streams indicated by the routing records in the first operator routing table, if it is determined that the traversed data stream is in a locked state, it continues to traverse the next data stream. If it is determined that the traversed data stream is in an unlocked state, according to the writing performance characteristics, routing record adjustment is performed for the traversed data stream until the traversal is completed, and a second operator routing table for the second cycle period is obtained.

[0253] Step 703: In response to the traversed data stream being in an unlocked state, the computer device determines a multi-cycle period to be statistically analyzed, where the multi-cycle period includes the first cycle period and at least one historical cycle period before the first cycle period; obtains the writing performance characteristics of each of the at least one historical cycle period; performs statistical analysis on the multi-cycle period according to the writing performance characteristics of each of the multi-cycle periods to obtain the writing performance metric values of each data stream written in the first cycle period; obtains the number of operators of the data writing operator on the data stream processing engine, and obtains the concurrent routing upper limit value of each data writing operator on the data stream processing engine; determines the concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value.

[0254] In some exemplary embodiments, the computer device can first determine the multi-cycle period to be statistically analyzed, and then obtain the writing performance characteristics of each of the at least one historical cycle period.

[0255] Further, after the computer device obtains the writing performance characteristics of the first cycle period and each of the at least one historical cycle period, it can perform statistical analysis on the multi-cycle period according to the writing performance characteristics of each of the multi-cycle periods to obtain the writing performance metric values of each data stream written in the first cycle period.

[0256] The computer device can also obtain the number of operators of the data writing operator on the data stream processing engine, and obtain the concurrent routing upper limit value of each data writing operator on the data stream processing engine.

[0257] Further, the computer device can determine the concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value. For each data stream written in the first cycle period, when the computer device determines that the writing performance metric value is greater than the first threshold, it also needs to determine whether the number of routing records in the first operator routing table is less than the concurrent routing constraint value.

[0258] Step 704: For each data stream written in the first cycle by the computer device, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value, and the write performance metric value is greater than the first threshold, and an expanded routing record indicating the data stream has been generated for the first cycle, then determine the timing order of this routing expansion in the continuous routing expansion operations; according to this timing order, determine the number of expanded routing records for the data stream; based on the existing routing records indicating the data stream in the first operator routing table of the first cycle, determine the candidate data write operators that have not established routing records with the data stream; for each such candidate data write operator, according to the routing records in the first operator routing table, count the number of routing records corresponding to the candidate data write operator; from these candidate data write operators, select, in ascending order of the number of routing records, the number of new data write operators that is the same as the number of expanded routing records; respectively establish routing records indicating each such new data write operator and the data stream to obtain the expanded routing record indicating the data stream; according to the existing routing records and the expanded routing record, generate the second operator routing table for the second cycle.

[0259] In some exemplary embodiments, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value, and the write performance metric value is greater than the first threshold, and an expanded routing record indicating the data stream has been generated for the first cycle. Then the timing order of this routing expansion in the continuous routing expansion operations can be determined.

[0260] Specifically, for each data stream written in the first cycle, if the computer device determines that the write performance metric value of the data stream is greater than the first threshold, it can first determine whether an expanded routing record indicating the data stream has been generated in the first cycle. If it has been generated, it can be determined as a continuous routing expansion operation, and then determine the timing order of this routing expansion in the continuous routing expansion operations. If it has not been generated, the timing order of this routing expansion can be determined to be 1.

[0261] Further, after the computer device determines the timing order, it can, according to this timing order, determine the number of expanded routing records for the data stream, and based on the existing routing records indicating the data stream in the first operator routing table of the first cycle, and in accordance with the number of expanded routing records, generate the expanded routing record indicating the data stream, and according to the existing routing records and the expanded routing record, generate the second operator routing table for the second cycle.

[0262] Step 705: The computer device determines the multi-cycle period to be counted, where the multi-cycle period includes the first cycle period and at least one historical cycle period before the first cycle period; obtains the write performance characteristics of each of the at least one historical cycle period; performs statistical analysis on the multi-cycle period according to the write performance characteristics of each of the multi-cycle period, and obtains the write performance metric values of each data stream written in the first cycle period; for each data stream written in the first cycle period, if the write performance metric value is less than or equal to the second threshold, determine the number of scaling routing records for the data stream according to a preset scaling strategy; according to the existing routing records indicating the data stream in the first operator routing table of the first cycle period, determine the candidate data write operators that have established routing records with the data stream; for each of the candidate data write operators, count the number of routing records corresponding to the candidate data write operator according to the routing records in the first operator routing table; select, from the candidate data write operators, the reduced data write operators with the number consistent with the number of scaling routing records in the order of priority of the routing record number in descending order, and obtain the scaling routing records indicating the data stream and the reduced data write operators; generate the second operator routing table of the second cycle period according to the existing routing records and the scaling routing records.

[0263] In some exemplary embodiments, the computer device may first determine the multi-cycle period to be counted, and then obtain the write performance characteristics of each of the at least one historical cycle period.

[0264] Further, after the computer device obtains the write performance characteristics of the first cycle period and each of the at least one historical cycle period, it may perform statistical analysis on the multi-cycle period according to the write performance characteristics of each of the multi-cycle period to obtain the write performance metric values of each data stream written in the first cycle period.

[0265] After receiving the write performance characteristics of the data stream processing engine in the first cycle period, the computer device may determine the write performance metric values of each data stream written in the first cycle period according to the write performance characteristics. Specifically, the computer device may use a preset algorithm and the write performance characteristics to determine the write performance metric values of each data stream written in the first cycle period.

[0266] Further, after the computer device determines the write performance metric values of each data stream written in the first cycle period, for each data stream written in the first cycle period, if the computer device determines that the write performance metric value of the data stream is less than or equal to the second threshold, it performs routing scaling based on the existing routing records indicating the data stream in the first operator routing table of the first cycle period to determine the scaling routing records indicating the data stream.

[0267] If the computer device determines that the write performance metric value of the data stream is greater than the second threshold, it does not perform route reduction based on the existing route record of the data stream indicated in the first operator routing table for the first cycle.

[0268] Furthermore, for each data stream written in the first cycle, the computer device can perform the above operations. After determining the reduced route record for each data stream, it can generate the second operator routing table for the second cycle based on the existing route record and the reduced route record.

[0269] Step 706: The computer device scans the data streams newly added by the data stream processing engine in the first cycle; determines the default number of route records for the newly added data stream according to the default operator allocation policy; counts the number of route records of each data write operator on the data stream processing engine according to the route records in the first operator routing table; selects the target data write operator with the number consistent with the default number of route records from the data write operators on the data stream processing engine in the ascending priority order of the number of route records; adds a target route record indicating the target data write operator and the data stream to the second operator routing table in the second cycle.

[0270] In some exemplary embodiments, the computer device can also scan the data streams newly added by the data stream processing engine in the first cycle and determine the default number of route records for the newly added data stream according to the default operator allocation policy.

[0271] Furthermore, after the computer device determines the default number of route records for the newly added data stream, it can count the number of route records of each data write operator on the data stream processing engine according to the route records in the first operator routing table and select the target data write operator with the number consistent with the default number of route records from the data write operators on the data stream processing engine in the ascending priority order of the number of route records.

[0272] Furthermore, after the computer device determines the target write operator, it can add a target route record indicating the target data write operator and the data stream to the second operator routing table in the second cycle.

[0273] In a specific application scenario, it is applicable to Internet of Things devices for real-time collection of various data. These devices first temporarily store the data in a cache, and the control node and the data stream processing engine are responsible for writing the data. Specifically, the data stream processing engine receives the first operator routing table of the first cycle from the control node through the data source operator. This routing table determines the writing paths of different types of Internet of Things device data (data streams), that is, through which data writing operators to write to the corresponding storage files, and these files can be used for long-term storage of device historical data or real-time analysis. The data source operator reads the data segments of the Internet of Things device data from the cache and distributes the data to the data writing operators according to the routing records indicated by the first operator routing table, so that the data writing operators can write the data segments. At the same time, for each device data stream, the data writing operator statistically obtains its performance characteristics during data writing in the first cycle and sends them to the control node.

[0274] After receiving the writing performance characteristics, the control node traverses the device data streams indicated by the first operator routing table. Due to hardware performance issues, some old devices have small amounts of data but unstable transmission and writing, and the processing takes a long time. The data streams of these devices are set to a locked state. During the traversal process, if the control node encounters a data stream in the locked state, it will skip it and continue to traverse the next one until it finds an unlocked device data stream.

[0275] For the unlocked device data streams, the control node determines the multi-cycle periods to be statistically analyzed, including the current first cycle and the historical cycle periods of several previous time periods. It obtains the writing performance characteristics of each historical cycle period, conducts statistical analysis, and obtains the writing performance metric values of each device data stream in the first cycle, such as the statistical value of data segment delay, the statistical value of data stream load, etc. At the same time, it obtains the number of data writing operators on the data stream processing engine and the concurrent routing upper limit value of each operator, and then determines the concurrent routing constraint value.

[0276] For each device data stream, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value, and the writing performance metric value is greater than the first threshold, and an expansion routing record for this device data stream has been generated previously, determine the timing order of this routing expansion in the continuous routing expansion operations. Then, select the new data writing operators in ascending order of the number of routing records from the candidate data writing operators that have not established routing records with this device data stream, establish new routing records, and generate the second operator routing table for the second cycle.

[0277] After completing the statistical analysis of multiple cycle periods, for the device data streams whose write performance metric values are less than or equal to the second threshold, determine the number of scaled-down routing records according to the preset scaling-down strategy. From the candidate data write operators that have established routing records with the device data stream, select the corresponding number of scaled-down data write operators in descending order of the number of routing records to obtain the scaled-down routing records, and generate the second operator routing table for the second cycle period.

[0278] With the update of IoT devices or new business requirements, some new devices may be added in the first cycle period, such as new environmental monitoring devices. After the control node scans the data streams of these new devices, according to the default operator allocation strategy, determine the default number of routing records for each new device data stream, and select the target data write operator from the data write operators in ascending order of the number of routing records, and add the target routing records for these new device data streams to the second operator routing table in the second cycle period.

[0279] In another specific application scenario, for the writing process of the log data of the application program. Specifically, assume that the log data is three data streams, namely data stream a, data stream b, and data stream c, and the data write operators configured in the data stream processing engine are three, namely Sink0, Sink1, and Sink2.

[0280] When the data stream processing engine starts to write the log data, the default operator allocation strategy can be adopted, that is, by default, each data stream is allocated a data write operator. After the data source operator consumes the data, the data stream is allocated to the data write operator according to the default operator allocation strategy. That is, data stream b - Sink0, data stream a - Sink1, data stream c - Sink2.

[0281] In a certain cycle period, the traffic of data stream a increases. The control node determines that its write performance metric value is greater than the expansion threshold (the first threshold) through the write performance characteristics, and then allocates another data write operator Sink - 2 with the fewest routing records to data stream a. Then the operator routing table for the next cycle period is data stream b - Sink0, data stream a - Sink1, data stream a - Sink2, data stream c - Sink2. In the next cycle period, the data processing engine will perform data writing for the data streams of the log data according to this operator routing table.

[0282] After at least one cycle, in another cycle, if the control node determines through the write performance characteristics that the write performance metric value of data stream a is greater than the expansion threshold (the first threshold) again, then allocate another data write operator Sink-0 with the fewest routing records to data stream a. Then the operator routing table for the next cycle is data stream b - Sink0, data stream a - Sink1, data stream a - Sink2, data stream a - Sink0, data stream c - Sink2. In the next cycle, the data processing engine will perform data writing for the data streams of the log data according to this operator routing table.

[0283] After at least one cycle, in another cycle, if the control node determines through the write performance characteristics that the write performance metric value of data stream c is also greater than the expansion threshold (the first threshold), then allocate a data write operator Sink-1 with the fewest routing records to data stream c as well. Then the operator routing table for the next cycle is data stream b - Sink0, data stream a - Sink1, data stream a - Sink2, data stream a - Sink0, data stream c - Sink2, data stream c - Sink1. In the next cycle, the data processing engine will perform data writing for the data streams of the log data according to this operator routing table.

[0284] After at least one cycle, in another cycle, if the control node determines through the write performance characteristics that the write performance metric values of data stream a and data stream c are less than the contraction threshold (the second threshold), then reduce the data write operators with more routing records for data stream c and data stream a. Then the operator routing table for the next cycle is data stream b - Sink0, data stream a - Sink2, data stream c - Sink1. In the next cycle, the data processing engine will perform data writing for the data streams of the log data according to this operator routing table.

[0285] To clearly describe the beneficial effects of the technical solution of this application, some prior arts will be explained below.

[0286] In some prior arts, a separate data stream processing engine is set for each data stream, and the number of data write operators in the corresponding data stream processing engine is configured according to the data volume of the data stream. Specifically, as shown in Figure 8 、 Figure 9 and Figure 10 shown, Figure 8 、 Figure 9 and Figure 10 respectively show setting different data stream processing engines for three data streams with different data volumes. For example, Figure 8The data stream in it is data stream a, and its data volume is large. Therefore, the number of data writing operators in the corresponding data stream processing engine is 3, namely data writing operator 0, data writing operator 1, and data writing operator 2. The data source operator first obtains data stream a from the cache, then distributes data stream a to data writing operator 0, data writing operator 1, and data writing operator 2, and then data writing operator 0, data writing operator 1, and data writing operator 2 write data stream a to the file system, namely file - data writing operator 0 - data stream a, file - data writing operator 1 - data stream a, file - data writing operator 2 - data stream a. Figure 9 The data stream in it is data stream b, and its data volume is small. Therefore, the number of data writing operators in the corresponding data stream processing engine is 1, namely data writing operator 0. The data source operator first obtains data stream b from the cache, then distributes data stream b to data writing operator 0, and then data writing operator 0 writes data stream b to the file system, namely file - data writing operator 0 - data stream b. Figure 10 The data stream in it is data stream c, and its data volume is smaller than that of data stream a and larger than that of data stream b. Therefore, the number of data writing operators in the corresponding data stream processing engine is 2, namely data writing operator 0 and data writing operator 1. The data source operator first obtains data stream c from the cache, then distributes data stream c to data writing operator 0 and data writing operator 1, and then data writing operator 0 and data writing operator 1 write data stream c to the file system, namely file - data writing operator 0 - data stream c, file - data writing operator 1 - data stream c.

[0287] However, although such a method realizes the resource independence of data streams, the cost of setting a data stream processing engine for each data stream is relatively high, especially in the scenario where the number of data streams is large.

[0288] Based on this, in some other existing technologies, first, all data writing operators in the data writing engine write the data stream to the intermediate file system. When each data writing operator writes the data stream to the intermediate file system, a corresponding small file will be generated, so that the same data stream is stored in the intermediate file system in the form of multiple small files. Then, the small files of the same data stream in the intermediate file system are integrated, and the large file corresponding to the integrated data stream is stored in the final file system. Such a method only needs to set one data stream processing engine. Specifically, as Figure 11 shown, in Figure 11In it, the data source operator in the data stream processing engine can obtain the data streams to be written from the cache, namely data stream a, data stream b, and data stream c, and allocate the data streams to the data writing operators, namely data writing operator 0, data writing operator 1, and data writing operator 2. Then, the data writing operators write the data streams to be written into the intermediate file system to generate multiple small files, namely file - data writing operator 0 - data stream a, file - data writing operator 0 - data stream b,......

[0289] Furthermore, the spark timing job can periodically obtain the small files of the same data stream from the intermediate file system, summarize them, and store the summarized large file in the final file system.

[0290] However, although this method solves the problem that the same data stream is stored in the form of multiple small files in the file system, resulting in poor data reading efficiency. However, setting up the intermediate file system and integrating the small files of the same data stream in the intermediate file system not only makes the writing cost of the data stream relatively high, but also makes the writing efficiency of the data stream relatively low, thus leading to poor writing performance of the data stream.

[0291] In an exemplary embodiment, as Figure 12 shown, a data stream writing control system is provided. The system includes a control node, a data stream processing engine, a file system, and a cache. The data stream processing engine includes a data writing operator and a data source operator;

[0292] The data source operator is used to receive the first operator routing table of the first cycle from the control node. The first operator routing table contains at least one routing record, and each routing record indicates a data stream to be written and a corresponding data writing operator; read the data segments of each data stream from the cache, and allocate the data segments to each data writing operator according to the routing records in the operator routing table; wherein, the writing performance characteristic is obtained and sent by the data writing operator.

[0293] The data writing operator is used to write the data stream to be written corresponding to the data writing operator indicated by the routing record into the file system, the file corresponding to the data writing operator and the data stream to be written; obtain the writing performance characteristic of the writing performed through the data writing operator during the first cycle; send the writing performance characteristic to the control node.

[0294] The control node is configured to receive the write performance characteristics of the data writing operator in the first cycle. The write performance characteristics characterize the write performance of the data stream processing engine for writing each data stream in the first cycle. According to the write performance characteristics, the first operator routing table in the first cycle is adjusted to obtain the second operator routing table in the second cycle. The first operator routing table includes routing records in the first cycle, and the second operator routing table includes routing records in the second cycle. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator. The second operator routing table is sent to the data source operator. The operator routing table is used to indicate that in the second cycle, the data stream processing engine uses the data writing operators indicated by each routing record in the second operator routing table to write the data streams to be written corresponding to the data writing operators into the file system, i.e., the files corresponding to the data writing operators and the data streams to be written.

[0295] In some exemplary embodiments, as Figure 12 shown, the data source operator can obtain data sources from the cache, namely data source a, data source b, and data source c, and allocate the data sources to data writing operator 0, data writing operator 1, and data writing operator 2 according to the operator routing table obtained from the control node. Then, data writing operator 0, data writing operator 1, and data writing operator 2 write data stream a, data source b, and data source c into the file system respectively, i.e., file - data writing operator 0 - data stream a, file - data writing operator 0 - data stream b, file - data writing operator 1 - data stream a, file - data writing operator 1 - data stream c, file - data writing operator 2 - data stream a, file - data writing operator 2 - data stream c. Meanwhile, the data writing operator also sends the write performance characteristics to the control node.

[0296] It should be understood that although the steps in the flowcharts involved in the above - described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0297] Based on the same inventive concept, an embodiment of the present application further provides a data stream writing control device for implementing the data stream writing control method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the data stream writing control device provided below can refer to the limitations on the data stream writing control method in the above text, and will not be elaborated here.

[0298] In an exemplary embodiment, as Figure 13 shown, a data stream writing control device 1300 is provided, including: a receiving module 1301, a first execution module 1302, and a sending module 1303, where:

[0299] The receiving module 1301 is configured to receive the writing performance characteristics of the data stream processing engine in the first cycle period, and the writing performance characteristics characterize the writing performance of the data stream processing engine for writing each data stream in the first cycle period;

[0300] The first execution module 1302 is configured to adjust the first operator routing table in the first cycle period according to the writing performance characteristics to obtain a second operator routing table in the second cycle period. The first operator routing table includes routing records in the first cycle period, and the second operator routing table includes routing records in the second cycle period. Each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator;

[0301] The sending module 1303 is configured to send the second operator routing table to the data stream processing engine. The second operator routing table is used to instruct the data stream processing engine to write the data stream to be written corresponding to each routing record in the second operator routing table into the file corresponding to the data writing operator and the data stream to be written in the second cycle period.

[0302] In an embodiment, the first execution module 1302 is specifically configured to determine the writing performance index values of each data stream written in the first cycle period according to the writing performance characteristics; for each data stream written in the first cycle period, if the writing performance index value is greater than a first threshold, perform routing expansion based on the existing routing record indicating the data stream in the first operator routing table in the first cycle period to generate an expanded routing record indicating the data stream; generate a second operator routing table in the second cycle period according to the existing routing record and the expanded routing record.

[0303] In one embodiment, the first execution module 1302 is further configured to obtain the number of operators of the data writing operator on the data stream processing engine, and obtain the concurrent routing upper limit value of each data writing operator on the data stream processing engine; determine the concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value; specifically, for each data stream written in the first cycle period, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value and the writing performance metric value is greater than the first threshold, the first execution module 1302 generates an expanded routing record indicating the data stream based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period.

[0304] In one embodiment, specifically, for each data stream written in the first cycle period, if the writing performance metric value is greater than the first threshold, the first execution module 1302 determines the number of expanded routing records for the data stream according to a preset expansion strategy; determines candidate data writing operators that have not established routing records with the data stream according to the existing routing records indicating the data stream in the first operator routing table of the first cycle period; for each candidate data writing operator, counts the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table; selects new data writing operators with a quantity consistent with the number of expanded routing records from the candidate data writing operators in ascending order of the number of routing records; and respectively establishes routing records indicating each new data writing operator and the data stream to obtain an expanded routing record indicating the data stream.

[0305] In one embodiment, specifically, for each data stream written in the first cycle period, if the writing performance metric value is greater than the first threshold and an expanded routing record indicating the data stream has been generated for the first cycle period, the first execution module 1302 determines the timing order of this routing expansion in consecutive routing expansion operations; determines the number of expanded routing records for the data stream according to the timing order, where the number of expanded routing records is positively correlated with the timing order; and generates an expanded routing record indicating the data stream based on the existing routing record indicating the data stream in the first operator routing table of the first cycle period and according to the number of expanded routing records.

[0306] In one embodiment, the first execution module 1302 is specifically configured to determine the write performance metric values of each data stream written in the first cycle according to the write performance characteristics; for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, perform routing capacity reduction based on the existing routing record indicating the data stream in the first operator routing table of the first cycle, and determine the reduced routing record indicating the data stream; generate the second operator routing table of the second cycle according to the existing routing record and the reduced routing record.

[0307] In one embodiment, the first execution module 1302 is specifically configured to, for each data stream written in the first cycle, if the write performance metric value is less than or equal to the second threshold, determine the number of reduced routing records for the data stream according to a preset capacity reduction policy; determine the candidate data write operators that have established routing records with the data stream according to the existing routing record indicating the data stream in the first operator routing table of the first cycle; for each candidate data write operator, count the number of routing records corresponding to the candidate data write operator according to the routing record in the first operator routing table; select the reduced data write operators with the same number as the number of reduced routing records from the candidate data write operators in the descending order of the priority of the number of routing records, and obtain the reduced routing record indicating the data stream and the reduced data write operators.

[0308] In one embodiment, the first execution module 1302 is specifically configured to determine the multi-cycle to be statistically analyzed, where the multi-cycle includes the first cycle and at least one historical cycle before the first cycle; obtain the write performance characteristics of each of the at least one historical cycle; perform statistical analysis of the multi-cycle according to the write performance characteristics of each of the multi-cycles, and obtain the write performance metric values of each data stream written in the first cycle.

[0309] In one embodiment, each data stream includes a plurality of data segments. The write performance characteristics include a publish timestamp and a consumption timestamp for each of the data segments. The publish timestamp represents the time when the data segment is written to the cache, and the consumption timestamp represents the time when the data segment read from the cache is written to the file. The write performance metric values of the data streams written in the first cycle period include the data segment latency statistical values for multiple cycle periods. The first execution module 1302 is specifically configured to determine the latency value of each data segment according to the publish timestamp and the consumption timestamp of each data segment written in each cycle period within the multiple cycle periods; for each data stream, determine the total latency value of the data stream according to the latency values of the data segments in the data stream written in the multiple cycle periods; for each data stream, count the total number of data segments in the data stream written in the multiple cycle periods; for each data stream, determine the data segment latency statistical value of the data stream in the multiple cycle periods according to the total latency value of the data stream and the total number of data segments in the data stream.

[0310] In one embodiment, the write performance characteristics include the amount of data written to each data stream in each cycle period. The write performance metric values of the data streams written in the first cycle period include the data stream load statistical values for multiple cycle periods. The first execution module 1302 is specifically configured to, for each data stream written in the first cycle period, count the total amount of data written to the data stream in the multiple cycle periods according to the amount of data written to the data stream in each cycle period; for each data stream written in the first cycle period, count the number of routing records indicating the data stream from the first operator routing table; for each data stream written in the first cycle period, determine the data stream load statistical value of the data stream in the multiple cycle periods according to the total amount of data, the number of routing records, and the number of cycle periods in the multiple cycle periods.

[0311] In one embodiment, the first execution module 1302 is further configured to scan the data streams newly added by the data stream processing engine in the first cycle period; determine the default number of routing records of the newly added data streams according to the default operator allocation policy; count the number of routing records of each data write operator on the data stream processing engine according to the routing records in the first operator routing table; select a target data write operator with a number consistent with the default number of routing records from the data write operators on the data stream processing engine in the ascending priority order of the number of routing records; add a target routing record indicating the target data write operator and the data stream to the second operator routing table in the second cycle period.

[0312] In one embodiment, the first execution module 1302 is specifically configured to traverse the data streams indicated by the routing records in the first operator routing table; in response to the traversed data stream being in a locked state, continue to traverse the next data stream; in response to the traversed data stream being in an unlocked state, adjust the routing record for the traversed data stream according to the write performance characteristic; until the traversal is complete, obtain the second operator routing table for the second cycle.

[0313] In an exemplary embodiment, as Figure 14 shown, a data stream write control device 1400 is provided, including: a first acquisition module 1401, a second execution module 1402, a second acquisition module 1403, and a sending module 1404, where:

[0314] The first acquisition module 1401 is configured to acquire a first operator routing table for a first cycle, where the first operator routing table includes at least one routing record, and each routing record indicates a data stream to be written and a corresponding data write operator;

[0315] The second execution module 1402 is configured to use the data write operator indicated by each routing record to write the data stream to be written corresponding to the data write operator into a file corresponding to the data write operator and the data stream to be written;

[0316] The second acquisition module 1403 is configured to, for each data stream to be written, acquire a write performance characteristic of writing through at least one data write operator within the first cycle;

[0317] The sending module 1404 is configured to send the write performance characteristic, and the sent write performance characteristic is used to indicate adjusting the routing record in the first operator routing table according to the write performance characteristic to obtain a second operator routing table for a second cycle, and the write performance characteristic characterizes the write performance of writing each data stream in the first cycle.

[0318] In one embodiment, the first acquisition module 1401 is specifically configured to use a data source operator to receive the first operator routing table for the first cycle from a control node; the first acquisition module is further configured to use the data source operator to read data segments of each data stream from a cache, and allocate the data segments to each data write operator according to the routing records in the first operator routing table; where the write performance characteristic is acquired and sent by the data write operator.

[0319] Each module in the above data stream writing control device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0320] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a data stream writing control method.

[0321] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 16 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC), or other technologies. When the computer program is executed by the processor, it implements a data stream writing control method.

[0322] Those skilled in the art can understand,Figure 15 and Figure 16 The structure shown in Figure 16 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0323] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned embodiments of the data stream writing control method are implemented.

[0324] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned embodiments of the data stream writing control method are implemented.

[0325] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above-mentioned embodiments of the data stream writing control method are implemented.

[0326] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0327] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0328] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A data stream writing control method, characterized in that: The method comprises: receiving a write performance characteristic of a data stream processing engine in a first cycle, wherein the write performance characteristic represents a write performance of the data stream processing engine in writing each data stream in the first cycle; According to the write performance characteristics, the first operator routing table of the first cycle is adjusted to obtain the second operator routing table of the second cycle, the first cycle and the second cycle are time periods of fixed length, the first operator routing table includes routing records within the first cycle, the second operator routing table includes routing records within the second cycle, each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data write operator; the first operator routing table of the first cycle is adjusted to include routing expansion and routing contraction; the number of expanded routing records of the routing expansion is determined according to the timing order of this routing expansion in the continuous routing expansion operation, and the number of expanded routing records is positively correlated with the timing order; after the continuous routing expansion operation is interrupted, the number of expanded routing records of the routing expansion performed again is the initial expanded routing number; the number of contracted routing records of the routing contraction is a preset number; The second operator routing table is sent to the data stream processing engine, where the second operator routing table is used to instruct the data stream processing engine to use the data write operator indicated by each routing record in the second operator routing table in the second cycle to write the data into the data stream to be written corresponding to the data write operator, and write the data into the file corresponding to the data write operator and the data stream to be written in the file system.

2. The method according to claim 1, characterized in that According to the write performance characteristics, adjusting the first operator routing table of the first cycle to obtain the second operator routing table of the second cycle includes: Determining, according to the write performance characteristics, respective write performance index values ​​of respective data streams written in the first cycle; For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, perform route expansion based on an existing route record indicating the data stream in the first operator routing table of the first cycle, and generate an expanded route record indicating the data stream; A second operator routing table of a second cycle is generated according to the existing routing record and the expanded routing record.

3. The method according to claim 2, characterized in that The method further comprises: Obtain the number of operators of the data writing operator on the data stream processing engine, and obtain the concurrent routing upper limit value of each data writing operator on the data stream processing engine; Determining a concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value; For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, performing route expansion based on an existing route record indicating the data stream in a first operator routing table of the first cycle, and generating an expanded route record indicating the data stream, including: For each data stream written in the first cycle, if the number of routing records in the first operator routing table is less than the concurrent routing constraint value, and the write performance indicator value is greater than the first threshold, then based on the existing routing records indicating the data stream in the first operator routing table of the first cycle, an expanded routing record indicating the data stream is generated.

4. The method according to claim 2, characterized in that For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, performing route expansion based on an existing route record indicating the data stream in a first operator routing table of the first cycle, and generating an expanded route record indicating the data stream, including: For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, determining the number of expansion routing records for the data stream according to a preset expansion strategy; Determine, according to the existing routing record indicating the data flow in the first operator routing table of the first cycle, a candidate data writing operator for which no routing record is established with the data flow; For each of the candidate data writing operators, counting the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table; From the candidate data writing operators, select new data writing operators whose number is the same as the number of the expanded routing records in ascending order of priority of the number of routing records; A routing record indicating each of the newly added data writing operators and the data flow is established respectively, and an expanded routing record indicating the data flow is obtained.

5. The method according to claim 2, characterized in that: For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, based on an existing routing record indicating the data stream in a first operator routing table of the first cycle, generating an expanded routing record indicating the data stream, including: For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold value, and an expansion route record indicating the data stream has been generated for the first cycle, determining a timing order of this route expansion in the continuous route expansion operation; Determining the number of expansion routing records for the data flow according to the timing sequence; Based on the existing routing records indicating the data flow in the first operator routing table of the first cycle period, and according to the number of the expanded routing records, an expanded routing record indicating the data flow is generated.

6. The method according to claim 1, characterized in that According to the write performance characteristics, adjusting the first operator routing table of the first cycle to obtain the second operator routing table of the second cycle includes: Determining, according to the write performance characteristics, respective write performance index values ​​of respective data streams written in the first cycle; For each data stream written in the first cycle, if the write performance indicator value is less than or equal to the second threshold, perform routing shrinkage based on an existing routing record indicating the data stream in the first operator routing table of the first cycle, and determine a shrinkage routing record indicating the data stream; A second operator routing table of a second cycle is generated according to the existing routing record and the reduced capacity routing record.

7. The method according to claim 6, characterized in that: For each data stream written in the first cycle, if the write performance indicator value is less than or equal to the second threshold, scaling down the route based on the existing route record indicating the data stream in the first operator routing table of the first cycle, and determining the scaled-down route record indicating the data stream, including: For each data stream written in the first cycle, if the write performance indicator value is less than or equal to a second threshold, determining the number of shrinking routing records for the data stream according to a preset shrinking strategy; Determine, according to the existing routing record indicating the data flow in the first operator routing table of the first cycle, a candidate data writing operator for which a routing record has been established with the data flow; For each of the candidate data writing operators, counting the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table; From the candidate data writing operators, in descending order of priority of the number of routing records, a reduction data writing operator having a number consistent with the number of the reduction routing records is selected to obtain a reduction routing record indicating the data flow and the reduction data writing operator.

8. The method according to any one of claims 2 to 7, characterized in that: Determining the write performance index value of each data stream written in the first cycle according to the write performance characteristic includes: Determine a multi-cycle period to be counted, wherein the multi-cycle period includes the first cycle period and at least one historical cycle period before the first cycle period; Obtaining a write performance characteristic of each of the at least one historical cycle; According to the write performance characteristics of each of the multiple cycle periods, a statistical analysis of the multiple cycle periods is performed to obtain the write performance index value of each data stream written in the first cycle period.

9. The method according to claim 8, characterized in that Each data stream includes a plurality of data segments, and the write performance characteristics include a publishing timestamp and a consumption timestamp of each data segment, wherein the publishing timestamp indicates a time when the data segment is written into a cache, and the consumption timestamp indicates a time when the data segment read from the cache is written into the file; The write performance index value of each data stream written in the first cycle period, including the data segment delay statistics of multiple cycle periods; According to the write performance characteristics of each of the multiple cycles, statistical analysis of the multiple cycles is performed to obtain the write performance index value of each of the data streams written in the first cycle, including: Determine a delay value of each data segment according to a publishing timestamp and a consumption timestamp of each data segment written in each cycle within the multi-cycle period; For each data stream, determining a total delay value of the data stream according to delay values ​​of each data segment in the data stream written in the multi-cycle period; For each data stream, counting the total number of data segments in the data stream written in the multi-cycle period; For each data flow, a data segment delay statistic value of the data flow in the multi-cycle period is determined according to a total delay value of the data flow and a total number of data segments of the data flow.

10. The method according to claim 8, characterized in that The write performance characteristics include the amount of data written to each data stream in each cycle; the write performance index values ​​of each data stream written in the first cycle, including the data stream load statistics of multiple cycles; According to the write performance characteristics of each of the multiple cycles, statistical analysis of the multiple cycles is performed to obtain the write performance index value of each of the data streams written in the first cycle, including: For each data stream written in the first cycle, according to the amount of data written to the data stream in each cycle, counting the total amount of data written to the data stream in the multiple cycles; For each data flow written in the first cycle, counting the number of routing records indicating the routing records of the data flow from the first operator routing table; For each data stream written in the first cycle, a data stream load statistic value of the data stream in the multi-cycle period is determined according to the total amount of data, the number of routing records and the number of periods in the multi-cycle period.

11. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Scan the data stream added by the data stream processing engine in the first cycle; Determine the number of default route records for the newly added data flow according to the default operator allocation strategy; According to the routing records in the first operator routing table, counting the number of routing records of each data writing operator on the data stream processing engine; From the data writing operators on the data stream processing engine, in order of priority in ascending order of the number of routing records, select a target data writing operator whose number is consistent with the number of the default routing records; In the second operator routing table of the second cycle, a target routing record indicating the target data write operator and the data flow is added.

12. The method according to any one of claims 1 to 7, characterized in that: According to the write performance characteristics, adjusting the first operator routing table of the first cycle to obtain the second operator routing table of the second cycle includes: Traversing the data stream indicated by the routing record in the first operator routing table; In response to the traversed data stream being in a locked state, continuing to traverse the next data stream; In response to the traversed data stream being in an unlocked state, adjusting a routing record for the traversed data stream according to the write performance characteristic; After the traversal is completed, the second operator routing table of the second cycle is obtained.

13. A data stream writing control method, characterized in that: The method comprises: Obtaining a first operator routing table of a first cycle, wherein the first operator routing table includes at least one routing record, each routing record indicating a data stream to be written and a corresponding data writing operator; Using the data writing operator indicated by each of the routing records, write the data stream to be written corresponding to the data writing operator into a file in the file system corresponding to the data writing operator and the data stream to be written; For each data stream to be written, obtaining a write performance characteristic of writing by at least one data writing operator in the first cycle; The write performance characteristic is sent, and the sent write performance characteristic is used to indicate adjusting the routing records in the first operator routing table according to the write performance characteristic, and obtaining the second operator routing table of the second cycle, wherein the write performance characteristic characterizes the write performance of writing each data stream in the first cycle, and the first cycle and the second cycle are time periods of fixed length; the first operator routing table for adjusting the first cycle includes routing expansion and routing contraction; the number of expanded routing records of the routing expansion is determined according to the timing order of this routing expansion in the continuous routing expansion operation, and the number of expanded routing records is positively correlated with the timing order; after the continuous routing expansion operation is interrupted, the number of expanded routing records of the routing expansion performed again is the initial number of expanded routing records; the number of contracted routing records of the routing contraction is determined according to a preset contraction policy, and the preset contraction policy indicates that the number of contracted routing records for each routing contraction is a fixed number.

14. The method according to claim 13, characterized in that The obtaining of the first operator routing table of the first cycle includes: Using the data source operator, receiving a first operator routing table of a first cycle from the control node; The method further comprises: Using the data source operator, read the data segments of each data stream from the cache, and distribute the data segments to each data write operator according to the routing records in the first operator routing table; The write performance characteristics are acquired and sent by the data write operator.

15. A data stream writing control system, characterized in that: The system includes a control node, a data stream processing engine, a file system and a cache, and the data stream processing engine includes a data writing operator and a data source operator; The data source operator is used to receive a first operator routing table of a first cycle from a control node, wherein the first operator routing table includes at least one routing record, each routing record indicating a data stream to be written and a corresponding data writing operator; read data segments of each data stream from the cache, and distribute the data segments to each data writing operator according to the routing records in the operator routing table; The data writing operator is used to write the data stream to be written corresponding to the data writing operator indicated by the routing record into the file system, and the file corresponding to the data writing operator and the data stream to be written; Acquire a write performance characteristic of writing by the data write operator in the first cycle; send the write performance characteristic to a control node; wherein the write performance characteristic is acquired and sent by the data write operator; The control node is used to receive the write performance characteristics of the data write operator in the first cycle, the write performance characteristics characterizing the write performance of the data stream processing engine in writing each data stream in the first cycle; according to the write performance characteristics, adjust the first operator routing table of the first cycle to obtain the second operator routing table of the second cycle, the first operator routing table includes routing records in the first cycle, the second operator routing table includes routing records in the second cycle, each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data write operator; send the second operator routing table to the data source operator, the operator routing table is used to instruct the data stream processing engine to use the data indicated by each routing record in the second operator routing table in the second cycle. According to the write operator, the data stream to be written corresponding to the data write operator is written into the file system, and the data write operator and the file corresponding to the data stream to be written are written; the first cycle period and the second cycle period are time periods of fixed length; the first operator routing table for adjusting the first cycle period includes routing expansion and routing contraction; the number of expanded routing records of the routing expansion is determined according to the timing order of this routing expansion in the continuous routing expansion operation, and the number of expanded routing records is positively correlated with the timing order; after the continuous routing expansion operation is interrupted, the number of expanded routing records of the routing expansion performed again is the initial number of expanded routing; the number of contracted routing records of the routing contraction is determined according to a preset contraction strategy, and the preset contraction strategy indicates that the number of contracted routing records for each routing contraction is a fixed number.

16. A data stream writing control device, characterized in that: The device comprises: A receiving module, configured to receive a write performance characteristic of a data stream processing engine in a first cycle, wherein the write performance characteristic represents a write performance of the data stream processing engine in writing each data stream in the first cycle; A first execution module is used to adjust the first operator routing table of the first cycle according to the write performance characteristics, and obtain the second operator routing table of the second cycle, wherein the first cycle and the second cycle are time periods of fixed duration, the first operator routing table includes routing records within the first cycle, and the second operator routing table includes routing records within the second cycle, and each routing record indicates a data stream to be written in the data stream processing engine and a corresponding data writing operator; the first operator routing table of the first cycle is adjusted to include routing expansion and routing contraction; the number of expanded routing records of the routing expansion is determined according to the timing order of the current routing expansion in the continuous routing expansion operation, and the number of expanded routing records is positively correlated with the timing order; after the continuous routing expansion operation is interrupted, the number of expanded routing records of the routing expansion performed again is the initial number of expanded routing records; the number of contracted routing records of the routing contraction is determined according to a preset contraction strategy, and the preset contraction strategy indicates that the number of contracted routing records for each routing contraction is a fixed number; A sending module is used to send the second operator routing table to the data stream processing engine, where the second operator routing table is used to instruct the data stream processing engine to use the data write operator indicated by each routing record in the second operator routing table in the second cycle period to write the data into the data stream to be written corresponding to the data write operator, and write the data into the file corresponding to the data write operator and the data stream to be written in the file system.

17. The device according to claim 16, characterized in that The first execution module is specifically configured to determine a write performance indicator value of each data stream written in the first cycle according to the write performance characteristics; For each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold, perform route expansion based on an existing route record indicating the data stream in the first operator routing table of the first cycle, and generate an expanded route record indicating the data stream; A second operator routing table of a second cycle is generated according to the existing routing record and the expanded routing record.

18. The device according to claim 17, characterized in that The first execution module is further used to obtain the number of operators of the data writing operator on the data stream processing engine, and obtain the concurrent routing upper limit value of each data writing operator on the data stream processing engine; Determining a concurrent routing constraint value of the data stream processing engine according to the number of operators and the concurrent routing upper limit value; The first execution module is specifically used to generate an expanded routing record indicating the data flow based on the existing routing record indicating the data flow in the first operator routing table of the first cycle period for each data flow written in the first cycle period if the number of routing records in the first operator routing table is less than the concurrent routing constraint value and the write performance indicator value is greater than the first threshold.

19. The device according to claim 17, characterized in that The first execution module is specifically configured to determine, for each data stream written in the first cycle, a number of expansion routing records for the data stream according to a preset expansion strategy if the write performance indicator value is greater than a first threshold; Determine, according to the existing routing record indicating the data flow in the first operator routing table of the first cycle, a candidate data writing operator for which no routing record is established with the data flow; For each of the candidate data writing operators, counting the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table; From the candidate data writing operators, select new data writing operators whose number is the same as the number of the expanded routing records in ascending order of priority of the number of routing records; A routing record indicating each of the newly added data writing operators and the data flow is established respectively, and an expanded routing record indicating the data flow is obtained.

20. The device according to claim 17, characterized in that The first execution module is specifically configured to determine, for each data stream written in the first cycle, if the write performance indicator value is greater than a first threshold value and an expansion route record indicating the data stream has been generated for the first cycle, a timing order of the current route expansion in the continuous route expansion operation; Determining the number of expansion routing records for the data flow according to the timing sequence; Based on the existing routing records indicating the data flow in the first operator routing table of the first cycle period, and according to the number of the expanded routing records, an expanded routing record indicating the data flow is generated.

21. The device according to claim 16, characterized in that The first execution module is specifically configured to determine a write performance indicator value of each data stream written in the first cycle according to the write performance characteristics; For each data stream written in the first cycle, if the write performance indicator value is less than or equal to the second threshold, perform routing shrinkage based on an existing routing record indicating the data stream in the first operator routing table of the first cycle, and determine a shrinkage routing record indicating the data stream; A second operator routing table of a second cycle is generated according to the existing routing record and the reduced capacity routing record.

22. The device according to claim 21, characterized in that The first execution module is specifically configured to determine, for each data stream written in the first cycle, a number of shrinkage routing records for the data stream according to a preset shrinkage strategy if the write performance indicator value is less than or equal to a second threshold; Determine, according to the existing routing record indicating the data flow in the first operator routing table of the first cycle, a candidate data writing operator for which a routing record has been established with the data flow; For each of the candidate data writing operators, counting the number of routing records corresponding to the candidate data writing operator according to the routing records in the first operator routing table; From the candidate data writing operators, in descending order of priority of the number of routing records, a reduction data writing operator having a number consistent with the number of the reduction routing records is selected to obtain a reduction routing record indicating the data flow and the reduction data writing operator.

23. The device according to any one of claims 17 to 22, characterized in that The first execution module is specifically used to determine a multi-cycle period to be counted, wherein the multi-cycle period includes the first cycle period and at least one historical cycle period before the first cycle period; Obtaining a write performance characteristic of each of the at least one historical cycle; According to the write performance characteristics of each of the multiple cycle periods, a statistical analysis of the multiple cycle periods is performed to obtain the write performance index value of each data stream written in the first cycle period.

24. The device according to claim 23, characterized in that The write performance index value of each data stream written in the first cycle period, including the data segment delay statistics of multiple cycle periods; The first execution module is specifically used to determine the delay value of each data segment according to the publishing timestamp and the consumption timestamp of each data segment written in each cycle within the multi-cycle period; For each data stream, determining a total delay value of the data stream according to delay values ​​of each data segment in the data stream written in the multi-cycle period; For each data stream, counting the total number of data segments in the data stream written in the multi-cycle period; For each data flow, a data segment delay statistic value of the data flow in the multi-cycle period is determined according to a total delay value of the data flow and a total number of data segments of the data flow.

25. The device according to claim 23, characterized in that The write performance characteristics include the amount of data written to each data stream in each cycle; the write performance index values ​​of each data stream written in the first cycle, including the data stream load statistics of multiple cycles; The first execution module is specifically configured to count the total amount of data written to the data stream in the multiple cycles according to the amount of data written to the data stream in each cycle for each data stream written in the first cycle; For each data flow written in the first cycle, counting the number of routing records indicating the routing records of the data flow from the first operator routing table; For each data stream written in the first cycle, a data stream load statistic value of the data stream in the multi-cycle period is determined according to the total amount of data, the number of routing records and the number of periods in the multi-cycle period.

26. The device according to any one of claims 16 to 22, characterized in that The first execution module is further used to scan the data stream added by the data stream processing engine in the first cycle; Determine the number of default route records for the newly added data flow according to the default operator allocation strategy; According to the routing records in the first operator routing table, counting the number of routing records of each data writing operator on the data stream processing engine; From the data writing operators on the data stream processing engine, in order of priority in ascending order of the number of routing records, select a target data writing operator whose number is consistent with the number of the default routing records; In the second operator routing table of the second cycle, a target routing record indicating the target data write operator and the data flow is added.

27. The device according to any one of claims 16 to 22, characterized in that The first execution module is specifically configured to traverse the data stream indicated by the routing record in the first operator routing table; In response to the traversed data stream being in a locked state, continuing to traverse the next data stream; In response to the traversed data stream being in an unlocked state, adjusting a routing record for the traversed data stream according to the write performance characteristic; After the traversal is completed, the second operator routing table of the second cycle is obtained.

28. A data stream writing control device, characterized in that: The device comprises: A first acquisition module, used to acquire a first operator routing table of a first cycle, wherein the first operator routing table includes at least one routing record, each routing record indicating a data stream to be written and a corresponding data writing operator; A second execution module is used to use the data writing operator indicated by each routing record to write the data stream to be written corresponding to the data writing operator into a file in the file system corresponding to the data writing operator and the data stream to be written; A second acquisition module, configured to acquire, for each data stream to be written, a write performance characteristic of writing by at least one data writing operator in the first cycle; A sending module is used to send the write performance characteristics, and the sent write performance characteristics are used to indicate adjusting the routing records in the first operator routing table according to the write performance characteristics to obtain the second operator routing table of the second cycle, wherein the write performance characteristics characterize the write performance of writing each data stream in the first cycle, and the first cycle and the second cycle are time periods of fixed length; the first operator routing table for adjusting the first cycle includes routing expansion and routing contraction; the number of expanded routing records of the routing expansion is determined according to the timing order of this routing expansion in the continuous routing expansion operation, and the number of expanded routing records is positively correlated with the timing order; after the continuous routing expansion operation is interrupted, the number of expanded routing records of the routing expansion performed again is the initial number of expanded routing records; the number of contracted routing records of the routing contraction is determined according to a preset contraction strategy, and the preset contraction strategy indicates that the number of contracted routing records for each routing contraction is a fixed number.

29. The device according to claim 28, characterized in that The first acquisition module is specifically configured to receive a first operator routing table of a first cycle from a control node using a data source operator; Using the data source operator, read the data segments of each data stream from the cache, and distribute the data segments to each data write operator according to the routing records in the first operator routing table; The write performance characteristics are acquired and sent by the data write operator.

30. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 14 are implemented.

31. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.

32. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.

Citation Information

Patent Citations

  • Self-adaptive data circulation method, system and device and storage medium

    CN117807272A