Cross-cloud region optimization scheduling method and system oriented to stream computing application

Through the optimization methods of operator topology reconstruction, cross-domain split execution and elastic scaling of Flink stream computing applications, the problems of low efficiency and high cost of cross-cloud area collaborative scheduling in the existing technology are solved, and the resource utilization rate and cloud cost saving are achieved.

CN120216104APending Publication Date: 2025-06-27SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311792284.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the existing Flink stream computing framework coordinates scheduling across cloud regions, there are problems such as resource sharing mechanism restrictions and inability to schedule to specified nodes, resulting in low computing efficiency and high cloud resource costs.

Method used

A cross-cloud area optimization scheduling method for Flink stream computing applications is proposed, and resource allocation and cost management are optimized through operator topology reconstruction, cross-domain split execution and elastic scaling. Specific steps include: operator topology analysis and performance sampling in the offline stage, load monitoring and performance model optimization in the online stage.

Benefits of technology

Through operator topology reconstruction, the CPU resource utilization rate is reduced to 55% to 65%; through cross-domain deployment, cloud resource cost is saved by 40%; the performance model can be adjusted to the optimal resource configuration in one step, achieving elastic scaling of Flink stream computing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216104A_ABST
    Figure CN120216104A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-cloud area optimization scheduling method and system oriented to a stream computing application, and the method comprises the steps: firstly combining a code static analysis technology with a construction rule of the Flink stream computing application in an offline stage, carrying out the analysis to obtain an operator topological structure of a current application and an operation statement of each operator, constructing a relation table between the operators and data fields, and carrying out the analysis of the relation table; after the real data dependency relationship between operators is identified, solving the equivalent operator topology by a method of adjusting the position of an operator without data dependency; then solving the load proportion and the communication bandwidth of each operator grouping division scheme from the operator topology by adopting a performance sampling method to obtain a division scheme adaptive to cross-domain execution and high in processing efficiency; carrying out gradient sampling on the performance of the operator groups, and training a performance model capable of predicting resource use conditions under different data processing rates and resource configurations; and monitoring the change condition of the load in the online stage, solving the optimal resource configuration of the current load through the trained performance model, selecting the cloud resource combination scheme with the optimal cost to carry out cross-cloud splitting deployment, and carrying out online optimization on the performance model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of distributed computing, specifically a cross-cloud region optimization scheduling method and system for stream computing applications. Background Art

[0002] Apache Flink is a currently popular high-performance open-source stream processing framework in the industry. Cloud computing can simplify its deployment and provide elastic resources. To utilize the resource and price advantages of different cloud regions and improve the scalability of Flink stream computing applications, the cross-cloud region collaborative Flink stream computing task scheduling has become an important breakthrough. However, the Flink stream computing framework adopts an operator-intermediate resource sharing mechanism and does not support scheduling operators to specified range nodes. Summary of the Invention

[0003] Aiming at the above deficiencies of the existing technology, the present invention proposes a cross-cloud region optimization scheduling method and system for stream computing applications, and proposes a complete cross-domain collaborative load transformation, split execution, and elastic scaling solution for the optimization reconstruction of the operator topology of Flink stream computing applications themselves, cloud cost optimization of cross-domain splitting and execution, and rapid resource response during load fluctuations. Experiments on various Flink stream computing applications show that for Flink stream computing applications after operator topology optimization reconstruction, the required CPU resources can be reduced to 55% - 65% of the original at most; after cross-domain deployment of Flink stream computing applications, the cost of cloud resources can be saved by 40%; at the same time, based on the performance model constructed by the present invention, the optimal resource configuration can be adjusted to in one step.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to a cross-cloud region optimization scheduling method for Flink stream computing applications. In the offline stage, first, in combination with code static parsing technology and the construction rules of Flink stream computing applications, the operator topology structure of the current application and the operation statements of each operator are parsed to construct a relationship table between operators and data fields. After identifying the true data dependency relationship between operators, the equivalent operator topology is solved by adjusting the positions of operators without data dependency; then, the method of performance sampling is used to solve the load ratio and communication bandwidth of each operator grouping division scheme from the operator topology to obtain a division scheme suitable for cross-domain execution and with high processing efficiency; then, gradient sampling is performed on the performance performance of operator groups to train a performance model that can predict the resource usage under different data processing rates and resource configurations; in the online stage, the change of the load is monitored and the optimal resource configuration of the current load is solved through the trained performance model, and the cloud resource combination scheme with the optimal cost is selected for cross-cloud splitting deployment and online optimization of the performance model.

[0006] The performance sampling of the operator topology mentioned above refers to: running the Flink stream computing applications corresponding to different operator topologies in a way that disables operator merging and operator exclusive resource slots, and collecting performance reports containing indicators such as the computing efficiency of each operator, the load ratio, and the communication volume between operators.

[0007] The performance models mentioned above include: a multi-layer perceptron regression model and a gradient boosting regression model.

[0008] The training mentioned above refers to: performing gradient sampling with a specific strategy, covering low, medium, and high load intervals, providing samples for different parts of the performance model, and then training the multi-layer perceptron regression model and the gradient boosting regression model according to different prediction targets.

[0009] The cross-cloud split deployment mentioned above refers to: calculating the optimal resource allocation plan according to the current data arrival rate and combining with the performance model; and in the multi-cloud resource manager, calculating the costs of different cloud resource combinations according to the fluctuations of the price board, and selecting the optimal one for cross-cloud deployment.

[0010] The online optimization mentioned above refers to: monitoring the resource usage of the Flink stream computing application, and updating the performance model in real time according to the data executed online. When the data arrival rate fluctuates, generating the optimal resource allocation for the current load again through the trained performance model to achieve the elastic scaling of the Flink stream computing application.

[0011] The present invention relates to a system for implementing the above method, including: an Application Manager and a Multi-Cloud Resource Manager, where: the Application Manager performs operator topology optimization and reconstruction according to the load structure characteristic information of the submitted Flink stream computing application, and after obtaining the equivalent operator topology and operator grouping division results adapted for cross-domain execution, constructs a performance model for each operator group; the Multi-Cloud Resource Manager solves the optimal cost cloud resource combination plan according to the target resource configuration and the real-time price information of cloud resources, and obtains the optimal cross-domain deployment plan for deployment.

[0012] The application manager described above includes: an operator topology reconstruction unit (TopologyReconstructor), a load-aware splitting unit (Load-Aware Splitter), and a resource predictor unit (Resource Predictor), where: The operator topology reconstruction unit performs operator topology reconstruction based on the parsed data dependency relationship information between operators to obtain an equivalent operator topology result; The load-aware splitting unit calculates the load ratio of different operator grouping division schemes and the inter-group communication volume according to the operator calculation load and the inter-operator communication volume during the operation of different operator topologies, and obtains a division scheme suitable for cross-domain splitting execution; The resource predictor unit predicts the resource usage under different parallelism configurations according to the throughput requirement of the current application and in combination with the performance model, and obtains the optimal resource configuration.

[0013] The multi-cloud resource manager described above includes: a price information collection unit (Price Tracker), a monitoring platform (Monitor), and a cross-region executor (Cross-Region Executor), where: The price information collection unit collects the real-time computing and bandwidth resource prices of bid instances and reserved instances of candidate models on different cloud platforms and cloud regions; The monitoring platform monitors the Flink operation efficiency and resource usage; The cross-region executor deploys the operator loads of each part to the corresponding cloud region instances in combination with the operator grouping scheme and the resource configuration scheme.

[0014] The operator topology reconstruction mentioned above refers to: by adjusting the positions of operators without data dependencies in the operator topology, the computing and communication efficiency is improved, specifically including:

[0015] Step ①: Obtain the computational flow graph (StreamGraph) of the current Flink stream computing application through the interface provided by the Flink framework. Operators can be obtained through the nodes (StreamNode) in the computational graph, and the data flow relationship between operators can be obtained through the edges (StreamEdge) of the nodes, thereby constructing the current operator topology.

[0016] The construction code of Flink has structural characteristics. For example: adjacent operators executed successively are connected by "."; The implementation of the internal logic of the operator needs to overload and implement the class methods of a specific structure. Based on these construction rules of the Flink stream computing application, the present invention uses the Java static parsing library JavaParser (https: / / github.com / javaparser / javaparser) to implement methods for extracting information such as operator operation statements and associated fields, and proposes a method for constructing a relationship table between operators and data fields. The relationship table contains the association information of operators, operator types, operation fields, new fields, and deleted fields.

[0017] Step ③: Based on the relationship table in Step ②, according to different types of operator operations, an operator topology with a real data dependency can be constructed for each part of the link (if operator A has a data dependency on operator B, then A must be executed after B). For example: The operator that operates on field A has a data dependency on the operator that generates field A; the operator that deletes field A has a data dependency on all operators that operate on field A.

[0018] Step ④: Based on the topology representing the real data dependency, after performing processing operations such as removing redundant dependencies between nodes, an operator structure with an adjustable execution order can be obtained. Their topological characteristics are that they have the same predecessor and successor nodes, but there is no data dependency between them.

[0019] Step ⑤: Adjust the order of the operators to construct equivalent loads for different operator topologies.

[0020] In addition, the operator topology reconstruction unit also explores the optimization effect of large language models on Flink stream computing applications. This unit focuses on using large language model generation optimization methods of prompt engineering and context engineering to improve the analysis and construction capabilities of large language models for Flink stream computing application code without modifying the weights and parameters of the models, and automatically generate reference operator orchestration optimized construction code. To improve the accuracy of the generated code and reduce the workload of secondary processing, the present invention has conducted multiple rounds of tests and optimizations, and summarized the prompts and contexts shown in Table 1, which have good and stable performance in the task of operator rearrangement:

[0021] Table 1

[0022] Solving the computational load ratio and inter-group communication volume of different operator grouping schemes during cross-domain splitting specifically includes:

[0023] Step ①: Execute the stream computing loads corresponding to different operator topologies in the disabled operator link mode and the mode where each operator has its own exclusive resource slot, and generate a performance sampling report, including the performance efficiency of each operator topology, the computational load of each operator, and the communication volume between operators.

[0024] Step ②: Using the graph traversal algorithm, starting from the data source of the Flink stream computing application, traverse different graph partitioning schemes (the operator graph is divided into upstream and downstream operator groups, and only upstream to downstream data transmission is allowed), and combined with the goal of cross-domain collaboration of the present invention (maximizing the load ratio of downstream operator groups while reducing communication between operator groups), prune the partitioning scheme (set thresholds for load ratio and data communication volume) to obtain an operator grouping scheme that is suitable for cross-domain execution, providing candidate schemes for subsequent load partitioning.

[0025] The resource configuration management unit in the scheduling system of the present invention limits each operator group to share the same degree of parallelism. Since Flink uses a resource slot sharing mechanism, the same degree of parallelism can ensure that there is an instance of each operator on each resource slot to prevent the "laggard" problem caused by load imbalance. At the same time, after analyzing the operation of various Flink stream computing applications, the results show that according to the relationship between data processing rate and resource usage, the load of Flink stream computing applications can be divided into the following parts:

[0026] a) Constant part: fixed basic overhead for each resource slot, such as the basic overhead of running the Flink stream computing framework and tasks;

[0027] b) Amortization: The overall overhead that is fixed or has an upper limit in Flink stream computing applications will be amortized to each resource slot. For example, in an operator that sorts the real-time sales of goods, since the total number of goods is fixed, the load of the sorting operator has an upper limit and will be amortized to each resource slot running the operator.

[0028] c) Quasi-linear part: According to existing work and experimental results, the computing resources required by Flink stream computing applications (CPU resource utilization, number of resource slots used) are positively correlated with the throughput.

[0029] d) Superlinear part: Within the resource slot, due to the overhead caused by resource competition, thread switching, garbage collection, etc., as the throughput increases, the additional overhead grows superlinearly, and the throughput that can be improved per unit CPU resource gradually decreases.

[0030] The resource rapid response specifically includes:

[0031] Step ①: Sample from historical execution records with a specific strategy, covering low, medium, and high load intervals, and provide samples for different parts of the performance model. First, determine the basic overhead of resource slot operation by using resource usage at different parallelisms at a lower data arrival rate; second, collect CPU usage at different gradients in medium and high load intervals at different parallelisms in a grid manner to grasp the rules of the changing parts.

[0032] Step ②: Group each operator to build a performance model. According to the target throughput and the number of resource slots, the average CPU usage rate and the total CPU usage rate of each resource slot can be predicted. That is, the goal of the performance model is to fit cpu total = f total (throughput, p slot ) and cpu average = f average (throughput, p slot ). The performance model is a multi-layer perceptron regression and gradient boosting regression model. When predicting online resources, the average value of the prediction results of the two models is used.

[0033] Step ③: Based on the above performance model, the resource configurations with the optimal total number of resource slots and the optimal total CPU utilization rate can be solved respectively. With the constraint that the average CPU usage rate of the resource slot is less than the resource upper limit of a single resource slot, and on the premise of meeting the current target throughput, when the goal is to optimize the total number of resource slots, the minimum number of resource slots needs to be found; when the goal is to optimize the total CPU utilization rate, the number of resource slots with the minimum total CPU usage rate needs to be found. Technical Effects

[0034] Through the operator topology reconstruction, operator group resource prediction, and operator group cross-domain execution technologies, the present invention deeply studies the internal mechanism of the Flink framework and the load characteristics of applications, and provides a complete solution for Flink stream computing applications to cross cloud platforms / cloud regions.

[0035] Compared with the prior art, the present invention can effectively discover the redundant parts of transmission and calculation in Flink, and through the way of operator graph reconstruction, improve the calculation and communication efficiency of Flink stream computing applications. For the reconstructed operator graph, the overall CPU usage rate can be reduced to 60% of the original, while the data transmission volume between operators is compressed, and the communication volume between operator groups during cross-domain splitting is greatly reduced; the present invention can accurately predict the resource usage under different throughputs and configurations, and can quickly find the resource configuration with the optimal number of resource slots or the optimal total CPU usage rate. Through experimental evaluation, based on the prediction results of the performance model, the resource configuration plan with the optimal number of resource slots can be found after at most one step adjustment; with the help of the elastic scaling strategy implemented by the performance model proposed by the present invention and the selection strategy of the cross-domain division scheme, the cross-domain splitting execution of Flink stream computing applications is realized. Through experimental evaluation, cross-domain execution can realize the cloud migration of Flink stream computing applications at the lowest cost, and the cloud migration cost can be saved by nearly 40% compared with other solutions. Description of the Drawings

[0036] Figure 1 It is a flow chart of the cross-cloud region optimization scheduling method of the present invention;

[0037] Figure 2 Schematic diagram of the steps for migrating the Flink streaming computing application to the cloud in the embodiment;

[0038] Figure 3 Schematic diagram of the operator topology structure in the embodiment;

[0039] Figure 4 Schematic diagram of reconstructing the operator topology according to data dependencies in the embodiment;

[0040] Figure 5 Schematic diagram of removing redundant edges from the operator topology in the embodiment;

[0041] Figure 6 Schematic diagram of the equivalent operator topology solution 1 in the embodiment;

[0042] Figure 7 Schematic diagram of the equivalent operator topology solution 2 in the embodiment;

[0043] Figure 8 Schematic diagram of the equivalent operator topology solution 3 in the embodiment;

[0044] Figure 9 Schematic diagram of the operator grouping and partitioning solution 1 in the embodiment;

[0045] Figure 10 Schematic diagram of the operator grouping and partitioning solution 2 in the embodiment;

[0046] Figure 11 Schematic diagram of the operator grouping and partitioning solution 3 in the embodiment;

[0047] Figure 12 Schematic diagram of the operator grouping and partitioning solution 4 in the embodiment;

[0048] Figure 13 Schematic diagram of the operator grouping and partitioning solution 5 in the embodiment;

[0049] Figure 14 Schematic diagram of the predicted results of the overall CPU usage on the performance model in the embodiment;

[0050] Figure 15 Schematic diagram of the predicted results of the average CPU usage on the performance model in the embodiment;

[0051] Figure 16 Schematic diagram of the region-aware task scheduling mechanism implemented by the present invention. Detailed implementation manners

[0052] The specific experimental environment of this embodiment is as follows: The operating system version is Ubuntu 20.04 LTS, configured with an Intel(R) Xeon(R) Gold 6230R CPU @ 2.10 GHz and 64 GB of DDR4 RAM. Among them, Flink is started through a docker container for resource management, and the Flink image number is flink: 1.13.0. The embodiment is a Flink stream computing application for web link click stream statistics, and its implemented function is: counting and sorting the number of clicks on specific web links in the past period of time. The operator topology structure is as Figure 3 shown.

[0053] As Figure 2 shown, the cross-cloud region optimization scheduling method for the Flink stream computing application in the above scenario involved in this embodiment includes:

[0054] Step 1, solving the equivalent operator topology: Combining code static analysis technology and the construction rules of the Flink stream computing application, after parsing the current operator topology and the operation statements of each operator, constructing a relationship table between operators and data fields, and parsing out the true data dependency relationship. Combining the structural characteristics of the Flink code and mature Java code static analysis technology, the fields operated inside each operator, the newly added fields, and the deleted fields can be parsed. The relationship between the operators and data fields in the current embodiment is shown in Table 2:

[0055] Table 2

[0056] From the true data dependency relationship between the above operators, an operator topology representing the true data dependency relationship can be reconstructed. The steps are as Figure 4 shown. After link merging, the three equivalent operator topologies obtained are as Figure 5 shown. In the figure, the marked candidate solutions (the <1>, <2>, and <3> on the edges indicate the candidate positions where Filter-1 can be inserted).

[0057] Perform performance tests on the loads of the above three different operator topologies, and allocate 1 CPU of computing resources to each resource slot. The results are shown in Table 3. Under two throughput loads, reconstructing the operator topology can save 35% - 45% of the original CPU usage rate.

[0058] Table 3

[0059] In this embodiment, by leveraging the semantic understanding and content generation capabilities of the large language model, the reconstruction of the optimal operator topology has also been successfully achieved. The prompt words and generated results for interacting with the large language model are shown in Table 4 (the codes in the table are only examples).

[0060] Table 4

[0061] Step 2: Solve the cross-domain load partitioning scheme for the Flink streaming computing application. By performing performance sampling on different splitting schemes, the performance index situation shown in Table 5 can be obtained.

[0062] Table 5

[0063] Step 3: Construct and train the performance model: In this embodiment, splitting scheme 3 is selected for cross-domain execution, and performance models are constructed for the upstream and downstream loads respectively. Taking the downstream load (including three operator operations: extracting timestamps, window grouping counting, and sorting) as an example, two regression models, namely the multi-layer perceptron and gradient boosting, are trained. The prediction results on the test data (the load and resource configuration do not overlap with the training set) are shown in Table 6:

[0064] Table 6

[0065] Multiple experiments have shown that the prediction results after averaging the above two models are closer to the real data. For any range of loads, the optimal resource configuration predicted by the performance model will not exceed the error of one resource slot, and it only takes at most one resource adjustment to reach the real optimal resource configuration. The overall average CPU usage prediction and total CPU usage prediction results are as Figure 14 and Figure 15 shown, and the prediction results of some nodes with load and resource configuration are shown in Table 7.

[0066] Table 7

[0067] Step 4: Select the optimal cloud adoption plan: In this embodiment, the Alibaba Cloud economy e instance is used as the available cloud model (the minimum available resource configuration is 2 CPUs / 2 GB of memory). Two Flink resource slots are deployed on each instance, and each resource slot can use 1 CPU resource. The charging standard of this model during the experiment is shown in Table 8.

[0068] Table 8

[0069] Step 5: Cross-cloud split deployment and online optimization of the performance model: Implement a task scheduling mechanism that is aware of the regions (nodes) of the Flink streaming computing application using a cross-domain executor. As Figure 16 shown, specifically, through the cross-domain executor combined with the operator grouping scheme and the resource configuration scheme, deploy the operator loads of each part to the corresponding cloud region instances, including:

[0070] 1) When creating resource slots, the Flink execution node adds the identification information of the region where it is located and also registers the region information with the Flink resource manager;

[0071] 2) When submitting a Flink streaming computing application, the Flink client divides the operators within different operator groups into the resource slot groups of the corresponding regions and extracts the resource slot requests containing region information;

[0072] 3) The Flink resource manager realizes the allocation of target region resources to the corresponding operator groups by matching the resource slot requests and the region information of the registered resource slots.

[0073] According to the actual measurement data, when the data throughput reaches 80k, for single execution, a total of 11 resource slots of 1CPU / 1GB are required; for cross-domain split execution, 2 resource slots of 1CPU / 1GB memory are required upstream and 10 resource slots of 1CPU / 1GB are required downstream. The data volume generated by the data source is 8.14MB / s, and the data transmission volume between the upstream and downstream for selection of partitioning scheme 3 is 1.97MB / s.

[0074] According to the charging standard, the hourly cost of different cloud adoption scenarios is shown in Table 9 (calculation method: reserved instance number * reserved instance price + bandwidth * reserved bandwidth price + spot instance number * spot instance price).

[0075] Table 9

[0076] The cloud adoption cost can save nearly 40% compared with local execution / unloading to remote execution.

[0077] In summary, this system provides a complete solution for cross-domain split execution of Flink streaming computing applications. First, both static analysis and large language models can discover equivalent and more efficient operator topologies; second, the constructed performance model can accurately predict the required resource situation and has been verified on multiple load benchmarks; finally, after cross-domain split execution of the load, the cloud cost is significantly reduced.

[0078] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments. All implementation solutions within its scope are subject to the present invention.

Claims

1. A cross-cloud region optimization scheduling method for Flink stream computing applications, characterized in that In the offline phase, first, combined with code static analysis technology and the construction rules of Flink stream computing applications, the operator topology structure of the current application and the operation statements of each operator are parsed, a relationship table between operators and data fields is constructed, and after identifying the true data dependencies between operators, the equivalent operator topology is solved by adjusting the positions of operators without data dependencies; then, a performance sampling method is used to solve the load ratio and communication bandwidth of each operator grouping division scheme from the operator topology, obtaining a division scheme that is suitable for cross-domain execution and has high processing efficiency; then, gradient sampling is performed on the performance of operator groups to train a performance model that can predict resource usage under different data processing rates and resource configurations; in the online phase, the change of the load is monitored and the optimal resource configuration for the current load is solved through the trained performance model, and the cloud resource combination scheme with the lowest cost is selected for cross-cloud splitting and deployment and online optimization of the performance model; The performance sampling of the operator topology mentioned above refers to: running Flink stream computing applications corresponding to different operator topologies in a way that disables operator merging and operator exclusive resource slots, and collecting performance reports including the computing efficiency, load ratio, and communication volume between operators; The performance model mentioned above includes: a multi-layer perceptron regression model and a gradient boosting regression model; The training mentioned above refers to: performing gradient sampling with a specific strategy, covering low, medium, and high load intervals, providing samples for different parts of the performance model, and then training the multi-layer perceptron regression model and the gradient boosting regression model according to different prediction targets.

2. The cross-cloud region optimization scheduling method for Flink stream computing applications according to claim 1, characterized in that The cross-cloud splitting and deployment mentioned above refers to: calculating the optimal resource configuration scheme according to the current data arrival rate, combined with the performance model; and in the multi-cloud resource manager, calculating the costs of different cloud resource combinations according to the fluctuations of the price board, and selecting the optimal scheme for cross-cloud deployment.

3. The cross-cloud region optimization scheduling method for Flink stream computing applications according to claim 1, characterized in that, The online optimization mentioned above refers to: monitoring the resource usage of the Flink stream computing application, and updating the performance model in real time according to the data executed online. When the data arrival rate fluctuates, the optimal resource configuration for the current load is generated again through the trained performance model, realizing the elastic scaling of the Flink stream computing application.

4. A cross-cloud region optimization scheduling system for Flink stream computing applications implementing the method described in any one of claims 1-3, characterized in that, Including: An application manager and a multi-cloud resource manager, where: the application manager optimizes and reconstructs the operator topology according to the load structure characteristic information of the submitted Flink stream computing application, and after obtaining the equivalent operator topology and operator grouping division results suitable for cross-domain execution, constructs a performance model for each operator group; the multi-cloud resource manager solves the cloud resource combination scheme with the lowest cost according to the target resource configuration and the real-time price information of cloud resources, and obtains the optimal cross-domain deployment scheme for deployment.

5. The cross-cloud region optimization scheduling system for Flink stream computing applications according to claim 4, characterized in that, The application manager described above includes: an operator topology reconstruction unit, a load-aware splitting unit, and a resource predictor unit, where: The operator topology reconstruction unit performs operator topology reconstruction based on the parsed data dependency relationship information between operators to obtain an equivalent operator topology result; The load-aware splitting unit calculates the load ratio of different operator grouping division schemes and the communication volume between groups according to the operator calculation load and the communication volume between operators during the operation of different operator topologies, and obtains a division scheme suitable for cross-domain splitting execution; The resource predictor unit predicts the resource usage under different parallelism configurations based on the throughput requirement of the current application and in combination with the performance model, and obtains the optimal resource configuration.

6. The cross-cloud region optimization scheduling system for Flink stream computing applications according to claim 4, characterized in that, The multi-cloud resource manager described above includes: a price information collection unit, a monitoring platform, and a cross-domain executor, where: The price information collection unit collects the real-time computing and bandwidth resource prices of bid instances and reserved instances of candidate models on different cloud platforms and cloud regions; The monitoring platform monitors the Flink operation efficiency and resource usage; The cross-domain executor deploys the operator loads of each part to the corresponding cloud region instances in combination with the operator grouping scheme and the resource configuration scheme.

7. The cross-cloud region optimization scheduling system for Flink stream computing applications according to claim 5, characterized in that, The operator topology reconstruction mentioned above refers to: by adjusting the positions of operators without data dependencies in the operator topology to improve the computing and communication efficiency, which specifically includes: Step ①: Obtain the computational flow graph of the current Flink stream computing application through the interface provided by the Flink framework, obtain operators through the nodes in the computational graph, and obtain the data flow relationship between operators through the edges of the nodes, so as to construct the current operator topology. Step ②: The construction code of Flink has the characteristics of being structured. For example, adjacent operators executed successively are connected by "."; The implementation of the internal logic of the operator needs to override and implement the class methods of a specific structure; Based on the construction rules of the Flink stream computing application, a Java static parsing library is used to extract operator operation statements and associated fields, and a construction method for the relationship table between operators and data fields is proposed. Step ③: Based on the relationship table in Step ②, according to different types of operator operations, an operator topology with a real data dependency relationship can be constructed for each part of the link. Step ④: Based on the topology representing the real data dependency relationship, after performing processing operations such as removing redundant dependencies between nodes, an operator structure with an adjustable execution order is obtained, that is, these operators have the same predecessor and successor nodes, but there is no data dependency between them. Step ⑤: Adjust the order of the operators to construct the equivalent load of different operator topologies.

8. The cross-cloud region optimization scheduling system for Flink stream computing applications according to claim 5, characterized in that The calculation of the load ratio of different operator grouping schemes and the communication volume between groups when solving cross-domain splitting specifically includes: Step ①: Execute the stream computing load corresponding to different operator topologies in the disabled operator chaining mode and the mode where each operator has its own exclusive resource slot, and generate a performance sampling report, including the performance efficiency of each operator topology, the computing load of each operator, and the communication volume between operators. Step ②: Using the graph traversal algorithm, starting from the data source of the Flink streaming computing application, traverse different graph partitioning schemes, and combine the cross-domain collaboration objectives of the present invention to prune the partitioning schemes, obtaining an operator grouping scheme adapted for cross-domain execution, providing a candidate scheme for subsequent load partitioning; The loads include: a) Constant part: The fixed basic overhead on each resource slot; b) Amortized part: The total overhead that is fixed or has an upper limit value in the Flink streaming computing application, which will be amortized to each resource slot; c) Quasi-linear part: According to existing work and experimental results, the computing resources required by the Flink streaming computing application are positively correlated with the throughput; d) Super-linear part: Inside the resource slot, due to resource competition, thread switching, and garbage collection, the overhead increases super-linearly as the throughput increases, and the throughput that can be improved by a unit of CPU resource gradually decreases.

9. The cross-cloud region optimization scheduling system for Flink stream computing applications according to claim 5, characterized in that, The rapid resource response specifically includes: Step ①: Sample from the historical execution records with a specific strategy, covering low, medium, and high load intervals, providing samples for different parts of the performance model. First, determine the basic overhead of the resource slot operation by the resource usage at different parallelism levels under a relatively low data arrival rate; secondly, collect the CPU usage at different gradients and different parallelism levels in the medium and high load intervals in a grid manner to master the variation pattern; Step ②: Group each operator to build a performance model. According to the target throughput and the number of resource slots, the average CPU usage rate and the total CPU usage rate of each resource slot can be predicted. That is, the goal of the performance model is to fit cpu total = f total (throughput, p slot ) and cpu average = f average (throughput, p slot ). The performance model is a multi-layer perceptron regression and gradient boosting regression model. When predicting online resources, the average of the prediction results of the two models is used; Step ③: Based on the above performance model, the resource configurations with the optimal total number of resource slots and the optimal total CPU utilization rate can be solved respectively. With the constraint that the average CPU usage rate of the resource slots is less than the resource upper limit of a single resource slot, when the goal is to optimize the total number of resource slots under the premise of meeting the current target throughput, the minimum number of resource slots needs to be found; when the goal is to optimize the total CPU utilization rate, the number of resource slots with the minimum total CPU usage rate needs to be found.