Flink cluster deployment management method and device based on k8s and medium

Through the K8s-based Flink cluster deployment management method, automated deployment and configuration management are solved, and the problem of complex traditional deployment methods and lack of dynamic adjustment is achieved, and efficient and automated Flink cluster management is achieved.

CN120215960APending Publication Date: 2025-06-27INSPUR SMART TECH (NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264177.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional way of deploying Flink clusters depends on a large number of manual steps, resulting in complex deployment parameters setting and lack of dynamic adjustment, which cannot meet the data processing needs in different business scenarios.

Method used

The Flink cluster deployment management method based on k8s is adopted, and the Flink cluster configuration template is generated by obtaining the cluster configuration requirement information input by the user and converting it into resource object definition information in the k8s container to realize automated deployment and dynamic adjustment.

Benefits of technology

The automated Flink cluster deployment and configuration management is realized, which reduces manual operations and ensures the standardization and consistency of configurations. The dynamic adjustment mechanism can automatically adjust the resource scale according to business load to avoid performance bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215960A_ABST
    Figure CN120215960A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a k8s-based Flink cluster deployment management method and device and a medium, and relates to the technical field of computers.The method comprises the steps that cluster configuration demand information input by a user is acquired, and an Flink cluster configuration template is generated based on the cluster configuration demand information; converting the Flink cluster configuration template into resource object definition information in a k8s container so as to create a plurality of component resources corresponding to the Flink cluster by using the resource object definition information; acquiring deployment environment information corresponding to the Flink cluster, performing pre-deployment verification on the plurality of component resources based on the deployment environment information, and after the verification is passed, deploying the plurality of component resources corresponding to the Flink cluster into a target operation environment; and monitoring a real-time performance index of the Flink cluster in the target operation environment, so as to adjust the Flink cluster according to a preset dynamic adjustment strategy and the real-time performance index, the adjustment including any one or more of adaptive scaling operation and adaptive configuration parameter adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a method, device, and medium for deploying and managing a Flink cluster based on k8s. Background Art

[0002] With the rapid development of big data technology, the powerful distributed computing ability of the Flink cluster enables it to efficiently process massive amounts of data. Whether it is offline batch processing tasks, such as the construction and data analysis of enterprise-level data warehouses, deeply mining large-scale historical data to obtain valuable business insights; or real-time stream processing scenarios, such as real-time risk monitoring in the financial field, instantly capturing abnormal transaction signals, and real-time recommendation systems in the Internet industry, accurately pushing personalized content based on users' real-time behaviors, etc., all reflect the key role of the Flink cluster in improving business efficiency and competitiveness.

[0003] However, in practical applications, the existing Flink cluster deployment methods have many drawbacks. Traditional deployment processes often rely on a large number of manual steps. Starting from the construction of the basic environment, operations and maintenance personnel need to install various dependent software required for Flink to run on each target server one by one, including precisely configuring the appropriate version of the Java runtime environment, as well as other system-level library files and tools. Subsequently, it enters a complex and cumbersome configuration stage. For the core components Job Manager and Task Manager in the Flink cluster, it is necessary to manually edit the configuration file to set numerous key parameters. Improper settings may affect the performance of the entire cluster.

[0004] In addition, the existing manual method of deploying the Flink cluster cannot achieve efficient dynamic allocation of resources. During peak business periods, task processing may be delayed or even fail due to insufficient resource allocation, affecting the real-time performance and accuracy of the business; while during off-peak business periods, resources cannot be automatically reduced, resulting in idle waste of resources and increasing unnecessary cost expenditures. Moreover, the configuration inconsistency problem brought about by manual deployment cannot be ignored. There are differences in the operating habits and technical levels of different operations and maintenance personnel, which may lead to different configurations of the Flink cluster on each node. Such differences are extremely likely to cause stability hazards and performance bottlenecks in a large-scale cluster environment, making fault troubleshooting and performance optimization work extremely complex and restricting the application and development of the Flink cluster in complex business scenarios.

[0005] Therefore, in the traditional method of deploying the Flink cluster, the setting of deployment parameters depends on a large number of manual steps, and there is no dynamic adjustment process after configuring the parameters, which cannot meet the data processing requirements in different business scenarios. Summary of the Invention

[0006] One or more embodiments of this specification provide a method, device, and medium for deploying and managing a Flink cluster based on k8s to solve the following technical problems: In the traditional method of deploying a Flink cluster, the setting of deployment parameters depends on a large number of manual steps, and there is no dynamic adjustment process after configuring the parameters, which cannot meet the data processing requirements in different business scenarios.

[0007] One or more embodiments of this specification adopt the following technical solutions:

[0008] One or more embodiments of this specification provide a method for deploying and managing a Flink cluster based on k8s. The method includes: obtaining the cluster configuration requirement information input by the user, and generating a Flink cluster configuration template based on the cluster configuration requirement information; converting the Flink cluster configuration template into resource object definition information in a k8s container, and using the resource object definition information to create multiple component resources corresponding to the Flink cluster; obtaining the deployment environment information corresponding to the Flink cluster, performing pre-deployment verification on the multiple component resources based on the deployment environment information, and after the verification passes, deploying the multiple component resources corresponding to the Flink cluster to the target running environment; monitoring the real-time performance metrics of the Flink cluster in the target running environment, and adjusting the Flink cluster according to the preset dynamic adjustment policy and the real-time performance metrics, where the adjustment includes any one or more of adaptive scaling operations and adaptive configuration parameter adjustments.

[0009] One or more embodiments of this specification provide a device for deploying and managing a Flink cluster based on k8s, including:

[0010] At least one processor; and,

[0011] A memory communicatively connected to the at least one processor; wherein,

[0012] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0013] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.

[0014] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Through the above technical solution, by obtaining the cluster configuration requirement information input by the user to generate a Flink cluster configuration template, the cumbersome process of manually writing complex configuration files from scratch in the traditional method is avoided. In addition, the method of generating a template according to the cluster configuration requirements input by the user can ensure the standardization and consistency of the configuration, and reduce configuration errors caused by human negligence or understanding differences; converting the Flink cluster configuration template into the resource object definition information in the Kubernetes (k8s) container realizes the deep integration of the Flink cluster and the powerful container orchestration ability of Kubernetes. Using the standardized resource management and deployment mechanism of Kubernetes, it is possible to create multiple component resources corresponding to the Flink cluster more conveniently and automatically, avoiding problems such as format errors and unreasonable resource mapping that may occur when manually configuring Kubernetes resources, and further improving the accuracy and efficiency of deployment; obtaining the deployment environment information corresponding to the Flink cluster and performing pre-deployment verification on multiple component resources based on this to discover potential configuration incompatibilities, resource conflicts, etc. in advance; by monitoring the real-time performance metrics of the Flink cluster in the target running environment and performing adaptive scaling operations according to the pre-set dynamic adjustment strategy, the Flink cluster can dynamically adjust the resource scale according to the actual business load; the dynamic adjustment mechanism can automatically adapt to changes in business logic, evolution of data characteristics, etc., ensuring that the Flink cluster continuously maintains good performance during long-term operation and avoiding performance bottleneck problems caused by fixed configurations being unable to adapt to business development; covering automated operations in multiple links from configuration generation, deployment to runtime dynamic adjustment, greatly reducing the manual operation workload of operation and maintenance personnel in the process of Flink cluster management. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0016] Figure 1 It is a schematic flowchart of a method for deploying and managing a Flink cluster based on k8s provided by an embodiment of this specification;

[0017] Figure 2 It is a schematic structural diagram of a device for deploying and managing a Flink cluster based on k8s provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0019] The embodiments of this specification provide a method for deploying and managing a Flink cluster based on k8s. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 It is a schematic flowchart of a method for deploying and managing a Flink cluster based on k8s provided by the embodiments of this specification, as Figure 1 shown, mainly including the following steps:

[0020] Step S101, obtain the cluster configuration requirement information input by the user, and generate a Flink cluster configuration template based on the cluster configuration requirement information.

[0021] In an embodiment of this specification, an interactive command-line interface or graphical user interface is designed in advance to enable the user to input cluster configuration requirement information, including the type of data processing (batch processing, stream processing, or hybrid processing), expected throughput, data latency requirements, concurrency of jobs, etc., and resource limit information, including the number of available CPU cores, total memory, disk space size, etc. Verify the data input by the user to ensure its rationality and integrity. For example, check whether the input throughput value is within a reasonable range and whether the resource limits meet the actual hardware conditions. At the same time, preprocess the data and convert it into an internally unified data format for subsequent analysis and calculation.

[0022] Based on the cluster configuration requirement information, a Flink cluster configuration template is generated, which specifically includes: collecting the historical cluster operation data of the Flink cluster in the actual operation environment, where the historical cluster operation data includes the corresponding historical configuration information and historical operation performance information under different business loads; analyzing the relationship between the cluster performance and configuration parameters of the Flink cluster through the historical cluster operation data to determine the performance parameter matching model corresponding to the Flink cluster, where the performance parameter matching model includes the relationship between each key performance indicator and at least one key configuration parameter; determining the corresponding Flink cluster resource quantity according to the cluster configuration requirement information and the performance parameter matching model, and generating a Flink cluster configuration template through the Flink cluster resource quantity, where the Flink cluster resource quantity includes parallelism parameters, memory requirement parameters, and CPU core allocation parameters.

[0023] In an embodiment of this specification, first, it is necessary to determine the key configuration parameters that affect the performance of the Flink cluster, such as parallelism, memory allocation (including Job Manager memory, Task Manager memory, and memory settings of each operator, etc.), the number of CPU cores, network bandwidth, checkpoint interval, and timeout time. At the same time, determine the key indicators to measure the cluster performance, such as throughput (the amount of data processed per unit time), latency (the time elapsed from when the data enters the system to when it is processed), and resource utilization (CPU, memory utilization, etc.). Relevant data can be obtained through various channels. In addition to the historical cluster operation data of the Flink cluster in the actual operation environment, it can also include conducting experimental tests on the Flink cluster in different configuration scenarios and recording the configuration parameters and corresponding performance indicator data for each test.

[0024] For the obtained historical cluster operation data or test data, analyze the relationship between the cluster performance and configuration parameters of the Flink cluster to determine the performance parameter matching model corresponding to the Flink cluster. It should be noted that the performance parameter matching model here refers to the corresponding relationship between the key performance indicators and the key configuration parameters, that is, the performance parameter matching model includes the relationship between each key performance indicator and at least one key configuration parameter. When determining the performance parameter matching model, various methods such as empirical formulas, experimental data fitting, or machine learning algorithm training can be used.

[0025] When determining the performance parameter matching model through empirical formulas, based on the internal working principle of Flink and the basic knowledge of relevant computer systems, the theoretical relationship between configuration parameters and performance indicators is deduced. For example, starting from the basic logic of data processing, when other conditions are relatively stable, the throughput is proportional to the parallelism. A simple formula such as "throughput = parallelism × single-task processing rate" can be initially constructed, where the single-task processing rate may be affected by factors such as CPU performance and memory read / write speed. Then, the association between these factors and configuration parameters is further refined to gradually improve the formula. In addition, the official Flink documentation usually provides some descriptions and suggestions on the impact of configuration parameters on performance. Combining the existing experience summaries in the industry, some common empirical formulas are sorted out. For example, for the relationship between memory allocation and task execution efficiency, there may be an empirical description like "the task execution efficiency increases with the increase of memory before the memory allocation reaches a certain threshold, and the improvement effect is not obvious or may even decrease after exceeding the threshold". Based on this, a corresponding piecewise function is constructed to describe this relationship.

[0026] When determining the performance parameter model through the experimental data fitting method, planned experiments are carried out for different combinations of configuration parameters. For example, fixing other parameters and only changing the parallelism to observe the changes in throughput and latency; then fixing the parallelism and changing the memory allocation amount, and recording the corresponding performance data, etc. Ensure that the experiments cover a wide enough range of parameter values and the experimental environment is relatively stable to reduce other interfering factors. After collecting the data obtained from the experiments, mathematical analysis tools (such as the numpy and scipy libraries in Python) are used for data fitting. If it is found that the throughput and parallelism show an approximately linear relationship, the linear regression method can be used to fit the specific linear function expression; if the relationship is a complex one, methods such as polynomial fitting and non-linear fitting (such as exponential functions, logarithmic functions, etc.) need to be adopted to find the function form that can best describe the data law, so as to establish a mathematical model between configuration parameters and performance indicators.

[0027] When training and determining a performance parameter matching model through machine learning algorithms, the determined configuration parameters are used as input features, and the performance metrics are used as the target output. The data is preprocessed, such as normalizing the data to bring parameters of different magnitudes to the same scale range; performing operations such as encoding conversion on some categorical data (such as whether the data processing type is stream processing or batch processing) to meet the input requirements of machine learning algorithms. Regression algorithms can be selected to build the model. For example, linear regression is used for modeling simple linear relationships. If the relationship is complex, non-linear regression algorithms such as decision tree regression, random forest regression, and gradient boosting regression tree (GBRT) can be tried. The collected experimental data or historical data from the production environment is divided into a training set and a test set. The selected algorithm is used to train the model with the training set, and by continuously adjusting the hyperparameters of the model, the accuracy and generalization ability of the model are evaluated on the test set until a model with better performance is obtained. This model can describe the relationship between cluster performance and configuration parameters.

[0028] After determining the performance parameter matching model, to ensure the accuracy of the model, the performance parameter matching model can be verified and optimized. An independent data set that did not participate in the model training is used to verify the established model, and the differences between the performance metrics predicted by the model and the actual data are compared. Metrics such as mean squared error (MSE) and mean absolute error (MAE) can be calculated to measure the accuracy of the model. If the error is large, it is necessary to re-examine the modeling process and check issues such as data quality and the rationality of the modeling method selection. As the Flink version is updated, the hardware environment changes, or the business scenario expands, new data is regularly collected, and the model is retrained and optimized to better reflect the actual situation. For example, the new Flink version may optimize some internal algorithms, resulting in a change in the relationship between configuration parameters and performance. In this case, the model needs to be updated to adapt to the new features.

[0029] Based on the cluster configuration requirement information and the performance parameter matching model, the corresponding Flink cluster resource quantities are determined, specifically including: determining the data processing type, required throughput parameter, data latency requirement parameter, and job concurrency parameter in the cluster configuration requirement information, and determining the parallelism matching sub-model, memory requirement matching sub-model, and CPU core number matching sub-model in the performance parameter matching model; determining the parallelism parameter according to the required throughput parameter and the parallelism matching sub-model; determining the corresponding memory requirement parameters for the JobManager and Task Manager according to the data processing type and the job concurrency parameter, based on the memory requirement matching sub-model; and determining the CPU core allocation parameter based on the required throughput parameter and the data latency requirement parameter, using the CPU core number matching sub-model.

[0030] In one embodiment of this specification, according to the business requirements and performance model input by the user, calculate the amount of Flink cluster resources required to meet these requirements. Extract the key performance indicator requirements from the business requirements input by the user, such as requiring a throughput of 100 GB of data processed per hour and an average latency controlled within 10 seconds, etc. At the same time, determine the data processing type of the business (whether it is real-time stream processing, batch processing, or hybrid processing) and the desired task concurrency, etc. According to the throughput requirement in the business requirements, combine the functional relationship between throughput and parallelism in the performance model to calculate the appropriate parallelism value. For example, if the throughput and parallelism in the performance model show a linear relationship (assumed to be "throughput = parallelism × 10 GB / hour"), when the required throughput is 100 GB of data processed per hour, the calculated parallelism should be 10. Then, based on the data processing type and concurrency, determine the memory requirements and CPU core allocation for the Job Manager and Task Manager. Consider factors such as data processing type and concurrency to determine the memory allocation. For the stream processing scenario, if there is a high concurrency and a fast data inflow rate, a larger memory needs to be allocated to the Task Manager to cache data and execute tasks to prevent problems such as out-of-memory. The appropriate memory size can be determined according to the relationship between memory and task execution efficiency, latency, and other indicators in the performance model. For example, if the model shows that when the memory allocation for each Task Manager reaches 4 GB, the latency can meet the requirements under the current concurrency, the memory of the Task Manager can be set to about 4 GB; at the same time, according to the responsibilities of the Job Manager (such as managing task scheduling, coordinating resources, etc.), allocate the corresponding memory according to a certain ratio (usually can refer to empirical values or obtained through model analysis), assumed to be 1 GB.

[0031] Allocate the number of CPU cores in combination with the relationship between CPU resources and throughput, latency in the performance model. For example, if the performance model indicates that each parallel task can achieve better performance when evenly allocated 0.5 CPU cores, and the currently calculated parallelism is 10, then a total of about 5 CPU cores need to be allocated for the task execution of the Flink cluster. Then, according to experience or further analysis, determine the specific number of cores allocated to the Job Manager and Task Manager. For example, allocate 1 core to the Job Manager and 4 cores to the Task Manager (assuming that the main computing tasks of the cluster are executed on the Task Manager). At the same time, considering system overhead and redundancy, appropriately adjust the calculated amount of resources.

[0032] Generate a Flink cluster configuration template based on the Flink cluster resource volume, specifically including: obtaining auxiliary configuration parameters, where the auxiliary configuration parameters include any one or more of network configuration parameters, checkpoint and fault tolerance configuration parameter intervals, and resource allocation priority parameters; generating a Flink cluster configuration template in YAML format according to the Flink cluster resource volume and the auxiliary configuration parameters.

[0033] In one embodiment of this specification, obtain auxiliary configuration parameters, where the auxiliary configuration parameters include any one or more of network configuration parameters, checkpoint and fault tolerance configuration parameter intervals, and resource allocation priority parameters; the network configuration parameters include network buffer size and transmission timeout. For example, parameters such as taskmanager.network.memory.fraction and taskmanager.network.memory.min for network buffer size control the proportion and minimum value of the memory used by the Task Manager for network data transmission. Reasonable settings can optimize network communication performance and prevent performance degradation caused by network data congestion. Especially in high-concurrency data transmission scenarios, appropriately increasing the network buffer can improve the efficiency and stability of data transmission. Transmission timeout parameters such as akka.ask.timeout and akka.tcp.timeout are used to set the timeout for network communication between Flink components. If the network environment is unstable or the task has high real-time requirements, these timeout times need to be adjusted according to the actual situation to avoid task failures or increased delays caused by timeouts. The checkpoint and fault tolerance configuration parameters include Checkpoint interval, Checkpoint timeout, and the maximum number of concurrent checkpoints. For example

[0034] execution.checkpointing.interval determines how often a checkpoint is performed during the execution of a Flink task. Frequent checkpoints will increase a certain performance overhead but provide better data fault tolerance; while a larger checkpoint interval will reduce the overhead but may lead to more data loss and increased recovery time in case of failures, and it needs to be set according to the trade-off between data consistency and performance of the business. The execution.checkpointing.timeout parameter defines the maximum allowed time for the checkpoint operation. If the checkpoint process does not complete within the timeout, it may be regarded as a failure and trigger the corresponding fault tolerance mechanism. Reasonable setting of this parameter can ensure that the checkpoint is completed within an acceptable time range and avoid unnecessary checkpoint failures caused by too short timeout settings.

[0035] execution.checkpointing.max-concurrent-checkpoints is the maximum number of concurrent checkpoints, which controls the maximum number of checkpoint operations that can occur simultaneously. For clusters with limited resources or tasks sensitive to performance, it is necessary to reasonably limit this number to prevent excessive checkpoint operations from competing for resources and affecting the normal execution of tasks. In addition, Flink supports setting resource allocation priorities for different tasks or task groups. For example, by using the resource allocation priority parameters slotSharingGroup and coLocationGroup in combination with priority strategies, it can be ensured that critical tasks can obtain the required resources first when resources are scarce, guaranteeing the performance and stability of critical business operations.

[0036] Based on the Flink cluster resource quantity and this auxiliary configuration parameter, generate a Flink cluster configuration template in YAML format. Assign a unique version number to each generated Flink cluster configuration template, and record metadata information such as the creation time, creator, and modification history of the template. A database (such as MySQL, SQLite, etc.) or a version control system (such as Git) can be used to store this information for subsequent traceability and management of template changes. Through the above steps, it is possible to automatically generate a relatively optimized Flink cluster configuration template according to the business requirements and resource limitations input by the user, improving the deployment efficiency and performance of the Flink cluster, while reducing the user's configuration difficulty and error risk.

[0037] In step S102, convert the Flink cluster configuration template into resource object definition information within the k8s container, and use this resource object definition information to create multiple component resources corresponding to the Flink cluster.

[0038] In an embodiment of this specification, use a programming language (such as the yaml library in Python) to open and read the Flink cluster configuration template file, and parse its content into a data structure of the corresponding programming language (usually in the form of dictionaries, lists, etc.) for subsequent data extraction and conversion operations. Analyze the parsed configuration template data structure, and extract key parameters related to the creation of Kubernetes resource objects from it, such as the resource allocation parameters (such as the number of CPU cores, memory size) of the Job Manager and Task Manager, the number of replicas, the configuration related to service exposure (port number, service type, etc.), and the content of the configuration files of various Flink components (such as flink-conf.yaml, etc., which need to be converted into a ConfigMap later).

[0039] Build the Kubernetes Deployment object definition to create basic metadata information for the Job Manager's Deployment, including the name (such as job-manager-deployment), the namespace it belongs to (which can be specified according to actual needs, for example, the flink-cluster namespace), labels (used to identify the Flink cluster and role to which this Deployment belongs, such as app:flink, component:job-manager), etc. Specify the Flink JobManager image used by the Job Manager container, usually obtained from the official image repository, such as flink:latest or specified according to a specific version, for example, flink:1.16.0. Set the resource requests (resources.requests) and resource limits (resources.limits) of the container according to the CPU and memory parameters extracted from the Flink configuration template. Configure the ports exposed by the Job anager container (such as the port for REST API access, etc.), and mount the Flink configuration files (such as flink-conf.yaml extracted from the template) to the corresponding directories inside the container so that the correct configuration can be loaded when the container starts. Obtain the number of replicas of the Job Manager from the Flink configuration template (usually 1), and add it to the spec.replicas field of the Deployment. The construction process of the Task Manager Deployment is similar to that of the Job Manager Deployment, creating information such as the Deployment metadata and container configuration of the TaskManager. The difference is that the number of replicas, resource allocation, and configuration parameters of the Task Manager vary according to the workload requirements of the Flink cluster and the settings in the configuration template. For example, set the number of replicas of the Task Manager to be dynamically scalable according to the template, and adjust the resource allocation according to the business scenario (assuming 2 CPU cores and 2Gi of memory).

[0040] Create a Kubernetes Service object definition, construct the metadata of the Job Manager Service, including the name (such as job-manager-service), the namespace it belongs to (flink-cluster), labels (matching the labels of the Job Manager Deployment for easy service discovery), etc. According to the deployment requirements of the Flink cluster, select an appropriate service type, such as ClusterIP (only accessible within the cluster), NodePort (accessible from the outside through the IP and specific port of the cluster nodes), or LoadBalancer (if there is external load balancer support for providing services externally). At the same time, configure the ports exposed by the service and the corresponding target ports (i.e., the ports inside the Job Manager container).

[0041] Generate a ConfigMap object definition, construct the metadata of the Job Manager ConfigMap, including the name (such as job-manager-config), the namespace it belongs to (flink-cluster), etc., and add the content of the Job Manager-related configuration files (such as flink-conf.yaml) extracted from the Flink configuration template to the data field of the ConfigMap in the form of key-value pairs. Create a ConfigMap object definition for the Task Manager in the same way, set the metadata and populate the configuration data related to the Task Manager (such as configuration content corresponding to parameters like memory allocation, network buffers, etc.).

[0042] Integrate the defined information of resource objects such as the above-mentioned Deployment, Service, and ConfigMap into an overall Kubernetes resource manifest dictionary, so that these resources can be created at once through the Kubernetes API. Through the above detailed steps, the key information in the Flink cluster configuration template can be converted into the defined information of resource objects within the Kubernetes container, and then the corresponding resources can be created with the help of the Kubernetes API to achieve the automated deployment of the Flink cluster in the Kubernetes environment.

[0043] In step S103, obtain the deployment environment information corresponding to the Flink cluster, perform pre-deployment verification on multiple component resources based on the deployment environment information, and after passing the verification, deploy the multiple component resources corresponding to the Flink cluster to the target running environment.

[0044] In one embodiment of this specification, before deploying the Flink cluster to the production environment, comprehensive verification is performed on the generated configuration to check the legality, integrity of the configuration parameters, and compatibility with other systems (such as data sources, data storage systems, etc.). In addition, the configuration verification tool provided by Flink can be used, combined with custom verification scripts, to perform static analysis and simulation run verification on the configuration to ensure that there are no errors in the actual operation of the configuration. Additionally, obtain the deployment environment information corresponding to the Flink cluster, and perform pre-deployment verification on multiple component resources based on the deployment environment information.

[0045] Perform pre-deployment verification on the multiple component resources based on the deployment environment information, specifically including: calculate the expected total resources corresponding to each component resource respectively according to the resource configuration parameters in the Flink cluster configuration template, where the resources include CPU resources and memory resources; perform node resource verification on the expected total resources corresponding to each component resource through the available resource information of the cluster nodes in the deployment environment information to determine the node resource verification result; use a pre-set test task to test the Flink cluster, collect the task execution results, and when the task execution results meet the preset requirements, monitor the metrics in the test task for the Flink cluster to collect test monitoring metrics; verify each test monitoring metric to determine the overall function verification result corresponding to the Flink cluster; determine the verification result of the pre-deployment verification according to the node resource verification result and the overall function verification result.

[0046] In one embodiment of this specification, the expected total CPU of the Job Manager and the Task Manager is calculated respectively according to the CPU requests and limits set for the Job Manager and the Task Manager in the Flink configuration template. For example, if the CPU request of the Job Manager is 1 core, the CPU request of the Task Manager is 2 cores, and the number of replicas is 3, then the expected total CPU of the Task Manager is 6 cores. The available CPU resource information of the cluster nodes is obtained through the Kubernetes API. Check whether the total CPU resources requested by the Flink components (Pods) running on each node exceed the allocable CPU resources of the node. In addition, it is also possible to verify whether the CPU limit setting is reasonable, that is, to ensure that the limit value is not lower than the request value and the limit value is within the tolerable range of the node resources, so as to avoid the Pod being killed due to frequent memory overflow caused by too low a limit or resource waste caused by too high a limit. Similar to the CPU resource verification, the expected total memory of the Job Manager and the Task Manager is calculated. For example, if the memory request of the Job Manager is 1 Gi, the memory request of the Task Manager is 2 Gi, and the number of replicas is 3, then the expected total memory of the Task Manager is 6 Gi. Check whether the total memory allocated to the Flink components on the node exceeds the available memory of the node, and at the same time verify the relationship between the memory limit and the request to ensure that the memory limit can meet the memory requirements of the components under normal and peak loads. If the total CPU resources requested by the Flink components running on each node do not exceed the allocable CPU resources of the node, and the total memory allocated to the Flink components does not exceed the available memory of the node, it is determined that the node resource verification result is passed.

[0047] Submit a simple test task, such as a simple Word Count example task, to the Flink cluster through the REST API of the Job Manager or the Kubernetes command-line tool. Observe whether the task can be successfully submitted, scheduled to the Task Manager, and executed normally, and check whether the execution result of the task meets the expectations. Use monitoring tools to check various monitoring metrics of the Flink cluster during the execution of the test task, such as CPU utilization, memory usage, task throughput, latency, etc. Verify whether the above metrics are within a reasonable range and whether they match the expected resource usage in the configuration template, so as to judge whether the overall performance of the cluster is normal. When metrics such as CPU utilization, memory usage, task throughput, and latency are within a reasonable range and match the expected resource usage in the configuration template, the verification result of the overall function is determined to be passed. When both the node resource verification result and the overall function verification result pass, it is determined that the pre-deployment verification process has passed. After the verification passes, deploy multiple component resources corresponding to the Flink cluster to the target running environment.

[0048] Step S104, monitor the real-time performance metrics of the Flink cluster in the target running environment, so as to adjust the Flink cluster according to the pre-set dynamic adjustment policy and real-time performance metrics.

[0049] Among them, the adjustment includes any one or more of adaptive scaling operations and adaptive configuration parameter adjustments.

[0050] After deploying the Flink cluster, use monitoring tools such as Prometheus and Grafana to collect various performance metrics of the Flink cluster, such as CPU utilization, memory usage, task throughput, latency, etc. At the same time, collect Kubernetes-related metrics, such as the resource usage of Pods and node loads, in order to comprehensively understand the running status of the cluster. According to the collected performance metrics, formulate a dynamic adjustment policy. If it is found that the cluster resource utilization is too high or too low, the Flink cluster can be dynamically expanded or contracted through the Horizontal Pod Autoscaler (HPA) of Kubernetes or manual adjustment of resource allocation. If it is found that the performance of some tasks is not ideal, the configuration parameters of Flink (such as parallelism, memory allocation, etc.) can be adjusted according to the performance data, and the configuration can be hot-updated by updating the ConfigMap, so that the cluster can be adaptively adjusted according to the actual workload and always maintain an efficient running state.

[0051] Adjust the Flink cluster according to the pre-set dynamic adjustment strategy and the real-time performance metric, specifically including: when the real-time performance metric is workload data, the workload data includes task execution time, resource utilization rate, and data input / output rate; obtain the application service information of the Flink cluster, where the application service information includes data traffic characteristics, data real-time characteristics, and service processing logic characteristics; according to the application service information of the Flink cluster, match the scaling metric corresponding to the Flink cluster, so as to determine the adaptive scaling strategy based on the scaling metric, where the scaling metric includes any one or more of resource utilization rate, task queue length, amount of data input per second, and data output; in the real-time performance metric, obtain the real-time metric parameter corresponding to the scaling metric, and compare the real-time metric parameter with the pre-set scaling metric threshold in the adaptive scaling strategy. When the real-time metric parameter meets the scaling metric threshold, trigger the pre-set Kubernetes horizontal Pod autoscaler to dynamically expand or contract the component resources of the Flink cluster.

[0052] In an embodiment of this specification, when the real-time performance metric is workload data, the workload data includes task execution time, resource utilization rate, and data input / output rate. Obtain the application service information of the Flink cluster, and the application service information includes data traffic characteristics, data real-time characteristics, and service processing logic characteristics. Generally, when the Flink cluster runs in a resource-constrained environment, such as a shared Kubernetes cluster, and multiple applications compete for resources, scaling is usually based on the utilization of conventional resources such as CPU and memory. For example, in a cluster that is hybrid-deployed with multiple big data processing tasks and other microservices, in order to ensure that Flink tasks do not over-occupy resources and affect other services, and at the same time avoid performance problems due to resource exhaustion, it is set to scale when the CPU utilization rate of the Task Manager reaches 70% or the memory usage rate reaches 80% (these thresholds can be adjusted according to the actual situation). In addition, for some Flink tasks that do not have obvious business peaks and valleys, the workload is relatively stable, but the data volume and computing intensity may fluctuate to a certain extent. For example, a Flink job that regularly processes log files, the size and complexity of the log files may vary, but the processing process is relatively standardized. In this case, dynamically adjusting the number of Task Managers according to the resource utilization rate can effectively respond to changes in resource requirements.

[0053] Therefore, if the Flink cluster is deployed in a resource-constrained scenario such as a shared Kubernetes cluster and needs to compete for resources with multiple applications, to ensure the normal operation of each application and avoid resource exhaustion problems itself, it is preferable to choose a scaling strategy based on conventional resource utilization (such as CPU and memory). It can be set to trigger scaling when the CPU utilization of the Task Manager reaches a specific percentage (such as 70%) or the memory usage reaches a corresponding ratio (such as 80%) according to the actual situation. For Flink tasks with relatively stable workloads, without obvious business peaks and valleys, although the data volume and computing intensity will have certain fluctuations, dynamically adjusting the number of Task Managers using conventional resource utilization can effectively handle changes in resource requirements and maintain the stable operation of the cluster. For example, jobs that process regular log files in a standardized process.

[0054] For Flink applications with extremely high real-time requirements, such as real-time risk monitoring in financial trading systems and real-time traffic analysis in telecommunications networks, the task queue length is a key business metric. If the task queue length continues to increase, it means that the data processing speed cannot keep up with the data input speed, which may lead to data latency and business risks. At this time, perform custom scaling based on the task queue length. When the queue length reaches a certain threshold, quickly increase the number of Task Managers to ensure that the data can be processed in a timely manner and meet the real-time requirements. For Flink applications with extremely high real-time requirements, such as real-time risk monitoring in financial trading systems and real-time traffic analysis in telecommunications networks, given that the task queue length is directly related to the timeliness of data processing and business risks, when the business is extremely sensitive to real-time, a custom scaling strategy should be based on the task queue length, and when it reaches the set threshold, the number of Task Managers should be increased in a timely manner to ensure that the data is processed in a timely manner.

[0055] In some scenarios of user behavior analysis in Internet applications, the amount of data input per second may fluctuate greatly with changes in user activity. For example, during promotional activities on e-commerce platforms, the volume of user browsing, purchasing, and other behavior data will increase sharply. Performing a custom scaling strategy based on this business metric of the amount of data input per second can expand capacity in advance when the data traffic peak arrives and automatically scale down during the traffic trough, thereby effectively utilizing resources and ensuring the timeliness of data processing. If there are large fluctuations in the data traffic in the business scenario where the Flink task is located, such as the sharp change in user behavior data volume during promotional activities on e-commerce platforms, to achieve efficient resource utilization and ensure the timeliness of data processing, a scaling strategy needs to be customized based on business metrics such as the amount of data input per second that can reflect the data traffic situation, realizing capacity expansion during peaks and capacity reduction during troughs.

[0056] When Flink tasks involve complex business logics, such as multi-level data processing pipelines, multi-source data fusion, etc., and there are specific requirements for the generation speed or processing quality of some intermediate results, specific business metrics become particularly important. For example, in a Flink task for data cleaning and feature extraction involving multiple data sources, it is required that the processing speed of a certain key feature extraction stage cannot be lower than a certain value. At this time, a scaling strategy can be customized according to business metrics such as the processing progress or data output volume of this stage to ensure the efficient operation of the entire task. Therefore, when Flink tasks cover complex business logics, such as the existence of multi-level data processing pipelines, multi-source data fusion, etc., and there are specific performance requirements for aspects such as the generation speed and processing quality of intermediate results, a scaling strategy should be customized according to specific business metrics (such as the processing progress and data output volume of the key feature extraction stage, etc.) that reflect these key requirements to ensure the efficient operation of the entire task.

[0057] According to the application business information of the Flink cluster and the above rules, match the corresponding scaling metrics for the Flink cluster. The scaling metrics include any one or more of resource utilization rate, task queue length, input data volume per second, and data output volume. Based on the scaling metrics, determine an adaptive scaling strategy. In this adaptive scaling strategy, there are preset scaling metric thresholds corresponding to the scaling metrics, which are used to trigger contraction or expansion actions. Among the real-time performance metrics, obtain the real-time metric parameters corresponding to the scaling metric, and compare the real-time metric parameters with the scaling metric threshold. When the real-time metric parameters meet the scaling metric threshold, trigger the preset Kubernetes horizontal Pod autoscaler to dynamically expand or contract the component resources of the Flink cluster. For example, for the horizontal Pod autoscaler (HPA) configuration based on CPU utilization, the expansion threshold is set to 70%, and the contraction threshold is set to 40%. Automatically monitor the CPU utilization of the Task Manager. When the average utilization exceeds the set 70%, the number of Task Manager replicas will be automatically increased to meet the workload requirements; when the utilization drops below 40%, the number of replicas will be reduced accordingly to release resources.

[0058] Through the above technical solutions, by monitoring workload data in real time (such as task execution time, resource utilization rate, data input / output rate, etc.) and matching corresponding scaling metrics based on the application business information of the Flink cluster, the cluster can accurately adjust resources according to the actual business needs; for different business scenarios, based on appropriate scaling metrics (such as scaling according to the amount of input data per second for businesses with large fluctuations in data input / output), resource adjustment is carried out, avoiding over-allocation or under-allocation of resources, keeping the cluster in a relatively efficient resource configuration state all the time, helping to optimize task execution time and improve data processing efficiency; dynamically adjusting the cluster scale according to the real-time performance metrics of the business can better cope with changes in business load; considering the application business information of the Flink cluster (including data traffic characteristics, data real-time characteristics, and business processing logic characteristics) to determine scaling metrics and adaptive scaling strategies enables the cluster to adapt to the unique requirements in different business scenarios; monitoring workload data in real time and triggering scaling operations based on preset thresholds can prevent risks such as system crashes and data loss caused by excessive business load in advance. For example, when the resource utilization rate approaches the critical value, resources are expanded in a timely manner to prevent tasks from being unable to execute normally due to resource exhaustion, thus ensuring business continuity and reducing the risk of business interruption caused by technical problems; the entire adjustment process is automatically triggered based on the pre-set dynamic adjustment strategy and real-time performance metrics, without frequent manual intervention, greatly reducing the workload of operation and maintenance personnel and lowering operation and maintenance costs.

[0059] According to the pre-set dynamic adjustment strategy and the real-time performance metric, adjust the Flink cluster, specifically including: when the real-time performance metric is a node processing metric, the node processing metric includes throughput parameters and data latency parameters; monitor the configuration change requirement information according to the real-time throughput parameter and the real-time data latency parameter, and when any one of the real-time throughput parameter and the real-time data latency parameter does not meet the corresponding preset performance condition, trigger a configuration change requirement; under the trigger of the configuration change requirement, use the performance metric corresponding to the performance condition as input, and determine the target configuration parameter combination through a pre-trained performance parameter matching model; adaptively adjust the configuration parameters of the Flink cluster with the target configuration parameter combination.

[0060] In one embodiment of this specification, when the real-time performance metric is the node processing metric, the node processing metric includes throughput parameters and data latency parameters; the configuration change requirement information is monitored according to the real-time throughput parameters and real-time data latency parameters in the real-time performance metric. When any one of the real-time throughput parameters and real-time data latency parameters does not meet the corresponding preset performance condition, a configuration change requirement is triggered. According to business requirements and empirical data, reasonable preset performance conditions are set for the throughput parameters and data latency parameters respectively in advance. For example, it is stipulated that the throughput rate should reach a certain amount of data processed per hour, and the latency should be controlled within a specific time range, etc. The above preset conditions fit the actual business scenario and can be continuously adjusted and optimized according to business development. By comparing the index data obtained in real time with the preset performance conditions, when it is found that any one does not meet, a corresponding configuration change requirement signal is triggered.

[0061] Using historical configuration parameter data and corresponding real-time throughput, latency and other performance metric data as training samples, a performance parameter matching model is constructed. According to the actual situation, a machine learning algorithm is selected, and the internal relationship between the configuration parameters and performance metrics is mined through learning a large amount of historical data, so as to train a performance parameter matching model that can predict the appropriate configuration parameter combination according to the input performance metrics. Under the trigger of the configuration change requirement, using the performance metrics corresponding to the performance conditions as the input, the target configuration parameter combination is output through the performance parameter matching model. In the Flink cluster, the configuration parameters are usually stored in a configuration file (such as flink-conf.yaml), and can be managed and mounted into the corresponding containers through mechanisms such as Kubernetes' ConfigMap. When the target configuration parameter combination is determined, using the Kubernetes API or the configuration update interface provided by Flink, the new configuration parameters can be updated to the corresponding configuration file, and the Flink component is notified to reload the configuration to achieve the adaptive adjustment of the configuration.

[0062] By monitoring throughput and latency parameters in real time and triggering configuration adjustments in a timely manner according to their comparison with preset performance conditions, the configuration parameters of the Flink cluster can be dynamically optimized, enabling the cluster to always operate close to the optimal performance state; different services have strict requirements for performance metrics such as throughput and latency. For example, real-time financial transaction monitoring requires extremely low data latency, and data processing during e-commerce big promotions requires high throughput. Through real-time monitoring and adaptive adjustment, this solution can ensure that the performance of the Flink cluster always meets the key performance requirements of the service, guarantee the stable operation of the service, and avoid service risks caused by performance problems; by monitoring performance metrics in real time and adjusting the configuration in a timely manner, problems can be detected and optimized at the initial stage of performance degradation, avoiding performance deterioration and subsequent failures caused by unreasonable configuration; and when there are some temporary performance fluctuations, it can quickly and automatically recover to a reasonable configuration state, improving the reliability and stability of the cluster, reducing the probability of failures and the time for fault repair.

[0063] Adjust the Flink cluster according to the preset dynamic adjustment strategy and the real-time performance metrics, specifically including: when the real-time performance metric is a state data metric, the state data metric includes state data size information and state data access frequency; according to the real-time state data size information and the real-time state data access frequency, match the corresponding state data storage type, where the storage type includes any one of high-performance solid-state drive storage, large-capacity hard disk drive storage, and distributed storage systems; monitor the real-time state data size information, and when the real-time state data size information meets the preset storage capacity adjustment threshold, adjust the requested storage capacity size corresponding to the state data storage type.

[0064] In an embodiment of this specification, when the real-time performance metric is a state data metric, the state data metric includes state data size information and state data access frequency. Comprehensive analysis is performed based on the collected state data size and access frequency information.

[0065] In general, for small-sized but highly accessed status data, such as the critical metadata frequently used for task scheduling decisions in the Job Manager, high-performance SSD storage can provide fast data read and write speeds, meet the access requirements with high real-time requirements, and reduce task scheduling delays caused by storage read and write latency. For large-sized and relatively low-access-frequency status data, such as some long-term accumulated historical task execution records in the Task Manager, large-capacity HDD storage is adopted, which can reduce storage costs while meeting the data storage requirements. For some situations that require high reliability, support multi-copy storage, and may involve sharing status data across large-scale clusters and nodes, such as the global status data that needs to be collaboratively accessed among multiple Job Managers or Task Managers in a distributed Flink cluster, choosing a distributed storage system (such as Ceph, GlusterFS, etc.) can provide data redundancy, high availability, and good horizontal scalability, facilitating the cluster to ensure the integrity and accessibility of status data during node expansion or fault recovery.

[0066] According to the above empirical data, storage type matching rules are preset. For example, it is set that when the access frequency of the status data is higher than 100 times per minute and the data size is less than 1GB, high-performance SSD storage is preferentially selected; when the access frequency is lower than 10 times per minute and the data size exceeds 10GB, large-capacity HDD storage is considered; and for data that needs to be accessed simultaneously across multiple nodes and has high requirements for data consistency and redundancy, regardless of its size and access frequency, a distributed storage system is selected. Monitor the real-time status data size information. When the real-time status data size information meets the preset storage capacity adjustment threshold, adjust the requested storage capacity size corresponding to the storage type of the status data. The storage capacity adjustment threshold here can be set to 80%. When the real-time status data size information approaches 80% of the current PVC capacity, trigger the dynamic adjustment operation of the PVC size. Through the Kubernetes API, write an automated script or use a dedicated storage management tool to modify the size parameters of the PVC to achieve dynamic expansion of the storage capacity. At the same time, when the business load drops and the amount of status data decreases significantly, if the PVC capacity utilization rate is too low (such as lower than 30%), the size of the PVC can be appropriately shrunk to release redundant storage resources and avoid resource waste.

[0067] In one embodiment of this specification, a unified monitoring platform deeply integrated with Kubernetes is built to seamlessly interface with common monitoring systems such as Prometheus and Grafana, comprehensively and deeply collecting and analyzing various metrics of the Flink cluster. In addition to regular resource metrics such as CPU, memory, and network, and execution metrics of Flink tasks (such as completion time, throughput, latency, etc.), the internal states of Flink jobs (such as the states of operators, the usage of data buffers, etc.) are also monitored in real time and visually displayed. Through customized Grafana dashboards, users can intuitively understand the overall operation status of the Flink cluster, quickly discover potential performance issues and fault hazards, and conduct in-depth fault troubleshooting and performance optimization through an interactive interface.

[0068] A log analysis engine based on machine learning is adopted to collect, parse, and analyze the log data of the Flink cluster in real time. The log analysis engine can automatically identify key information in the logs (such as error messages, warning messages, log events related to performance bottlenecks, etc.), and classify and aggregate them according to predefined rules. At the same time, by continuously learning the log data over a long period, a log pattern library is established, which can automatically discover abnormal log patterns, predict potential faults in advance, and promptly send alerts to administrators. In addition, full-text search and visual analysis of log data are also supported, facilitating users to quickly locate and solve problems and improving the efficiency of fault troubleshooting.

[0069] In one embodiment of this specification, when establishing a secure communication channel based on two-way TLS / SSL between Flink components, an innovative certificate management and key distribution mechanism is adopted. Utilizing the Secrets function of Kubernetes, certificates and keys are securely stored and managed, and an automated certificate update and key rotation system is developed to ensure the security and reliability of communication. At the same time, strict control is exerted over the network access of Flink components. Through Network Policies, network isolation and access control between components are achieved, allowing only legitimate communication traffic to be transmitted between components to prevent malicious attacks and data leakage.

[0070] Based on the Role-Based Access Control (RBAC) mechanism of Kubernetes, further refine the access control policy of the Flink cluster. Define fine-grained permission sets for different user roles (such as administrators, developers, operators, etc.), including not only the access permissions to Flink cluster resources (such as creating, deleting, modifying cluster configurations, etc.), but also the operation permissions to Flink jobs (such as submitting jobs, stopping jobs, viewing job status, etc.). At the same time, establish a perfect compliance audit system to record and audit the operation behaviors of all users in detail, ensuring that the operations comply with the requirements of relevant regulations such as GDPR and HIPAA. When a violation is detected, an alarm can be issued in a timely manner and corresponding measures can be taken (such as blocking the operation, recording the violation event, etc.) to ensure the security and compliance of the cluster.

[0071] Through the above technical solutions, by obtaining the cluster configuration requirement information input by the user to generate the Flink cluster configuration template, it avoids the cumbersome process of manually writing complex configuration files from scratch in the traditional way. In addition, the method of generating the template according to the cluster configuration requirements input by the user can ensure the standardization and consistency of the configuration, reducing configuration errors caused by human negligence or differences in understanding; converting the Flink cluster configuration template into the resource object definition information within the Kubernetes (k8s) container realizes the deep integration of the Flink cluster and the powerful container orchestration ability of Kubernetes. Using the standardized resource management and deployment mechanism of Kubernetes, it is possible to create multiple component resources corresponding to the Flink cluster more conveniently and automatically, avoiding problems such as format errors and unreasonable resource mapping that may occur when manually configuring Kubernetes resources, and further improving the accuracy and efficiency of deployment; obtaining the deployment environment information corresponding to the Flink cluster and performing pre-deployment verification on multiple component resources based on this to discover potential configuration incompatibilities, resource conflicts, etc. in advance; by monitoring the real-time performance metrics of the Flink cluster in the target running environment and performing adaptive scaling operations according to the pre-set dynamic adjustment strategy, the Flink cluster can dynamically adjust the resource scale according to the actual business load; the dynamic adjustment mechanism can automatically adapt to changes in business logic, evolution of data characteristics, etc., ensuring that the Flink cluster continuously maintains good performance during long-term operation and avoiding performance bottleneck problems caused by fixed configurations being unable to adapt to business development; covering automated operations in multiple links from configuration generation, deployment to runtime dynamic adjustment, greatly reducing the manual operation workload of operators in the process of Flink cluster management.

[0072] The embodiment of this specification also provides a Flink cluster deployment management device based on k8s, such as Figure 2As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0073] An embodiment of this specification also provides a non-volatile computer storage medium storing computer-executable instructions configured to: execute the above method.

[0074] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0075] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0076] The device and medium provided by the embodiments of this specification correspond one-to-one with the method. Therefore, the device and medium also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.

[0077] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0078] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0081] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0082] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.

[0083] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0084] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0085] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.

Claims

1. A Flink cluster deployment management method based on k8s, characterized in that: The method comprises: Obtain cluster configuration requirement information input by the user, and generate a Flink cluster configuration template based on the cluster configuration requirement information; Convert the Flink cluster configuration template into resource object definition information in the k8s container, so as to create multiple component resources corresponding to the Flink cluster by using the resource object definition information; Obtaining deployment environment information corresponding to the Flink cluster, performing pre-deployment verification on the multiple component resources based on the deployment environment information, and after the verification passes, deploying the multiple component resources corresponding to the Flink cluster to the target operating environment; Monitor the real-time performance indicators of the Flink cluster in the target operating environment to adjust the Flink cluster according to a preset dynamic adjustment strategy and the real-time performance indicators, wherein the adjustment includes any one or more of an adaptive scaling operation and an adaptive configuration parameter adjustment.

2. According to a k8s-based Flink cluster deployment management method according to claim 1, it is characterized in that: Based on the cluster configuration requirement information, a Flink cluster configuration template is generated, which specifically includes: Collect historical cluster operation data of the Flink cluster in the actual operation environment, where the historical cluster operation data includes historical configuration information and historical operation performance information corresponding to different business loads; Analyze the relationship between the cluster performance and configuration parameters of the Flink cluster through the historical cluster operation data, and determine a performance parameter matching model corresponding to the Flink cluster, wherein the performance parameter matching model includes a relationship between each key performance indicator and at least one key configuration parameter; According to the cluster configuration requirement information and the performance parameter matching model, a corresponding Flink cluster resource amount is determined, and a Flink cluster configuration template is generated through the Flink cluster resource amount, wherein the Flink cluster resource amount includes a parallelism parameter, a memory requirement parameter, and a CPU core allocation parameter.

3. According to a k8s-based Flink cluster deployment management method according to claim 2, it is characterized in that: According to the cluster configuration requirement information and the performance parameter matching model, the corresponding Flink cluster resource amount is determined, specifically including: Determine the data processing type, required throughput parameter, data delay requirement parameter and job concurrency parameter in the cluster configuration requirement information, and determine the parallelism matching sub-model, memory requirement matching sub-model and CPU core number matching sub-model in the performance parameter matching model; Determining the parallelism parameter according to the required throughput parameter and the parallelism matching sub-model; Determine the memory requirement parameters corresponding to the Job Manager and Task Manager according to the data processing type and the job concurrency parameter and the memory requirement matching sub-model; Based on the required throughput parameter and the data delay requirement parameter, the CPU core allocation parameter is determined using the CPU core number matching sub-model.

4. According to a k8s-based Flink cluster deployment management method according to claim 1, it is characterized in that: Performing pre-deployment verification on the plurality of component resources based on the deployment environment information specifically includes: According to the resource configuration parameters in the Flink cluster configuration template, the expected total amount of resources corresponding to each of the component resources is calculated respectively, wherein the resources include CPU resources and memory resources; Perform node resource verification on the expected total amount of resources corresponding to each of the component resources through the available resource information of the cluster nodes in the deployment environment information, and determine the node resource verification result; Using a pre-set test task, testing the Flink cluster, collecting task execution results, and when the task execution results meet preset requirements, monitoring the indicators in the test task executed by the Flink cluster to collect test monitoring indicators; Verify each of the test monitoring indicators to determine the overall functional verification result corresponding to the Flink cluster; The verification result of the pre-deployment verification is determined according to the node resource verification result and the overall function verification result.

5. According to a k8s-based Flink cluster deployment management method according to claim 1, it is characterized in that: According to the preset dynamic adjustment strategy and the real-time performance indicator, the Flink cluster is adjusted, specifically including: When the real-time performance indicator is workload data, the workload data includes task execution time, resource utilization, and data input and output rate; Obtain application business information of the Flink cluster, wherein the application business information includes data flow characteristics, data real-time characteristics, and business processing logic characteristics; According to the application business information of the Flink cluster, a scaling indicator corresponding to the Flink cluster is matched to determine an adaptive scaling strategy based on the scaling indicator, wherein the scaling indicator includes any one or more of resource utilization, task queue length, input data volume per second, and data output volume; In the real-time performance indicator, the real-time indicator parameter corresponding to the scaling indicator is obtained, and the real-time indicator parameter is compared with the scaling indicator threshold preset in the adaptive scaling strategy. When the real-time indicator parameter meets the scaling indicator threshold, the preset Kubernetes horizontal Pod autoscaler is triggered to dynamically expand or shrink the component resources of the Flink cluster.

6. According to a k8s-based Flink cluster deployment management method according to claim 1, it is characterized in that: According to the preset dynamic adjustment strategy and the real-time performance indicator, the Flink cluster is adjusted, specifically including: When the real-time performance indicator is a node processing indicator, the node processing indicator includes a throughput parameter and a data delay parameter; Monitoring configuration change requirement information according to the real-time throughput parameter and the real-time data delay parameter in the real-time performance indicator, and triggering the configuration change requirement when any one of the real-time throughput parameter and the real-time data delay parameter does not meet the corresponding preset performance condition; Under the triggering of the configuration change requirement, the performance indicator corresponding to the performance condition is used as input, and a target configuration parameter combination is determined through a pre-trained performance parameter matching model; Adaptively adjust the configuration parameters of the Flink cluster using the target configuration parameter combination.

7. According to a k8s-based Flink cluster deployment management method according to claim 1, it is characterized in that: According to the preset dynamic adjustment strategy and the real-time performance indicator, the Flink cluster is adjusted, specifically including: When the real-time performance indicator is a state data indicator, the state data indicator includes state data size information and state data access frequency; According to the real-time status data size information and the real-time status data access frequency in the real-time performance indicator, a corresponding status data storage type is matched, wherein the storage type includes any one of high-performance solid-state hard disk storage, large-capacity hard disk drive storage and distributed storage system; The real-time status data size information is monitored, and when the real-time status data size information meets a preset storage capacity adjustment threshold, the requested storage capacity size corresponding to the status data storage type is adjusted.

8. According to a k8s-based Flink cluster deployment management method according to claim 2, it is characterized in that: Generate a Flink cluster configuration template based on the Flink cluster resource quantity, which specifically includes: Acquire auxiliary configuration parameters, wherein the auxiliary configuration parameters include any one or more of a network configuration parameter, a checkpoint and fault tolerance configuration parameter interval, and a resource allocation priority parameter; A Flink cluster configuration template is generated in YAML format according to the Flink cluster resource amount and the auxiliary configuration parameters.

9. A Flink cluster deployment management device based on k8s, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.