Flink multi-cluster secure authentication data synchronization method and system

By establishing secure connection channels between Flink clusters and optimizing incremental encoding compression algorithms, the security and efficiency issues of data synchronization between Flink clusters are resolved, enabling secure and reliable data transmission and state recovery, and improving resource utilization and data consistency.

CN120710799BActive Publication Date: 2025-11-18北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188137.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-18
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

The existing Flink cluster inter-data synchronization has an imperfect security authentication mechanism, lacks a unified authentication standard and secure connection channel, which leads to security risks during data transmission, lacks dynamic optimization capabilities for resource allocation, affects data synchronization efficiency and data quality control, and cannot guarantee data consistency and integrity.

Method used

By receiving data synchronization requests from the source Flink cluster, obtaining access tokens based on authentication information and establishing a secure connection channel, using incremental coding compression algorithms to dynamically optimize data processing operators, generating fine-grained task fragments, and performing data quality assessment and state recovery in the target Flink cluster, and dynamically allocating computing resources, the security and efficiency of data transmission are improved.

Benefits of technology

It enables secure data synchronization between multiple Flink clusters, improves the security and reliability of data transmission, enhances data processing efficiency and resource utilization, ensures the integrity and consistency of the data synchronization process, and strengthens the system's fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120710799B_ABST
    Figure CN120710799B_ABST
Patent Text Reader

Abstract

The application provides a Flink multi-cluster security authentication data synchronization method and system, relates to the technical field of distributed computing, and comprises the following steps: establishing a secure connection channel, adopting an incremental encoding compression algorithm to optimize an execution plan, setting an adaptive computing load threshold, dynamically allocating computing resources, writing state checkpoint information, performing data quality evaluation and writing the evaluation data into a distributed ledger, realizing target cluster state recovery and starting a synchronization task, improving data synchronization security and efficiency, enhancing computing resource utilization, and ensuring data quality and synchronization consistency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed computing, and in particular to a Flink multi-cluster secure authentication data synchronization method and system. BACKGROUND

[0002] With the growth of big data processing needs, Apache Flink as a stream processing framework has been widely used in real-time data processing scenarios, Flink supports low-latency, high-throughput data stream processing, while providing state management, fault tolerance mechanism and exactly-once semantics, etc. In enterprise-level applications, data synchronization between different Flink clusters is often required to meet data backup, disaster recovery, multi-region deployment and other needs;

[0003] Traditional Flink inter-cluster data synchronization usually uses an intermediate storage system as a bridge, data needs to be written from the source cluster to the intermediate storage first, and then read by the target cluster, which increases the delay and complexity of data synchronization. Direct inter-cluster data transmission has become a kind of demand that cannot be ignored. However, the existing technology still has the problems of imperfect security authentication mechanism, lack of unified authentication standard and secure connection channel, resulting in security risks in data transmission process, lack of dynamic optimization ability for task scheduling and resource allocation, inability to adaptively adjust resource allocation according to actual computing load, low resource utilization, affecting data synchronization efficiency, lack of effective data quality control mechanism, inability to real-time evaluate and verify data quality in the synchronization process, and difficulty in guaranteeing data consistency and integrity.

[0004] Therefore, there is an urgent need for a solution to solve the problems in the prior art. SUMMARY

[0005] The Flink multi-cluster secure authentication data synchronization method and system provided by the embodiments of the present application can at least solve some of the problems in the prior art.

[0006] In a first aspect, the present application provides a Flink multi-cluster secure authentication data synchronization method, comprising:

[0007] Receiving a data synchronization request of a source Flink cluster, obtaining an access token based on authentication information in the data synchronization request and establishing a secure connection channel;

[0008] In the source Flink cluster, a incremental encoding compression algorithm is used to perform dynamic optimization on data processing operators and generate an execution plan, the execution plan is processed in a distributed parallel manner to obtain fine-grained task shards, an adaptive computing load threshold is set for each task shard and computing resources are dynamically allocated, and state checkpoint information is written in the task shard;

[0009] sending the task shards to the target Flink cluster through the secure connection channel;

[0010] After receiving the task shards in the target Flink cluster, performing data quality evaluation on the task shards, writing evaluation data into a distributed ledger, judging whether data quality meets a preset threshold according to the evaluation data, allocating computing nodes from an elastic computing resource pool when the preset threshold is met, distributing the task shards to each computing node for parallel reorganization, and restoring a data processing state according to the state checkpoint information;

[0011] After restoring the data processing state, starting a data synchronization task in the target Flink cluster and returning confirmation information.

[0012] In an optional implementation,

[0013] receiving a data synchronization request of a source Flink cluster, obtaining an access token based on authentication information in the data synchronization request, and establishing a secure connection channel include:

[0014] receiving a data synchronization request sent by a source Flink cluster, the data synchronization request including authentication information and target Flink cluster address information;

[0015] generating a temporary session key by performing a two-way authentication protocol between an authentication center and the source Flink cluster, and obtaining an access token by encrypting the authentication information using the temporary session key;

[0016] establishing a secure connection channel according to the access token and the target Flink cluster address information, and configuring an end-to-end encrypted transmission mechanism based on the temporary session key in the secure connection channel.

[0017] In an optional implementation,

[0018] In the source Flink cluster, performing dynamic optimization on a data processing operator using an incremental encoding compression algorithm and generating an execution plan, processing the execution plan in a distributed parallel manner to obtain fine-grained task shards, setting an adaptive computing load threshold for each task shard and dynamically allocating computing resources, and writing state checkpoint information in the task shards include:

[0019] In the source Flink cluster, performing timing analysis on a data processing operator, writing a set corresponding to the data processing operator and a pre-acquired data stream dependency relationship into an operator dependency graph, obtaining an incremental change amount of a data stream in the data stream dependency relationship based on the operator dependency graph and setting a weight coefficient, collecting execution times of operators in the operator dependency graph, substituting the incremental change amount, the weight coefficient, and the execution times into an encoding function of the incremental encoding compression algorithm and minimizing the encoding function to obtain an execution plan.

[0020] build a set of plan nodes and a set of hierarchical relationships in the execution plan, collect memory consumption, data dependency and input / output overhead of each plan node in the set of plan nodes and calculate a complexity evaluation value, and when the complexity evaluation value exceeds a preset threshold, process the execution plan in a distributed parallel manner to generate fine-grained task shards;

[0021] collect processor usage, memory occupancy and network transmission rate of the fine-grained task shards and calculate an adaptive computing load threshold, and dynamically allocate computing resources based on the proportion of the adaptive computing load threshold in the total sum of adaptive computing load thresholds of all task shards in the fine-grained task shards;

[0022] write state identifiers, state data and timestamps into the fine-grained task shards to build state checkpoint information, calculate the difference between the state checkpoint information corresponding to two adjacent timestamps to obtain incremental checkpoint information and write the incremental checkpoint information into the fine-grained task shards.

[0023] In an optional implementation,

[0024] collect memory consumption, data dependency and input / output overhead of each plan node in the set of plan nodes and calculate a complexity evaluation value, and when the complexity evaluation value exceeds a preset threshold, process the execution plan in a distributed parallel manner to generate fine-grained task shards, including:

[0025] extract a time sequence feature sequence corresponding to the execution plan, calculate a hidden state output according to time sequence correlation of historical data in a specified time window in the time sequence feature sequence, calculate an attention score for the hidden state output to obtain a feature weight coefficient, multiply the time sequence feature sequence by the feature weight coefficient and sum to obtain a task feature vector, and perform recursive iteration calculation on the task feature vector to obtain a predicted complexity;

[0026] calculate a topology feature matrix between plan nodes, calculate a propagation influence value of data flow between nodes based on the topology feature matrix, perform graph embedding operation on node degree distribution value and edge weight distribution value to obtain data dependency strength, and calculate a resource competition degree matrix to obtain resource competition degree;

[0027] build a state transition probability matrix based on execution states, reward values and action sequences in the pre-acquired historical execution data, calculate an optimal balance coefficient based on the state transition probability matrix, multiply the predicted complexity, data dependency strength and resource competition degree by the balance coefficient respectively and perform joint optimization to obtain a complexity evaluation value;

[0028] When the complexity evaluation value is greater than the preset threshold value, a task similarity matrix is calculated based on data dependency strength, and a spectral clustering operation is performed to obtain a task cooperative execution group, a chromosome coding sequence is constructed for tasks in the task cooperative execution group according to calculation complexity and resource demand, an optimal parallelism configuration scheme is obtained through iterative optimization of a genetic algorithm, and the task cooperative execution group is distributed to different computing nodes according to the parallelism configuration scheme to perform parallel processing to obtain fine-grained task shards.

[0029] In an optional implementation,

[0030] After the target Flink cluster receives the task shards, data quality evaluation is performed on the task shards, evaluation data is written into a distributed ledger, and it is determined whether the data quality meets a preset threshold value according to the evaluation data. When the preset threshold value is met, a computing node is allocated from an elastic computing resource pool, the task shards are distributed to each computing node for parallel reorganization, and the data processing state is recovered according to the state checkpoint information, including:

[0031] After the target Flink cluster receives the task shards, the ratio of the number of missing values in each field in the task shards to the total number of records is calculated to obtain a data missing rate, the entropy value of the data value occurrence probability is calculated to obtain a data consistency score, and the standard deviation of the numerical data relative to the mean value is calculated to obtain an outlying degree. The evaluation data is constructed based on the data missing rate, the data consistency score, and the outlying degree.

[0032] The evaluation data is written into a hierarchical storage structure of a distributed ledger, and it is determined whether the data quality meets a preset threshold value according to the evaluation data. If the preset threshold value is met, the task complexity, memory demand, and network bandwidth demand are determined, and the resource demand amount is calculated. A multi-objective optimization function is constructed by combining resource utilization, load balancing degree, and communication overhead, and a resource allocation scheme is obtained by solving. Computing nodes are allocated from a pre-set elastic computing resource pool according to the resource allocation scheme.

[0033] The task shards are distributed to each computing node for parallel reorganization, the state item identifier sequence in the state checkpoint information is extracted, the dependency relationship between state items is calculated by traversing the state item identifier sequence to obtain a state dependency matrix, the state transition cost corresponding to the state items is calculated by analyzing and calculating the state dependency matrix using a longest common subsequence algorithm, a state recovery path is constructed based on the state transition cost, the total cost of the state recovery path is taken as an incremental recovery cost, and the data processing state is recovered according to the incremental recovery cost.

[0034] In an optional implementation,

[0035] allocating the task fragments to the computing nodes for parallel reconstruction, extracting a state item identification sequence from the state checkpoint information, calculating state transition costs corresponding to the state items using a longest common subsequence algorithm, and constructing a state recovery path based on the state transition costs include:

[0036] allocating the task fragments to the computing nodes for parallel reconstruction, extracting a state item identification sequence from the state checkpoint information, traversing the state item identification sequence to calculate a dependency relationship between the state items, constructing a forward dependency chain and a backward dependency chain based on the dependency relationship, and storing analysis results of the forward dependency chain and the backward dependency chain in a state dependency matrix;

[0037] simultaneously starting from a start end and a terminal end of the state item identification sequence, calculating using a longest common subsequence algorithm, and calculating state transition costs according to a forward dependency number, a backward dependency number, and a distance between the state items in the state dependency matrix;

[0038] calculating a dependency strength index based on the forward dependency number and the backward dependency number of the state items, dynamically adjusting a search direction according to a comparison result of the dependency strength index and a preset dependency strength threshold, generating a forward search result and a backward search result, calculating a ratio of the forward dependency number of each state item to a total dependency amount of the state item to obtain a dynamic merging weight, and performing weighted merging on the forward search result and the backward search result according to the dynamic merging weight to obtain a state transition sequence;

[0039] constructing a state recovery path based on the state transition costs and the state transition sequence, obtaining the state recovery path by minimizing a total sum of the state transition costs between adjacent state items in the state item identification sequence, and recovering a data processing state according to the state recovery path.

[0040] In an optional implementation,

[0041] after the data processing state is recovered, starting a data synchronization task in the target Flink cluster and returning confirmation information includes:

[0042] obtaining a recovery completion state of the data processing state from the target Flink cluster, verifying whether a recovery result of the data processing state satisfies a preset running condition, initializing a running environment of the data synchronization task in the target Flink cluster if the preset running condition is satisfied, loading configuration parameters required by the data synchronization task, and starting the data synchronization task;

[0043] monitoring a starting process of the data synchronization task, obtaining running state information of the data synchronization task, and returning the running state information as the confirmation information.

[0044] In a second aspect, the present application provides a Flink multi-cluster secure authentication data synchronization system, comprising:

[0045] A first unit is configured to receive a data synchronization request of a source Flink cluster, obtain an access token based on authentication information in the data synchronization request, and establish a secure connection channel;

[0046] A second unit is configured to perform dynamic optimization on a data processing operator in the source Flink cluster by using an incremental encoding compression algorithm, generate an execution plan, process the execution plan in a distributed parallel manner to obtain fine-grained task shards, set an adaptive computing load threshold for each task shard and dynamically allocate computing resources, and write state checkpoint information in the task shard;

[0047] A third unit is configured to send the task shard to a target Flink cluster through the secure connection channel;

[0048] A fourth unit is configured to perform data quality evaluation on the task shard after the target Flink cluster receives the task shard, write the evaluation data into a distributed ledger, determine whether the data quality meets a preset threshold according to the evaluation data, allocate computing nodes from an elastic computing resource pool when the preset threshold is met, recombine the task shard in parallel on each computing node, and restore the data processing state according to the state checkpoint information;

[0049] A fifth unit is configured to start a data synchronization task in the target Flink cluster after the data processing state is restored and return confirmation information.

[0050] In a third aspect, the present application provides an electronic device, comprising:

[0051] A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0052] In a fourth aspect, the present application provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the method described above.

[0053] In the present application, data is transmitted through a secure connection channel based on authentication information, realizing secure data synchronization between multiple Flink clusters, effectively preventing data leakage and unauthorized access, improving the security and reliability of data transmission, dynamically optimizing data processing operators using incremental encoding compression algorithms, and combining fine-grained task partitioning and adaptive computing load threshold mechanisms to significantly improve data processing efficiency and resource utilization, reduce system resource overhead, adapt to different scales of data synchronization requirements, introduce data quality evaluation mechanisms and distributed ledger records in the target cluster, and realize precise recovery of data processing state in combination with state checkpoint information, ensuring the integrity and consistency of the data synchronization process, and improving the system fault tolerance and reliability of data synchronization. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 A flowchart of the Flink multi-cluster secure authentication data synchronization method of the embodiment of the present application is shown in

[0055] Figure 2 A state recovery optimization flowchart of the Flink multi-cluster secure authentication data synchronization method of the embodiment of the present application is shown in DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0057] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and some embodiments may not be described again for the same or similar concepts or processes.

[0058] Figure 1 A flowchart of the Flink multi-cluster secure authentication data synchronization method of the embodiment of the present application is shown in Figure 1 As shown, the method comprises:

[0059] Receiving a data synchronization request of a source Flink cluster, obtaining an access token based on authentication information in the data synchronization request and establishing a secure connection channel;

[0060] In the source Flink cluster, a dynamic optimization is performed on a data processing operator by using an incremental encoding compression algorithm, and an execution plan is generated, the execution plan is processed in a distributed parallel manner to obtain fine-grained task shards, an adaptive computing load threshold is set for each task shard and computing resources are dynamically allocated, and state checkpoint information is written in the task shard;

[0061] The task shard is sent to the target Flink cluster through the secure connection channel;

[0062] After the target Flink cluster receives the task shard, data quality evaluation is performed on the task shard, evaluation data is written into a distributed ledger, and it is judged whether the data quality meets a preset threshold according to the evaluation data. When the preset threshold is met, a computing node is allocated from an elastic computing resource pool, the task shard is distributed to each computing node for parallel reconstruction, and the data processing state is recovered according to the state checkpoint information;

[0063] After the data processing state is recovered, a data synchronization task is started in the target Flink cluster and confirmation information is returned.

[0064] In an optional implementation,

[0065] A data synchronization request of a source Flink cluster is received, an access token is obtained based on authentication information in the data synchronization request, and a secure connection channel is established, comprising:

[0066] A data synchronization request sent by a source Flink cluster is received, and the data synchronization request includes authentication information and target Flink cluster address information;

[0067] A temporary session key is generated by performing a two-way authentication protocol between an authentication center and the source Flink cluster, and the authentication information is encrypted by using the temporary session key to obtain an access token;

[0068] A secure connection channel is established according to the access token and the target Flink cluster address information, and an end-to-end encryption transmission mechanism based on the temporary session key is configured in the secure connection channel.

[0069] The authentication center receives the data synchronization request sent by the source Flink cluster, and the data synchronization request contains two core information: authentication information and target Flink cluster address information. The authentication information includes the unique identifier of the source Flink cluster, the digital signature and the timestamp. The source Flink cluster identifier can be in the UUID format, such as "f47ac10b-58cc-4372-a567-0e02b2c3d479", which is used to uniquely identify the request initiator. The digital signature is generated by using the SHA-256 algorithm to digest the request content and encrypting it using the RSA-2048 private key, and the timestamp is in the ISO-8601 format such as "2025-07-29T08:15:30Z", which is used to prevent replay attacks. The target Flink cluster address information includes the IP address such as "192.168.1.100" and the port number such as "8081", and can also include the target cluster name such as "target-flink-cluster-001".

[0070] The authentication center and the source Flink cluster perform a two-way authentication protocol, using a two-way verification mechanism based on TLS1.3. The specific process of two-way authentication is that the authentication center sends a challenge string to the source Flink cluster, and the challenge string is a 32-byte random number, such as "8f7d6b5a4c3b2a1908172635475665748"; the source Flink cluster receives the challenge string, signs it using its private key, and returns the signature result to the authentication center; the authentication center uses the pre-stored public key of the source Flink cluster to verify whether the signature is valid. Then, the source Flink cluster sends a challenge string to the authentication center, and the authentication center signs it using its own private key and returns it, and the source Flink cluster verifies the signature using the public key of the authentication center. After two-way authentication is completed, each party generates a random number, the source Flink cluster generates a 128-bit random number R1 such as "a1b2c3d4e5f6g7h8i9j0k1l2m3n4o5p6", and the authentication center generates a 128-bit random number R2 such as "q7r8s9t0u1v2w3x4y5z6a7b8c9d0e1f2", and the two parties exchange these random numbers securely through a key exchange protocol such as the Diffie-Hellman key exchange algorithm. The temporary session key is generated by performing a bitwise XOR operation on R1 and R2, resulting in a 256-bit temporary session key.

[0071] When the authentication center encrypts the authentication information using the temporary session key, the AES-256-GCM encryption algorithm is used. In the encryption process, the source Flink cluster identifier, permission level (such as "admin" or "read-only"), validity period (such as "3600 seconds"), resource access scope (such as "jobs / read, savepoints / create"), and other information are constructed into a JSON format authentication payload, for example {"cluster_id":"f47ac10b-58cc-4372-a567-0e02b2c3d479","permissions":"admin","expires_in":3600,"scope":"jobs / read,savepoints / create","issued_at":"2025-07-29T08:15:30Z"}. The JSON is AES-256-GCM encrypted using the temporary session key to generate encrypted binary data. The encrypted binary data is Base64 encoded to obtain the access token, for example "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkZsaW5rIENsdXN0ZXIiLCJpYXQiOjE1MTYyMzkwMjJ9.4Adcg3Z60OJN0B4Rgh5oXO-74V_7vsWaGpX".

[0072] According to the access token and the target Flink cluster address information, a secure connection channel is established, and the authentication center verifies the availability of the target Flink cluster by sending an HTTPS probe request to the target cluster address to confirm its status. The authentication center establishes a TLS 1.3 secure channel with the target Flink cluster and includes the access token in the HTTP request header in the format "Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...". After receiving the connection request, the target Flink cluster verifies the validity of the access token, including checking whether the token is expired, the signature is valid, and the permissions are met. After verification, the target Flink cluster returns a connection confirmation message containing a session identifier such as "session-id: 67890abcdef12345" and a transmission parameter such as "transfer-mode: stream", indicating that the secure connection channel is successfully established.

[0073] In the secure connection channel, a temporary session key-based end-to-end encryption transmission mechanism is configured, and a block encryption transmission strategy is adopted. The source data is divided into fixed-size data blocks, each data block being 4 MB in size. Each data block is independently encrypted, and a temporary session key is used as a base key in the encryption process. A block encryption key is derived by combining the data block sequence number (such as "block-1", "block-2") and the block encryption key. The encryption is performed using the AES-256-CTR mode. Each encrypted data block includes the following fields: block sequence number (4 bytes), data size (4 bytes), initialization vector (16 bytes), encrypted data content (variable length), and authentication tag (16 bytes). During data transmission, the breakpoint resume function is implemented, and by maintaining a list of successfully transmitted data blocks, the transmission can be continued from the breakpoint when the transmission is interrupted. To ensure transmission efficiency, an adaptive flow control mechanism is implemented, and the transmission window size is dynamically adjusted according to the network conditions, with a window size range of 16 KB to 8 MB. After transmission is completed, the receiver performs integrity verification on all data blocks, calculates the SHA-256 digest value of the data blocks, and compares it with the original digest value provided by the source cluster to ensure that the data has not been tampered with.

[0074] In this embodiment, a temporary session key is generated by performing a two-way authentication protocol between the source Flink cluster and the authentication center, avoiding the security risks that may be caused by fixed keys. The authentication information is encrypted using the temporary session key to obtain an access token, achieving dynamic encryption protection of the authentication information, ensuring that the authentication information cannot be stolen or tampered with during transmission. A temporary session key-based end-to-end encryption transmission mechanism is configured in the secure connection channel, achieving full-process encryption of data transmission.

[0075] In an alternative embodiment,

[0076] In the source Flink cluster, an incremental encoding compression algorithm is used to perform dynamic optimization on data processing operators and generate an execution plan. The execution plan is processed in a distributed parallel manner to obtain fine-grained task shards. An adaptive computing load threshold is set for each task shard, and computing resources are dynamically allocated. State checkpoint information is written in the task shard, including:

[0077] In the source Flink cluster, a timing analysis is performed on the data processing operators. The set corresponding to the data processing operators and the pre-acquired data stream dependency relationship are written into an operator dependency graph. Based on the operator dependency graph, the incremental change amount of the data stream in the data stream dependency relationship is obtained and a weight coefficient is set. The execution time of the operators in the operator dependency graph is collected. The incremental change amount, the weight coefficient, and the execution time are substituted into the encoding function of the incremental encoding compression algorithm, and the encoding function is minimized to obtain an execution plan.

[0078] A set of plan nodes and a set of hierarchical relationships are constructed in the execution plan, memory consumption, data dependency and input / output overhead of each plan node in the set of plan nodes are collected, and a complexity evaluation value is calculated, and when the complexity evaluation value exceeds a preset threshold, the execution plan is processed in a distributed parallel manner to generate fine-grained task shards;

[0079] The processor usage, memory occupancy and network transmission rate of the fine-grained task shards are collected, and an adaptive computing load threshold is calculated, and the computing resources are dynamically allocated based on the proportion of the adaptive computing load threshold in the total sum of adaptive computing load thresholds of all task shards in the fine-grained task shards;

[0080] State checkpoint information is constructed by writing state identifiers, state data and time stamps in the fine-grained task shards, the difference between the state checkpoint information corresponding to two adjacent time stamps is calculated to obtain incremental checkpoint information, and the incremental checkpoint information is written into the fine-grained task shards.

[0081] In the source Flink cluster, the operator timing analysis is performed to collect the execution information of the data processing operators. The Flink API is used to obtain all the operator nodes in the JobGraph, each of which contains operator types, input and output data formats, parallelism, and other attributes. For example, the operator set can be represented as {MapOperator-1, FilterOperator-2, WindowOperator-3, JoinOperator-4}. The pre-acquired data flow dependency describes the data transmission path between operators, which can be represented as {(MapOperator-1→FilterOperator-2), (FilterOperator-2→WindowOperator-3), (MapOperator-1→JoinOperator-4), (WindowOperator-3→JoinOperator-4)}. The operator set and data flow dependency are written into the operator dependency graph, which uses an adjacency list storage structure, with nodes storing operator information and edges representing data flow directions. Based on the operator dependency graph, the data flow increment change is obtained by comparing the data flow differences in adjacent time windows. For example, the data flow from FilterOperator-2 to WindowOperator-3 is 500MB at T1 and 520MB at T2, so the increment change is 20MB. A weight coefficient is set for the data flow increment change, which is determined according to the importance and urgency of the data flow, ranging from 0.1 to 1.0. For critical business data flow, the weight coefficient is set to 0.9; for non-critical data flow, the weight coefficient is set to 0.3. The execution time of the operators in the operator dependency graph is collected through the Flink index collection system, and the time required for each operator to process a unit of data is recorded, such as 10ms for MapOperator-1 to process 1MB of data and 5ms for FilterOperator-2 to process 1MB of data. The increment change, weight coefficient, and execution time are substituted into the encoding function of the increment encoding compression algorithm, which considers the balance between data processing cost and data transmission cost. The encoding function is minimized through an iterative optimization method, which stops when the difference between the results of two adjacent iterations is less than 0.001 or the maximum number of iterations is reached, which is 100, to obtain the optimized execution plan.

[0082] A set of plan nodes and a set of hierarchy relations are constructed in the execution plan, the set of plan nodes contains all operation nodes in the execution plan, each node has a unique identifier, an operation type and configuration parameters. For example, the plan nodes can be represented as {Node-1(Map, parallelism=4), Node-2(Filter, parallelism=2), Node-3(Window, parallelism=4), Node-4(Join, parallelism=8)}. The set of hierarchy relations describes the dependency relations and execution order between nodes, represented by a directed acyclic graph, for example, {Level-1:[Node-1], Level-2:[Node-2, Node-3], Level-3:[Node-4]}. Performance indicators are collected for each plan node in the set of plan nodes, including memory consumption, data dependency and input / output overhead. Memory consumption is obtained by sampling the heap memory and off-heap memory usage of the node at runtime, such as Node-2 memory consumption of 256MB; data dependency is determined by calculating the number of input edges of the node, such as Node-4 data dependency of 2; input / output overhead is determined by measuring the data serialization and deserialization time, such as Node-3 input / output overhead of 15ms / MB. The complexity evaluation value is calculated according to the collected indicators, and the complexity evaluation value is the weighted sum of memory consumption, data dependency and input / output overhead. When the complexity evaluation value exceeds a preset threshold (such as 1000), the execution plan is processed in a distributed parallel manner to generate fine-grained task shards. In the task shard process, large operators are split into multiple sub-tasks, each sub-task processes part of the data, such as Node-4(Join) is split into 4 sub-tasks {SubTask-4-1, SubTask-4-2, SubTask-4-3, SubTask-4-4}, each sub-task is responsible for processing a specific data partition.

[0083] The resource usage of fine-grained task shards is collected, including processor usage, memory occupancy, and network transmission rate. Processor usage represents the average load of CPU cores, such as the processor usage of SubTask-4-1 being 65%; memory occupancy represents the proportion of allocated memory that is in use, such as the memory occupancy of SubTask-4-2 being 78%; network transmission rate represents the amount of data transmitted per unit time, such as the network transmission rate of SubTask-4-3 being 50 MB / s. Based on the collected indicators, an adaptive computing load threshold is calculated, which takes into account the combined influence of processor usage, memory occupancy, and network transmission rate. For example, the adaptive computing load threshold of SubTask-4-1 is 0.72, that of SubTask-4-2 is 0.81, that of SubTask-4-3 is 0.65, that of SubTask-4-4 is 0.58, and the total is 2.76. According to the proportion of each task shard's adaptive computing load threshold in the total, computing resources are dynamically allocated, such as SubTask-4-1 obtaining 26.1% of the resources (0.72 / 2.76) and being allocated 5 CPU cores (total number of cores 20 26. 1%), memory is 2.6 GB (total memory 10 GB 26.1%)).

[0084] State checkpoint information is constructed in fine-grained task shards, writing three types of key data: state identifier, state data, and timestamp. The state identifier is a unique identifier for the checkpoint, formatted as "TaskID-CheckpointID", such as "SubTask-4-1-CP001"; the state data contains a snapshot of the task execution state, such as processing progress, intermediate results, and cache data; the timestamp records the creation time of the checkpoint, using a millisecond-level UNIX timestamp, such as 1627547631000. The difference between the state checkpoint information corresponding to two adjacent timestamps is calculated to obtain incremental checkpoint information. For example, the state data size of SubTask-4-1 at timestamp 1627547631000 is 150 MB, and at timestamp 1627547641000 it is 155 MB. By comparing the difference between the two state data, a binary difference algorithm is used to generate incremental checkpoint information with a size of 8 MB. The incremental checkpoint information is written to the fine-grained task shard, reducing the data transmission volume and storage overhead during state recovery. The incremental checkpoint information contains the set of key-value pairs that have changed, the operation type (add, modify, or delete), and metadata, so that the complete state can be accurately reconstructed during fault recovery.

[0085] In this embodiment, by performing time series analysis on the data processing operator and constructing an operator dependency graph, combining the incremental change amount and weight coefficient of the data stream, and using an incremental encoding compression algorithm to optimize the execution plan, the memory consumption, data dependency degree and input / output overhead of the collection plan node are calculated to evaluate the complexity, the adaptive division of task shards is realized, the adaptive computing load threshold is dynamically calculated based on the processor usage rate, memory occupancy rate and network transmission rate, and the computing resources are proportionally allocated according to the actual load condition, the resource utilization efficiency is improved, the resource waste is avoided, the state checkpoint information is constructed in the task shard and the incremental checkpoint information is calculated, the fine-grained state management and fault recovery mechanism are realized, the state storage overhead is significantly reduced while the reliability and consistency of data processing are ensured.

[0086] In an optional implementation,

[0087] The memory consumption, data dependency degree and input / output overhead of each plan node in the set of plan nodes are collected, and a complexity evaluation value is calculated. When the complexity evaluation value exceeds a preset threshold, the execution plan is processed in a distributed parallel manner to generate fine-grained task shards.

[0088] The time series feature sequence corresponding to the execution plan is extracted, the historical data in a specified time window in the time series feature sequence is calculated according to the time series correlation to obtain a hidden state output, the attention score of the hidden state output is calculated to obtain a feature weight coefficient, the time series feature sequence and the feature weight coefficient are multiplied and summed to obtain a task feature vector, and the task feature vector is recursively iterated to obtain a predicted complexity.

[0089] A topology feature matrix between plan nodes is calculated, a propagation influence value of data flow between nodes is calculated based on the topology feature matrix, a data dependency strength is obtained by graph embedding operation of node degree distribution value and edge weight distribution value, and a resource competition degree matrix between nodes is calculated to obtain a resource competition degree.

[0090] A state transition probability matrix is constructed based on the execution state, reward value and action sequence in the pre-acquired historical execution data, an optimal balance coefficient is calculated based on the state transition probability matrix, the predicted complexity, data dependency strength and resource competition degree are multiplied by the balance coefficient respectively and jointly optimized to obtain a complexity evaluation value.

[0091] When the complexity evaluation value is greater than a preset threshold, a task similarity matrix is calculated based on the data dependency strength, and a spectral clustering operation is performed to obtain a task cooperative execution group. Chromosome encoding sequences are constructed for the tasks in the task cooperative execution group according to the computational complexity and resource demand, and an optimal parallelism configuration scheme is obtained by iterative optimization of a genetic algorithm. The task cooperative execution group is distributed to different computing nodes for parallel processing according to the parallelism configuration scheme to obtain fine-grained task shards.

[0092] The time sequence feature sequence corresponding to the extraction execution plan is formed by monitoring the running indicators of each operator in the Flink execution plan. The time sequence feature sequence includes CPU usage, memory occupation, data processing volume, and operator throughput, and other key indicators. The key indicators are sampled and stored at fixed time intervals (such as 5 seconds). For example, for a MapOperator operator, the CPU usage sequence in a period of time may be [45%, 48%, 52%, 47%, 50%]. The historical data of a specified time window (such as the last 30 minutes) is intercepted from the time sequence feature sequence for analysis, and a long short-term memory network is used to calculate the hidden state output. The network includes three control units of an input gate, a forgetting gate, and an output gate. The input gate controls the degree of new information entering, the forgetting gate controls the degree of historical information retention, and the output gate controls the degree of current hidden state output. The network input is the time sequence feature sequence, which is processed through a four-layer network structure (16 nodes in the input layer, 32 nodes in each of the two hidden layers, and 16 nodes in the output layer) to obtain the hidden state output with a dimension of 16. When calculating the attention score of the hidden state output, a self-attention mechanism is used to evaluate the importance of data at different time points. The attention score is calculated by the similarity between the hidden state vectors. The higher the similarity, the stronger the relevance. For example, the attention score distribution at a certain time is [0.02, 0.03, 0.15, 0.25, 0.40, 0.15], indicating that the weight of the most recent data point is 0.40. After normalizing the attention score, a feature weight coefficient is obtained. The feature values at each time point in the time sequence feature sequence are multiplied by the corresponding feature weight coefficient and summed to obtain a 16-dimensional task feature vector. For example, the task feature vector of MapOperator may be [0.48, 0.53, 0.62, 0.45, 0.58, 0.49, 0.51, 0.55, 0.47, 0.52, 0.56, 0.50, 0.54, 0.49, 0.53, 0.51]. The task feature vector is input into a recurrent neural network, and a predicted complexity value is calculated through 5 rounds of iteration. The value ranges between 0 and 1, such as the predicted complexity of MapOperator being 0.65, indicating a moderate computing complexity.

[0093] The topological feature matrix between nodes in the execution plan is calculated, representing the node relationships in the execution plan as an adjacency matrix. For an execution plan containing n nodes, an n×n topological feature matrix is ​​constructed, where the matrix element values ​​represent the association strength between nodes. For example, the topological feature matrix of a four-node execution plan might be [[0, 1, 0, 0], [0, 0, 1, 1], [0, 0, 0, 1], [0, 0, 0, 0]], where 1 indicates the presence of data flow and 0 indicates the absence of data flow. Based on the topological feature matrix, the propagation impact value of data flow between nodes is calculated, and an information propagation model is used to analyze the diffusion effect of data flow between nodes. The influence radius and influence strength of each node are calculated; for example, the propagation impact value of node 2 is 0.78, indicating that it has a strong influence on downstream nodes. When calculating the node degree distribution value, the in-degree and out-degree of each node are counted; for example, node 1 has an in-degree of 0 and an out-degree of 1, and node 2 has an in-degree of 1 and an out-degree of 2. The edge weight distribution values ​​are determined based on the data traffic volume. For example, the edge weight from node 1 to node 2 is 500 MB / s, and the edge weight from node 2 to node 3 is 300 MB / s. The node degree distribution values ​​and edge weight distribution values ​​are input into a graph embedding algorithm. A 16-dimensional embedding vector is generated through random walks and jumps. A data dependency strength matrix is ​​obtained by calculating cosine similarity; the element values ​​in this matrix range from 0 to 1, with larger values ​​indicating stronger dependencies. When calculating the resource contention matrix between nodes, the competition for resources such as CPU, memory, and network bandwidth is analyzed. For example, when nodes 2 and 3 are running simultaneously, the CPU contention is 0.65, the memory contention is 0.42, and the network contention is 0.38, resulting in a total resource contention of 0.55.

[0094] A state transition probability matrix is ​​constructed based on historical execution data, collecting information such as execution status, reward value, and action sequence. Execution status includes resource usage and task execution progress; reward value reflects execution effect, such as increased throughput or reduced latency; and action sequence records resource allocation adjustment operations. The states in the historical data are divided into multiple discrete intervals, such as resource utilization divided into four intervals: [0-0.3], [0.3-0.6], [0.6-0.9], and [0.9-1.0]. A probability matrix is ​​constructed based on the state transition data, where matrix element P(i,j) represents the probability of transitioning from state i to state j. For example, the probability of transitioning from a resource utilization of 0.5 to 0.7 is 0.35. The optimal balance coefficient is calculated based on the state transition probability matrix, and the optimal strategy is solved using a value iteration algorithm, yielding balance coefficients of 0.4, 0.35, and 0.25 for prediction complexity, data dependency strength, and resource competition, respectively. Prediction complexity, data dependency strength, and resource competition are multiplied by their respective balance coefficients, and a weighted summation is used for joint optimization to obtain a comprehensive complexity evaluation value. For example, if the prediction complexity of a node is 0.65, the data dependency strength is 0.78, and the resource contention degree is 0.55, the calculated complexity evaluation value is 0.4×0.65+0.35×0.78+0.25×0.55=0.67.

[0095] When the complexity assessment value exceeds a preset threshold (e.g., 0.6), task sharding optimization is required. A task similarity matrix is ​​calculated based on data dependency strength, and cosine similarity is used to measure the similarity between tasks. For example, a similarity of 0.85 between task A and task B indicates highly similar data processing characteristics. Spectral clustering is performed on the similarity matrix, dividing the tasks into multiple collaborative execution groups through eigenvalue decomposition and K-means clustering. For example, 10 tasks might be divided into 3 collaborative execution groups: {Task1, Task3, Task7}, {Task2, Task5, Task8}, and {Task4, Task6, Task9, Task10}. Chromosome encoding sequences are constructed for the tasks in each collaborative execution group according to computational complexity and resource requirements. Each chromosome represents a parallelism configuration scheme. The chromosome length equals the number of tasks, and the gene value represents the parallelism. For example, chromosome [4, 2, 8, 4, 6, 3, 5, 2, 4, 3] represents the parallelism configuration of 10 tasks. Parallelism configuration is optimized using a genetic algorithm, with a population size of 100, a crossover probability of 0.8, a mutation probability of 0.1, and 50 iterations. The fitness function comprehensively considers execution time, resource utilization, and data transfer overhead. After 500 iterations, the optimal chromosome [6, 3, 9, 5, 7, 4, 8, 3, 5, 4] is obtained, representing the optimal parallelism configuration. Task co-execution groups are then distributed to different computing nodes according to the optimal parallelism configuration. For example, the three tasks in task co-execution group 1, configured with parallelism [6, 9, 8], are distributed across four computing nodes. Allocation is based on task characteristics and the principle of data locality, ensuring that tasks with strong data dependencies are preferentially assigned to the same node to reduce data transfer overhead. Fine-grained task fragments are generated through parallel processing, with each fragment responsible for processing a portion of the data, achieving efficient data synchronization across multiple Flink clusters.

[0096] In this embodiment, by using temporal feature sequence analysis and attention mechanism, the temporal correlation features of task execution can be captured, achieving accurate extraction of task features and making the predicted complexity more accurate and forward-looking. Through topological feature matrix analysis and graph embedding operation, the data dependency relationship and resource competition between nodes are comprehensively evaluated, enabling a deep understanding of the correlation between tasks and resource demand characteristics, providing a reliable basis for subsequent optimization. By introducing reinforcement learning ideas and constructing a state transition probability matrix based on historical execution data, the adaptive calculation of the optimal balance coefficient is realized, improving the accuracy of complexity assessment.

[0097] In one alternative implementation,

[0098] After the target Flink cluster receives the task shard, it performs a data quality assessment on the task shard, writes the assessment data into the distributed ledger, determines whether the data quality meets a preset threshold based on the assessment data, and if it does, allocates computing nodes from the elastic computing resource pool. The task shard is then distributed to each computing node for parallel reassembly. The data processing state is restored based on the state checkpoint information, including:

[0099] After the target Flink cluster receives the task shards, the data missing rate is obtained by calculating the ratio of the number of missing values ​​in each field of the task shard to the total number of records. The data consistency score is obtained by calculating the entropy value of the probability of the occurrence of data values. The outlier degree is obtained by calculating the standard deviation of numerical data relative to the mean. Evaluation data is constructed based on the data missing rate, data consistency score and outlier degree.

[0100] The evaluation data is written into the hierarchical storage structure of the distributed ledger. Based on the evaluation data, it is determined whether the data quality meets the preset threshold. If it does, the task complexity, memory requirements, and network bandwidth requirements are determined and the resource requirements are calculated. A multi-objective optimization function is constructed by combining resource utilization, load balancing, and communication overhead, and the resource allocation scheme is obtained by solving it. Computing nodes are allocated from the pre-set elastic computing resource pool according to the resource allocation scheme.

[0101] The task is divided and distributed to each computing node for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The state item identifier sequence is traversed to calculate the dependency relationship between state items to obtain the state dependency matrix. The longest common subsequence algorithm is used to analyze and calculate the state dependency matrix to obtain the state transition cost corresponding to the state item. A state recovery path is constructed based on the state transition cost. The total cost of the state recovery path is used as the incremental recovery cost. The data processing state is restored according to the incremental recovery cost.

[0102] After receiving task shards, the target Flink cluster performs a quality assessment on the data within each shard. The missing value rate is calculated by dividing the number of missing values ​​in each field of the task shard by the total number of records. This is done by iterating through each field in the data table, counting the number of null or empty values, and dividing by the total number of records. For example, a task shard might contain a user behavior data table where the `user_id` field has 5000 records with 120 missing values, resulting in a missing value rate of 0.024; the `event_time` field has 5000 records with 80 missing values, resulting in a missing value rate of 0.016; and the `device_type` field has 5000 records with 350 missing values, resulting in a missing value rate of 0.07. When calculating the entropy value (the probability of data values ​​occurring) to obtain the data consistency score, the frequency of different values ​​for each category field is counted, and the information entropy is calculated. For example, for the `device_type` field, its value distribution is {"mobile": 2800, "desktop": 1500, "tablet": 350}, the calculated information entropy is 1.23, and the normalized data consistency score is 0.82. When calculating the outlier level by the standard deviation of numerical data relative to the mean, a Z-score is calculated for each numerical field, and the proportion of data points exceeding 3 standard deviations is counted. For example, for the user dwell time field, the average is 120 seconds, the standard deviation is 30 seconds, and there are 15 data points exceeding 3 standard deviations, accounting for 0.003, with an outlier level of 0.003. Based on the above three indicators, evaluation data is constructed, and a weighted average method is used with weights of 0.4, 0.35, and 0.25, respectively, to calculate the comprehensive evaluation score. For example, the overall evaluation score of a certain task segment is 0.4×(1-0.024)+0.35×0.82+0.25×(1-0.003)=0.866.

[0103] Evaluation data is written into a hierarchical storage structure of the distributed ledger, which includes a metadata layer, a metric layer, and a content layer. The metadata layer stores basic information about task shards, such as shard ID, creation time, and data source. The metric layer stores evaluation data, including data missing rate, consistency score, and outlier rate. The content layer stores the raw data and processing results. The distributed ledger uses a hash chain-based structure, where each block contains the hash value of the previous block, a timestamp, data content, and a digital signature. For example, the block content might be {"blockId": "b6f23", "prevHash": "a1c5e", "timestamp": 1627892465, "data": {"shardId": "shard-42", "quality": 0.866}, "signature": "3a7bd"}. The evaluation data is used to determine if the data quality meets a preset threshold of 0.8; a value higher than this indicates good data quality. If the data quality meets the preset threshold, determine the task complexity, memory requirements, and network bandwidth requirements, and calculate the resource requirements. Task complexity is estimated using operator computation and data size; for example, the complexity of MapOperator is O(n), and the complexity of window aggregation is O(n log n). Memory requirements are predicted using historical execution data and a linear regression model of the data size; for example, processing 5000 records requires 512MB of memory. Network bandwidth requirements are calculated using data transmission volume and processing time; for example, if the task's data chunk size is 200MB and the expected transmission time is 10 seconds, the bandwidth requirement is 20MB / s. A multi-objective optimization function is constructed by combining resource utilization, load balancing, and communication overhead. Resource utilization represents the efficiency of resource allocation, with the objective of maximizing it; load balancing represents the degree of load balance among nodes, with the objective of minimizing the standard deviation; and communication overhead represents the amount of data transmitted between nodes, with the objective of minimizing it. A genetic algorithm is used to solve the multi-objective optimization problem, with a population size of 50, 100 iterations, a crossover probability of 0.8, and a mutation probability of 0.1. After optimization, a resource allocation scheme is obtained, such as allocating 4 CPU cores, 2GB of memory, and 50MB / s bandwidth to each task shard. Computing nodes are allocated from a pre-configured elastic computing resource pool according to the resource allocation scheme. The resource pool contains nodes of different specifications, such as small nodes (2 cores, 4GB), medium nodes (4 cores, 8GB), and large nodes (8 cores, 16GB). The appropriate node type is selected based on task requirements; for example, large nodes are selected for compute-intensive tasks, and medium nodes are selected for I / O-intensive tasks.

[0104] The task is sharded and distributed to various computing nodes for parallel reassembly. The sharded data is distributed to each node according to a data partitioning strategy. For example, hash partitioning is used to distribute data with user IDs as hash keys to four nodes, ensuring that data from the same user is assigned to the same node. The state item identifier sequence is extracted from the state checkpoint information. The state item identifier format is "operatorID-stateType-timestamp", such as "op3-window-1627892460". The state item identifier sequence is traversed to calculate the dependencies between state items, resulting in a state dependency matrix. The matrix elements represent the dependency strength between state items. For example, the dependency strength between state item A and state item B is 0.8, indicating that recovering state item B strongly depends on state item A. The longest common subsequence algorithm is used to analyze the state dependency matrix and find the optimal recovery order between state items. This algorithm compares two state sequences, finds their longest common part, and calculates the similarity and difference. For example, the longest common subsequence of the sequences ["op1-map-1627892450", "op2-filter-1627892455", "op3-window-1627892460"] and ["op1-map-1627892450", "op3-window-1627892460", "op4-join-1627892465"] is ["op1-map-1627892450", "op3-window-1627892460"], with a length of 2. Calculate the state transition cost corresponding to each state item. The cost includes both computational cost and I / O cost. The computational cost is related to the processing complexity of the state item, while the I / O cost is related to the amount of state data read and written. For example, the computational cost of transitioning from state item A to state item B is 150ms, the I / O cost is 200ms, and the total cost is 350ms. A state recovery path is constructed based on the state transition cost, and a dynamic programming algorithm is used to find the path with the minimum total cost. For example, the optimal path from the initial state to the target state is [State0, StateA, StateC, StateF], with corresponding transition costs of [0, 350ms, 280ms, 420ms], respectively. The total cost of the state recovery path is used as the incremental recovery cost, which is 1050ms in this embodiment. The data processing state is recovered based on the incremental recovery cost, achieving an efficient transition from the checkpoint to the target state.

[0105] During the recovery process, the basic checkpoint data is loaded first, and then incremental changes are applied sequentially according to the recovery path. For example, the basic checkpoint data with timestamp 1627892450 (100MB in size) is loaded first, and then three incremental changes (15MB, 12MB, and 18MB respectively) are applied in sequence to restore the state to timestamp 1627892465. Throughout the recovery process, the data transfer volume was reduced from 145MB to 45MB, improving efficiency by 69%. After recovery, task fragments can continue execution from the breakpoint, achieving seamless data processing.

[0106] In this embodiment, a comprehensive data quality assessment system is established by comprehensively evaluating data missing rate, consistency score, and outlier degree. This system can accurately identify data anomalies, ensure the quality of data in subsequent processing, and improve the reliability of the system. By comprehensively considering resource utilization, load balancing, and communication overhead through a multi-objective optimization function, the optimal allocation of computing resources is achieved, which ensures processing efficiency and avoids resource waste. By analyzing the state item identifier sequence to construct a state dependency matrix and applying the longest common subsequence algorithm to calculate the state transition cost, the overhead of state recovery is significantly reduced.

[0107] In one alternative implementation,

[0108] The task is fragmented and distributed to various computing nodes for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The longest common subsequence algorithm is used to calculate the state transition cost corresponding to the state item. Based on the state transition cost, a state recovery path is constructed, including:

[0109] The task is divided and allocated to computing nodes for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The state item identifier sequence is traversed to calculate the dependency relationship between state items. Based on the dependency relationship, the forward dependency chain and the backward dependency chain are constructed. The analysis results of the forward dependency chain and the backward dependency chain are stored in the state dependency matrix.

[0110] Starting simultaneously from the beginning and end of the state item identifier sequence, the longest common subsequence algorithm is used to calculate the state transition cost based on the number of forward dependencies, the number of backward dependencies, and the distance between state items in the state dependency matrix.

[0111] The dependency strength index is calculated based on the number of forward dependencies and backward dependencies of state items. The search direction is dynamically adjusted according to the comparison result of the dependency strength index and the preset dependency strength threshold. Forward search results and backward search results are generated. The ratio of the number of forward dependencies of each state item to its total dependencies is calculated to obtain the dynamic merging weight. The forward search results and the backward search results are weighted and merged according to the dynamic merging weight to obtain the state transition sequence.

[0112] A state recovery path is constructed based on the state transition cost and the state transition sequence. The state recovery path is obtained by minimizing the sum of the state transition costs between adjacent state items in the state item identifier sequence. The data processing state is then restored according to the state recovery path.

[0113] Task fragments are allocated to compute nodes and reassembled in parallel. Based on the resource allocation scheme, task fragments are mapped to specific compute nodes. Each task fragment contains processing logic, data partitions, and state information. For example, task fragment TS001 is allocated to compute node N1 to process partition 1 of user behavior data. State item identifier sequences are extracted from the state checkpoint information, which is stored in a distributed file system and contains a complete state snapshot. State item identifiers use the format "operator ID-state type-state name-timestamp," such as "map-1-window-counts-1627890000," which represents the window-counts state of operator ID map-1 at time 1627890000. The extracted state item identifier sequence might be ["map-1-window-counts-1627890000", "filter-2-keyed-threshold-1627890010", "window-3-aggregate-sum-1627890020", "join-4-state-mapping-1627890030", "sink-5-buffer-queue-1627890040"]. The state item identifier sequence is traversed to calculate the dependencies between state items. These dependencies are determined by analyzing the data flow graph and state access patterns. Based on these dependencies, forward and backward dependency chains are constructed. The forward dependency chain indicates which previous state items the current state item depends on, and the backward dependency chain indicates which subsequent state items the current state item depends on. For example, the forward dependency of state item "window-3-aggregate-sum-1627890020" is ["map-1-window-counts-1627890000", "filter-2-keyed-threshold-1627890010"], and the backward dependency is ["join-4-state-mapping-1627890030"]. The analysis results of the forward and backward dependency chains are stored in a state dependency matrix with dimensions n×n, where n is the number of state items. The matrix element (i, j) represents the degree of dependency of state item i on state item j, ranging from 0 to 1, where 0 indicates no dependency and 1 indicates a strong dependency. For example, the value of element (1, 3) in the dependency matrix is ​​0.8, indicating that state item 1 has a strong dependency on state item 3.

[0114] Starting simultaneously from the beginning and end of the state item identifier sequence, the similarity between state items is calculated using the longest common subsequence algorithm. This algorithm uses dynamic programming to find the longest common part between two sequences, with a time complexity of O(m×n), where m and n are the lengths of the two sequences, respectively. For example, a two-dimensional table is constructed to record the solutions to the subproblem. Table element (i, j) represents the length of the longest common subsequence of the first i elements of sequence 1 and the first j elements of sequence 2. For instance, the longest common subsequence of sequences A["map-1", "filter-2", "window-3"] and B["map-1", "join-4", "window-3"] is ["map-1", "window-3"], with a length of 2. The state transition cost is calculated based on the number of forward dependencies, the number of backward dependencies, and the distance between state items in the state dependency matrix. The number of forward dependencies indicates the number of state items that need to be recovered to restore the current state item. For example, the number of forward dependencies for the state item "window-3-aggregate-sum-1627890020" is 2. The number of backward dependencies indicates the number of state items that can be recovered after restoring the current state item. For example, the number of backward dependencies for this state item is 1. The distance between state items indicates the interval between two state items in the time series. For example, the distance from "filter-2-keyed-threshold-1627890010" to "window-3-aggregate-sum-1627890020" is 10 seconds. The formula for calculating the state transition cost is: the number of forward dependencies multiplied by the forward dependency weight, plus the distance between state items multiplied by the distance weight, minus the number of backward dependencies multiplied by the backward dependency weight. The weight parameters are determined through optimization using historical execution data. For example, the forward dependency weight is 0.4, the distance weight is 0.3, and the backward dependency weight is 0.5.

[0115] The dependency strength index is calculated based on the number of forward and backward dependencies of a state item. The dependency strength index is the sum of the number of forward and backward dependencies divided by the total number of state items. For example, the dependency strength index of the state item "window-3-aggregate-sum-1627890020" is (2+1) / 5=0.6. The search direction is dynamically adjusted based on the dependency strength index compared to a preset dependency strength threshold (e.g., 0.5). When the dependency strength is higher than the threshold, dependencies are considered first; when the dependency strength is lower than the threshold, the distance between state items is considered first. A bidirectional search strategy is adopted, searching backward from the starting state item to obtain forward search results, and searching forward from the ending state item to obtain backward search results. During the forward search, the next state item with the lowest transformation cost is selected at each step to generate a forward path; the backward search is similar, generating a backward path. For example, the forward search path is ["map-1-window-counts-1627890000", "filter-2-keyed-threshold-1627890010", "window-3-aggregate-sum-1627890020"], and the backward search path is ["sink-5-buffer-queue-1627890040", "join-4-state-mapping-1627890030", "window-3-aggregate-sum-1627890020"]. The dynamic merge weight is obtained by calculating the ratio of the number of forward dependencies to the total number of dependencies for each state item. For example, the dynamic merge weight for the state item "window-3-aggregate-sum-1627890020" is 2 / (2+1) = 0.67. The forward and backward search results are weighted and merged according to dynamic merging weights to obtain a state transition sequence. During the merging process, overlapping parts of the search results are processed to ensure that the final sequence does not contain duplicate state items. For example, if the forward search weight is 0.67 and the backward search weight is 0.33, the merged state transition sequence is ["map-1-window-counts-1627890000", "filter-2-keyed-threshold-1627890010", "window-3-aggregate-sum-1627890020", "join-4-state-mapping-1627890030", "sink-5-buffer-queue-1627890040"].

[0116] A state recovery path is constructed based on state transition costs and state transition sequences. The path is obtained by minimizing the sum of state transition costs between adjacent state items in the state item identifier sequence. A dynamic programming algorithm is used to solve the shortest path problem, defining the subproblem as finding the minimum-cost path from the initial state to the current state. The state transition cost is represented as a cost matrix, where each element C(i,j) represents the cost of transitioning from state i to state j. For example, the transition cost from "map-1-window-counts-1627890000" to "filter-2-keyed-threshold-1627890010" is 25, and the transition cost from "filter-2-keyed-threshold-1627890010" to "window-3-aggregate-sum-1627890020" is 40. The path with the minimum total cost is found by filling in a dynamic programming table. The optimized state recovery path might be ["map-1-window-counts-1627890000", "window-3-aggregate-sum-1627890020", "join-4-state-mapping-1627890030", "filter-2-keyed-threshold-1627890010", "sink-5-buffer-queue-1627890040"], with a total cost of 105, a 25% reduction compared to the original total cost of 140. The data processing state is restored according to the state recovery path, and state item data is loaded in the optimized order. During the recovery process, the basic data for each state item is loaded first, and then incremental changes are applied. For example, load the base data (50MB) and incremental changes (5MB) of "map-1-window-counts-1627890000", then load the data of "window-3-aggregate-sum-1627890020". For state items with strong dependencies, a parallel loading strategy is used to improve recovery efficiency. For example, the dependency between "join-4-state-mapping-1627890030" and "filter-2-keyed-threshold-1627890010" is weak, so they can be loaded in parallel.

[0117] In this embodiment, by constructing a bidirectional analysis mechanism of forward and backward dependency chains, the system comprehensively captures the dependencies between state items. This not only improves the accuracy of dependency identification but also provides complete dependency information support for subsequent state recovery optimization. The longest common subsequence algorithm is used in combination with the number of dependencies and distance of state items to calculate the state transition cost, achieving accurate evaluation of state transition overhead. By dynamically calculating the dependency strength index and adaptively adjusting the search direction, efficient state recovery path search is achieved. By minimizing the sum of transition costs between adjacent state items to optimize the state recovery path, a globally optimal state recovery strategy is achieved, ensuring both the efficiency of the recovery process and maintaining the continuity of state transitions.

[0118] Figure 2 This is a flowchart illustrating the state recovery optimization process of the Flink multi-cluster security authentication data synchronization method according to an embodiment of the present invention.

[0119] In one alternative implementation,

[0120] After resuming data processing, a data synchronization task is initiated in the target Flink cluster and confirmation information is returned, including:

[0121] Obtain the data processing status recovery completion status from the target Flink cluster, verify whether the data processing status recovery result meets the preset running conditions, and if the preset running conditions are met, initialize the running environment of the data synchronization task in the target Flink cluster, load the configuration parameters required for the data synchronization task, and start the data synchronization task.

[0122] Monitor the startup process of the data synchronization task, obtain the running status information of the data synchronization task, and return the running status information as confirmation information.

[0123] To obtain the recovery completion status of the data processing state from the target Flink cluster, you need to use Flink's JobManager API to retrieve the recovery progress and results. For example, construct a REST API request to obtain the recovery status, with the request path " / jobs / {jobId} / checkpoints / details / {checkpointId}", where jobId is the task identifier (e.g., "job_7e9581a708de43e9a74a10a5579c4a69") and checkpointId is the checkpoint identifier (e.g., "checkpoint_39a685b0a52a4a2c9ec37125675f6c98"). The request header contains authentication information and the request type. The authentication information uses a JWT token format, such as "Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...". The request returns a JSON response containing the recovery status, including the recovery progress percentage, the number of recovered states, the total number of states, and the recovery time. For example, the return value might be {"id": 39685, "status": "completed", "is_savepoint": false, "trigger_timestamp": 1627892465000, "completion_timestamp": 1627892470000, "state_size": 268435456, "duration": 5000, "aligned": true, "num_subtasks": 16, "num_acknowledged_subtasks": 16, "tasks": {...}}. The recovery completion status is determined by parsing the response content, checking if the "status" field value is "completed" and if "num_subtasks" is equal to "num_acknowledged_subtasks".

[0124] Verify that the data processing state recovery results meet the preset operating conditions, ensuring the integrity and consistency of the state recovery through multi-dimensional checks. The preset operating conditions include three aspects: state integrity, consistency, and performance. State integrity requires that all necessary state items have been successfully recovered without loss or damage. Integrity is verified by checking the state item count and checksum. For example, if the task requires recovering the state of 5 operators, check if the number of recovered state items is 5 and if the MD5 checksum of each state item matches the expected value. Consistency requires that related state items maintain data consistency, especially those with dependencies. Consistency is checked by verifying the data reference relationships between state items. For example, if the state of operator A references data from operator B, verify if the reference is valid. Performance requires that the post-recovery processing capacity meets the expected level. Performance is verified by testing processing latency and throughput. For example, the post-recovery processing latency should be less than 200 milliseconds, and the throughput should reach 10,000 records / second. All verification results are recorded in the state verification report, in the format {"completeness": true, "consistency": true, "performance": true, "details": {...}}. If the preset running conditions are met, the running environment for the data synchronization task will be initialized in the target Flink cluster.

[0125] Initializing the runtime environment includes two parts: preparing execution resources and configuring the runtime environment. When preparing execution resources, TaskManager nodes are allocated according to task requirements, and the number of slots and memory configuration are set. For example, four TaskManager nodes are allocated for a data synchronization task, each node provides four slots, and each slot is allocated 4GB of memory. When configuring the runtime environment, parameters such as the state backend type, checkpoint interval, and fault recovery strategy are set. For example, RocksDB is selected as the state backend, the checkpoint interval is set to 60 seconds, the fault recovery strategy is set to "fixed-delay", the number of retries is 3, and the retry interval is 10 seconds. The configuration parameters required for the data synchronization task are loaded. These configuration parameters are stored in a distributed configuration center and read through parameter key-value pairs. The configuration parameters include four categories: data source configuration, data target configuration, transformation rule configuration, and performance tuning configuration. The data source configuration specifies the data source and connection method, such as {"source.type": "kafka", "source.bootstrap.servers": "broker1:9092, broker2:9092", "source.topic": "data-sync-topic", "source.group.id": "data-sync-group"}; the data target configuration specifies the data writing target, such as {"sink.type": "jdbc", "sink.url": "jdbc:postgresql: / / db-server:5432 / data_warehouse", "sink.table": "user_b {"ehavior", "sink.batch.size": 1000}; Transformation rule configuration defines the data transformation logic, such as {"transform.fields": ["user_id", "event_time", "event_type", "product_id"], "transform.filter": "event_type!='heartbeat'"}; Performance tuning configuration sets the execution parameters, such as {"performance.parallelism": 16, "performance.buffer.timeout": 100, "performance.checkpoint.interval": 60000}.

[0126] Initiate the data synchronization task by submitting it for execution via Flink's JobManager API. Construct a task submission request with the path " / jars / {jarId} / run", where jarId is the task's JAR package identifier, such as "jar_b1a11ece35aa40bd95dfd8a4a33d3028". The request body contains task execution parameters, such as {"entryClass": "com.example.DataSyncJob", "parallelism": 16, "programArgs": "--configconfig.properties"}. The request header contains authentication information and content type, with the content type set to "application / json". After sending the request, JobManager initiates the task execution process: allocating resources, deploying the task, establishing the data flow topology, and starting data processing. The task startup process consists of five stages: CREATED, SCHEDULED, DEPLOYING, RUNNING, and FINISHED. Each stage records the corresponding timestamp and status information. For example, the timestamp of the CREATED stage is 1627892480000, and the status is "task created, waiting for scheduling".

[0127] The monitoring process of the data synchronization task startup is achieved by periodically polling the task status API to obtain the latest status. The polling interval uses an exponential backoff strategy, with an initial interval of 100 milliseconds, a maximum interval of 5 seconds, and a maximum monitoring time of 5 minutes. Each poll sends a GET request with the request path " / jobs / {jobId} / status", returning a JSON response containing the task status. For example, the return value is {"id": "job_7e9581a708de43e9a74a10a5579c4a69", "status": "RUNNING", "start-time": 1627892485000, "end-time": -1, "duration": 15000, "now": 1627892500000}. The system retrieves the running status information of the data synchronization task, including four parts: task status, subtask status distribution, resource usage, and performance metrics. The task status indicates the overall execution status, such as "RUNNING"; the subtask status distribution displays the status statistics of each subtask, such as {"RUNNING": 16, "SCHEDULED": 0, "FINISHED": 0, "CANCELED": 0, "FAILED": 0}; resource usage includes CPU utilization, memory usage, and network I / O, such as {"cpu_usage": 0.65, "memory_usage": 0.72, "network_input_rate": 15000, "network_output_rate": 12000}; performance metrics include processing latency and throughput, such as {"latency_avg": 45, "latency_p95": 120, "throughput": 25000}. The running status information is returned to the source Flink cluster as confirmation information in JSON format, containing the status code, status description, and detailed status data. For example, {"code": 200, "message": "Data sync task started successfully", "data": {"job_id": "job_7e9581a708de43e9a74a10a5579c4a69", "status": "RUNNING", "start_time": "2023-08-02T10:14:45Z", "metrics": {"throughput": 25000, "latency": 45}}}.

[0128] In this embodiment, the verification mechanism for the data processing status recovery completion status ensures that the data synchronization task starts on the correct status. The step-by-step task startup process ensures the standardization and integrity of the task startup process, reduces the risk of task startup failure, and improves the stability of the system. By monitoring the startup process of the data synchronization task in real time and obtaining the running status information, the entire task startup process is observable, which can promptly detect and respond to abnormal situations during the startup process, thereby improving the controllability of the system.

[0129] A second aspect of this invention provides a Flink multi-cluster secure authentication data synchronization system, comprising:

[0130] The first unit is used to receive data synchronization requests from the source Flink cluster, obtain an access token based on the authentication information in the data synchronization request, and establish a secure connection channel.

[0131] The second unit is used to dynamically optimize the data processing operator and generate an execution plan in the source Flink cluster using an incremental encoding compression algorithm, process the execution plan in a distributed parallel manner to obtain fine-grained task partitions, set an adaptive computing load threshold for each task partition and dynamically allocate computing resources, and write state checkpoint information into the task partitions.

[0132] The third unit is used to send the task fragments to the target Flink cluster through the secure connection channel;

[0133] The fourth unit is used to perform data quality assessment on the task fragment after the target Flink cluster receives the task fragment, write the assessment data into the distributed ledger, determine whether the data quality meets the preset threshold based on the assessment data, allocate computing nodes from the elastic computing resource pool when the preset threshold is met, allocate the task fragment to each computing node for parallel reassembly, and restore the data processing state based on the state checkpoint information.

[0134] The fifth unit is used to start a data synchronization task in the target Flink cluster and return confirmation information after restoring the data processing state.

[0135] A third aspect of the present invention provides an electronic device, comprising:

[0136] A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0137] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0138] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Flink multi-cluster secure authentication data synchronization method, characterized in that, include: Receive a data synchronization request from the source Flink cluster, obtain an access token based on the authentication information in the data synchronization request, and establish a secure connection channel. In the source Flink cluster, an incremental encoding compression algorithm is used to dynamically optimize the data processing operators and generate an execution plan. The execution plan is processed in a distributed parallel manner to obtain fine-grained task partitions. An adaptive computing load threshold is set for each task partition and computing resources are dynamically allocated. Status checkpoint information is written into the task partition. The task fragments are sent to the target Flink cluster via the secure connection channel; After the target Flink cluster receives the task shard, it performs a data quality assessment on the task shard, writes the assessment data into the distributed ledger, determines whether the data quality meets a preset threshold based on the assessment data, and allocates computing nodes from the elastic computing resource pool when the preset threshold is met. The task shard is then allocated to each computing node for parallel reassembly, and the data processing state is restored based on the state checkpoint information. After resuming data processing, a data synchronization task is started in the target Flink cluster and a confirmation message is returned.

2. The method according to claim 1, characterized in that, Receiving a data synchronization request from the source Flink cluster, obtaining an access token based on the authentication information in the data synchronization request, and establishing a secure connection channel includes: Receive a data synchronization request sent by the source Flink cluster, the data synchronization request including authentication information and target Flink cluster address information; A two-way authentication protocol is executed between the authentication center and the source Flink cluster to generate a temporary session key. The authentication information is then encrypted using the temporary session key to obtain an access token. A secure connection channel is established based on the access token and the target Flink cluster address information, and an end-to-end encrypted transmission mechanism based on the temporary session key is configured in the secure connection channel.

3. The method according to claim 1, characterized in that, In the source Flink cluster, an incremental encoding compression algorithm is used to dynamically optimize the data processing operators and generate an execution plan. The execution plan is then processed in a distributed parallel manner to obtain fine-grained task partitions. An adaptive computational load threshold is set for each task partition, and computational resources are dynamically allocated. Status checkpoint information is written into each task partition, including: In the source Flink cluster, time series analysis is performed on the data processing operators. The set corresponding to the data processing operators and the pre-acquired data flow dependencies are written into the operator dependency graph. Based on the operator dependency graph, the incremental change of the data flow in the data flow dependencies is obtained and a weight coefficient is set. The execution time of the operators in the operator dependency graph is collected. The incremental change, the weight coefficient and the execution time are substituted into the encoding function of the incremental encoding compression algorithm and the encoding function is minimized to obtain the execution plan. In the execution plan, a set of plan nodes and a set of hierarchical relationships are constructed. The memory consumption, data dependency, and input / output overhead of each plan node in the set of plan nodes are collected and a complexity evaluation value is calculated. When the complexity evaluation value exceeds a preset threshold, the execution plan is processed in a distributed parallel manner to generate fine-grained task shards. The processor utilization, memory usage, and network transmission rate of the fine-grained task fragments are collected, and an adaptive computing load threshold is calculated. Computing resources are dynamically allocated based on the proportion of the adaptive computing load threshold to the sum of the adaptive computing load thresholds of all task fragments in the fine-grained task fragments. State identifiers, state data, and timestamps are written into the fine-grained task slice to construct state checkpoint information. The difference between the state checkpoint information corresponding to two adjacent timestamps is calculated to obtain incremental checkpoint information, which is then written into the fine-grained task slice.

4. The method according to claim 3, characterized in that, The memory consumption, data dependency, and input / output overhead of each plan node in the plan node set are collected, and a complexity evaluation value is calculated. When the complexity evaluation value exceeds a preset threshold, the execution plan is processed in a distributed parallel manner to generate fine-grained task shards, including: Extract the temporal feature sequence corresponding to the execution plan, calculate the hidden state output of the historical data within the specified time window in the temporal feature sequence according to the temporal correlation, calculate the attention score of the hidden state output to obtain the feature weight coefficient, multiply the temporal feature sequence with the feature weight coefficient and sum them to obtain the task feature vector, and recursively iterate the task feature vector to obtain the prediction complexity. Calculate the topological feature matrix between the planned nodes, calculate the propagation impact value of data traffic between the nodes based on the topological feature matrix, perform graph embedding operation on the node degree distribution value and edge weight distribution value to obtain the data dependency strength, and calculate the resource competition degree matrix between the nodes to obtain the resource competition degree. A state transition probability matrix is ​​constructed based on the execution status, reward value, and action sequence in the pre-acquired historical execution data. The optimal balance coefficient is calculated based on the state transition probability matrix. The prediction complexity, data dependency strength, and resource competition degree are multiplied by the balance coefficient and jointly optimized to obtain the complexity evaluation value. When the complexity assessment value is greater than the preset threshold, the task similarity matrix is ​​calculated based on the data dependency strength and spectral clustering is performed to obtain the task collaborative execution group. Chromosome coding sequences are constructed for the tasks in the task collaborative execution group according to the computational complexity and resource requirements. The optimal parallelism configuration scheme is obtained through iterative optimization by genetic algorithm. The task collaborative execution group is allocated to different computing nodes according to the parallelism configuration scheme to perform parallel processing and obtain fine-grained task fragmentation.

5. The method according to claim 1, characterized in that, After the target Flink cluster receives the task shard, it performs a data quality assessment on the task shard, writes the assessment data into the distributed ledger, determines whether the data quality meets a preset threshold based on the assessment data, and if it does, allocates computing nodes from the elastic computing resource pool. The task shard is then distributed to each computing node for parallel reassembly. The data processing state is restored based on the state checkpoint information, including: After the target Flink cluster receives the task shards, the data missing rate is obtained by calculating the ratio of the number of missing values ​​in each field of the task shard to the total number of records. The data consistency score is obtained by calculating the entropy value of the probability of the occurrence of data values. The outlier degree is obtained by calculating the standard deviation of numerical data relative to the mean. Evaluation data is constructed based on the data missing rate, data consistency score and outlier degree. The evaluation data is written into the hierarchical storage structure of the distributed ledger. Based on the evaluation data, it is determined whether the data quality meets the preset threshold. If it does, the task complexity, memory requirements, and network bandwidth requirements are determined and the resource requirements are calculated. A multi-objective optimization function is constructed by combining resource utilization, load balancing, and communication overhead, and the resource allocation scheme is obtained by solving it. Computing nodes are allocated from the pre-set elastic computing resource pool according to the resource allocation scheme. The task is divided and distributed to each computing node for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The state item identifier sequence is traversed to calculate the dependency relationship between state items to obtain the state dependency matrix. The longest common subsequence algorithm is used to analyze and calculate the state dependency matrix to obtain the state transition cost corresponding to the state item. A state recovery path is constructed based on the state transition cost. The total cost of the state recovery path is used as the incremental recovery cost. The data processing state is restored according to the incremental recovery cost.

6. The method according to claim 5, characterized in that, The task is fragmented and distributed to various computing nodes for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The longest common subsequence algorithm is used to calculate the state transition cost corresponding to the state item. Based on the state transition cost, a state recovery path is constructed, including: The task is divided and allocated to computing nodes for parallel reassembly. The state item identifier sequence is extracted from the state checkpoint information. The state item identifier sequence is traversed to calculate the dependency relationship between state items. Based on the dependency relationship, the forward dependency chain and the backward dependency chain are constructed. The analysis results of the forward dependency chain and the backward dependency chain are stored in the state dependency matrix. Starting simultaneously from the beginning and end of the state item identifier sequence, the longest common subsequence algorithm is used to calculate the state transition cost based on the number of forward dependencies, the number of backward dependencies, and the distance between state items in the state dependency matrix. The dependency strength index is calculated based on the number of forward dependencies and backward dependencies of state items. The search direction is dynamically adjusted according to the comparison result of the dependency strength index and the preset dependency strength threshold. Forward search results and backward search results are generated. The ratio of the number of forward dependencies of each state item to its total dependencies is calculated to obtain the dynamic merging weight. The forward search results and the backward search results are weighted and merged according to the dynamic merging weight to obtain the state transition sequence. A state recovery path is constructed based on the state transition cost and the state transition sequence. The state recovery path is obtained by minimizing the sum of the state transition costs between adjacent state items in the state item identifier sequence. The data processing state is then restored according to the state recovery path.

7. The method according to claim 1, characterized in that, After resuming data processing, a data synchronization task is initiated in the target Flink cluster and confirmation information is returned, including: Obtain the data processing status recovery completion status from the target Flink cluster, verify whether the data processing status recovery result meets the preset running conditions, and if the preset running conditions are met, initialize the running environment of the data synchronization task in the target Flink cluster, load the configuration parameters required for the data synchronization task, and start the data synchronization task. Monitor the startup process of the data synchronization task, obtain the running status information of the data synchronization task, and return the running status information as confirmation information.

8. A Flink multi-cluster secure authentication data synchronization system, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to receive data synchronization requests from the source Flink cluster, obtain an access token based on the authentication information in the data synchronization request, and establish a secure connection channel. The second unit is used to dynamically optimize the data processing operator and generate an execution plan in the source Flink cluster using an incremental encoding compression algorithm, process the execution plan in a distributed parallel manner to obtain fine-grained task partitions, set an adaptive computing load threshold for each task partition and dynamically allocate computing resources, and write state checkpoint information into the task partitions. The third unit is used to send the task fragments to the target Flink cluster through the secure connection channel; The fourth unit is used to perform data quality assessment on the task fragment after the target Flink cluster receives the task fragment, write the assessment data into the distributed ledger, determine whether the data quality meets the preset threshold based on the assessment data, allocate computing nodes from the elastic computing resource pool when the preset threshold is met, allocate the task fragment to each computing node for parallel reassembly, and restore the data processing state based on the state checkpoint information. The fifth unit is used to start a data synchronization task in the target Flink cluster and return confirmation information after restoring the data processing state.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unstructured data synchronization method and system based on Flink

    CN120407294A

  • Persistent Flink job file loading and displaying method and system

    CN120408687A