Multi-stage collaborative reconfiguration method based on fine-grained state transition, electronic equipment and storage medium

By building a formal model of multi-level collaborative reconfiguration optimization problem on the stream computing platform and designing a multi-level collaborative reconfiguration method based on second-order heuristic algorithms, the problem of large reconfiguration overhead in edge environments is solved, and system performance improvement and delay reduction is achieved.

CN120110907AActive Publication Date: 2025-06-06HARBIN INST OF TECH

Patent Information

Application Number
CN202510262789.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

When reconfiguring in an edge environment, the existing stream computing platform lacks a collaborative optimization strategy that comprehensively considers the three levels of partitioning, scheduling and elasticity, resulting in excessive reconfiguration overhead and deterioration of system performance.

Method used

A multi-level collaborative reconfiguration method based on fine-grained state migration is proposed. By building a formal model of multi-level collaborative reconfiguration optimization problem, a multi-level collaborative reconfiguration method based on second-order heuristic algorithm is designed, and a reconfiguration scheme at three levels of partitioning, scheduling and elasticity are optimized.

Benefits of technology

It effectively reduces the reconfiguration overhead and improves system performance. The experimental results show that in the mobile edge scenario, the end-to-end delay of the application has dropped by 52.64%, with an average execution delay of 339ms, and has online execution capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110907A_ABST
    Figure CN120110907A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level collaborative reconfiguration method based on fine-grained state transition, electronic equipment and a storage medium, and belongs to the technical field of stream computing platform design. In order to avoid unnecessary reconfiguration overhead, the method comprises the following steps of: constructing a formalized model of a multi-stage collaborative reconfiguration optimization problem based on a built-in reconfiguration mechanism of a stream computing platform by combining characteristics of an edge computing environment and overhead of three levels of partitioning, scheduling and elasticity in a reconfiguration process; on the basis of the constructed formal model of the multi-stage collaborative reconfiguration optimization problem, constructing a granularity adjustment method for fine granularity state transition, and independently solving the selection of the optimal transition state granularity in the multi-stage collaborative reconfiguration optimization problem; and designing a multi-stage collaborative reconfiguration method based on a second-order heuristic algorithm, obtaining a multi-stage collaborative reconfiguration solution integrating three levels of partitioning, scheduling and elasticity, and obtaining an approximate optimal solution in polynomial time. According to the invention, the number of reconfiguration times can be effectively reduced, and unnecessary reconfiguration overhead is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of stream computing platform design, and specifically relates to a multi-level collaborative reconfiguration method based on fine-grained state migration, an electronic device and a storage medium. Background Art

[0002] The built-in reconfiguration mechanism in the existing stream computing platform Flink includes fine-grained state migration and stream computing reconfiguration mechanism.

[0003] When stream computing reconfiguration involves stateful operators, state migration becomes an important part of the reconfiguration process, which may cause a lot of reconfiguration overhead and even deteriorate system performance. Fine-grained state migration is proposed to alleviate the negative impact of state migration. Its core is to divide the state into multiple key groups and use the key group as the smallest migration unit. Key-value databases are usually used to save state key groups and are responsible for forwarding state key groups between source instances and target instances.

[0004] The stream computing reconfiguration mechanism is designed to cope with dynamically changing network conditions and limited computing resources in edge environments. The stream processing system must have a flexible runtime reconfiguration mechanism. The reconfiguration strategy should minimize system overhead while ensuring efficient performance, including partitioning, scheduling, and elasticity:

[0005] Partitioning is used to adjust the input load of operators. It adjusts the load allocated to downstream operators by modifying the routing strategy of upstream operators to achieve load balancing or unload instances with excessive load. Therefore, partitioning is mainly used to deal with problems such as load surges or load tilt.

[0006] Scheduling involves modifying the entire job graph structure. Since scheduling occurs at runtime, consistency is ensured by introducing an intermediate topology. For example, when scheduling instance a from node A to node B, you first need to start a new instance on node B, modify the routing of the upstream operator and the state of the migration operator, cancel the node, and finally delete the old instance on node A.

[0007] Elasticity is achieved by adjusting resource sets to adapt to sudden increases or decreases in load. When expanding a node, increase the parallelism of the bottleneck operator and expand the instance to the newly added node, then update the routing strategy and migrate the state. Conversely, when reclaiming a node, if there are stateful operators on the node, the operator state needs to be migrated to other nodes, then adjust the routing of the upstream operator, and finally release the node.

[0008] Since the node resources in the edge environment are usually heterogeneous and the network transmission speed between nodes is unstable, the static runtime reconfiguration mechanism cannot meet the needs of edge streaming processing in some scenarios. Among them, most of the existing reconfiguration mechanisms only optimize performance from one of the levels of partitioning, scheduling or elasticity, and lack a collaborative optimization strategy that integrates the three levels. In addition, some reconfiguration mechanisms also introduce fine-grained migration to reduce the overhead caused by state migration during reconfiguration. However, the prior art does not specifically discuss the benefits of state granularity on migration performance and the additional overhead caused. Summary of the invention

[0009] The problem to be solved by the present invention is to avoid unnecessary reconfiguration overhead, and a multi-level collaborative reconfiguration method based on fine-grained state migration, an electronic device and a storage medium are proposed.

[0010] To achieve the above object, the present invention is implemented through the following technical solutions:

[0011] A multi-level collaborative reconfiguration method based on fine-grained state migration includes the following steps:

[0012] S1. Based on the built-in reconfiguration mechanism of the stream computing platform, combined with the characteristics of the edge computing environment and the overhead of partitioning, scheduling and elasticity in the reconfiguration process, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed;

[0013] S2. Based on the formal model of the multi-level collaborative reconfiguration optimization problem constructed in step S1, a granularity adjustment method for fine-grained state migration is constructed, and the selection of the optimal migration state granularity in the multi-level collaborative reconfiguration optimization problem is solved separately;

[0014] S3. Design a multi-level collaborative reconfiguration method based on a second-order heuristic algorithm to obtain a multi-level collaborative reconfiguration solution at the three levels of comprehensive partitioning, scheduling, and flexibility to obtain an approximate optimal solution in polynomial time.

[0015] Furthermore, the specific implementation method of step S1 includes the following steps:

[0016] S1.1. Construct a migration state key group model;

[0017] Define the set of migration state key groups key as K, where K i A key set representing the migration status of the i-th batch;

[0018] Set k as the key of a certain transition state, S(k) represents the state corresponding to k, I o k is the source instance that sends the migration tuple, I t k is the target instance that receives the migration tuple, then the i-th batch state set S sent by instance j iso (i, j), and the partial state set S sent to instance x o (i,j,x) is represented as:

[0019]

[0020] The i-th batch state set S received by instance j t (i, j) and its partial state set S from instance x t (i,j,x) are represented as:

[0021]

[0022] S1.2. Analyze the factors affecting the reconfiguration overhead of the migration state, including the time T for aligning the migration instructions align , the time of state transition T mig and the time T to return to normal processing after successful migration res , the sum of the three is the total migration time T;

[0023] State transition time T mig The time T serialized by the state srl , state transfer time T trans and the time T for state deserialization desrl is composed of the following expressions:

[0024] T srl +T desrl =α*(max{S o (i,j)}+max{S t (i,j)})#(2-5)

[0025] Among them, α is the slope of the linear function of the consumption time on the state size;

[0026] Let B(x,j) be the network transmission rate between instances x and j. Then the state transmission time T for instance j to send the state to the target instance in the i-th batch is o (i, j), and the state transfer time T for receiving the state from the source instance t (i,j) are represented as:

[0027]

[0028] Based on the fact that when instance j receives the state of the i-th batch of migration, it blocks the tuples coming from the upstream that will affect the state recovery, and defines M i,j is the number of tuples blocked by instance j during the i-th batch migration, expressed as:

[0029] M i,j=f(k)*λ i,j *(T align +T mig )#(2-8)

[0030] Among them, λ i,j represents the input flow intensity of instance j during the i-th batch migration, and f(k) represents the probability of k appearing during the i-th batch migration;

[0031] Calculate the time T to return to normal processing after the i-th batch of state migration is successful res (i), the expression is:

[0032]

[0033] Among them, μ j is the processing capability of instance j;

[0034] S1.3. Based on the factors affecting the reconfiguration overhead of the migration state, a state migration overhead model is constructed, and the expression is:

[0035]

[0036] T peak (m,B,P)=T align +T mig #(2-11)

[0037] Among them, T span (m,B,P) is the total time of state migration of m batches, T peak (m, B, P) is the peak latency of m batches of state migration, P is a triple used to represent the multi-level reconfiguration strategy, each element in P represents the partition, scheduling and elasticity strategy, and B is the network transmission rate;

[0038] S1.4.Formulate the partitioning, scheduling and elasticity processes based on the state migration overhead model;

[0039] S1.4.1.Formulate the partitioning process based on the state migration cost model:

[0040] The upstream operator calculates the route based on the key of the tuple and distributes the data to the downstream operator, using the mapping relationship f p :k→i' represents a partition operation, where i' represents the index of the downstream instance;

[0041] Divide the total key group K into n sub-key groups kg = {kg 0 ,kg 1 ,…,kg n-1},have With matrix M pRecord the routing relationship between the subkey set kg and the instance set V, and the value of each element is defined as follows:

[0042]

[0043] Among them, M p [i'][j] is the subkey group kg i' and instance V j The routing mapping relationship, V j is the jth instance;

[0044] Reconfiguration requires adjustment of the partitions. After the partition operation, the routing relationship between the partitioned sub-key groups and instances is represented by M' p ; then the repartitioning strategy is expressed as ΔM p =M' p -M p , ΔM p Each element takes the value of -1, 0 or 1, representing the subkey group kg i' From the example V j Migrated out, not migrated, and migrated to instance V j Three strategies;

[0045] Suppose the partitioned multi-level reconfiguration strategy triple It is a placeholder empty set, indicating that scheduling strategy and elasticity strategy are not involved here;

[0046] According to the repartitioning strategy, calculate the state quantity that needs to be migrated for a repartitioning. First, define the partition migration state vector Used to save the state size corresponding to each sub-key group. Next, define the partition adjustment instance index set A collection of indices representing instances that sent status during the partitioning process;

[0047] Then we get the state quantity involved in the migration during the partition process.

[0048] Combining formula (2-10) and formula (2-11), we can finally get the total time T for partition operation. span (m,B,P p ) and the peak latency of partition state migration T peak (m,B,P p );

[0049] S1.4.2.Formulate the scheduling process based on the state transition overhead model;

[0050] Definition M s , M' s and ΔM sThey represent the instance deployment before reconfiguration, the instance deployment after reconfiguration, and the scheduling strategy, respectively. s The value of the element is 0 or 1, indicating the instance V i' Whether to deploy on node N j On, ΔM s The value of the [i'][j] element is -1, 0 or 1, indicating the instance V i' From Node N j Node N that has been migrated, not migrated, or migrated j , let the triple of the multi-level reconfiguration strategy of scheduling be

[0051] Define the scheduling adjustment instance index set represents the index of the instance that sends the status during the scheduling process, then the number of migration instances is V mig =|I s |;

[0052] The time overhead related to instance management before and after migration is expressed as ρ*V mig ,ρ is the instance management overhead coefficient;

[0053] For the migration overhead caused by scheduling, define the scheduling migration state vector Used to save the state size corresponding to each stateful instance, combined with ΔM s That is, the state quantity involved in the scheduling process is obtained

[0054] according to The total time T associated with the state transition caused by B being scheduled span (m,B,P s ) and the peak delay of scheduling state transition T peak (m,B,P s ), the total time of the scheduling phase and the peak delay of the scheduling state migration are ρ*V mig +T span (m,B,P s ) and T peak (m,B,P s );

[0055] S1.4.3.Formulate modeling of elasticity process based on state migration overhead model;

[0056] Defining vectors Indicates the node usage. Each element in the vector takes a value of 0 or 1, representing node N. j After an elastic operation, the usage of node resources is readjusted to The elastic strategy is expressed as Its element values ​​are -1, 0, or 1, which represent that the node is released, the node status is unchanged, and the node is removed from the resource pool. Considering that the elastic operation only adjusts the resource set, the additional overhead caused by the elastic operation is expressed as β is the overhead coefficient corresponding to the adjusted resource set;

[0057] S1.5. Constructing a multi-level reconfiguration strategy triple representing the three mechanisms of comprehensive partitioning, scheduling, and elasticity Perform overhead analysis of multi-level reconfiguration strategies;

[0058] Set the time T for elastic adjustment of node resource consumption elastic is a constant value, and elastic operations will not cause migration blocking; the time taken to deploy an instance in the stage is T deploy , then the total deployment time of all related instances is T deploy *V mig ; Both scheduling and partitioning involve state migration and there is migration blocking and delay peak. When the state migration batch is m, the total reconfiguration time and the delay peak corresponding to the reconfiguration strategy P are calculated by the following two formulas:

[0059] T 1 (m,B,P)=δ p *T span (m,B,P p )+δ S *(T deploy *V mig +T span (m,B,P s ))+δ E *T elastic #(2-13)

[0060] T 2 (m,B,P)=max{T peak (m,B,P p ),T peak (m,B,P s )}#(2-14)

[0061] Among them, δ p , δ S and δ E Indicates whether the reconfiguration strategy involves partitioning, scheduling, or elastic operations. A value of 0 indicates no operation is involved, and a value of 1 indicates that an operation is involved.

[0062] The reconfiguration time overhead is defined as The overhead associated with the latency peak is expressed as Where L 1 is the average end-to-end delay of the application before reconfiguration;

[0063] Finally, the total reconfiguration cost is calculated according to the following formula:

[0064]

[0065] Wherein, a and b are the first trade-off coefficient and the second trade-off coefficient respectively;

[0066] Benefits of computing reconfiguration strategy:

[0067]

[0068] Wherein, c and d are the third trade-off coefficient and the fourth trade-off coefficient respectively;

[0069] S1.6. Based on the cost analysis of the multi-level reconfiguration strategy in step S1.5, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed, which is subject to constraints (1)-(6).

[0070] maximize Objective=benefit-cost#(2-17)

[0071]

[0072] Constraint (1) ensures that partition operations only migrate key groups from instances that have this key group; constraint (2) ensures that scheduling operations only occur from nodes that have this instance; constraints (3) and (4) limit partition and scheduling operations to being effectively transferred between instances and nodes, thereby ensuring consistency of operations; constraint (5) ensures that the load on each node after reconfiguration does not exceed its processing capacity; constraint (6) limits the number of loads on a node after reconfiguration to not exceed the number of its physical resource slots.

[0073] Furthermore, the specific implementation method of step S2 includes the following steps:

[0074] S2.1. Set the goal to minimize the reconfiguration overhead and output the corresponding state granularity. Rewrite equation (2-11) according to the following process to obtain the delay peak T of the rewritten m batches of state migration peak (m,B,P)' expression is:

[0075]

[0076] S2.2. For the total reconfiguration time, rewrite equation (2-10) from equation (2-9) to obtain the expression:

[0077]

[0078] Among them, Q is the sum of the state transition time of each batch, Q = α*(|So |+|S t |)+T o +T t , in the case of uniform particle size division, it can be considered

[0079] Then simplifying, we get

[0080] For the simplified T span Take the derivative and find T span The value m corresponding to the minimum * for The expression for the value range of m is:

[0081]

[0082] Furthermore, the specific implementation method of step S3 includes the following steps:

[0083] S3.1. Set the first stage to generate two secondary reconfiguration solutions, namely, a partition-first-then-scheduling secondary reconfiguration solution and a scheduling-first-then-partition secondary reconfiguration solution;

[0084] S3.2. Based on the simulated annealing algorithm, the second-level reconfiguration solution obtained in step S3.1 is searched in the domain to obtain the final second-level reconfiguration solution.

[0085] Furthermore, the specific implementation method of step S3.1 includes the following steps:

[0086] For the partition-first-then-scheduling two-level reconfiguration solution, an incremental repartitioning method is used. The instances with higher than average load are sorted from large to small according to load, and then their key groups are migrated to the instance with the lowest load in sequence until the load of the instance is lower than the average load after the key group migration;

[0087] For the two-level reconfiguration solution of scheduling first and partitioning later, the nodes are first sorted according to computing capabilities, and then the operators are scheduled according to network communication performance, with network communication priority. If a node cannot carry an operator pair, the scheduling is extended to the instance level, and the instance pair with the largest communication volume among the operator pairs is placed first until the node resources are unavailable, and then the next node is moved to continue placing instances.

[0088] In the case where node resources are not fully utilized in the two-level reconfiguration solution of scheduling first and partitioning later, the load of the operator instance on the node with fragmented resources is increased until the load reaches the upper limit that the node resources can bear.

[0089] Furthermore, the specific implementation method of step S3.2 is to partition first and then schedule the two-level reconfiguration solution ΔM obtained in step S3.1 p,s The solution of scheduling first and partitioning second-level reconfiguration ΔM s,p , using simulated annealing algorithm to search <ΔM p,s ,ΔM s,p >Nearby area, set a probability threshold p, if the solution obtained by the current cycle search is not better than the temporary optimal solution, and the exploration probability p e >p, the temporary optimal solution is updated to this solution to expand the search space; when the number of cycles of the simulated annealing algorithm reaches the maximum value, the optimal solution obtained is the final secondary reconfiguration solution.

[0090] Furthermore, the load predictor is designed to adjust the reconfiguration interval time by using the exponential smoothing method. The specific implementation method is to collect the load time series data X from the end of the last reconfiguration execution to the current moment, and use the exponential smoothing method to process the data items of X one by one. i and the previous data item x i-1 The corresponding residual Calculate the corresponding residual The calculation formula is:

[0091]

[0092] Get the residual corresponding to the data item at the current moment The time interval for performing the next reconfiguration.

[0093] Furthermore, the load predictor is integrated into Flink, and the load history data required by the load predictor comes from the indicator collector of TaskManager and is collected and summarized by the resource monitor of JobManager.

[0094] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a multi-level collaborative reconfiguration method based on fine-grained state migration when executing the computer program.

[0095] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a multi-level collaborative reconfiguration method based on fine-grained state migration.

[0096] Beneficial effects of the present invention:

[0097] The multi-level collaborative reconfiguration method based on fine-grained state migration described in the present invention designs a multi-level collaborative reconfiguration algorithm that integrates partitioning, scheduling, and elasticity for edge computing environments. First, a reconfiguration benefit and cost model is proposed based on the characteristics of the edge computing environment and the overhead of the three levels in the reconfiguration process, which realizes the formal description of the reconfiguration optimization problem. Then, a second-order heuristic algorithm based on simulated annealing is designed to solve the optimization problem to efficiently obtain a multi-level reconfiguration solution.

[0098] The present invention discloses a multi-level collaborative reconfiguration method based on fine-grained state migration. The present invention reduces the reconfiguration overhead based on fine-grained state migration, comprehensively considers the three-level reconfiguration mechanism of partitioning, scheduling and elasticity in stream computing, and models and describes the reconfiguration-related overhead and multi-level reconfiguration optimization problems. After determining the appropriate state migration granularity, a partitioning, scheduling and elasticity solution that optimizes system performance is generated through a second-order heuristic algorithm based on a simulated annealing algorithm. Experiments verify that when facing mobile edge scenarios, the proposed algorithm can be used for reconfiguration to reduce the end-to-end delay of the application by 52.64%, and the average execution delay is 339ms, indicating that the algorithm has online execution capabilities.

[0099] The multi-level collaborative reconfiguration method based on fine-grained state migration described in the present invention collects system load time series data, uses exponential smoothing method to calculate the residual corresponding to each data item of the time series data in each iteration of the loop, and uses the residual to calculate a new residual in the next iteration. The residual finally obtained is the time interval for triggering the reconfiguration next time, so that the system has the ability to actively perform reconfiguration and reduce the time overhead caused by reconfiguration. Experiments have shown that the use of prediction methods in mobile edge scenarios can effectively reduce the number of reconfiguration occurrences and avoid unnecessary reconfiguration overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 It is a structural schematic diagram of a heat-insulating outer cover for protecting the temperature of an elevator traction machine from low temperatures according to the present invention;

[0101] Figure 2 It is a block diagram of the multi-level collaborative reconfiguration mechanism of the present invention;

[0102] Figure 3 This is a functional framework diagram of Flink reconfiguration after the present invention is extended;

[0103] Figure 4 The load fluctuation trend diagram in the mobile edge environment simulated by the present invention;

[0104] Figure 5 It is a trend diagram of the serialization and deserialization time of the present invention as the amount of state data changes;

[0105] Figure 6 The latency peak graphs generated by reconfiguration at different migration granularities, where (a) shows the relationship between latency peak and migration granularity when NexMark Q4 is used as the application load, (b) shows the relationship between latency peak and migration granularity when NexMark Q8 is used as the application load, and (c) shows the relationship between latency peak and migration granularity when smart transportation is used as the application load;

[0106] Figure 7 The total time consumed for reconfiguration at different migration granularities, where (a) is the relationship between reconfiguration time and migration granularity when NexMark Q4 is used as the application load, (b) is the relationship between reconfiguration time and migration granularity when NexMark Q8 is used as the application load, and (c) is the relationship between reconfiguration time and migration granularity when smart transportation is used as the application load;

[0107] Figure 8 The following is a comparison chart of the end-to-end delay change trends of NexMark Q4 and smart transportation in scenario 1, where (a) shows the end-to-end delay change trends under different reconfiguration algorithms when NexMark Q4 is used as the application load, and (b) shows the end-to-end delay change trends under different reconfiguration algorithms when smart transportation is used as the application load;

[0108] Fig. 9 The following are the trends of the end-to-end delay of the application after using different reconfiguration algorithms to handle back pressure in scenario 1, where (a) is the trend of the end-to-end delay when using MLR-algorithm to handle back pressure, (b) is the trend of the end-to-end delay when using DR to handle back pressure, (c) is the trend of the end-to-end delay when using GT-Scheduler to handle back pressure, and (d) is the trend of the end-to-end delay when using MC-Stream to handle back pressure.

[0109] Fig.10 This is a comparison chart of the average response delay of the system after using different reconfiguration algorithms to handle back pressure in scenario 1;

[0110] Fig.11 The following are the trends of the end-to-end delay when reconfiguration is triggered periodically by different reconfiguration algorithms in scenario 2. (a) shows the trend of the end-to-end delay when reconfiguration is triggered periodically by MLR-algorithm, (b) shows the trend of the end-to-end delay when reconfiguration is triggered periodically by DR, (c) shows the trend of the end-to-end delay when reconfiguration is triggered periodically by GT-Scheduler, and (d) shows the trend of the end-to-end delay when reconfiguration is triggered periodically by MC-Stream.

[0111] Fig.12The end-to-end delay variation trends of smart transportation applications when using MLR-algorithm and Mc-Stream respectively, where (a) is the end-to-end delay variation trend of smart transportation applications when using MLR-algorithm, and (b) is the end-to-end delay variation trend of smart transportation applications when using Mc-Stream;

[0112] Fig.13 The figure is a comparison of the average execution delays of the four reconfiguration algorithms. DETAILED DESCRIPTION

[0113] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0114] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0115] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 -Attached Fig.13 The detailed instructions are as follows:

[0116] Embodiment 1:

[0117] A multi-level collaborative reconfiguration method based on fine-grained state migration includes the following steps:

[0118] S1. Based on the built-in reconfiguration mechanism of the stream computing platform, combined with the characteristics of the edge computing environment and the overhead of partitioning, scheduling and elasticity in the reconfiguration process, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed;

[0119] When a stream processing system is reconfigured at runtime, if the operator instances involved are stateful, state migration needs to be enabled to ensure state consistency. This process must block the instance from running, which will introduce additional migration overhead. Fine-grained migration can reduce tuple migration time and the total migration time. Considering the granularity of the split state, formulating and modeling the state migration overhead will help to evaluate the impact of state migration on reconfiguration overhead at different granularities.

[0120] Furthermore, the specific implementation method of step S1 includes the following steps:

[0121] S1.1. Construct a migration state key group model;

[0122] Define the set of migration state key groups key as K, where K i A key set representing the migration status of the i-th batch;

[0123] Set k as the key of a certain transition state, S(k) represents the state corresponding to k, I o k is the source instance that sends the migration tuple, I t k is the target instance that receives the migration tuple, then the i-th batch state set S sent by instance j is o (i, j), and the partial state set S sent to instance x o (i,j,x) is represented as:

[0124]

[0125]

[0126] The i-th batch state set S received by instance j t (i, j) and its partial state set S from instance x t (i,j,x) are represented as:

[0127]

[0128] S1.2. Analyze the factors affecting the reconfiguration overhead of the migration state, including the time T for aligning the migration instructions align , the time of state transition T mig and the time T to return to normal processing after successful migration res , the sum of the three is the total migration time T;

[0129] State transition time T mig The time T serialized by the state srl , state transfer time T trans and the time T for state deserialization desrl is composed of the following expressions:

[0130] T srl +T desrl =α*(max{S o (i,j)}+max{S t (i,j)})#(2-5)

[0131] Among them, α is the slope of the linear function of the consumption time on the state size;

[0132] Due to T align It is usually a small constant value and can be ignored when analyzing migration overhead. For state serialization and deserialization, both are performed simultaneously and there is a linear positive correlation between the time consumed and the state size. Therefore, the total time consumed by the serialization and deserialization process can be obtained by the following formula;

[0133] Let B(x,j) be the network transmission rate between instances x and j. Then the state transmission time T for instance j to send the state to the target instance in the i-th batch is o (i, j), and the state transfer time T for receiving the state from the source instance t (i,j) are represented as:

[0134]

[0135] Based on the fact that when instance j receives the state of the i-th batch of migration, it blocks the tuples coming from the upstream that will affect the state recovery, and defines M i,j is the number of tuples blocked by instance j during the i-th batch migration, expressed as:

[0136] M i,j =f(k)*λ i,j *(T align +T mig )#(2-8)

[0137] Among them, λ i,j represents the input flow intensity of instance j during the i-th batch migration, and f(k) represents the probability of k appearing during the i-th batch migration;

[0138] Calculate the time T to return to normal processing after the i-th batch of state migration is successful res (i), the expression is:

[0139]

[0140] Among them, μ j is the processing capability of instance j;

[0141] Affected by migration blocking, the end-to-end delay of the blocked tuple will be higher than the end-to-end delay under normal operation. From the perspective of delay statistics, this will cause a significant peak in the delay curve, and the peak is called the delay peak. Let latency be the time taken by the blocked tuple from arriving at the instance to being processed, that is, the blocking time. Then, for the t(0<=t<=T,T=T align +T mig ) The tuple of the time when Among them, there is an implicit condition λ<μ, which is a prerequisite for queue stability. Therefore, when t=0, the latency is maximum and the latency is equal to T align With T mig The end-to-end delay of the tuple arriving at this moment is the peak delay of the performance fluctuation during this migration period.

[0142] At this point, the total time T for m batches of state transitions can be calculated according to the following formula: span (m,B,P) and peak delay T peak (m, B, P) is used to analyze the negative impact of state migration on reconfiguration. Among them, P is a triple used to represent a multi-level reconfiguration strategy, and each element represents a partition, scheduling, and elasticity strategy.

[0143] S1.3. Based on the factors affecting the reconfiguration overhead of the migration state, a state migration overhead model is constructed, and the expression is:

[0144]

[0145] T peak (m,B,P)=T align +T mig #(2-11)

[0146] Among them, T span (m,B,P) is the total time of state migration of m batches, T peak (m, B, P) is the peak latency of m batches of state migration, P is a triple used to represent the multi-level reconfiguration strategy, each element in P represents the partition, scheduling and elasticity strategy, and B is the network transmission rate;

[0147] Stream processing reconfiguration optimizes the system load rate, resource utilization, and throughput through three mechanisms: partitioning, scheduling, and elasticity. Since reconfiguration requires structural adjustments to the system's resource topology and application topology, these three mechanisms all rely on state migration to maintain state consistency before and after reconfiguration. In order to evaluate the reconfiguration overhead, the partitioning, scheduling, and elasticity processes are formulated and modeled based on the state migration overhead model, which facilitates further analysis of the negative impact of state migration on stream processing performance.

[0148] S1.4.Formulate the partitioning, scheduling and elasticity processes based on the state migration overhead model;

[0149] S1.4.1.Formulate the partitioning process based on the state migration cost model:

[0150] The upstream operator calculates the route based on the key of the tuple and distributes the data to the downstream operator, using the mapping relationship f p :k→i' represents a partition operation, where i' represents the index of the downstream instance;

[0151] Divide the total key group K into n sub-key groups kg = {kg 0 ,kg 1 ,…,kg n-1},have With matrix M p Record the routing relationship between the subkey set kg and the instance set V, and the value of each element is defined as follows:

[0152]

[0153] Among them, M p [i'][j] is the subkey group kg i' and instance V j The routing mapping relationship, V j is the jth instance;

[0154] Reconfiguration requires adjustment of the partitions. After the partition operation, the routing relationship between the partitioned sub-key groups and instances is represented by M' p ; then the repartitioning strategy is expressed as ΔM p =M' p -M p , ΔM p Each element takes the value of -1, 0 or 1, representing the subkey group kg i' From the example V j Migrated out, not migrated, and migrated to instance V j Three strategies;

[0155] Suppose the partitioned multi-level reconfiguration strategy triple It is a placeholder empty set, indicating that scheduling strategy and elasticity strategy are not involved here;

[0156] According to the repartitioning strategy, calculate the state quantity that needs to be migrated for a repartitioning. First, define the partition migration state vector Used to save the state size corresponding to each sub-key group. Next, define the partition adjustment instance index set A collection of indices representing instances that sent status during the partitioning process;

[0157] Then we get the state quantity involved in the migration during the partition process.

[0158] Combining formula (2-10) and formula (2-11), we can finally get the total time T for partition operation. apan (m,B,P p ) and the peak latency of partition state migration T peak (m,B,P p );

[0159] S1.4.2.Formulate the scheduling process based on the state transition overhead model;

[0160] Unlike partitioning operations, scheduling operations require adjusting operator parallelism and broadcasting control information to downstream for alignment before migration in addition to state migration. After migration, resources of old instances need to be recycled and cleaned up. Both steps will incur additional overhead.

[0161] Definition M s , M' s and ΔM s They represent the instance deployment before reconfiguration, the instance deployment after reconfiguration, and the scheduling strategy, respectively. s The value of the element is 0 or 1, indicating the instance V i' Whether to deploy on node N j On, ΔM s The value of the [i'][j] element is -1, 0 or 1, indicating the instance V i' From Node N j Node N that has been migrated, not migrated, or migrated j , let the triple of the multi-level reconfiguration strategy of scheduling be

[0162] Define the scheduling adjustment instance index set represents the index of the instance that sends the status during the scheduling process, then the number of migration instances is V mig =|I s |;

[0163] The time overhead related to instance management before and after migration is expressed as ρ*V mig ,ρ is the instance management overhead coefficient;

[0164] For the migration overhead caused by scheduling, define the scheduling migration state vector Used to save the state size corresponding to each stateful instance, combined with ΔM s That is, the state quantity involved in the scheduling process is obtained

[0165] according to The total time T associated with the state transition caused by B being scheduled span (m,B,P s ) and the peak delay of scheduling state transition T peak (m,B,p s ), the total time of the scheduling phase and the peak delay of the scheduling state migration are ρ*V mig +T span (m,B,P s ) and T peak (m,B,P s );

[0166] S1.4.3.Formulate modeling of elasticity process based on state migration overhead model;

[0167] When a sudden load peak occurs in a stream processing system, neither partitioning nor scheduling operations can cope with it, and elastic operations must be used to handle it. Elastic operations take out available node sets or resource sets from the resource pool, similar to scheduling operations to scale bottleneck operators. The difference is that elastic operations do not need to reclaim the resources of the original instance after scaling.

[0168] Defining vectors Indicates the node usage. Each element in the vector takes a value of 0 or 1, representing node N. j After an elastic operation, the usage of node resources is readjusted to The elastic strategy is expressed as Its element values ​​are -1, 0, or 1, which represent that the node is released, the node status is unchanged, and the node is removed from the resource pool. Considering that the elastic operation only adjusts the resource set, the additional overhead caused by the elastic operation is expressed as β is the overhead coefficient corresponding to the adjusted resource set;

[0169] S1.5. Constructing a multi-level reconfiguration strategy triple representing the three mechanisms of comprehensive partitioning, scheduling, and elasticity Perform overhead analysis of multi-level reconfiguration strategies;

[0170] P reflects the changes in the mapping relationship between key groups and instances, the mapping relationship between instances and nodes, and the set of computational nodes;

[0171] Set the time T for elastic adjustment of node resource consumption elastic is a constant value, and elastic operations will not cause migration blocking; the time taken to deploy an instance in the stage is T deploy , then the total deployment time of all related instances is T deploy *V mig; Both scheduling and partitioning involve state migration and there is migration blocking and delay peak. When the state migration batch is m, the total reconfiguration time and the delay peak corresponding to the reconfiguration strategy P are calculated by the following two formulas:

[0172] T 1 (m,B,P)=δ p *T span (m,B,P p )+δ S *(T deploy *V mig +T span (m,B,P s ))+δ E *T elastic #(2-13)

[0173] T 2 (m,B,P)=max{T peak (m,B,P p ),T peak (m,B,P s )}#(2-14)

[0174] Among them, δ p , δ S and δ E Indicates whether the reconfiguration strategy involves partitioning, scheduling, or elastic operations. A value of 0 indicates no operation is involved, and a value of 1 indicates that an operation is involved.

[0175] In addition to the overhead caused by elastic operations, the reconfiguration overhead is also related to the reconfiguration time and the latency peak. a To evaluate the timeliness of reconfiguration optimization;

[0176] The reconfiguration time overhead is defined as The overhead associated with the latency peak is expressed as Where L 1 is the average end-to-end delay of the application before reconfiguration;

[0177] Finally, the total reconfiguration cost is calculated according to the following formula:

[0178]

[0179] Wherein, a and b are the first trade-off coefficient and the second trade-off coefficient respectively;

[0180] Benefits of computing reconfiguration strategy:

[0181]

[0182] Wherein, c and d are the third trade-off coefficient and the fourth trade-off coefficient respectively;

[0183] S1.6. Based on the cost analysis of the multi-level reconfiguration strategy in step S1.5, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed, which is subject to constraints (1)-(6).

[0184] maximize Objective=benefit-cost#(2-17)

[0185]

[0186] Constraint (1) ensures that partition operations only migrate key groups from instances that have this key group; constraint (2) ensures that scheduling operations only occur from nodes that have this instance; constraints (3) and (4) limit partition and scheduling operations to being effectively transferred between instances and nodes, thereby ensuring consistency of operations; constraint (5) ensures that the load on each node after reconfiguration does not exceed its processing capacity; constraint (6) limits the number of loads on a node after reconfiguration to not exceed the number of its physical resource slots.

[0187] S2. Based on the formal model of the multi-level collaborative reconfiguration optimization problem constructed in step S1, a granularity adjustment method for fine-grained state migration is constructed, and the selection of the optimal migration state granularity in the multi-level collaborative reconfiguration optimization problem is solved separately;

[0188] Fine-grained migration can reduce the additional overhead introduced by state migration during reconfiguration. However, equations (2-10) and (2-11) show that the choice of state granularity is closely related to the total state migration time T. span and the peak delay T peak There is a connection between the state granularity and the multi-level reconfiguration optimization problem, so the change of state granularity affects the reconfiguration cost. The variability of state granularity makes the solution space of the reconfiguration optimization problem extremely large and the problem solving process is more complicated. Obviously, it is unrealistic to rely on manual determination of the granularity size and achieve reconfiguration optimization. Therefore, a granularity adjustment algorithm for fine-grained state migration is proposed, which separates the selection of state granularity from the multi-level reconfiguration optimization problem and solves it separately, simplifying the difficulty of solving the multi-level reconfiguration optimization problem. It aims to minimize the reconfiguration cost and output the corresponding state granularity.

[0189] Furthermore, the specific implementation method of step S2 includes the following steps:

[0190] S2.1. Set the goal to minimize the reconfiguration overhead and output the corresponding state granularity. Rewrite equation (2-11) according to the following process to obtain the delay peak T of the rewritten m batches of state migration peak (m,B,P)' expression is:

[0191]

[0192] S2.2. For the total reconfiguration time, rewrite equation (2-10) from equation (2-9) to obtain the expression:

[0193]

[0194] Among them, Q is the sum of the state transition time of each batch, Q = α*(|S o |+|S t |+T o +T t , in the case of uniform particle size division, it can be considered

[0195] Then simplifying, we get

[0196] For the simplified T span Take the derivative and find T span The value m corresponding to the minimum * for The expression for the value range of m is:

[0197]

[0198] Furthermore, it should be noted that the prerequisite must be satisfied. Therefore, the final value of m is determined by formula (2-20).

[0199] S3. Design a multi-level collaborative reconfiguration method based on a second-order heuristic algorithm to obtain a multi-level collaborative reconfiguration solution at the three levels of comprehensive partitioning, scheduling, and flexibility to obtain an approximate optimal solution in polynomial time.

[0200] After solving the problem of selecting the optimal migration state granularity, how to obtain a multi-level reconfiguration solution with comprehensive partitioning, scheduling, and elasticity remains a complex problem. Obviously, this problem is an NP-hard problem. Therefore, a multi-level cooperative reconfiguration algorithm based on a second-order heuristic algorithm is designed to obtain an approximate optimal solution to this problem in polynomial time.

[0201] The specific implementation method of step S3 includes the following steps:

[0202] S3.1. Set the first stage to generate two secondary reconfiguration solutions, namely, a partition-first-then-scheduling secondary reconfiguration solution and a scheduling-first-then-partition secondary reconfiguration solution;

[0203] Furthermore, the specific implementation method of step S3.1 includes the following steps:

[0204] For the partition-first-then-scheduling two-level reconfiguration solution, an incremental repartitioning method is used. The instances with higher than average load are sorted from large to small according to load, and then their key groups are migrated to the instance with the lowest load in sequence until the load of the instance is lower than the average load after the key group migration;

[0205] For the two-level reconfiguration solution of scheduling first and partitioning later, the nodes are first sorted according to computing capabilities, and then the operators are scheduled according to network communication performance, with network communication priority. If a node cannot carry an operator pair, the scheduling is extended to the instance level, and the instance pair with the largest communication volume among the operator pairs is placed first until the node resources are unavailable, and then the next node is moved to continue placing instances.

[0206] If node resources are not fully utilized in the two-level reconfiguration solution of scheduling first and partitioning later, the load of the operator instance on the node with fragmented resources is increased until the load reaches the upper limit of the node resources.

[0207] S3.2. Based on the simulated annealing algorithm, the second-level reconfiguration solution obtained in step S3.1 is searched in the domain to obtain the final second-level reconfiguration solution.

[0208] Furthermore, the specific implementation method of step S3.2 is to partition first and then schedule the two-level reconfiguration solution ΔM obtained in step S3.1 p,s The solution of scheduling first and partitioning second-level reconfiguration ΔM s,p , using simulated annealing algorithm to search <ΔM p,s ,ΔM s,p >Nearby area, set a probability threshold p, if the solution obtained by the current cycle search is not better than the temporary optimal solution, and the exploration probability p e >p, the temporary optimal solution is updated to this solution to expand the search space; when the number of cycles of the simulated annealing algorithm reaches the maximum value, the optimal solution obtained is the final secondary reconfiguration solution.

[0209] Finally, in order to introduce elasticity mechanism to complete the multi-level collaborative reconfiguration process, it is selected to detect whether the system resources are sufficient and expand the nodes according to the actual situation before the first stage of the second-order heuristic algorithm starts; and release the completely unused nodes after the second stage. So far, the above method can obtain a multi-level reconfiguration solution combining the three reconfiguration mechanisms of partitioning, scheduling and elasticity. Based on the above second-order heuristic algorithm, the proposed multi-level collaborative reconfiguration optimization problem with comprehensive partitioning, scheduling and elasticity can obtain an approximate optimal solution in polynomial time.

[0210] Furthermore, due to the real-time requirements, the passively triggered reconfiguration algorithm cannot be applied to the edge environment because reconfiguration often occurs when the system load bursts. Even if the multi-level collaborative reconfiguration algorithm we proposed can be executed within a limited time, passive reconfiguration will inevitably generate a lot of time overhead when facing a burst load. Therefore, a load predictor is designed to predict the load change trend in the future and actively trigger reconfiguration before the load burst to reduce the impact of reconfiguration overhead on system performance. Considering the limited computing power of edge devices, the load predictor is designed and implemented based on the exponential smoothing method.

[0211] Furthermore, the load predictor is designed to adjust the reconfiguration interval time by using the exponential smoothing method. The specific implementation method is to collect the load time series data X from the end of the last reconfiguration execution to the current moment, and use the exponential smoothing method to process the data items of X one by one. i and the previous data item x i-1 The corresponding residual Calculate the corresponding residual The calculation formula is:

[0212]

[0213] Get the residual corresponding to the data item at the current moment The time interval for performing the next reconfiguration.

[0214] Furthermore, the load predictor is integrated into Flink. The load history data required by the load predictor comes from the indicator collector of TaskManager and is collected and summarized by the resource monitor of JobManager. The following are the specific steps that the system goes through from the data monitoring stage to the load prediction stage:

[0215] Step 1. Each TaskManager collects information about the load, input rate, network bandwidth and latency between other TaskManagers, and current idle resources of each instance since the last reconfiguration, and sends them to the JobManager;

[0216] Step 2. JobManager collects all monitoring data from TaskManager, including network bandwidth and transmission delay between cluster nodes, maintains load information of each node, and input rate of job graph;

[0217] Step 3. JobManager delivers the above summarized information to the exponential smoothing load predictor for calculation to obtain the interval time for the next reconfiguration execution;

[0218] Step 4. Use the reconfiguration interval as the amortized time and combine it with other indicator information as the input data of the reconfigurator to generate the reconfiguration strategy.

[0219] Based on the method of this embodiment, the method for experimental verification and the experimental verification data results are as follows:

[0220] The method of this embodiment is integrated into Flink and used as the underlying stream processing platform of the system to create a heterogeneous resource cluster consisting of 1 master node and 13 worker nodes. The cluster equipment includes 3 RK3588 and 11 RK3568. Each worker node can carry two operator instances. Redis cluster and Kafka cluster are deployed on three RK3588 nodes with sufficient resources to ensure distributed state management and reliable data source simulation. In the cluster network topology, the three RK3588 nodes communicate through wired connections, and the other 11 RK3568 nodes are interconnected by wireless networks.

[0221] The following three application workloads are selected to evaluate the performance of the proposed fine-grained migration algorithm and multi-level collaborative reconfiguration algorithm:

[0222] (1) NexMark Q4: Calculates the average auction price.

[0223] (2) NexMark Q8: Counts the number of active auctions over a period of time.

[0224] (3) Intelligent transportation application (CEP): Detect traffic jams and vehicle anomalies based on defined rules.

[0225] In addition, the following baseline methods are selected as comparison algorithms for the proposed fine-grained state migration algorithm and the MLR-algorithm:

[0226] (1) Batch migration (state migration): The state is sent to the target instance for synchronization at one time. The target instance is blocked until all state synchronization is restored.

[0227] (2)DR (reconfiguration): Only the partitioning mechanism is used for optimization.

[0228] (3)GT-Scheduler (reconfiguration): Only the scheduling mechanism is used for optimization.

[0229] (4) Mc-Stream (reconfiguration): It uses three-level mechanisms of partitioning, scheduling, and elasticity for optimization. However, it only optimizes each level of the mechanism separately, and does not consider the joint optimization of multiple levels of mechanisms.

[0230] In order to verify that the MLR algorithm has better performance than other methods in different scenarios, two representative scenarios are designed. The difference between them lies in the fluctuation trend of system load:

[0231] (1) Scenario 1: The data source input parameters such as content and rate are always unchanged, and the initial topology and running time of the application workflow are consistent after each scenario restart, so that the load fluctuation trend during each scenario operation is at a relatively stable level. After the running time reaches the set deadline, the system reconfiguration is triggered.

[0232] (2) Scenario 2: In order to simulate the dynamic factors in the mobile edge environment, the input rate of the data source is controlled to follow a random distribution within a period of time, so that the system load fluctuates with different amplitudes during this period of time. In this way, the obtained system load changes follow the following Figure 4 From the trend curve shown, we can see that the load has several peaks and valleys when the system is running.

[0233] Correctness verification of fine-grained state migration algorithm: The granularity of the state determined by the fine-grained migration algorithm depends on the hyperparameter α in formula (2-18). In order to determine the parameter value, NexMark is run in the system for a period of time before triggering the migration, and the migration state size and serialization and deserialization time are recorded in the log. Figure 5 As shown, the relationship between the amount of state data and the serialization and deserialization time can be approximated as a linear relationship, so α is set to 30.

[0234] After determining the value of the hyperparameter, it is necessary to further verify the correctness of the fine-grained migration algorithm, that is, whether the state granularity output by the algorithm is the optimal migration granularity. Run three loads, NexMark Q4, Q8, and smart transportation applications, respectively, and then perform state migration at different granularities and record the increased latency, limiting the state granularity to the interval [1,128].

[0235] When facing three types of loads, the impact of different state granularities on the latency peak generated during the migration process is as follows: Figure 6 As shown. From the results, we can analyze that the smaller the granularity, the smaller the delay peak. However, when the granularity is reduced to a certain extent, the negative effects it produces will offset or even exceed the benefits it brings. Table 1 shows the delay peak estimation values ​​output by the fine-grained migration algorithm at various granularities and the actual values ​​obtained from the experimental results, proving that the fine-grained migration algorithm has the ability to determine the optimal migration granularity.

[0236] Table 1 Comparison of the estimated delay peak and the actual delay peak of the fine-grained migration algorithm at different state granularities

[0237]

[0238] In addition to the correctness of the algorithm, it is also necessary to evaluate the impact of fine-grained migration on reconfiguration time. Prior to this, it is necessary to determine the hyperparameter μ in equation (2-19). By running the operator on a single machine without generating back pressure, the maximum input intensity when running the three loads is 4800, 9200, and 2200, respectively, and it is used to represent the processing capacity μ of the server when running the load. Similarly, run the three loads separately and trigger reconfiguration after a period of time. According to the migration state size and total reconfiguration time recorded in the log, the following can be obtained: Figure 7 The relationship between the state granularity and the total reconfiguration time is shown in Figure 2. From the experimental results, it can be seen that the theoretical optimal value output by the algorithm is roughly the same as the actual optimal value, which once again verifies the correctness of the fine-grained migration algorithm.

[0239] The batch migration method and the fine-grained migration algorithm are executed at the optimal migration granularity, and the effects of the two on the reconfiguration overhead are compared. The comparison results are shown in Table 2. The results show that the fine-grained migration algorithm can optimize the increased latency of migration without increasing the total reconfiguration time. At the same time, the latency peak can be stably reduced by more than 50% in the experimental environment. Finally, the above experimental results prove that the fine-grained migration algorithm effectively optimizes the negative impact of state migration on reconfiguration.

[0240] Table 2 Comparison of optimization degree of reconfiguration between batch migration and fine-grained migration

[0241]

[0242] Reconfiguration algorithm performance comparison: In order to verify the applicability of the multi-level collaborative reconfiguration algorithm and the benefits it brings to the system, DR, GT-scheduler and Mc-Stream were selected as comparison methods, and NexMark Q4 and smart transportation applications were selected as loads. Based on these reconfiguration algorithms and application loads, experiments were performed in scenarios 1 and 2, and the end-to-end average delay during system operation was recorded every 600ms to evaluate the optimization effect of different reconfiguration algorithms on application performance. In addition, the hyperparameters a, b, c, and d that the reconfiguration algorithm depends on are set to 0.5, 0.1, 1, and 0.1, respectively. These parameters can be customized according to user preferences.

[0243] from Figure 8It can be seen that when facing scenario 1, the MLR-algorithm can significantly reduce the reconfiguration delay by benefiting from fine-grained migration, and further reduces the end-to-end delay after reconfiguration. Among them, the MLR-algorithm has a smaller delay peak than GT-scheduler and Mc-Stream and brings greater benefits. Although the use of DR can also greatly reduce the negative impact caused by reconfiguration, DR cannot bring obvious benefits to the system through reconfiguration. In addition, Table 3 shows the end-to-end delay results of the system after optimization by four reconfiguration algorithms. It can be seen that the MLR-algorithm has the greatest degree of optimization for end-to-end delay, which is at least 7.48% higher than other reconfiguration algorithms.

[0244] Table 3 Comparison of the optimization degree of end-to-end delay of four reconfiguration algorithms

[0245]

[0246] After the system has been running for a while, the input rate of the data source is increased from 2000 tuples / second to 7000 tuples / second to trigger system reconfiguration to simulate burst load and evaluate the ability of the reconfiguration algorithm to handle back pressure. Fig. 9 The results shown in Figure 2 show that the DR method that relies solely on partitions cannot handle back pressure. GT-scheduler and Mc-Stream can solve the back pressure problem, but the latency peaks they produce show that they cannot meet real-time requirements. MLR-algorithm can not only handle back pressure, but also control the latency peak to the second level. The average end-to-end latency of the system after reconfiguration is shown in Figure 2. Fig.10 As shown in the figure, it is proved that MLR-algorithm can achieve the best optimization effect on the system when facing the back pressure problem compared with other methods.

[0247] Finally, in order to evaluate the ability of the four reconfiguration algorithms to cope with the mobile edge environment, the smart transportation application is used as the load, and the system reconfiguration is periodically triggered based on the background of scenario 2. Since the load prediction method based on exponential smoothing depends on a fixed-length time window, the window length is set to 30s, and the trigger period of the four reconfiguration algorithms in the experiment is uniformly set to this value. Fig.11 The variation of end-to-end delay during system operation is demonstrated. It is clear that the use of the MLR algorithm can minimize the average end-to-end delay during reconfiguration and further improve the degree of optimization of system performance by reconfiguration.

[0248] It is worth noting that Fig.12As shown in Figure 1, the MLR algorithm can effectively avoid unnecessary reconfiguration. In addition to the advantages of the algorithm itself, this is also due to the designed load predictor. It can actively trigger system reconfiguration to extract and cope with possible burst loads in the future, avoiding potential large reconfiguration overheads, while ensuring that the system end-to-end delay remains at a low level even if no reconfiguration is performed.

[0249] Fig.13 The average execution delay of the four reconfiguration algorithms is presented. The results show that the average execution delay of the MLR algorithm is 339ms, which is the lowest among the four reconfiguration algorithms, indicating that it has the ability to execute online and is suitable for mobile edge environments with high real-time requirements.

[0250] Embodiment 2:

[0251] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a multi-level collaborative reconfiguration method based on fine-grained state migration described in Example 1 when executing the computer program.

[0252] The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit, etc. Moreover, the processor is used to implement the steps of the above-mentioned multi-level collaborative reconfiguration method based on fine-grained state migration when executing the computer program stored in the memory.

[0253] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0254] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0255] Embodiment 3:

[0256] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a multi-level collaborative reconfiguration method based on fine-grained state migration as described in Example 1.

[0257] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned multi-level collaborative reconfiguration method based on fine-grained state migration can be implemented.

[0258] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0259] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0260] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A multi-level collaborative reconfiguration method based on fine-grained state migration, characterized in that: The steps include: S1. Based on the built-in reconfiguration mechanism of the stream computing platform, combined with the characteristics of the edge computing environment and the overhead of partitioning, scheduling and elasticity in the reconfiguration process, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed; S2. Based on the formal model of the multi-level collaborative reconfiguration optimization problem constructed in step S1, a granularity adjustment method for fine-grained state migration is constructed, and the selection of the optimal migration state granularity in the multi-level collaborative reconfiguration optimization problem is solved separately; S3. Design a multi-level collaborative reconfiguration method based on a second-order heuristic algorithm to obtain a multi-level collaborative reconfiguration solution at the three levels of comprehensive partitioning, scheduling, and flexibility to obtain an approximate optimal solution in polynomial time.

2. A multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.

1. Construct a migration state key group model; Define the set of migration state key groups key as K, where K i A key set representing the migration status of the i-th batch; Set k as the key of a certain transition state, S(k) represents the state corresponding to k, I o k is the source instance that sends the migration tuple, I t k is the target instance that receives the migration tuple, then the i-th batch state set S sent by instance j is o (i, j), and the partial state set S sent to instance x o (i,j,x) is represented as: The i-th batch state set S received by instance j t (i, j) and its partial state set S from instance x t (i,j,x) are represented as: S1.

2. Analyze the factors affecting the reconfiguration overhead of the migration state, including the time T for aligning the migration instructions align , the time of state transition T mig and the time T to return to normal processing after successful migration res , the sum of the three is the total migration time T; State transition time T mig The time T serialized by the state srl , state transfer time T trans and the time T for state deserialization desrl is composed of the following expressions: T srl +t desrl =α*(max{S o (i,j)}+max{S t (i,j)})#(2-5) Among them, α is the slope of the linear function of the consumption time on the state size; Let B(x,j) be the network transmission rate between instances x and j. Then the state transmission time T for instance j to send the state to the target instance in the i-th batch is o (i, j), and the state transfer time T for receiving the state from the source instance t (i,j) are represented as: Based on the fact that when instance j receives the state of the i-th batch of migration, it blocks the tuples coming from the upstream that will affect the state recovery, and defines M i,j is the number of tuples blocked by instance j during the i-th batch migration, expressed as: M i,j =f(k)*λ i,j *(T align +T mig )#(2-8) Among them, λ i,j represents the input flow intensity of instance j during the i-th batch migration, and f(k) represents the probability of k appearing during the i-th batch migration; Calculate the time T to return to normal processing after the i-th batch of state migration is successful res (i), the expression is: Among them, μ j is the processing capability of instance j; S1.

3. Based on the factors affecting the reconfiguration overhead of the migration state, a state migration overhead model is constructed, and the expression is: T peak (m,B,P)=T align +T mig #(2-11) Among them, T span (m,B,P) is the total time of state migration of m batches, T peak (m, B, P) is the peak latency of m batches of state migration, P is a triple used to represent the multi-level reconfiguration strategy, each element in P represents the partition, scheduling and elasticity strategy, and B is the network transmission rate; S1.4.Formulate the partitioning, scheduling and elasticity processes based on the state migration overhead model; S1.4.1.Formulate the partitioning process based on the state migration cost model: The upstream operator calculates the route based on the key of the tuple and distributes the data to the downstream operator, using the mapping relationship f p :k→i' represents a partition operation, where i' represents the index of the downstream instance; Divide the total key group K into n sub-key groups kg = {kg0, kg1, ..., kg n-1 },have With matrix M p Record the routing relationship between the subkey set kg and the instance set V, and the value of each element is defined as follows: Among them, M p [i'][j] is the subkey group kg i' and instance V j The routing mapping relationship, V j is the jth instance; Reconfiguration requires adjustment of the partitions. After the partition operation, the routing relationship between the partitioned sub-key groups and instances is represented by M' p ; then the repartitioning strategy is expressed as ΔM p =M' p -M p , ΔM p Each element takes the value of -1, 0 or 1, representing the subkey group kg i' From the example V j Migrated out, not migrated, and migrated to instance V j Three strategies; Suppose the partitioned multi-level reconfiguration strategy triple It is a placeholder empty set, indicating that scheduling strategy and elasticity strategy are not involved here; According to the repartitioning strategy, calculate the state quantity that needs to be migrated for a repartitioning. First, define the partition migration state vector Used to save the state size corresponding to each sub-key group. Next, define the partition adjustment instance index set A collection of indices representing instances that sent status during the partitioning process; Then we get the state quantity involved in the migration during the partition process. Combining formula (2-10) and formula (2-11), we can finally get the total time T for partition operation. span (m,B,P p ) and the peak latency of partition state migration T peak (m,B,P p ); S1.4.2.Formulate the scheduling process based on the state transition overhead model; Definition M s , M' s and ΔM s They represent the instance deployment before reconfiguration, the instance deployment after reconfiguration, and the scheduling strategy, respectively. s The value of the element is 0 or 1, indicating the instance V i' Whether to deploy on node N j On, ΔM s The value of the [i'][j] element is -1, 0 or 1, indicating the instance V i' From Node N j Node N that has been migrated, not migrated, or migrated j , let the triple of the multi-level reconfiguration strategy of scheduling be Define the scheduling adjustment instance index set represents the index of the instance that sends the status during the scheduling process, then the number of migration instances is V mig =|I s |; The time overhead related to instance management before and after migration is expressed as ρ*V mig ,ρ is the instance management overhead coefficient; For the migration overhead caused by scheduling, define the scheduling migration state vector Used to save the state size corresponding to each stateful instance, combined with ΔM s That is, the state quantity involved in the scheduling process is obtained according to The total time T associated with the state transition caused by B being scheduled span (m,B,P s ) and the peak delay of scheduling state transition T peak (m,B,P s ), the total time of the scheduling phase and the peak delay of the scheduling state migration are ρ*V mig +T span (m,B,P s ) and T peak (m,B,P s ); S1.4.3.Formulate modeling of elasticity process based on state migration overhead model; Defining vectors Indicates the node usage. Each element in the vector takes a value of 0 or 1, representing node N. j After an elastic operation, the usage of node resources is readjusted to The elastic strategy is expressed as Its element values ​​are -1, 0, or 1, which represent that the node is released, the node status is unchanged, and the node is removed from the resource pool. Considering that the elastic operation only adjusts the resource set, the additional overhead caused by the elastic operation is expressed as β is the overhead coefficient corresponding to the adjusted resource set; S1.

5. Constructing a multi-level reconfiguration strategy triple representing the three mechanisms of comprehensive partitioning, scheduling, and elasticity Perform overhead analysis of multi-level reconfiguration strategies; Set the time T for elastically adjusting node resource consumption elastic is a constant value, and elastic operations will not cause migration blocking; the time taken to deploy an instance in the stage is T deploy , then the total deployment time of all related instances is T deploy *V mig ; Both scheduling and partitioning involve state migration and there is migration blocking and delay peak. When the state migration batch is m, the total reconfiguration time and the delay peak corresponding to the reconfiguration strategy P are calculated by the following two formulas: T1(m,B,P)=δ p *T span (m,B,P p )+δ S *(T deploy *V mig +T span (m,B,P s ))+δ E *T elastic #(2-13) T2(m,B,P)=max{T peak (m,B,P p ),T peak (m,B,P s )}#(2-14) Among them, δ p , δ S and δ E Indicates whether the reconfiguration strategy involves partitioning, scheduling, or elastic operations. A value of 0 indicates no operation is involved, and a value of 1 indicates that an operation is involved. The reconfiguration time overhead is defined as The overhead associated with the peak latency is expressed as Where L1 is the average end-to-end delay of the application before reconfiguration; Finally, the total reconfiguration cost is calculated according to the following formula: Wherein, a and b are the first trade-off coefficient and the second trade-off coefficient respectively; Benefits of computing reconfiguration strategy: Wherein, c and d are the third trade-off coefficient and the fourth trade-off coefficient respectively; S1.

6. Based on the cost analysis of the multi-level reconfiguration strategy in step S1.5, a formal model of the multi-level collaborative reconfiguration optimization problem is constructed, which is subject to constraints (1)-(6). maximize Objective=benefit-cos#(2-17) Constraint (1) ensures that partition operations only migrate key groups from instances with this key group; constraint (2) ensures that scheduling operations only occur from nodes with this instance; constraints (3) and (4) limit partition and scheduling operations to be effectively transferred between instances and nodes, thereby ensuring the consistency of operations; constraint (5) ensures that the load of each node after reconfiguration does not exceed its processing capacity; constraint (6) limits the number of loads on the node after reconfiguration to not exceed the number of its physical resource slots.

3. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 2, characterized in that: The specific implementation method of step S2 includes the following steps: S2.

1. Set the goal to minimize the reconfiguration overhead and output the corresponding state granularity. Rewrite equation (2-11) according to the following process to obtain the delay peak T of the rewritten m batches of state migration peak (m,B,P)' expression is: S2.

2. For the total reconfiguration time, rewrite equation (2-10) from equation (2-9) to obtain the expression: Among them, Q is the sum of the state transition time of each batch, Q = α*(|S o |+|S t |)+T o +T t , in the case of uniform particle size division, it can be considered Then simplifying, we get For the simplified T span Take the derivative and find T span The value m corresponding to the minimum * for The expression for the value range of m is:

4. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 3 is characterized in that: The specific implementation method of step S3 includes the following steps: S3.

1. Set the first stage to generate two secondary reconfiguration solutions, namely, a partition-first-then-scheduling secondary reconfiguration solution and a scheduling-first-then-partition secondary reconfiguration solution; S3.

2. Based on the simulated annealing algorithm, the second-level reconfiguration solution obtained in step S3.1 is searched in the domain to obtain the final second-level reconfiguration solution.

5. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 4 is characterized in that: The specific implementation method of step S3.1 includes the following steps: For the partition-first-then-scheduling two-level reconfiguration solution, an incremental repartitioning method is used. The instances with higher than average load are sorted from large to small according to load, and then their key groups are migrated to the instance with the lowest load in sequence until the load of the instance is lower than the average load after the key group migration; For the two-level reconfiguration solution of scheduling first and partitioning later, the nodes are first sorted according to computing capabilities, and then the operators are scheduled according to network communication performance, with network communication priority. If a node cannot carry an operator pair, the scheduling is extended to the instance level, and the instance pair with the largest communication volume among the operator pairs is placed first until the node resources are unavailable, and then the next node is moved to continue placing instances. In the case where node resources are not fully utilized in the two-level reconfiguration solution of scheduling first and partitioning later, the load of the operator instance on the node with fragmented resources is increased until the load reaches the upper limit that the node resources can bear.

6. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 5, characterized in that: The specific implementation method of step S3.2 is to partition first and then schedule the two-level reconfiguration solution ΔM obtained in step S3.1 p,s The solution of scheduling first and partitioning second-level reconfiguration ΔM s,p , using simulated annealing algorithm to search <ΔM p,s ,ΔM s,p >Nearby area, set a probability threshold p, if the solution obtained by the current cycle search is not better than the temporary optimal solution, and the exploration probability p e >p, the temporary optimal solution is updated to this solution to expand the search space; when the number of cycles of the simulated annealing algorithm reaches the maximum value, the optimal solution obtained is the final secondary reconfiguration solution.

7. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 1, characterized in that: The load predictor is designed to adjust the reconfiguration interval time by using the exponential smoothing method. The specific implementation method is to collect the load time series data X from the end of the last reconfiguration execution to the current moment, and use the exponential smoothing method to process the data items of X one by one. Based on the current data item x in each iteration i and the previous data item x i-1 The corresponding residual Calculate the corresponding residual The calculation formula is: Get the residual corresponding to the data item at the current moment The time interval for performing the next reconfiguration.

8. The multi-level collaborative reconfiguration method based on fine-grained state migration according to claim 7, characterized in that: The load predictor is integrated into Flink. The load history data required by the load predictor comes from the indicator collector of TaskManager and is collected and summarized by the resource monitor of JobManager.

9. An electronic device, characterized in that: It comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of a multi-level collaborative reconfiguration method based on fine-grained state migration as described in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-level collaborative reconfiguration method based on fine-grained state migration described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Virtual network function dynamic migration method based on deep belief network resource demand forecasting

    CN108900358A

  • Cross-border cloud internal geographic boundary discovery system and method

    CN115134251A

  • Flow processing job capacity expansion and contraction scheduling method based on priority state transition

    CN115168006A

  • Flink-based multi-level collaborative reconfiguration stream processing system and processing method thereof

    CN115412501A

  • Active assurance of network slicing

    CN115918139A

Cited By

  • Edge flow system fault tolerance method based on parallel recovery

    CN121523799A