Gateway data scheduling method based on cuckoo search and Fibonacci Hash mapping

Through the combination of cuckoo search and Fibonacci hash mapping, scheduling and mapping strategies are dynamically adjusted, the problems of node overload and storage imbalance in gateway data scheduling are solved, more efficient data task allocation and storage are achieved, and the system's operating stability and resource utilization efficiency are improved.

CN120263853AInactive Publication Date: 2025-07-04HUNAN UNIV OF SCI & ENG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510418517.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing gateway data scheduling methods lack real-time perception of the dynamic resource state of nodes, resulting in uneven task allocation, node overload and abnormal growth in response time; hash mapping technology is difficult to achieve balanced mapping, and there are problems with high mapping conflict rate, local data accumulation and single-point hotspots, which affect system efficiency and stability.

Method used

The method of combining cuckoo search and Fibonacci hash mapping is adopted to construct the scheduling state association matrix and fitness function, dynamically adjust the search path and hash mapping rules, realize adaptive scheduling and mapping optimization, integrate task size, node load and response delay factors, and build a cross-module coupling optimization strategy.

Benefits of technology

Effectively avoid task blockage caused by high-latency nodes, realize the hash mapping of semantic correlation between data tasks and storage units, improve the system's scheduling robustness and data processing completeness, and reduce the risk of mapping conflicts and storage hotspots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263853A_ABST
    Figure CN120263853A_ABST
Patent Text Reader

Abstract

The invention discloses a gateway data scheduling method based on cuckoo search and Fibonacci Hash mapping. The method comprises the following steps: S1, forming a structured data set; s2, constructing a resource state data set; s3, iteratively generating a preliminary scheduling strategy; s4, inputting task numbers and scheduling instructions generated in the data distribution module into a Fibonacci Hash mapping module, and uniformly mapping the data tasks to a distributed storage unit according to a preset Hash mapping rule; s5, carrying out balance and conflict detection on a storage distribution result output by the Fibonacci Hash mapping module, and constructing a self-adaptive scheduling and mapping optimization mechanism; and S6, applying the dynamically adjusted scheduling strategy and the mapping rule to each processing node in the gateway. The method supports strategy adaptive updating under complex conditions of task type high aggregation and node delay abrupt change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gateway data scheduling, and in particular, to a gateway data scheduling method based on cuckoo search and Fibonacci hash mapping. Background Art

[0002] With the wide application of Internet of Things and edge computing technologies, various intelligent terminal devices continuously access the gateway system, and the generated data traffic shows the characteristics of high concurrency, heterogeneity, and multiple types. How to efficiently schedule and reasonably store the large-scale multi-source data accessed by the gateway has become a key factor affecting system performance and resource utilization efficiency. The existing gateway data scheduling methods mainly rely on scheduling mechanisms based on polling, weighted shortest job first, or static threshold strategies to allocate data tasks to gateway nodes according to preset rules. However, they usually ignore the dynamic resource state changes of processing nodes, lack real-time perception of factors such as the current load, processing capacity, and response latency of nodes, and are prone to problems such as uneven task allocation, node overload, or abnormal increase in response time.

[0003] In addition, existing data storage mapping technologies mostly use consistent hashing, MD5 hashing, or simple round-robin methods to map the data tasks after scheduling to storage units. In the face of drastic fluctuations in the scale of data streams or highly uneven loads of storage units, it is often difficult to achieve balanced mapping, and there are problems such as high mapping conflict rate, local data accumulation, and single-point hot spots, which seriously affect the data access efficiency and overall load stability of the system. Especially in the case of high-frequency write scenarios, traditional hash algorithms lack the ability of context perturbation and cannot flexibly adjust mapping rules according to task types, sizes, or node states, resulting in a static and rigid characteristic of storage distribution.

[0004] In summary, there are significant defects in the existing technologies for gateway multi-source data scheduling and distributed storage mapping, specifically manifested as follows: First, the scheduling strategy is static and has poor adaptability, and it is difficult to dynamically perceive the node resource state and make optimal allocation; second, the hash mapping mechanism lacks the ability to perceive data characteristics and resource states, resulting in unbalanced mapping and high conflict rate. There is an urgent need for an intelligent scheduling and mapping method with global optimization ability and dynamic feedback mechanism to effectively improve the operation efficiency and resource utilization level of the system. Summary of the Invention

[0005] An object of the present invention is to propose a gateway data scheduling method based on cuckoo search and Fibonacci hash mapping, and the present invention supports the strategy adaptive update in the case of highly aggregated task types and complex mutations of node delays.

[0006] A gateway data scheduling method based on cuckoo search and Fibonacci hash mapping according to an embodiment of the present invention includes the following steps:

[0007] S1. Collect the multi-source data stream information sets of each access device in the gateway, preprocess the multi-source data stream information sets to form a structured data set;

[0008] S2. Monitor the real-time resource status of each processing node in the gateway, collect the current load, processing capacity, response delay operation parameters of each node, and construct a resource status data set;

[0009] S3. According to the structured data set and the resource status data set, initialize the parameters of the global scheduling optimization, set the initial scheduling population and candidate scheduling solutions, use the cuckoo search algorithm to globally search for the candidate scheduling solutions, and use the real-time resource status of the nodes as the evaluation index to iteratively generate a preliminary scheduling strategy;

[0010] S4. Transmit the data task allocation information in the preliminary scheduling strategy to the data allocation module, input the task numbers and scheduling instructions generated in the data allocation module into the Fibonacci hash mapping module, and evenly map the data tasks to the distributed storage units according to the predetermined hash mapping rules;

[0011] S5. Perform balance and conflict detection on the storage distribution results output by the Fibonacci hash mapping module, obtain the distribution status and conflict status of data access in real time, and dynamically adjust the search parameters of the cuckoo search algorithm and the mapping rules of the Fibonacci hash mapping according to the detection results to construct an adaptive scheduling and mapping optimization mechanism;

[0012] S6. Apply the dynamically adjusted scheduling strategy and mapping rules to each processing node in the gateway.

[0013] Optionally, the S1 includes the following steps:

[0014] S11. Obtain the multi-source data stream information generated by each access device in the gateway, set a unified data collection time window, capture the data streams of all access devices within each time window, and construct a multi-source data stream information set D raw , defined as follows:

[0015]

[0016] where d i represents the i-th multi-source data stream information, N is the total number of data collected within the time window Δt, id i represents the access terminal identifier that generates the multi-source data stream, s i represents the size of the multi-source data stream, t i represents the collection timestamp of the multi-source data stream, c i represents the data type code corresponding to the multi-source data stream, which is used to identify the business category to which the data belongs;

[0017] S12. Clean and structure the data in the multi-source data stream information set, removing records with missing key fields, invalid timestamps, or redundant duplicates to form a structured data set D struct , where d' in the structured data set i is the cleaned and valid multi-source data stream information, satisfying the structural integrity verification rule valid(d i ) = True, that is, all fields are non-empty, and s i > 0, t i is legal, and id i is unique.

[0018] Optionally, S2 includes obtaining the identification set of all processing nodes in the gateway, setting a unified status monitoring period, collecting the real-time operation parameters of all processing nodes within each monitoring period, and constructing a resource status data set R node :

[0019]

[0020] where r j represents the resource status record of the jth processing node, M is the total number of processing nodes participating in scheduling in the gateway, nid j represents the unique identification number of the processing node, l j represents the current load rate of the processing node, defined as the ratio of the currently used processing capacity to the maximum available processing capacity, p j represents the available processing capacity of the processing node, and τ j represents the average response delay of the processing node within the current monitoring period.

[0021] Optionally, S3 includes the following steps:

[0022] S31. Based on the structured data set D struct and the resource status data set R node construct a scheduling status association matrix The scheduling status association matrix is used to measure the scheduling adaptation degree between data tasks and processing nodes:

[0023]

[0024] where Q i,j is the scheduling cost between the ith task and the jth node, is the average arrival time of the node task queue, is a type matching function, and λ1 to λ4 are fusion coefficients;

[0025] S32. According to the scheduling status association matrix Q, construct an initial population of cuckoo search Each individual represents a scheduling scheme, where represents assigning the i-th task to the -th processing node;

[0026] S33. During the cuckoo search iteration process, define a fitness function F adaptive (x k , t) to evaluate the quality of the scheduling solution. The fitness function considers the node load degree, delay increase, and task matching caused by the scheduling scheme:

[0027]

[0028] where represents the overall resource ratio of the tasks scheduled to node j in the current iteration, represents the increase in node response delay, represents the degree of adaptation violation between the task and the node capability type. φ1, φ2, and φ3 are the weight coefficients in the fitness function, and t represents the iteration round of the current cuckoo search algorithm;

[0029] S34. In each update of the cuckoo search solution, construct an adaptive perturbation step size by dynamically adjusting the search step size in combination with the node response delay

[0030]

[0031] and perform solution update:

[0032]

[0033] where represents the normalized response delay of node x k,i , α is the step size base coefficient, γ is the step size decay coefficient, Levy(β) is the Levy distribution perturbation factor, dynamically adjusting the search amplitude to avoid high-delay nodes, represents the index of the processing node to which the i-th task is assigned by the k-th individual in the t-th iteration, represents the index of the processing node to which the i-th task is assigned by the k-th individual in the (t - 1)-th iteration;

[0034] S35. Continuously iterate to perform fitness calculation and scheduling solution update. Combine the current optimal solution individual and the global solution replacement strategy to perform population evolution until the maximum iteration number T max or the scheduling convergence condition is satisfied;

[0035] S36. Define the optimal scheduling solution obtained from the final iteration as the preliminary scheduling strategy, indicating that the i-th data task in the structured dataset is processed by the Processed by a processing node.

[0036] Optionally, S4 includes the following steps:

[0037] S41. Based on the preliminary scheduling strategy x * Construct a data task allocation mapping table T map , where the data task allocation mapping table assigns the i-th task d' in the preliminary scheduling strategy i to the corresponding processing node and generates a task number tid for each task i , and the set of task numbers is denoted as T id ;

[0038] S42. Define a task hash perturbation vector to optimize the distribution balance and context awareness ability of the mapping process:

[0039]

[0040] where δ i represents the perturbation coefficient of the i-th task, and are the average size of all tasks and the average latency of all processing nodes respectively, represents the semantic deviation function between the data type c i and the target node processing preference, and κ1 to κ3 are perturbation fusion weight coefficients;

[0041] S43. Input the task number tid i and the perturbation coefficient δ i into the Fibonacci hash mapping function together to form a perturbation mapping function h ′ (tid i ):

[0042]

[0043] where, N s is the total number of storage units;

[0044] S44. Construct a mapping set from data tasks to storage units:

[0045]

[0046] where, represents the target storage unit number after mapping by the task tid i through the hash perturbation function;

[0047] S45. Detect the conflict rate of the hash mapping result and count the task write density c of each storage unit within a unit time windowj , define the storage unit conflict density vector ρ:

[0048]

[0049] where n j is the current number of write tasks for the j-th storage unit, is the average number of write tasks for all units. If ρ j > θ c , then mark this area as a high-conflict area, and θ c is the set conflict threshold;

[0050] S46. Adaptively correct the mapping results of high-conflict units to construct a conflict repair function h * (tid i ):

[0051]

[0052] where ζ is the perturbation correction amplification factor, and the corrected hash address replaces the original mapping address to form the final perturbed mapping table

[0053] Optionally, the S5 includes the following steps:

[0054] S51. Perform storage distribution statistics on the final perturbed mapping table to obtain the current task write density c of all storage units j , and calculate the global write density variance

[0055]

[0056] where c j represents the number of task writes for the j-th storage unit, is the average number of writes for all storage units.

[0057] S52. Combine the storage unit conflict density vector ρ to construct a conflict detection threshold interval Θ = [θ min , θ max , and make judgments on the balance and conflict status based on the following three types of conditions:

[0058] If the write density of a certain storage unit is higher than the conflict upper limit preset by the system, that is, max(ρ) > θ max , it means that there is serious single-point data accumulation in the system;

[0059] If the fluctuation of the write density of the overall storage distribution is greater than the preset value, that is, the global variance of the write density exceeds the variance threshold set by the system, then it satisfies It shows that the data tasks in the system are extremely unevenly distributed among multiple storage units;

[0060] If the difference between the maximum and minimum write densities among different storage units exceeds the acceptable range, i.e., |max(ρ)-min(ρ)|>θ range , it indicates that there are problems of structural skew or insufficient perturbation in the current mapping mechanism;

[0061] If any of the above conditions holds, the system will determine that there is an abnormality in the distribution result of the current Fibonacci hash mapping module, and it is necessary to trigger the subsequent dynamic adjustment mechanism;

[0062] Among them, θ var is the tolerance threshold of the write density variance set by the system, and θ range is the acceptable difference range between the maximum and minimum densities;

[0063] S53. If it is determined that there is a mapping abnormality, the dynamic perturbation parameters of the Fibonacci hash mapping module are adjusted, and the perturbation vector fusion coefficients κ1, κ2, κ3 and the perturbation correction amplification coefficient ζ are updated;

[0064] S54. Synchronously evaluate the effectiveness of the fitness function and search strategy parameters in the cuckoo search algorithm. If the delay of the corresponding processing nodes in the high-conflict area continues to increase, trigger the search strategy adjustment mechanism;

[0065] S55. Form the output rules of the adaptive scheduling and mapping optimization mechanism.

[0066] Optionally, the dynamic perturbation parameter adjustment rules are as follows:

[0067] If the global write density variance continues to rise, increase κ2, decrease κ1, and enhance the delay sensitivity;

[0068] If the high-conflict nodes are concentrated in a specific storage area, increase ζ to expand the hash perturbation correction range;

[0069] If the deviation between the task type and the node processing ability is large, enhance κ3 to increase the type adaptation constraint weight.

[0070] Optionally, the search strategy adjustment mechanism includes:

[0071] Update the priority of the fitness function weight φ2 to enhance the influence weight of the delay index in the search optimization;

[0072] Adjust the step size decay coefficient γ to enhance the repulsive ability of the search space convergence direction to high-delay nodes;

[0073] If multiple nodes enter the high-load state at the same time, reduce the Levy perturbation ratio α to shrink the search perturbation range.

[0074] Optionally, the output rules of the adaptive scheduling and mapping optimization mechanism include:

[0075] Scheduling disturbance priority type: If the storage mapping conflict is caused by a sudden change in node load, the fitness function will prioritize the delay term, and the step size disturbance will tend to avoid the loaded node;

[0076] Mapping perturbation priority: If the conflict is caused by task type concentration or hash imbalance, the Fibonacci perturbation weight and repair strategy are adjusted to achieve active guidance on the mapping side;

[0077] Linkage adjustment type: If the feedback on both the scheduling and mapping ends is abnormal, the cuckoo search and Fibonacci hash parameter linkage update mechanism are started at the same time to form cross-module coupling adjustment.

[0078] The beneficial effects of the present invention are:

[0079] (1) The present invention constructs a scheduling state association matrix and fitness function that integrates multiple factors in the cuckoo search algorithm, introduces multiple dynamic factors such as task size, node load, response delay and type matching, and comprehensively quantifies the adaptation cost between data tasks and processing nodes. In each iteration round, the standardized dynamic step size perturbation constructed based on the node response delay effectively avoids the task blocking problem caused by high-latency nodes. The mechanism can dynamically adjust the search path and individual migration direction, and achieve better global search convergence performance in the high-dimensional task scheduling space.

[0080] (2) The present invention constructs a perturbation vector that integrates task size, target node delay and type semantic deviation, and introduces it into the Fibonacci hash function to generate a perturbation mapping address, thereby realizing a hash mapping relationship with more semantic relevance between data tasks and storage units. At the same time, the hash function has nonlinear perturbation characteristics based on the golden section constant, has good distribution uniformity and high cycle coverage, and can achieve smoother write distribution without increasing storage overhead.

[0081] (3) The present invention introduces the conflict density vector and the global write variance as dynamic feedback indicators between the cuckoo search and Fibonacci hash modules, triggers the linkage adjustment mechanism of the disturbance factor and the search strategy, constructs a cross-module coupling optimization strategy, and supports the adaptive update of strategies under the situations of highly clustered task types and complex node delay mutations. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0083] Figure 1Flowchart of a gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping proposed by the present invention. Detailed implementation manners

[0084] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0085] Refer to Figure 1 , a gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping, includes the following steps:

[0086] S1. Collect the multi-source data stream information sets of each access device in the gateway, preprocess the multi-source data stream information sets to form a structured data set;

[0087] S2. Monitor the real-time resource status of each processing node in the gateway, collect the current load, processing capacity, response delay operation parameters of each node, and construct a resource status data set;

[0088] S3. According to the structured data set and the resource status data set, initialize the parameters of the global scheduling optimization, set the initial scheduling population and candidate scheduling solutions, use the cuckoo search algorithm to globally search for the candidate scheduling solutions, and use the real-time resource status of the nodes as the evaluation index to iteratively generate a preliminary scheduling strategy;

[0089] S4. Transmit the data task allocation information in the preliminary scheduling strategy to the data allocation module, input the task numbers and scheduling instructions generated in the data allocation module into the Fibonacci hashing mapping module, and evenly map the data tasks to the distributed storage units according to the predetermined hashing mapping rules;

[0090] S5. Perform balance and conflict detection on the storage distribution results output by the Fibonacci hashing mapping module, obtain the distribution situation and conflict status of data access and storage in real time, and dynamically adjust the search parameters of the cuckoo search algorithm and the mapping rules of the Fibonacci hashing mapping according to the detection results to construct an adaptive scheduling and mapping optimization mechanism;

[0091] S6. Apply the dynamically adjusted scheduling strategy and mapping rules to each processing node in the gateway.

[0092] In this embodiment, S1 includes the following steps:

[0093] S11. Obtain the multi-source data stream information generated by each access device in the gateway, set a unified data collection time window, capture the data streams of all access devices within each time window, and construct a multi-source data stream information set D raw , defined as follows:

[0094]

[0095] Among them, d i represents the i-th multi-source data stream information, N is the total number of data collected within the time window Δt, and id i represents the access terminal identifier that generates the multi-source data stream, and s i represents the size of the multi-source data stream, and t i represents the collection timestamp of the multi-source data stream, and c i represents the data type code corresponding to the multi-source data stream, which is used to identify the business category to which the data belongs;

[0096] S12. Clean and structure the data in the multi-source data stream information set, eliminate the records with missing key fields, invalid timestamps or redundant repetitions, and form a structured data set D struct , and d' i in the structured data set is the valid multi-source data stream information after cleaning, satisfying the structural integrity verification rule valid(d i ) = True, that is, all fields are non-empty, and s i > 0, t i is legal, and id i is unique.

[0097] In this embodiment, S2 includes obtaining the identifier set of all processing nodes in the gateway, setting a unified status monitoring period, collecting the real-time operation parameters of all processing nodes within each monitoring period, and constructing a resource status data set R node :

[0098]

[0099] Among them, r j represents the resource status record of the j-th processing node, M is the total number of processing nodes participating in scheduling in the gateway, nid j represents the unique identification number of the processing node, l j represents the current load rate of this processing node, which is defined as the ratio of the currently used processing capacity to the maximum available processing capacity, and p j represents the available processing capacity of this processing node, and τ j represents the average response delay of this processing node within the current monitoring period.

[0100] In this embodiment, S3 includes the following steps:

[0101] S31. Based on the structured data set D struct and the resource status data set R node construct a scheduling status association matrix The scheduling status correlation matrix is used to measure the scheduling adaptation degree between data tasks and processing nodes:

[0102]

[0103] Among them, Q i,j is the scheduling cost between the i-th task and the j-th node, is the average arrival time of the node task queue, is the type matching function, and λ1 to λ4 are fusion coefficients;

[0104] S32. Construct the initial population of the cuckoo search according to the scheduling status correlation matrix Q Each individual represents a scheduling scheme, where represents allocating the i-th task to the th processing node;

[0105] S33. During the cuckoo search iteration process, define the fitness function F adaptive (x k , t) to evaluate the quality of the scheduling solution. The fitness function considers the node load degree, delay increase, and task matching caused by the scheduling scheme:

[0106]

[0107] Among them, represents the total resource occupancy ratio of the tasks scheduled to node j in the current iteration, represents the increase in node response delay, represents the degree of adaptation violation between the task and the node capability type. φ1, φ2, and φ3 are the weight coefficients in the fitness function, and t represents the current iteration round of the cuckoo search algorithm;

[0108] S34. In each round of cuckoo search solution update, dynamically adjust the search step size in combination with the node response delay to construct an adaptive perturbation step size

[0109]

[0110] And perform solution update:

[0111]

[0112] Among them, represents the normalized response delay of node x k,i . α is the step size base coefficient, γ is the step size decay coefficient, Levy(β) is the Levy distribution perturbation factor, and dynamically adjust the search amplitude to avoid high-delay nodes, Denote the index of the processing node to which the \(i\)-th task is assigned by the \(k\)-th individual in the \(t\)-th iteration. Denote the index of the processing node to which the \(i\)-th task is assigned by the \(k\)-th individual in the \((t - 1)\)-th iteration;

[0113] S35. Continuously iterate to execute fitness calculation and scheduling solution update, and combine the current optimal solution individual and the global solution replacement strategy to perform population evolution until the maximum iteration number \(T\) max or the scheduling convergence condition is satisfied;

[0114] S36. Define the optimal scheduling solution obtained from the final iteration as the preliminary scheduling strategy, indicating that the \(i\)-th data task in the structured dataset is processed by the -th processing node.

[0115] In this embodiment, S4 includes the following steps:

[0116] S41. Based on the preliminary scheduling strategy \(x\) * Construct a data task assignment mapping table \(T\) map , where the data task assignment mapping table assigns the \(i\)-th task \(d'\) in the preliminary scheduling strategy i to the corresponding processing node and generates a task number \(tid\) for each task i , and the set of task numbers is denoted as \(T\) id ;

[0117] S42. Define a task hash perturbation vector to optimize the distribution balance and context awareness ability of the mapping process:

[0118]

[0119] where \(\delta\) i represents the perturbation coefficient of the \(i\)-th task, and are the average size of all tasks and the average latency of all processing nodes respectively, represents the semantic deviation function between the data type \(c\) i and the processing preference of the target node, and \(\kappa_1\) to \(\kappa_3\) are perturbation fusion weight coefficients;

[0120] S43. Input the task number \(tid\) i and the perturbation coefficient \(\delta\) i into the Fibonacci hash mapping function together to form a perturbation mapping function \(h\) ′ (tid i ):

[0121]

[0122] Among them, N s is the total number of storage units;

[0123] S44. Construct a mapping set from data tasks to storage units:

[0124]

[0125] Among them, represents the target storage unit number after the task tid i is mapped by the hash perturbation function;

[0126] S45. Detect the conflict rate of the hash mapping result, and count the task write density c of each storage unit within the unit time window j , and define the storage unit conflict density vector ρ:

[0127]

[0128] Among them, n j is the current number of write tasks of the j-th storage unit, is the average number of write tasks of all units. If ρ j > θ c , then mark this area as a high-conflict area, and θ c is the set conflict threshold;

[0129] S46. Adaptively correct the mapping result of the high-conflict unit, and construct a conflict repair function h * (tid i ):

[0130]

[0131] Among them, ζ is the perturbation correction amplification factor, and the corrected hash address replaces the original mapped address to form the final perturbation mapping table

[0132] In this embodiment, S5 includes the following steps:

[0133] S51. Perform storage distribution statistics on the final perturbation mapping table to obtain the current task write density c of all storage units j , and calculate the global write density variance

[0134]

[0135] Among them, c j represents the number of task writes of the j-th storage unit, is the average number of writes of all storage units.

[0136] S52. Construct a conflict detection threshold interval Θ = [θ min , θ max by combining the storage unit conflict density vector ρ, and determine the balance and conflict status based on the following three types of conditions:

[0137] If the write density of a certain storage unit is higher than the conflict upper limit preset by the system, that is, max(ρ) > θ max , it means that there is serious single-point data accumulation in the system;

[0138] If the fluctuation of the write density of the overall storage distribution is greater than the preset value, that is, the global variance of the write density exceeds the variance threshold set by the system, then it indicates that the data tasks in the system are extremely unbalanced among multiple storage units;

[0139] If the difference between the maximum and minimum write densities among different storage units exceeds the acceptable range, that is, |max(ρ) - min(ρ)| > θ range , it indicates that there is a problem of structural skew or insufficient perturbation in the current mapping mechanism;

[0140] If any of the above conditions is satisfied, the system will determine that the distribution result of the current Fibonacci hash mapping module is abnormal and trigger the subsequent dynamic adjustment mechanism;

[0141] Among them, θ var is the variance tolerance threshold of the write density set by the system, and θ range is the acceptable difference range between the maximum and minimum densities;

[0142] S53. If it is determined that there is a mapping abnormality, adjust the dynamic perturbation parameters of the Fibonacci hash mapping module, and update the perturbation vector fusion coefficients κ1, κ2, κ3 and the perturbation correction amplification factor ζ;

[0143] S54. Synchronously evaluate the effectiveness of the fitness function and search strategy parameters in the cuckoo search algorithm. If the delay of the processing node corresponding to the high-conflict area continues to increase, trigger the search strategy adjustment mechanism;

[0144] S55. Form an output rule for the adaptive scheduling and mapping optimization mechanism.

[0145] In this embodiment, the adjustment rules for the dynamic perturbation parameters are as follows:

[0146] If the global write density variance continues to rise, increase κ2, decrease κ1, and enhance the delay sensitivity;

[0147] If the high-conflict nodes are concentrated in a specific storage area, increase ζ and expand the hash perturbation correction range;

[0148] If the task type deviates greatly from the node processing capacity, κ3 is enhanced to increase the type adaptation constraint weight.

[0149] In this implementation, the search strategy adjustment mechanism includes:

[0150] Update the fitness function weight φ2 priority to enhance the influence of the delay indicator in search optimization;

[0151] Adjust the step size attenuation coefficient γ to enhance the search space convergence direction's ability to exclude high-latency nodes;

[0152] If multiple nodes enter a high-load state at the same time, reduce the Levy disturbance ratio α and shrink the search disturbance range.

[0153] In this implementation, the output rules of the adaptive scheduling and mapping optimization mechanism include:

[0154] Scheduling disturbance priority type: If the storage mapping conflict is caused by a sudden change in node load, the fitness function will prioritize the delay term, and the step size disturbance will tend to avoid the loaded node;

[0155] Mapping perturbation priority: If the conflict is caused by task type concentration or hash imbalance, the Fibonacci perturbation weight and repair strategy are adjusted to achieve active guidance on the mapping side;

[0156] Linkage adjustment type: If the feedback on both the scheduling and mapping ends is abnormal, the cuckoo search and Fibonacci hash parameter linkage update mechanism are started at the same time to form cross-module coupling adjustment.

[0157] Embodiment 1:

[0158] In December 2024, in the industrial park of City A, a smart manufacturing company deployed an edge intelligent gateway system to process the real-time data streams generated by a large number of heterogeneous devices in the workshop. The system consists of 10 processing nodes and 28 distributed storage units, and runs high-frequency data scheduling and distribution tasks.

[0159] During a trial run, the company's technical team found that between 10:30 and 11:30 a.m. every day, the device data processing system frequently encountered problems such as "single-node task accumulation", "surge in storage conflict rate", and "partial data packet loss and retransmission". For example, on the morning of December 6, 2024, node N7 had seven consecutive task queue times exceeding 80ms in just 20 minutes, and the number of hash collisions between task write units #13 and #14 was as high as 56 times / minute, far exceeding the system design threshold (20 times / minute), resulting in delayed writing of some camera video clips, triggering a cache overflow on the device side and causing data frame loss.

[0160] After the problem occurred, the technical staff decided to replace the original fixed-weight scheduling + consistent hashing mapping mechanism with the present invention, and a control deployment was carried out from December 7th to December 10th.

[0161] At around 10:30 am on December 8th, the system began to collect the data task flows of each device in real time. During the implementation process, the gateway detected that the upload frequency from the terminal device "AGV_TRANSPORT_09" was significantly higher than that of other terminals. On average, 86 control feedback data were sent per minute, and the data size was between 32KB - 49KB. The type code was "C3" (indicating status control data). At this time, there were also stable data inputs from the terminals "TEMP_NODE_17" and "CAMERA_12" in the gateway.

[0162] The system entered the scheduling stage and first formed the following original data task sample set:

[0163]

[0164] Meanwhile, the system monitored the resource status of the processing nodes in real time as follows:

[0165] Node ID Current Load (%) Remaining Processing Capacity (GFLOPS) Average Response Delay (ms) N1 68.2 114.2 59.4 N2 45.1 132.6 48.9 N7 89.5 88.7 102.6 N9 56.4 107.1 60.7

[0166] Under the traditional scheduling scheme, T10982 - T10986 were continuously assigned to the N7 node with stronger processing ability but gradually increasing latency. Since it had just processed a batch of high-definition video streams and the load was not fully digested, the system latency soared rapidly, and finally, the task processing of T10997, T10998, and T11000 failed during the period from 10:31:06 to 10:31:12, triggering timeout readjustment.

[0167] After the deployment of the method of the present invention, during the scheduling process of the same batch of tasks, the cuckoo search algorithm dynamically analyzed the latency risk of N7 and automatically adjusted the step size perturbation coefficients (α = 1.2, γ = 0.65), making the search direction tend to the low-latency nodes N2 and N1. Finally, T10982 was assigned to N2, and T10984 and T10985 were split and assigned to N1 and N9 respectively, effectively avoiding processing latency.

[0168] After the scheduling was completed, the system entered the hashing mapping stage. Taking task T10984 as an example, its task number tid = 10984, and the perturbation coefficient δ10984 was calculated as follows:

[0169] δ10984 = 0.35×(226 / 120)+0.4×(60.7 / 68.9)+0.25×η(V1,N9)

[0170] ≈0.658;

[0171] The mapping address generated by the Fibonacci perturbation hash function is as follows:

[0172] h′(10984) = 28×(10984 + 0.658)×1 / φ mod 1 ≈ 13;

[0173] The target unit for storing the mapping task is S13.

[0174] Subsequently, the system discovers that within the next 2 - minute window, unit #13 continuously receives 13 large - scale video frame tasks. The write density is 2.63 times the average value, and the conflict density ρ13 > θc (the system threshold is 1.75), triggering the adaptive mapping correction mechanism. ζ is dynamically increased to 1.4, and h*(10984) switches the target mapping to unit #17 to relieve the hot - spot pressure.

[0175] The data recorded continuously for 3 days from December 8th to 10th is summarized as follows:

[0176] Indicator Original Traditional Method Method of the Present Invention Improvement Situation Peak Scheduling Delay (ms) 119.7 67.3 ↓43.8% Average Response Failure Rate 3.7% 0.9% ↓75.7% Single-Point Node Load Concentration Rate 27.2% 10.8% ↓60.3% Hash Collision Density ρmax 2.81 1.29 ↓54.1% Number of Mapping Abnormality Adjustments 0 (Not Supported) 7 times per hour Implement Dynamic Feedback Write Data Completion Rate 94.3% 99.2% ↑5.2%

[0177] Through the scheduling mechanism of the present invention, in a real high - concurrency environment, the system can sensitively perceive the node load and response latency, and dynamically avoid "hot - spot" nodes through the search strategy, effectively avoiding task accumulation and timeout re - adjustment problems. At the same time, the Fibonacci perturbation hash mechanism combined with semantic deviation perception makes the data task distribution more balanced, and actively triggers the hash address repair logic in the scenario with high write conflicts, greatly improving the data - processing integrity rate and scheduling robustness of the system.

[0178] In summary, in the real - world industrial - level edge gateway scheduling application, the present invention solves the problems of "scheduling blind area", "storage hot - spot", and "lack of feedback adjustment" in the traditional method, and has good practical deployment effects and general promotion value.

[0179] The present invention constructs a scheduling state correlation matrix and a fitness function that integrate multiple factors in the cuckoo search algorithm, introduces multiple dynamic factors such as task size, node load, response latency, and type matching, comprehensively quantifies the adaptation cost between data tasks and processing nodes. In each iteration round, the standardized dynamic step - size perturbation constructed based on the node response latency effectively avoids the task - blocking problem caused by high - latency nodes. The mechanism can dynamically adjust the search path and the individual migration direction, and achieve better global search convergence performance in the high - dimensional task - scheduling space.

[0180] The present invention constructs a perturbation vector that combines the task size, the target node delay, and the type semantic deviation degree, and introduces it into the Fibonacci hash function to generate a perturbed mapping address, realizing a more semantically relevant hash mapping relationship between the data task and the storage unit. At the same time, based on the non-linear perturbation characteristics of the golden ratio constant, the hash function has good distribution uniformity and high cycle coverage, and can achieve a smoother write distribution without increasing the storage overhead.

[0181] The present invention introduces a conflict density vector and a global write variance as dynamic feedback indicators between the cuckoo search and the Fibonacci hash module, triggers a linkage adjustment mechanism for the perturbation factor and the search strategy, constructs a cross-module coupling optimization strategy, and supports the adaptive update of the strategy in the case of highly aggregated task types and complex mutations in node delays.

[0182] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.

Claims

1. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping, characterized in that It includes the following steps: S1. Collect the multi-source data stream information sets of each access device in the gateway, preprocess the multi-source data stream information sets to form a structured data set; S2. Monitor the real-time resource status of each processing node in the gateway, collect the current load, processing capacity, response delay operation parameters of each node, and construct a resource status data set; S3. According to the structured data set and the resource status data set, initialize the parameters of global scheduling optimization, set the initial scheduling population and candidate scheduling solutions, use the cuckoo search algorithm to globally search for the candidate scheduling solutions, and use the real-time resource status of the nodes as the evaluation index to iteratively generate a preliminary scheduling strategy; S4. Transmit the data task allocation information in the preliminary scheduling strategy to the data allocation module, input the task numbers and scheduling instructions generated in the data allocation module into the Fibonacci hash mapping module, and evenly map the data tasks to the distributed storage units according to the predetermined hash mapping rules; S5. Perform balance and conflict detection on the storage distribution results output by the Fibonacci hash mapping module, obtain the distribution situation and conflict status of data access in real time, and dynamically adjust the search parameters of the cuckoo search algorithm and the mapping rules of the Fibonacci hash mapping according to the detection results to construct an adaptive scheduling and mapping optimization mechanism; S6. Apply the dynamically adjusted scheduling strategy and mapping rules to each processing node in the gateway.

2. The gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 1, characterized in that, The S1 includes the following steps: S11. Obtain the multi-source data stream information generated by each access device in the gateway, set a unified data collection time window, capture the data stream of all access devices within each time window, and construct a multi-source data stream information set D raw , which is defined as follows: Among them, d i represents the i-th multi-source data stream information, N is the total number of data collected within the time window Δt, and id i represents the access terminal identifier that generates the multi-source data stream, s i represents the size of the multi-source data stream, t i represents the collection timestamp of the multi-source data stream, c i represents the data type code corresponding to the multi-source data stream, which is used to identify the business category to which the data belongs; S12. Clean and structure the data in the multi-source data stream information set, removing records with missing key fields, invalid timestamps, or redundant duplicates to form a structured data set D struct , where d' in the structured data set i is the cleaned and valid multi-source data stream information, satisfying the structural integrity verification rule valid(d i ) = True, that is, all fields are non-empty, and s i > 0, t i is legal, and id i is unique.

3. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 1, characterized in that, The S2 includes obtaining the identity set of all processing nodes in the gateway, setting a unified status monitoring period, collecting the real-time operation parameters of all processing nodes within each monitoring period, and constructing a resource status data set R node : Among them, r j represents the resource status record of the j-th processing node, M is the total number of processing nodes participating in scheduling in the gateway, and nid j represents the unique identification number of the processing node, and l j represents the current load rate of the processing node, defined as the ratio of the currently used processing capacity to the maximum available processing capacity, and p j represents the available processing capacity of the processing node, and τ j represents the average response delay of the processing node during the current monitoring period.

4. A gateway data scheduling method based on cuckoo search and Fibonacci hash mapping according to claim 3, characterized in that The S3 includes the following steps: S31. Based on the structured data set D struct and the resource status data set R node construct a scheduling status association matrix The scheduling status association matrix is used to measure the scheduling adaptation degree between data tasks and processing nodes: Among them, Q i,j is the scheduling cost of the i-th task and the j-th node, is the average arrival time of the node task queue, is the type matching function, and λ1 to λ4 are fusion coefficients; S32. Construct the initial population of the cuckoo search according to the scheduling status association matrix Q Each individual represents a scheduling scheme, where represents that the i-th task is assigned to the th processing node; S33. During the iteration process of the cuckoo search, define a fitness function F for evaluating the quality of the scheduling solution adaptive (x k , t), where the fitness function takes into account the node load, delay increase, and task matching caused by the scheduling scheme: Among them, represents the total resource occupancy ratio of the tasks scheduled to node j in the current iteration, represents the increase in node response delay, represents the degree of adaptation violation between the task and the node capability type. φ1, φ2, and φ3 are the weight coefficients in the fitness function, and t represents the iteration round of the current cuckoo search algorithm; S34. In each round of cuckoo search solution update, the search step size is dynamically adjusted in combination with the node response delay to construct an adaptive perturbation step size And perform solution update: Among them, represents the normalized response delay of node x k,i , α is the step base coefficient, γ is the step decay coefficient, Levy(β) is the Levy distribution perturbation factor, dynamically adjusting the search amplitude to avoid high-delay nodes, represents the processing node index to which the i-th task is assigned in the t-th iteration of the k-th individual, represents the processing node index to which the i-th task is assigned in the (t-1)-th iteration of the k-th individual; S35. Continuously and iteratively perform fitness calculation and scheduling solution update. Combine the current optimal solution individual and the global solution replacement strategy to perform population evolution until the maximum number of iterations T max or the scheduling convergence condition is met; S36. Define the optimal scheduling solution obtained from the final iteration as the preliminary scheduling strategy, indicating that the i-th data task in the structured dataset is processed by the th processing node.

5. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 4, characterized in that, The S4 includes the following steps: S41. Based on the preliminary scheduling policy x * Construct a data task allocation mapping table T map , the data task allocation mapping table allocates the i-th task d′ in the preliminary scheduling policy i to the corresponding processing node and generates a task number tid for each task i , the set of task numbers is denoted as T id ; S42. Define the task hash perturbation vector Optimize the distribution balance and context awareness of the mapping process: Among them, δ i represents the perturbation coefficient of the i-th task, and are respectively the average size of all tasks and the average latency of all processing nodes, represents the semantic deviation function between data type c i and the processing preference of the target node, and κ1 to κ3 are perturbation fusion weight coefficients; S43. Input the task number tid i and the perturbation coefficient δ i into the Fibonacci hash mapping function together to form the perturbation mapping function h ′ (tid i ): Among them, N s is the total number of storage units; S44. Construct a mapping set from data tasks to storage units: Among them, represents the task tid i The target storage unit number after being mapped by the hash perturbation function; S45. Detect the collision rate of the hash mapping result, and count the task write density c of each storage unit within the unit time window j , and define the storage unit collision density vector ρ: where n j is the current number of write tasks for the j-th storage unit, is the average number of write tasks for all units. If ρ j > θ c , then mark this area as a high-conflict area, and θ c is the set conflict threshold; S46. Adaptively correct the mapping result of high-conflict units to construct a conflict repair function h * (tid i ): Among them, ζ is the amplification coefficient of perturbation correction. The corrected hash address replaces the original mapped address to form the final perturbed mapping table 6. The gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 5, characterized in that The S5 includes the following steps: S51. Store distribution statistics for the final reflection perturbation mapping table Perform storage distribution statistics to obtain the current task write density c of all storage units j , and calculate the global write density variance Among them, c j represents the number of task writes to the j-th storage unit, and is the average number of writes for all storage units. S52. Combine the storage unit conflict density vector ρ to construct a conflict detection threshold interval Θ = [θ min , θ max , and determine the balance and conflict status based on the following three types of conditions: If the write density of a certain storage unit is higher than the conflict upper limit preset by the system, i.e., max(ρ)>θ max , it indicates that there is serious single-point data accumulation in the system; If the write density fluctuation of the overall storage distribution is greater than the preset value, that is, the global variance of the write density exceeds the variance threshold set by the system, then it satisfies indicating that the data tasks in the system are extremely unevenly distributed among multiple storage units; If the difference between the maximum and minimum write densities among different memory cells exceeds the acceptable range, i.e., |max(ρ)-min(ρ)|>θ range , it indicates that there are problems with structural skew or insufficient perturbation in the current mapping mechanism; If any of the above conditions is satisfied, the system will determine that there is an abnormality in the distribution result of the current Fibonacci hash mapping module, and it is necessary to trigger the subsequent dynamic adjustment mechanism; where θ var is the tolerance threshold of the write density variance set by the system, and θ range is the acceptable difference range between the maximum and minimum densities; S53. If it is determined that there is a mapping abnormality, adjust the dynamic perturbation parameters of the Fibonacci hash mapping module, and update the perturbation vector fusion coefficients κ1, κ2, κ3 and the perturbation correction amplification coefficient ζ; S54. Synchronously evaluate the effectiveness of the fitness function and search strategy parameters in the cuckoo search algorithm. If the delay of the processing nodes corresponding to the high-conflict area continues to increase, trigger the search strategy adjustment mechanism; S55. Form the output rules of the adaptive scheduling and mapping optimization mechanism.

7. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 6, characterized in that, The adjustment rules of the dynamic perturbation parameters are as follows: If the global write density variance continues to rise, increase κ2, decrease κ1, and enhance delay sensitivity; If the high-conflict nodes are concentrated in a specific storage area, increase ζ and expand the hash perturbation correction range; If the deviation between the task type and the node processing capacity is large, enhance κ3 and increase the type adaptation constraint weight.

8. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 6, characterized in that, The search strategy adjustment mechanism includes: Update the priority of the fitness function weight φ2 to enhance the influence weight of the delay index in the search optimization; Adjust the step size attenuation coefficient γ to enhance the repulsive ability of the search space convergence direction to high-delay nodes; If multiple nodes enter the high-load state at the same time, reduce the Lévy perturbation ratio α and shrink the search perturbation range.

9. A gateway data scheduling method based on cuckoo search and Fibonacci hashing mapping according to claim 6, characterized in that, The output rules of the adaptive scheduling and mapping optimization mechanism include: Scheduling perturbation priority type: If the storage mapping conflict is caused by a sudden change in node load, the fitness function preferentially strengthens the delay term, and the step size perturbation tends to avoid the load nodes; Mapping perturbation priority type: If the conflict stems from the task type set or hash imbalance, adjust the Fibonacci perturbation weight and repair strategy to achieve active guidance at the mapping end; Linkage adjustment type: If the feedback from both the scheduling and mapping ends is abnormal, simultaneously start the cuckoo search and Fibonacci hash parameter linkage update mechanism to form cross-module coupling adjustment.

Citation Information

Cited By

  • NetCDF multi-source ionosphere data reconstruction method and device

    CN120541388A