Novel spatio-temporal stream data distributed computing load balancing method and system, terminal and storage medium
By building an adaptive quadtree and scalable ring time wheel, combining the learnable spatio-temporal embedding layer and multi-head attention mechanism, high-precision future distribution prediction results of spatio-temporal stream data are generated, and computing resource allocation is allocated based on the adversarial network, the problem of data skew in spatio-temporal stream data processing is solved and the computing processing efficiency is improved.
Patent Information
- Application Number
- CN202510482114.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
When processing spatiotemporal stream data, the prior art fails to effectively consider the distribution rules of spatiotemporal stream data in time, resulting in data tilt and reducing calculation and processing efficiency.
By constructing an adaptive quadtree structure for spatial dimension division, combining scalable ring time wheel for time dimension modeling, bit crossover operation generates space-time unique identifiers, and uniformly maps the spatiotemporal stream data to the distributed computing node. At the same time, a learnable spatio-temporal embedding layer and multi-head attention mechanism are introduced to build a timing prediction model to generate high-precision future spatio-temporal flow distribution prediction results, and generate optimal computing resource allocation requirements based on the adversarial network to achieve optimal data division and elastic scheduling of computing resources.
It effectively alleviates the data tilt phenomenon, improves the computing and processing efficiency of spatiotemporal stream data, and realizes dynamic load balancing of computing nodes.
Smart Images

Figure CN120011084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a novel space-time stream data distributed computing load balancing method, system, terminal and computer-readable storage medium. Background Art
[0002] At present, widely distributed observation and perception facilities generate massive spatiotemporal flow data with spatial attributes and flow characteristics such as sequential arrival and infinite increase every moment. These spatiotemporal flow data have the characteristics of huge data volume and uneven spatial distribution. How to perform efficient real-time calculation on spatiotemporal flow data and quickly extract information from huge data has gradually become a hot issue in current research.
[0003] Using a distributed computing framework to distribute huge amounts of data to different computing nodes for processing is currently the mainstream method for improving the processing speed of big data. However, due to the uneven spatial distribution of spatiotemporal data, the use of a distributed processing framework can easily lead to a large difference in the amount of data processed by different computing nodes, which is the phenomenon of data skew. This will greatly reduce the overall computing efficiency of the cluster and increase the latency of cluster computing. At present, the main methods for dealing with data skew are: first, extract samples from existing data and balance the server load according to the sample situation; second, rebuild the index according to the density of the partition index, improve the balance of the index to improve cluster performance; third, let the idle computing nodes capture and process some of the data that is not processed by the busy nodes to improve the overall efficiency of the cluster. However, the load balancing implemented based on the above methods does not take into account the temporal distribution law of spatiotemporal streaming data, and cannot effectively alleviate the data skew phenomenon.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a new type of distributed computing load balancing method, system, terminal and computer-readable storage medium for spatiotemporal stream data, aiming to solve the problem that the existing technology achieves load balancing but does not take into account the temporal distribution law of spatiotemporal streaming data, cannot effectively alleviate the data skew phenomenon, and leads to low efficiency in spatiotemporal stream data computing and processing.
[0006] To achieve the above object, the present invention provides a novel space-time stream data distributed computing load balancing method, the novel space-time stream data distributed computing load balancing method comprising the following steps: Construct an adaptive quadtree structure for spatial dimension partitioning, combine with a scalable ring time wheel for time dimension modeling, use bit crossover operation to generate a unique spatiotemporal identifier, and evenly map the spatiotemporal flow data to distributed computing nodes based on the unique spatiotemporal identifier; The spatiotemporal stream data is converted into a multi-dimensional feature tensor, a learnable spatiotemporal embedding layer is introduced to capture the spatiotemporal correlation, and a multi-head attention mechanism is used to aggregate multi-scale features to construct a time series prediction model to generate a high-precision prediction result of the future distribution of the spatiotemporal stream data; Based on the adversarial network combined with the high-precision prediction results, the optimal computing resource allocation requirements are generated, the spatiotemporal flow data is optimally divided, an elastic scheduling strategy for distributed computing resources is developed, and distributed node status data is collected in real time. The data partitioning and task allocation strategies are dynamically adjusted to form a closed-loop optimization system to continuously optimize the load balancing performance.
[0007] In addition, to achieve the above-mentioned purpose, the present invention also provides a novel space-time stream data distributed computing load balancing system, wherein the novel space-time stream data distributed computing load balancing system comprises: The module for adaptive network partitioning and distributed node mapping of spatiotemporal stream data is used to construct an adaptive quadtree structure for spatial dimension partitioning, combine a scalable annular time wheel for time dimension modeling, generate a spatiotemporal unique identifier using bit crossover operation, and evenly map the spatiotemporal stream data to distributed computing nodes based on the spatiotemporal unique identifier; A spatiotemporal stream data feature deep modeling and distribution prediction module, which is used to convert the spatiotemporal stream data into a multi-dimensional feature tensor, introduce a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and use a multi-head attention mechanism to aggregate multi-scale features to build a time series prediction model to generate high-precision prediction results for the future distribution of spatiotemporal stream data; The distributed computing resource demand deduction and load balancing implementation module is used to generate optimal computing resource allocation requirements based on the adversarial network combined with the high-precision prediction results, optimally divide the spatiotemporal flow data, develop a flexible scheduling strategy for distributed computing resources, and collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, and form a closed-loop optimization system to continuously optimize load balancing performance.
[0008] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a new type of spatiotemporal stream data distributed computing load balancing program stored in the memory and run on the processor, and when the new type of spatiotemporal stream data distributed computing load balancing program is executed by the processor, the steps of the new type of spatiotemporal stream data distributed computing load balancing method as described above are implemented.
[0009] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a new type of spatiotemporal stream data distributed computing load balancing program, and when the new type of spatiotemporal stream data distributed computing load balancing program is executed by a processor, the steps of the new type of spatiotemporal stream data distributed computing load balancing method as described above are implemented.
[0010] In the present invention, an adaptive quadtree structure is constructed to perform spatial dimension division, a scalable annular time wheel is combined to perform time dimension modeling, a space-time unique identifier is generated by bit crossover operation, and the space-time flow data is evenly mapped to distributed computing nodes based on the space-time unique identifier; the space-time flow data is converted into a multidimensional feature tensor, a learnable space-time embedding layer is introduced to capture the space-time correlation, and a multi-head attention mechanism is used to aggregate multi-scale features, and a time series prediction model is constructed to generate a high-precision prediction result of the future distribution of space-time flow data; based on the adversarial network combined with the high-precision prediction result, the optimal computing resource allocation requirement is generated, the space-time flow data is optimally divided, and a flexible scheduling strategy for distributed computing resources is developed, and distributed node status data is collected in real time, and data partitioning and task allocation strategies are dynamically adjusted to form a closed-loop optimization system to continuously optimize load balancing performance. The present invention can effectively mine future rules from space-time flow data, realize dynamic load balancing for computing nodes, effectively alleviate data tilt phenomenon, and improve the efficiency of space-time flow data calculation and processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a flow chart of a preferred embodiment of the novel time-space stream data distributed computing load balancing method of the present invention; Figure 2 It is a schematic diagram of the overall process of realizing load balancing in a preferred embodiment of the novel time-space stream data distributed computing load balancing method of the present invention; Figure 3 It is a schematic diagram of the adaptive network division of space-time stream data in a preferred embodiment of the novel space-time stream data distributed computing load balancing method of the present invention; Figure 4 It is a schematic diagram of deep modeling of spatiotemporal stream data features in a preferred embodiment of the novel spatiotemporal stream data distributed computing load balancing method of the present invention; Figure 5 It is a flow chart of distributed computing resource demand deduction and load balancing implementation in a preferred embodiment of the novel time-space stream data distributed computing load balancing method of the present invention; Figure 6 It is a structural diagram of a preferred embodiment of the novel time-space stream data distributed computing load balancing system of the present invention; Figure 7 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0012] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0013] The present invention models and predicts the timing characteristics of spatiotemporal data through a deep learning model, thereby achieving efficient deduction of future spatiotemporal distribution and optimal allocation of computing resources. The present invention mainly includes three core steps: first, an adaptive quadtree structure is constructed through an improved Hilbert space filling curve to realize spatial dimension division, and a scalable ring time wheel is combined to complete the time dimension modeling. At the same time, a bit crossing operation is used to generate a 64-bit spatiotemporal unique identifier to support efficient data positioning and retrieval. Finally, a hash function is designed based on the identifier to evenly map the data to distributed computing nodes. Compared with the traditional Hilbert curve, in the spatial dimension, the improved curve of the present invention realizes adaptive division under non-uniform data distribution through a dynamic adjustment mechanism, overcoming the defects of traditional fixed-order curves in over-division in sparse areas or insufficient division in dense areas; in terms of spatiotemporal coding, the optimized space filling path combined with the bit crossover operation significantly improves the local characteristics of the identifier, so that adjacent spatial entities still maintain numerical proximity after encoding, and this feature is much better than the spatial jump problem caused by traditional curves in high-dimensional mapping; at the same time, through the collaborative design with the scalable time wheel, the decoupled coding of the spatiotemporal dimension is realized, solving the dimensional coupling problem faced by traditional methods in spatiotemporal joint query; finally, the identifier generated based on the improved curve has better discrete characteristics, which provides a better load balancing basis for distributed hash mapping and avoids the hot spot problem of traditional methods in data skew scenarios. Secondly, the spatiotemporal flow data is converted into a multi-dimensional feature tensor, a learnable spatiotemporal embedding layer is introduced to capture the spatiotemporal correlation, and multi-scale features are aggregated through a multi-head attention mechanism to build a time series prediction model to generate high-precision future spatiotemporal flow distribution prediction results; finally, the optimal computing resource requirements are deduced based on the spatiotemporal conditions, and the improved K-means++ algorithm is used for dynamic data partitioning, and a flexible scheduling strategy is designed to achieve sub-millisecond task migration. A load balancing feedback controller is developed to adjust the task allocation strategy in real time, forming a closed-loop optimization system to continuously improve load balancing performance and system stability. This method effectively alleviates the data skew problem in distributed systems and significantly improves the computing resource utilization and overall computing efficiency in large-scale spatiotemporal flow data scenarios.
[0014] The novel time-space stream data distributed computing load balancing method described in the preferred embodiment of the present invention is as follows: Figure 1 and Figure 2As shown in the (overall process diagram), the novel time-space stream data distributed computing load balancing method includes the following steps: Step S10, construct an adaptive quadtree structure for spatial dimension division, combine with a scalable ring time wheel for time dimension modeling, use bit crossover operation to generate a spatiotemporal unique identifier, and map the spatiotemporal stream data evenly to distributed computing nodes based on the spatiotemporal unique identifier (adaptive network division of spatiotemporal stream data and distributed node mapping).
[0015] Specifically, step S10 of the present invention mainly includes: S11. Spatial dimension division: According to the spatial scope of the study area, an improved Hilbert space filling curve is used to construct an adaptive quadtree structure. The data distribution density is monitored in real time according to the density adaptive quadtree division, and the grid resolution is dynamically adjusted to adapt to the change of data density. S12, time dimension modeling: build a scalable ring time wheel, set basic time segments, and dynamically adjust the time segment granularity according to the data flow rate; S13, generation of spatiotemporal unique identifier: the spatial Hilbert code and the time wheel position are bit-interleaved to generate a 64-bit spatiotemporal unique identifier, which supports fast data location and retrieval with O(1) complexity (O(1) complexity means that the time complexity of the algorithm is at a constant level, that is, the execution time of the algorithm is fixed regardless of the input scale); S14, distributed node mapping: Based on the unique spatiotemporal identifier, a hash function is designed to evenly map the spatiotemporal flow data to distributed computing nodes as the initial state of the prediction.
[0016] Among them, S11 (spatial dimension division) specifically includes: Spatial dimension division is the basis for efficient spatiotemporal data management. Through the improved Hilbert space filling curve and adaptive quadtree structure, the grid resolution can be dynamically adjusted to adapt to the changes in data distribution density, thereby optimizing storage and computing efficiency.
[0017] First, an improved Hilbert space filling curve is generated. The Hilbert curve is a fractal curve that can map a two-dimensional space into a one-dimensional sequence while maintaining spatial locality. Assume that the spatial range of the study area is a two-dimensional plane [0, L x ]×[0,L y ], where L x and L y Respectively represent the length and width of the study area on the two-dimensional plane, and divide the two-dimensional plane into an N×N uniform grid, where N represents the number of rows and columns of the grid, and N=2 k, k is a positive integer, the recursive construction formula of the Hilbert curve is as follows: ; in, Indicate point The one-dimensional mapping value on the Hilbert curve, Represents a recursive call to divide the current quadrant into 4 sub-quadrants. , , and Represent four sub-areas respectively; the core advantage of the Hilbert curve is that it can maintain the continuity of spatially adjacent points on the curve to the greatest extent, which is very beneficial for subsequent data retrieval and partitioning.
[0018] Then, an adaptive quadtree is constructed. An adaptive quadtree is a dynamically adjusted tree structure that can dynamically adjust the grid resolution according to the data distribution density. The specific implementation process is as follows: 1) During initialization, the entire study area is divided into a root node.
[0019] 2) Define a threshold value θ, and control the maximum number of data points in each grid according to the threshold value θ. If the number of data points in a grid exceeds the threshold value θ, the grid is further subdivided into 4 subgrids; otherwise, it remains unchanged.
[0020] 3) Use a recursive algorithm to construct an adaptive quadtree.
[0021] In practical applications, the threshold θ can be dynamically adjusted according to system resources and performance requirements to balance computational complexity and accuracy.
[0022] Finally, real-time monitoring and dynamic adjustment are performed. In the spatiotemporal flow data scenario, the data distribution density may change over time. Therefore, a real-time monitoring mechanism needs to be introduced to dynamically adjust the grid resolution. The specific implementation is as follows: 1) Use sliding window technology to count the number of data points of each grid in the current time period in real time.
[0023] 2) If the number of data points in a grid continues to exceed the threshold θ, the grid is further subdivided; if the number of data points is lower than the threshold, the adjacent low-density grids are merged.
[0024] 3) The dynamic adjustment process can be triggered by event-driven, for example, every T w The time interval determines the data distribution based on the prediction results and updates the quadtree structure as needed.
[0025] Among them, S12 (time dimension modeling) specifically includes: Time dimension modeling aims to accurately describe the time characteristics of spatiotemporal flow data in order to support efficient time series analysis and prediction. By building a scalable ring time wheel and dynamically adjusting the time segment granularity, it is possible to flexibly respond to load requirements under different data flow rates.
[0026] First, define the ring time wheel, which is a cyclic data structure used to organize information in the time dimension. Set the basic time segment to , the ring time wheel is composed of slots, and the time interval corresponding to each slot is , the update rules of the time wheel are as follows: ; in, Indicates that the current slot stores the data flow information in the current time period (that is, the slot number pointed to by the current time wheel). Indicates a modulo operation to ensure that the slot number is between 0 and The advantage of the ring time wheel is that it can efficiently manage periodic and non-periodic time series data.
[0027] Then, dynamically adjust the time segment granularity. Data flow rate It is the key factor in determining the granularity of time segments. In order to achieve dynamic adjustment, define the target load balancing factor , which is used to measure the ideal amount of data in each time segment. Indicates the adjusted time segment granularity. represents the average amount of data, Indicates the data flow rate, and the adjustment rules are as follows: ; When the data flow rate is too high, the time segment granularity is reduced to improve the time resolution. When the data flow rate is low, the time segment granularity is increased to reduce the computing overhead. The adjustment process can be achieved through a timer or event trigger mechanism to ensure the real-time response capability of the system.
[0028] Finally, the time wheel is expanded or compressed (that is, the time wheel is expanded or compressed according to the situation of the slots); in extreme cases (such as sudden high or low traffic), the number of slots in the time wheel can be expanded or compressed To further optimize performance. For example: Expansion: When all slots are close to full, increase the number of slots to relieve pressure.
[0029] Compression: When most of the slots are empty, reduce the number of slots to save storage space.
[0030] See the schematic diagram of the adaptive network partitioning process for spatiotemporal flow data. Figure 3.
[0031] Among them, S13 (generation of spatiotemporal unique identifier) specifically includes: Spatiotemporal unique identifiers are the core tools for efficient data location and retrieval. By combining spatial Hilbert coding with temporal wheel position coding, a 64-bit identifier is generated, which can support query operations with O(1) complexity (O(1) complexity means that the time complexity of the algorithm is constant, that is, no matter how the input scale changes, the execution time of the algorithm is fixed).
[0032] First, perform spatial Hilbert encoding, such as Figure 3 As shown, assuming that the spatial grid number is , The corresponding spatial Hilbert encoding is , a spatial Hilbert code is generated by a recursive construction method. The spatial Hilbert code is a fixed-length binary string used to represent the spatial position.
[0033] Then, the time wheel position encoding is performed. Figure 3 As shown, if the timing wheel position is ,Will Convert the time wheel position code into binary form For example, if the time wheel has =16 slots, then The value range is [0, 15], and the corresponding binary code length is log2( )=4 digits.
[0034] Then, a bit crossover operation is performed. The spatial Hilbert code and the time wheel position code are bit-crossed to generate a 64-bit spatiotemporal unique identifier ( ). For example, for a 32-bit Hilbert code and a 32-bit temporal code: ; in, Represents Hilbert encoding from 32-bit space Extracted from Bit, Indicates the position encoding from the 32-bit time wheel Extracted from The core idea of bit interleaving operation is to interweave spatial and temporal information to form a compact and unique identifier.
[0035] Finally, the mapping relationship between the spatiotemporal unique identifier and the data block is stored in a hash table to support query operations with O(1) complexity. The specific implementation is as follows: 1) The key of the hash table is a unique identifier in time and space, and the value is the corresponding data block pointer.
[0036] 2) When querying, the hash value is calculated directly through the identifier and the target data block is found, avoiding the high time complexity of traditional linear search.
[0037] Among them, S14 (distributed node mapping) specifically includes: The goal of distributed node mapping is to evenly distribute spatiotemporal stream data blocks to multiple computing nodes to achieve load balancing and efficient computing.
[0038] First, define the hash function , convert the 64-bit unique spatial and temporal identifier Mapped to distributed node numbers. Assuming there are nodes, the hash function is: ; The hash function distributes the identifiers evenly to each node through modulo operation to ensure load balancing of data blocks.
[0039] Then, in order to further optimize the load balancing performance, the consistent hashing algorithm is introduced. The core idea of consistent hashing is to reduce the data migration cost through the virtual node mechanism. The specific implementation is as follows: 1) Evenly distribute Nvirtual virtual nodes on the hash ring, and each virtual node corresponds to a physical node.
[0040] 2) The data block obtains the physical node to which the nearest virtual node belongs clockwise according to the hash value for storage.
[0041] 3) When adding or removing nodes, only the affected data blocks need to be reallocated to avoid large-scale data migration.
[0042] Finally, in practical applications, distributed systems may need to dynamically expand or reduce the number of nodes. To this end, a dynamic expansion mechanism is designed: 1) When adding a new node, assign a new virtual node to it and migrate some data blocks to the new node.
[0043] 2) When a node is removed, the data blocks it is responsible for are redistributed to other nodes.
[0044] 3) The entire process is coordinated by the load balancing controller to ensure the stability and efficiency of the system.
[0045] Step S20: convert the spatiotemporal flow data into a multi-dimensional feature tensor, introduce a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and use a multi-head attention mechanism to aggregate multi-scale features to build a time series prediction model to generate high-precision prediction results of the future distribution of spatiotemporal flow data (deep modeling and distribution prediction of spatiotemporal flow data features).
[0046] Specifically, Figure 4 As shown, step S20 of the present invention mainly includes: S21, feature tensor conversion: converting the spatiotemporal stream data into a multidimensional feature tensor, extracting geographic location code and time period feature, and performing tensor fusion of the geographic location code and the time period feature to form a high-dimensional spatiotemporal feature; S22. Design of spatiotemporal embedding layer: introduce a learnable spatiotemporal embedding layer to jointly model the geographic location code and the time period feature, capture the spatiotemporal correlation and enhance the feature expression capability; S23, multi-scale feature aggregation: use the multi-head attention mechanism to aggregate spatiotemporal features of different scales to generate a unified spatiotemporal stream data feature representation; S24. Time series prediction model construction: Establish a time series prediction model based on the feature representation of the spatiotemporal flow data, define the loss function and optimizer, learn the distribution characteristics of the spatiotemporal flow data through a deep learning algorithm, and generate high-precision prediction results for the future distribution of the spatiotemporal flow data.
[0047] Among them, S21 (feature tensor conversion) specifically includes: Converting raw spatiotemporal stream data into multidimensional feature tensors is the first step to achieve deep modeling. By extracting geographic location codes and time period features and performing tensor fusion, a high-dimensional spatiotemporal feature representation can be formed, providing a basis for further feature learning.
[0048] First, perform geographic location coding. Geographic location coding is a numerical representation of spatial information. Assuming that the study area is divided into N×N grids, the location of each grid is represented by a two-dimensional coordinate system. In order to enhance the model's perception of spatial locality, sine and cosine functions can be introduced to generate position encoding: ; ; ; ; in, represents the dimension of the position encoding, represents the dimension index, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension.
[0049] Then, the time period characteristics are clarified. The time period characteristics are used to capture the regularity in the time dimension (such as daily cycle, weekly cycle, monthly cycle, etc.). Time codes can be generated by sine and cosine functions, for example: ; ; in, Indicates the timestamp, represents the length of the period (e.g. a day or a week), Indicates the time code The value of the dimension, Indicates the time code The value of the dimension.
[0050] Finally, the geographic location code and time period features are fused into a high-dimensional feature tensor. The specific method is as follows: Assume that the geographic location is encoded as ∈ , the time period characteristics are ∈ , then the fused feature tensor for: ; in, represents the dimension of the geolocation encoding, The dimension representing the time period characteristics, represents the geographic location encoding space, represents the time period feature space, It represents the space after the fusion of geographic location encoding space and time period feature space.
[0051] If you need to further enhance the feature expression capability, you can introduce nonlinear transformation: ; in, Represents the feature tensor after enhancing the feature expression capability, represents the activation function (such as ReLU or Sigmoid), represents the weight matrix, Represents the first bias vector.
[0052] Among them, S22 (temporal and spatial embedding layer design) specifically includes: The spatiotemporal embedding layer is a learnable neural network layer that is used to jointly model geographic location and time period features, capture spatiotemporal correlations, and enhance feature expression capabilities.
[0053] First, design the embedding layer structure. The core idea of the spatiotemporal embedding layer is to map discrete spatial and temporal features into a continuous low-dimensional vector space. Assume that the input feature is , the output of the embedding layer is , the embedding process is expressed as: ; in, represents the embedding weight matrix, represents the second bias vector.
[0054] Then, for the fused feature tensor , joint modeling is achieved by sharing parameters: ; ; ; in, , and Respectively represent the weight matrix of spatial features, the weight matrix of temporal features, and the weight matrix of spatiotemporal features, , and They represent the bias term of spatial features, the bias term of temporal features, and the bias term of spatiotemporal features respectively. and Represent the embedding vector of spatial features and the embedding vector of temporal features respectively, represents the joint spatiotemporal embedding vector.
[0055] Finally, in order to capture the complex correlation between spatiotemporal features, an interaction module can be introduced after the embedding layer. For example, the dot product operation is used to calculate the similarity of spatial and temporal features. : ; in, represents the function for calculating the similarity of spatial and temporal features, express Embedding vector of transposed temporal features.
[0056] Among them, S23 (multi-scale feature aggregation) specifically includes: Multi-scale feature aggregation aims to integrate spatiotemporal features of different scales to generate a unified and robust feature representation. The multi-head attention mechanism can effectively capture the correlation between global and local features.
[0057] First, the multi-head attention mechanism calculates the importance weights of features independently through multiple attention heads, and then weights the sum of the results. If the input feature is X (There are multiple elements in the input features), the calculation of the attention mechanism is: ; in, represents the attention mechanism, , and denote query, key, and value matrices respectively, represents the transposed and fused feature tensor, Indicates the dimension of the key.
[0058] Then, in the multi-head attention mechanism, the input feature X is linearly transformed into queries, keys, and values in multiple subspaces: ; ; ; in, Indicates The query weight matrix of the attention heads, Indicates The key weight matrix of the attention head, No. The value weight matrix of the attention head, , and Respectively represent The query, key, and value matrices for each attention head.
[0059] The output of each attention head for: ; The final output is the concatenation of all attention heads. : ; in, It means concatenating the outputs of all attention heads. represents the number of attention heads, represents the output of the first attention head, represents the output of the second attention head, Indicates The output of an attention head is represents the output weight matrix; Finally, in order to integrate features of different scales, convolution operations are introduced on the basis of the multi-head attention mechanism. For example, convolution kernels of different sizes are used to extract local features and global features, and the local features and global features are concatenated and sent to the multi-head attention module.
[0060] Among them, S24 (construction of time series prediction model) specifically includes: After the above steps, a shape of A three-dimensional matrix, where is the timestamp, and Represents the division of space in longitude and latitude. The element values of this matrix describe the time period Inner geospatial area The distribution characteristics of spatiotemporal flow records. In order to further predict the spatial distribution of spatiotemporal flow in future time periods, the present invention adopts a deep learning spatiotemporal data time series prediction model to analyze and deduce the feature matrix. The core architecture of the prediction model is composed of a spatiotemporal feature analysis layer composed of multiple layers of stacked long short-term memory network units and a multi-layer fully connected layer. Among them, the spatiotemporal feature analysis layer focuses on learning the time series patterns of the input data and capturing complex time dependencies and dynamic correlation features; while the fully connected layer linearly maps and dimensionalizes the hidden states output by the spatiotemporal feature analysis layer to generate the spatiotemporal distribution results of the predicted time period.
[0061] Before applying the prediction model, the spatiotemporal feature data needs to be strictly preprocessed and converted into a specific shape to meet the input requirements. Specifically, the spatiotemporal feature data needs to be constructed into a three-dimensional matrix of [batch_size, time_step, input_sizeension], where batch_size represents the number of samples processed by the model at one time. Its value can be flexibly set according to the size of the data set to balance computational efficiency and memory consumption. Time_step represents the time dimension, which is the first dimension of the historical spatiotemporal feature data. , input_sizeension represents the input feature dimension, which is the spatial distribution feature in this invention, corresponding to the second and third dimensions in the historical spatiotemporal feature data and To achieve the above transformation, the original feature matrix needs to be reshaped. The specific operation is as follows: First, the spatial feature dimension and Merge the three-dimensional matrix into a single dimension Compressed to , and add the batch dimension Upgrade to , in order to meet the normative requirements of deep learning models for input data formats.
[0062] Before model prediction, in order to ensure the accuracy of the prediction results and the generalization ability of the model, the data set needs to be reasonably divided into two parts: training set and test set. Usually, the data set is divided into training set and test set in a 4:1 ratio of the total data set, that is, 80% of the data is used for model training and the remaining 20% of the data is used for performance verification, so as to construct the shape of The training set and shape The test set.
[0063] Subsequently, in the model optimization phase, the appropriate loss function and optimizer must be carefully selected to achieve effective gradient updates and fast convergence. The choice of loss function should be based on the task objectives and data characteristics to ensure accurate quantification of errors, while the optimizer must take into account both training efficiency and stability.
[0064] When the time series forecasting model starts forecasting, the time series forecasting model will initialize the long-term memory With short-term status , respectively recording long-term memory and short-term information, at time step , the spatial feature recognition unit receives the long-term memory from the previous time step Short-term status And the input features of the current time step , through the nonlinear calculation of the forget gate, input gate and output gate, dynamically adjust the retention and update of information to generate the long-term memory of the current time step and short-term status , and recursively uses it as input to the next time step. The calculation process is described as follows: ; ; in, represents the weight matrix that maps input features and short-term states to candidate memory states, represents the bias term, and ⊙ represents the Hadamard product.
[0065] Forget Gate Function : ; Input gate function : ; Output gate function : ; in, , and Both represent learning weight parameters, , and Both represent bias parameters.
[0066] Activation Function Functions and The function expressions are: ; ; in, Indicates the input parameters; After the recursive calculation is completed in the spatiotemporal feature analysis layer, the result is passed to the fully connected layer, which adjusts the input dimension through linear transformation and generates a high-precision prediction result of the future distribution of the spatiotemporal flow data.
[0067] Step S30: Based on the adversarial network and the high-precision prediction results, generate the optimal computing resource allocation requirements, optimally divide the spatiotemporal flow data, develop a flexible scheduling strategy for distributed computing resources, and collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, and form a closed-loop optimization system to continuously optimize load balancing performance (distributed computing resource demand deduction and load balancing implementation).
[0068] Specifically, Figure 5 As shown, step S30 of the present invention mainly includes: S31, generating optimal computing resource requirements: generating an adversarial network based on spatiotemporal conditions, combining the high-precision prediction results, generating optimal computing resource allocation requirements, and improving resource utilization; S32, data dynamic partition optimization: using the improved K-means++ algorithm based on spatiotemporal constraints to optimally partition the spatiotemporal flow data to ensure that the data partition meets the spatiotemporal consistency requirements; S33, elastic scheduling strategy design: Develop elastic scheduling strategies for distributed computing resources (improve the overall efficiency of computing clusters), achieve sub-millisecond task transfer through dynamic load migration technology, reduce system latency and improve response speed; S34. Load balancing feedback control: Design a load balancing feedback controller to collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies as a continuous incremental state of load balancing, form a closed-loop optimization system, continuously optimize load balancing performance, and ensure system stability and efficiency.
[0069] Among them, S31 (generation of optimal computing resource requirements) specifically includes: The optimal computing resource demand generation is based on the spatiotemporal conditional generative adversarial network and the spatiotemporal flow data distribution prediction results, dynamically deducing the computing resource allocation requirements of each node in the distributed system. This method can significantly improve resource utilization and reduce redundancy.
[0070] First, a spatiotemporal conditional generative adversarial network is designed. The adversarial network includes a generator and a discriminator. The generator is used to generate a computing resource requirement plan that meets the spatiotemporal distribution characteristics, and the discriminator is used to evaluate the authenticity and rationality of the generated plan. Generator input: includes high-precision prediction results, current system status (such as CPU utilization, memory usage, etc.) and historical resource allocation records.
[0071] Generator output: The computing resource allocation requirement vector Rt=[r t,1 , r t,2 ,…,r t,N’ ], where r t,1 represents the resource requirement of the first node at time t, r t,2 represents the resource requirement of the second node at time t, r t,N’ It represents the resource requirement of the N'th node at time t, and N' represents the number of nodes.
[0072] Discriminator function: receiving the output of the generator and real historical resource allocation data, judging whether the generation scheme is reasonable, and providing feedback to optimize the generator.
[0073] Then, define the loss function and the loss function of the generator And the loss function of the discriminator To optimize model performance: ; ; in, represents the output of the generator, represents random noise, represents the output of the discriminator, represents the real input data, represents the expectation of noise, Represents real data expectations, Represents the discriminator to generate data The discrimination result indicates the probability that the generated data is judged as real data.
[0074] Finally, the gradient descent method is used to alternately optimize the parameters of the generator and the discriminator until the adversarial network converges and generates the optimal computing resource allocation requirements; the resource allocation requirements finally generated can meet the needs of spatiotemporal flow data distribution prediction to the greatest extent.
[0075] Among them, S32 (data dynamic partition optimization) specifically includes: The purpose of data dynamic partition optimization is to improve the computing efficiency of distributed systems by optimally partitioning spatiotemporal stream data to ensure that data partitioning meets spatiotemporal consistency requirements.
[0076] First, based on the traditional K-means++ algorithm (K-means++ is an improved version of the K-means clustering algorithm, which aims to improve the clustering effect and convergence speed through a more intelligent selection of initial center points. The standard K-means algorithm is very sensitive to the selection of initial center points, which may lead to suboptimal clustering results. K-means++ selects the initial center points through a probabilistic method, reducing this sensitivity), time and space constraints are introduced to optimize data partitioning.
[0077] Initialize center point selection: randomly select the first center point, and the probability of selecting subsequent center points is proportional to the distance to other selected center points.
[0078] Distance metric: Based on the traditional Euclidean distance, the time dimension weight is added and spatial dimension weights : ; in, Indicates data points, Indicates Cluster centers, Represents data points The spatial characteristics of Represents data points The time characteristics of Represents the cluster center The spatial characteristics of Represents the cluster center The time characteristics of Represents the integrated time and space distance.
[0079] Then, in the clustering process, we ensure that the data in the same spatiotemporal region are divided into the same group as much as possible. To this end, we introduce constraints into the objective function: ; in, represents the objective function, which is used to represent the total cost of clustering. represents the number of cluster centers, represents the constraint weight, Represents a constraint item, which is used to ensure that data in the same spatiotemporal area are divided into the same group as much as possible.
[0080] Finally, the number of cluster centers is dynamically adjusted according to the real-time data distribution density and computing resource requirements. , time dimension weight and spatial dimension weights , to meet the division requirements in different scenarios.
[0081] Among them, S33 (elastic scheduling strategy design) specifically includes: The elastic scheduling strategy uses dynamic load migration technology to achieve sub-millisecond task transfer, reduce system latency and improve response speed. This strategy can maintain efficient operation of the system when resource demand fluctuates greatly. First, a priority-based load migration algorithm is designed to prioritize the migration of tasks on high-load nodes to low-load nodes. The specific steps are as follows: 1) Task priority calculation: Calculate the task priority based on the task's computational complexity, remaining execution time, and dependencies: ; in, Indicates the task priority. represents the computational complexity of the task, Indicates the remaining execution time. represents the task dependency weight, , and Both represent weight coefficients.
[0082] 2) Migration decision: for high-load nodes N h , select the task with the highest priority T max Migrate to a less loaded node N l , and update the node load status.
[0083] Then, the task migration overhead is minimized through efficient serialization and deserialization techniques, for example, using zero-copy technology and asynchronous I / O operations to speed up the data transfer process.
[0084] Finally, before migration, sufficient computing resources are reserved for the target node to avoid task failure due to insufficient resources. The amount of reserved resources can be dynamically adjusted based on historical load statistics and forecast results.
[0085] Among them, S34 (load balancing feedback control) specifically includes: The load balancing feedback controller collects distributed node status data in real time, dynamically adjusts data partitioning and task allocation strategies, forms a closed-loop optimization system, and continuously optimizes load balancing performance.
[0086] First, the node status data of distributed nodes is collected in real time, including indicators such as CPU utilization, memory usage, network bandwidth, and disk I / O. The collection frequency can be dynamically adjusted according to the system load, usually once per second.
[0087] Then, the node status data is evaluated using machine learning or deep learning models to generate a load balancing score. , for example, a scoring model based on a multilayer perceptron (MLP, Multilayer Perceptron, a feedforward neural network): ; in, represents a deep learning model based on a multi-layer perceptron. , , and Respectively represent The CPU utilization, memory usage, network bandwidth, and disk I / O of each node.
[0088] Then, the data partitioning and task allocation strategies are dynamically adjusted according to the node status score: 1) If the score of a node is lower than the threshold, some tasks will be migrated to other nodes.
[0089] 2) If the score of a node is higher than the threshold, its task allocation is increased to make full use of resources.
[0090] Finally, a closed-loop feedback control system is constructed to regularly evaluate the load balancing performance and adjust the control parameters. For example, a PID controller (proportional-integral-derivative controller, a feedback control mechanism widely used in industrial control systems) is used to optimize the weight coefficients in the migration decision: ; in, represents the weight coefficient in the migration decision, represents the load error, represents the error signal, , and Both represent controller parameters.
[0091] The present invention proposes an innovative distributed computing load balancing method for spatiotemporal stream data. By constructing a deep learning-driven spatiotemporal data time series prediction model, the present invention can accurately identify and analyze the spatial distribution characteristics of historical spatiotemporal stream data, and obtain spatiotemporal stream distribution predictions for several future time periods based on the prediction results. Furthermore, based on these prediction data, the present invention scientifically and rationally allocates the distributed computing cluster resources of spatiotemporal stream data, thereby achieving load balancing in the distributed spatiotemporal stream data processing process, and optimizing the utilization efficiency of computing resources and the overall performance of the system. The present invention can achieve efficient computing resource allocation, data partitioning and load balancing in a distributed system, significantly improving the stability and performance of the system.
[0092] Furthermore, if Figure 6 As shown, based on the above novel space-time stream data distributed computing load balancing method, the present invention also provides a novel space-time stream data distributed computing load balancing system, wherein the novel space-time stream data distributed computing load balancing system includes: The spatiotemporal stream data adaptive network partitioning and distributed node mapping module 51 is used to construct an adaptive quadtree structure for spatial dimension partitioning, combine a scalable annular time wheel for time dimension modeling, generate a spatiotemporal unique identifier using a bit crossover operation, and evenly map the spatiotemporal stream data to distributed computing nodes based on the spatiotemporal unique identifier; The spatiotemporal stream data feature deep modeling and distribution prediction module 52 is used to convert the spatiotemporal stream data into a multi-dimensional feature tensor, introduce a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and use a multi-head attention mechanism to aggregate multi-scale features to build a time series prediction model to generate a high-precision prediction result of the future distribution of the spatiotemporal stream data; The distributed computing resource demand deduction and load balancing implementation module 53 is used to generate the optimal computing resource allocation demand based on the adversarial network combined with the high-precision prediction results, optimally divide the spatiotemporal flow data, develop a flexible scheduling strategy for distributed computing resources, and collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, and form a closed-loop optimization system to continuously optimize load balancing performance.
[0093] Furthermore, if Figure 7 As shown, based on the novel space-time stream data distributed computing load balancing method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 7 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0094] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a new type of space-time stream data distributed computing load balancing program 40 is stored on the memory 20, and the new type of space-time stream data distributed computing load balancing program 40 can be executed by the processor 10, thereby realizing the new type of space-time stream data distributed computing load balancing method in the present application.
[0095] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the novel spatiotemporal stream data distributed computing load balancing method.
[0096] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, the memory 20, and the display 30 of the terminal communicate with each other via a system bus.
[0097] In one embodiment, when the processor 10 executes the novel spatiotemporal stream data distributed computing load balancing program 40 in the memory 20, the steps of the novel spatiotemporal stream data distributed computing load balancing method described above are implemented.
[0098] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a new type of space-time stream data distributed computing load balancing program, and when the new type of space-time stream data distributed computing load balancing program is executed by a processor, the steps of the new type of space-time stream data distributed computing load balancing method as described above are implemented.
[0099] In summary, the present invention provides a new type of distributed computing load balancing method, system, terminal and computer-readable storage medium for spatiotemporal stream data, the method comprising: constructing an adaptive quadtree structure for spatial dimension division, combining a scalable annular time wheel for time dimension modeling, generating a spatiotemporal unique identifier using a bit crossover operation, and uniformly mapping the spatiotemporal stream data to distributed computing nodes based on the spatiotemporal unique identifier; converting the spatiotemporal stream data into a multidimensional feature tensor, introducing a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and using a multi-head attention mechanism to aggregate multi-scale features, constructing a time series prediction model to generate a high-precision prediction result of the future distribution of spatiotemporal stream data; based on an adversarial network combined with the high-precision prediction result, generating the optimal computing resource allocation requirement, optimally dividing the spatiotemporal stream data, developing a flexible scheduling strategy for distributed computing resources, and collecting distributed node status data in real time, dynamically adjusting data partitioning and task allocation strategies, and forming a closed-loop optimization system to continuously optimize load balancing performance. The present invention can effectively mine future rules from spatiotemporal stream data, realize dynamic load balancing for computing nodes, effectively alleviate data tilt, and improve the computing and processing efficiency of spatiotemporal stream data.
[0100] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0101] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0102] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A new type of distributed computing load balancing method for spatiotemporal stream data, characterized in that: The novel time-space stream data distributed computing load balancing method includes: Construct an adaptive quadtree structure for spatial dimension partitioning, combine with a scalable ring time wheel for time dimension modeling, use bit crossover operation to generate a unique spatiotemporal identifier, and evenly map the spatiotemporal flow data to distributed computing nodes based on the unique spatiotemporal identifier; The spatiotemporal stream data is converted into a multi-dimensional feature tensor, a learnable spatiotemporal embedding layer is introduced to capture the spatiotemporal correlation, and a multi-head attention mechanism is used to aggregate multi-scale features to construct a time series prediction model to generate a high-precision prediction result of the future distribution of the spatiotemporal stream data; Based on the adversarial network combined with the high-precision prediction results, the optimal computing resource allocation requirements are generated, the spatiotemporal flow data is optimally divided, an elastic scheduling strategy for distributed computing resources is developed, and distributed node status data is collected in real time. The data partitioning and task allocation strategies are dynamically adjusted to form a closed-loop optimization system to continuously optimize the load balancing performance.
2. According to the novel space-time stream data distributed computing load balancing method described in claim 1, it is characterized in that: The method comprises: constructing an adaptive quadtree structure for spatial dimension division, combining a scalable ring time wheel for time dimension modeling, using a bit crossover operation to generate a spatiotemporal unique identifier, and uniformly mapping the spatiotemporal flow data to distributed computing nodes based on the spatiotemporal unique identifier. Specifically, the method comprises: According to the spatial scope of the study area, an improved Hilbert space filling curve is used to construct an adaptive quadtree structure. The data distribution density is monitored in real time according to the division of the density adaptive quadtree, and the grid resolution is dynamically adjusted to adapt to the change of data density. Build a scalable ring time wheel, set basic time segments, and dynamically adjust the time segment granularity according to the data flow rate; The spatial Hilbert code and the time wheel position are bit-interleaved to generate a 64-bit spatiotemporal unique identifier, which is used for data location and retrieval; Based on the unique spatiotemporal identifier, a hash function is designed to evenly map the spatiotemporal flow data to distributed computing nodes as an initial state for prediction.
3. The novel time-space stream data distributed computing load balancing method according to claim 2 is characterized in that: The process of converting the spatiotemporal stream data into a multi-dimensional feature tensor, introducing a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and using a multi-head attention mechanism to aggregate multi-scale features, constructing a time series prediction model to generate a high-precision prediction result of the future distribution of the spatiotemporal stream data, specifically includes: Converting the spatiotemporal flow data into a multidimensional feature tensor, extracting geographic location code and time period features, and performing tensor fusion of the geographic location code and the time period features to form a high-dimensional spatiotemporal feature; Introducing a learnable spatiotemporal embedding layer to jointly model the geographic location encoding and the time period feature to capture spatiotemporal correlation; Use the multi-head attention mechanism to aggregate spatiotemporal features of different scales and generate a unified spatiotemporal stream data feature representation; A time series prediction model based on the feature representation of the spatiotemporal stream data is established, a loss function and an optimizer are defined, and the distribution characteristics of the spatiotemporal stream data are learned through a deep learning algorithm to generate high-precision prediction results for the future distribution of the spatiotemporal stream data.
4. The novel time-space stream data distributed computing load balancing method according to claim 3 is characterized in that: The adversarial network is combined with the high-precision prediction results to generate optimal computing resource allocation requirements, optimally divide the spatiotemporal flow data, develop a flexible scheduling strategy for distributed computing resources, and collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, and form a closed-loop optimization system to continuously optimize load balancing performance, specifically including: Generate an optimal computing resource allocation requirement based on a spatiotemporal conditional generative adversarial network and the high-precision prediction results; Using an improved K-means++ algorithm based on spatiotemporal constraints to optimally divide the spatiotemporal flow data; Develop elastic scheduling strategies for distributed computing resources as a continuous incremental state of load balancing, and achieve sub-millisecond task transfer through dynamic load migration technology; Design a load balancing feedback controller to collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, form a closed-loop optimization system, and continuously optimize load balancing performance.
5. The novel time-space stream data distributed computing load balancing method according to claim 4 is characterized in that: According to the spatial scope of the study area, an improved Hilbert space filling curve is used to construct an adaptive quadtree structure, data distribution density is monitored in real time according to the division of the density adaptive quadtree, and the grid resolution is dynamically adjusted to adapt to the change of data density, specifically including: Generate an improved Hilbert space filling curve. If the spatial range of the study area is a two-dimensional plane [0, L x ]×[0,L y ], L x and L y Respectively represent the length and width of the study area on the two-dimensional plane, and divide the two-dimensional plane into an N×N uniform grid, where N represents the number of rows and columns of the grid, and N=2 k , k is a positive integer, the recursive construction formula of the Hilbert curve is as follows: ; in, Indicate point The one-dimensional mapping value on the Hilbert curve, Represents a recursive call to divide the current quadrant into 4 sub-quadrants. , , and Represent four sub-areas respectively; During initialization, the entire study area is divided into a root node, a threshold θ is defined, the maximum number of data points in each grid is controlled according to the threshold θ, and an adaptive quadtree is constructed using a recursive algorithm; Use the sliding window technology to count the number of data points of each grid in the current time period in real time. If the number of data points of a grid continues to exceed the threshold θ, the grid is subdivided. If the number of data points is lower than the threshold θ, the adjacent low-density grids are merged. The construction of a scalable annular time wheel, setting basic time segments, and dynamically adjusting the time segment granularity according to the data flow rate specifically includes: Define a circular time wheel and set the basic time segment as , the ring time wheel is composed of slots, and the time interval corresponding to each slot is , the update rules of the time wheel are as follows: ; in, Indicates that the current slot stores the data flow information in the current time period. Indicates a modulo operation to ensure that the slot number is between 0 and Cycle between Defining the target load balancing factor , used to measure the ideal amount of data in each time segment, Indicates the adjusted time segment granularity. represents the average amount of data, Indicates the data flow rate, and the adjustment rules are as follows: ; When the data flow rate is too high, the time segment granularity is reduced to improve the time resolution. When the data flow rate is low, the time segment granularity is increased to reduce the computational overhead. Expand or compress the time wheel according to the slot situation; The generating of a 64-bit spatiotemporal unique identifier by performing a bit crossover operation on the spatial Hilbert code and the time wheel position specifically includes: If the spatial grid number is , The corresponding spatial Hilbert encoding is , a spatial Hilbert code is generated by a recursive construction method. The spatial Hilbert code is a fixed-length binary string used to represent the spatial position; If the time wheel position is ,Will Convert the time wheel position code into binary form ; The spatial Hilbert code and the time wheel position code are bit-by-bit interleaved to generate a 64-bit unique spatial and temporal identifier. For 32-bit Hilbert code and 32-bit time code: ; in, Represents Hilbert encoding from 32-bit space Extracted from Bit, Indicates the position encoding from the 32-bit time wheel Extracted from Bit; The mapping relationship between the spatiotemporal unique identifier and the data block is stored in a hash table. The key of the hash table is the spatiotemporal unique identifier, and the value is the corresponding data block pointer. When querying, the hash value is directly calculated through the identifier and the target data block is found; Based on the spatiotemporal unique identifier, a hash function is designed to evenly map the spatiotemporal stream data to the distributed computing nodes, specifically including: Defining a hash function , convert the 64-bit spatiotemporal unique identifier Mapped to distributed node number, if there are nodes, the hash function is: ; The hash function distributes the identifiers evenly to each node through modulo operation to ensure load balancing of data blocks; Evenly distribute Nvirtual virtual nodes on the hash ring. Each virtual node corresponds to a physical node. Data blocks are stored in the physical node to which the nearest virtual node belongs clockwise according to the hash value. When adding or removing a node, only the affected data blocks need to be reallocated. Depending on the actual application, the distributed system dynamically expands or reduces the number of nodes.
6. The novel time-space stream data distributed computing load balancing method according to claim 5 is characterized in that: The converting the spatiotemporal stream data into a multidimensional feature tensor, extracting geographic location codes and time period features, and performing tensor fusion of the geographic location codes and the time period features to form high-dimensional spatiotemporal features specifically includes: If the study area is divided into N×N grids, the position of each grid is represented by the two-dimensional coordinates Indicates that the sine and cosine functions are introduced to generate position encoding: ; ; ; ; in, represents the dimension of the position encoding, represents the dimension index, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension, Indicates the position code The value of the dimension; Clearly define the time period characteristics and generate time codes through sine and cosine functions: ; ; in, Indicates the timestamp, represents the cycle length, Indicates the time code The value of the dimension, Indicates the time code The value of the dimension; If the geographic location code is ∈ , the time period characteristics are ∈ , then the fused feature tensor for: ; in, represents the dimension of the geolocation encoding, The dimension representing the time period characteristics, represents the geographic location coding space, represents the time period feature space, It represents the space after the geographic location coding space and the time period feature space are fused; If you need to enhance the feature expression capability, introduce nonlinear transformation: ; in, Represents the feature tensor after enhancing the feature expression capability, represents the activation function, represents the weight matrix, represents the first bias vector; The introduction of a learnable spatiotemporal embedding layer to jointly model the geographic location code and the time period feature to capture spatiotemporal correlation specifically includes: If the input feature is , the output of the embedding layer is , the embedding process is expressed as: ; in, represents the embedding weight matrix, represents the second bias vector; For the fused feature tensor , joint modeling is achieved by sharing parameters: ; ; ; in, , and Respectively represent the weight matrix of spatial features, the weight matrix of temporal features, and the weight matrix of spatiotemporal features, , and They represent the bias term of spatial features, the bias term of temporal features, and the bias term of spatiotemporal features respectively. and Represent the embedding vector of spatial features and the embedding vector of temporal features respectively, represents the joint spatiotemporal embedding vector; To capture the complex correlation between spatiotemporal features, an interaction module is introduced after the embedding layer, and the dot product operation is used to calculate the similarity of spatial and temporal features. : ; in, represents the function for calculating the similarity of spatial and temporal features, express The embedding vector of the transposed temporal features; The multi-head attention mechanism is used to aggregate spatiotemporal features of different scales to generate a unified spatiotemporal stream data feature representation, specifically including: The importance weights of the features are calculated independently by multiple attention heads, and then the results are weighted summed. If the input feature is X , the calculation of the attention mechanism is: ; in, represents the attention mechanism, , and denote query, key, and value matrices respectively, represents the transposed fused feature tensor, Represents the dimension of the key; In the multi-head attention mechanism, the input features X is linearly transformed into queries, keys, and values in multiple subspaces: ; ; ; in, Indicates The query weight matrix of the attention heads, Indicates The key weight matrix of the attention head, No. The value weight matrix of the attention head, , and Respectively represent The query, key, and value matrices for each attention head; The output of each attention head for: ; The final output is the concatenation of all attention heads. : ; in, It means concatenating the outputs of all attention heads. represents the number of attention heads, represents the output of the first attention head, represents the output of the second attention head, Indicates The output of an attention head is represents the output weight matrix; On the basis of the multi-head attention mechanism, the convolution operation is introduced, and the local features and global features are extracted using convolution kernels of different sizes. The local features and global features are then concatenated and sent to the multi-head attention module. The method of establishing a time series prediction model based on the feature representation of the spatiotemporal stream data, defining a loss function and an optimizer, learning the distribution features of the spatiotemporal stream data through a deep learning algorithm, and generating a high-precision prediction result of the future distribution of the spatiotemporal stream data specifically includes: The spatiotemporal feature data is constructed as batch_size, time_step, input_sizeension], where batch_size represents the number of samples processed by the model at one time to balance computational efficiency and memory consumption, time_step represents the time dimension, input_sizeension represents the input feature dimension, and the spatial feature dimension and Merge the three-dimensional matrix into a single dimension Compress to , and add the batch dimension Upgrade to ; The data set is divided into a training set and a test set according to the 4:1 ratio of the total data set, and the shapes are constructed respectively. The training set and shape The test set; Select the loss function based on the task objectives and data characteristics, and select the optimizer based on training efficiency and stability; When the time series forecasting model starts forecasting, the time series forecasting model will initialize the long-term memory With short-term status , respectively recording long-term memory and short-term information, at time step , the spatial feature recognition unit receives the long-term memory from the previous time step Short-term status And the input features of the current time step , through the nonlinear calculation of the forget gate, input gate and output gate, dynamically adjust the retention and update of information to generate the long-term memory of the current time step and short-term status , and recursively uses it as input to the next time step. The calculation process is described as follows: ; ; in, represents the weight matrix that maps input features and short-term states to candidate memory states, represents the bias term, ⊙ represents the Hadamard product; Forget Gate Function : ; Input gate function : ; Output gate function : ; in, , and Both represent learning weight parameters, , and All represent bias parameters; Activation Function Functions and The function expressions are: ; ; in, Indicates the input parameters; After the recursive calculation is completed in the spatiotemporal feature analysis layer, the result is passed to the fully connected layer, which adjusts the input dimension through linear transformation and generates a high-precision prediction result of the future distribution of the spatiotemporal flow data.
7. The novel time-space stream data distributed computing load balancing method according to claim 6 is characterized in that: The generation of an adversarial network based on spatiotemporal conditions, combined with the high-precision prediction results, generates optimal computing resource allocation requirements, specifically including: Design a spatiotemporal conditional generative adversarial network, the adversarial network comprising a generator and a discriminator, the generator is used to generate a computing resource requirement plan that conforms to the spatiotemporal distribution characteristics, and the discriminator is used to evaluate the authenticity and rationality of the generated plan; The high-precision prediction results, current system status and historical resource allocation records are input into the generator, which outputs the computing resource allocation requirement vector Rt=[r t,1 r t,2 … r t,N’ ], where r t,1 represents the resource requirement of the first node at time t, r t,2 represents the resource requirement of the second node at time t, r t,N’ represents the resource demand of the N'th node at time t, where N' represents the number of nodes; The discriminator receives the output of the generator and real historical resource allocation data, determines whether the generation scheme is reasonable, and provides feedback to optimize the generator; Define the loss function and the loss function of the generator And the loss function of the discriminator To optimize model performance: ; ; in, represents the output of the generator, represents random noise, represents the output of the discriminator, represents the real input data, represents the expectation of noise, Represents real data expectations, Represents the discriminator to generate data The discrimination result indicates the probability that the generated data is judged as real data; Use the gradient descent method to alternately optimize the parameters of the generator and the discriminator until the adversarial network converges and generates the optimal computing resource allocation requirements; The improved K-means++ algorithm based on spatiotemporal constraints is used to optimally divide the spatiotemporal flow data, specifically including: Based on the traditional K-means++ algorithm, time and space constraints are introduced to optimize data partitioning, the first center point is randomly selected, and the time dimension weight is added on the basis of the traditional Euclidean distance. and spatial dimension weights : ; in, Indicates data points, Indicates Cluster centers, Represents data points The spatial characteristics of Represents data points The time characteristics of Represents the cluster center The spatial characteristics of Represents the cluster center The time characteristics of Represents the integrated time and space distance; In the clustering process, constraints are introduced into the objective function: ; in, represents the objective function, which is used to represent the total cost of clustering. represents the number of cluster centers, represents the constraint weight, Represents a constraint item, which is used to ensure that data in the same spatiotemporal region are divided into the same group as much as possible; Dynamically adjust the number of cluster centers based on real-time data distribution density and computing resource requirements , time dimension weight and spatial dimension weights , to meet the division requirements in different scenarios; The development of a flexible scheduling strategy for distributed computing resources implements sub-millisecond task transfer through dynamic load migration technology, specifically including: Design a priority-based load migration algorithm to prioritize the migration of tasks on high-load nodes to low-load nodes. Calculate the task priority based on the computational complexity, remaining execution time, and dependencies of the task: ; in, Indicates the task priority. represents the computational complexity of the task, Indicates the remaining execution time. represents the task dependency weight, , and All represent weight coefficients; For high load nodes N h , select the task with the highest priority T max Migrate to a less loaded node N l , and update the node load status; Minimize task migration overhead through efficient serialization and deserialization technology, and dynamically adjust the amount of resources reserved for the target node based on historical load statistics and prediction results; The load balancing feedback controller is designed to collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, form a closed-loop optimization system, and continuously optimize load balancing performance, specifically including: Collect node status data of distributed nodes in real time; Use machine learning or deep learning models to evaluate node status data and generate load balancing scores : ; in, represents a deep learning model based on a multi-layer perceptron. , , and Respectively represent CPU utilization, memory usage, network bandwidth, and disk I / O of each node; Dynamically adjust data partitioning and task allocation strategies based on node status scores; Build a closed-loop feedback control system to regularly evaluate load balancing performance and adjust control parameters, and use a PID controller to optimize the weight coefficients in migration decisions: ; in, represents the weight coefficient in the migration decision, represents the load error, represents the error signal, , and Both represent controller parameters.
8. A new type of distributed computing load balancing system for spatiotemporal stream data, characterized in that: The novel space-time stream data distributed computing load balancing system includes: The module for adaptive network partitioning and distributed node mapping of spatiotemporal stream data is used to construct an adaptive quadtree structure for spatial dimension partitioning, combine a scalable annular time wheel for time dimension modeling, generate a spatiotemporal unique identifier using bit crossover operation, and evenly map the spatiotemporal stream data to distributed computing nodes based on the spatiotemporal unique identifier; A spatiotemporal stream data feature deep modeling and distribution prediction module, which is used to convert the spatiotemporal stream data into a multi-dimensional feature tensor, introduce a learnable spatiotemporal embedding layer to capture spatiotemporal correlation, and use a multi-head attention mechanism to aggregate multi-scale features to build a time series prediction model to generate high-precision prediction results for the future distribution of spatiotemporal stream data; The distributed computing resource demand deduction and load balancing implementation module is used to generate optimal computing resource allocation requirements based on the adversarial network combined with the high-precision prediction results, optimally divide the spatiotemporal flow data, develop a flexible scheduling strategy for distributed computing resources, and collect distributed node status data in real time, dynamically adjust data partitioning and task allocation strategies, and form a closed-loop optimization system to continuously optimize load balancing performance.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a new type of spatiotemporal stream data distributed computing load balancing program stored in the memory and executable on the processor. When the new type of spatiotemporal stream data distributed computing load balancing program is executed by the processor, the steps of the new type of spatiotemporal stream data distributed computing load balancing method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a new type of space-time stream data distributed computing load balancing program, and when the new type of space-time stream data distributed computing load balancing program is executed by a processor, the steps of the new type of space-time stream data distributed computing load balancing method as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Traffic flow prediction system based on deep learning and dynamic network analysis and application method thereof
CN118675324A
AI-based big data distributed computing task automatic optimization method and system
CN119576507A
Method for dynamically allocating resources in an SDN / NFV network based on load balancing
US20190182169A1
Node load-based dynamic data partitioning system
WO2021073083A1
Cited By
Distribution box coordination control method and system for distributed energy access
CN120784972A
Distributed intelligent deduction system and method for multi-source data fusion and dynamic scheduling
CN121301004A
Distributed intelligent inference system and method for multi-source data fusion and dynamic scheduling
CN121301004B
Efficient big data processing system based on distributed computing architecture
CN121326581A