Unstructured database federated learning collaborative method and system based on swarm intelligence
Through the unstructured database federated learning collaborative method based on group intelligence, the query execution scheme and resource allocation are dynamically optimized, and the problem of inefficient unstructured data query is solved, and efficient resource utilization and adaptive optimization are achieved.
Patent Information
- Application Number
- CN202510186996.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing technology is difficult to efficiently process unstructured data queries, and the lack of global resource coordination mechanism in a distributed environment, resulting in low query efficiency and low resource utilization.
Using an unstructured database federated learning collaboration method based on group intelligence, a non-structured database federated learning collaborative method is used to collect query characteristics and resource indicators of multiple database nodes, an initial training sample set is constructed and input into a deep neural network for training, and a sub-model of query pattern recognition, path planning and resource allocation is generated. A hybrid group intelligent algorithm combining ant colony path optimization mechanism and particle swarm parameter optimization mechanism is used to dynamically optimize the query execution plan, and distributed resource scheduling is performed through resource allocation sub-model.
It improves query efficiency, optimizes resource utilization, and realizes adaptive query optimization, improving the overall performance of the system.
Smart Images

Figure CN119669433B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to structured data technology, and in particular to a collaborative method and system for unstructured database federated learning based on swarm intelligence. Background Art
[0002] Unstructured databases, such as document databases and graph databases, have significant advantages in storing and managing complex data and are widely used in various fields. With the continuous growth of data scale and the increasing complexity of query requirements, how to efficiently process queries has become a key challenge. Traditional query optimization methods usually rely on predefined patterns or indexes, which are difficult to adapt to the flexibility and diversity of unstructured data. In addition, in a distributed environment, how to coordinate the resources of multiple database nodes to minimize query latency and maximize resource utilization is also a problem that needs to be solved.
[0003] First, traditional query optimization methods are difficult to effectively handle queries on unstructured data. Since unstructured data lacks a fixed pattern, traditional optimization strategies based on indexes or statistical information are difficult to apply, resulting in low query efficiency.
[0004] Secondly, existing distributed query processing methods usually lack a global resource coordination mechanism. Each database node executes query tasks independently, which easily causes resource competition and load imbalance, thus affecting the overall query performance.
[0005] Finally, traditional query optimization methods are usually static and difficult to adapt to changes in query load. As data scale and query patterns change, pre-set optimization strategies may no longer be effective, and query execution plans need to be dynamically adjusted to meet new requirements. Summary of the invention
[0006] The embodiments of the present invention provide a collaborative method and system for federated learning of unstructured databases based on swarm intelligence, which can solve the problems in the prior art.
[0007] According to a first aspect of the embodiments of the present invention,
[0008] Provides a collaborative method for federated learning of unstructured databases based on swarm intelligence, including:
[0009] Collect query statement feature vectors, query resource consumption indicators, query response time indicators, and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; input the initial training sample set into a preset deep neural network, train the query pattern recognition sub-model, path planning sub-model, and resource allocation sub-model through a federated learning framework, and generate corresponding initial model parameters; establish a swarm intelligence collaborative network based on the initial model parameters, and set an initial path convergence time threshold and a resource utilization threshold as a network evaluation benchmark;
[0010] The query request input by the user is sent to the query pattern recognition sub-model for semantic analysis, query semantic features are extracted, and candidate query paths are generated through the path planning sub-model according to the initial path convergence time threshold; the candidate query path is input into a hybrid swarm intelligence algorithm that integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, and a query execution plan is generated based on the resource utilization threshold dynamic optimization; the resource allocation sub-model performs distributed resource scheduling according to the query execution plan to form a final query execution plan; the final query execution plan is executed, and the actual path convergence time and resource occupancy are recorded;
[0011] The actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold are differentially analyzed to generate a performance optimization target; based on the performance optimization target, the local model parameters of each node are synchronously updated, and the updated local model parameters are fused through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; the optimized global model parameters are distributed to the swarm intelligent collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization.
[0012] Inputting the candidate query path into a hybrid swarm intelligence algorithm that integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, and dynamically optimizing and generating a query execution plan based on the resource utilization threshold comprises:
[0013] Construct an initial candidate query path set, map the execution order of each path in the initial candidate query path set into a path code in the form of an integer sequence, map the resource allocation strategy of the query operation into a parameter code in the form of a real vector, and establish a two-layer path representation model of path structure and resource allocation; calculate the query cost ratio of each candidate path based on the two-layer path representation model to generate an initial path evaluation index;
[0014] The ratio of the initial path evaluation index to the preset benchmark threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligence optimization model, and the hybrid swarm intelligence optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; based on the path transfer probability matrix, a path structure optimization sequence is generated by an ant colony algorithm;
[0015] The evaluation score of the path structure optimization sequence is used as the pheromone increment, and the pheromone increment is converted into the fitness function of the particle swarm algorithm; based on the fitness function, the particle swarm algorithm is guided to iteratively optimize the resource allocation parameters to generate a parameter optimization vector; system resource utilization data is collected, and when the resource utilization data exceeds a preset benchmark threshold, the value space of the parameter optimization vector is adaptively adjusted according to the degree of excess; the path structure optimization sequence and the parameter optimization vector are combined to form a complete query execution plan.
[0016] The ratio of the initial path evaluation index to the preset benchmark threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligence optimization model, and the hybrid swarm intelligence optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; based on the path transfer probability matrix, the path structure optimization sequence is generated by the ant colony algorithm, including:
[0017] Calculating a ratio between the initial path evaluation index and a preset reference threshold to obtain path heuristic information, wherein the initial path evaluation index includes a path length index, a path energy consumption index, and a path time index, and the preset reference threshold includes a reference path length threshold, a reference path energy consumption threshold, and a reference path time threshold;
[0018] The path heuristic information and the initial pheromone concentration are combined to obtain a path transfer probability matrix, wherein the initial pheromone concentration adopts a fixed initial value, and the calculation result of the path transfer probability matrix includes a pheromone concentration weight coefficient and a heuristic information weight coefficient;
[0019] Inputting the path transfer probability matrix into a hybrid swarm intelligent optimization model, wherein the hybrid swarm intelligent optimization model integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, wherein the ant colony path optimization mechanism is responsible for path structure search, and the particle swarm parameter optimization mechanism is responsible for optimizing the pheromone concentration weight coefficient and the heuristic information weight coefficient;
[0020] Generate candidate paths through an ant colony algorithm based on the path transfer probability matrix, calculate the path evaluation score corresponding to the candidate path, determine the pheromone release amount according to the path evaluation score, and update the path transfer probability matrix based on the pheromone release amount and a preset pheromone volatility coefficient;
[0021] The path evaluation score is used as the fitness value of the particle swarm algorithm, the individual optimal solution and the global optimal solution of the particle swarm are updated based on the fitness value, and the particle speed update amount and the position update amount are calculated according to the individual optimal solution and the global optimal solution;
[0022] The updated particle positions are used as weight coefficient configurations to optimize the path transfer probability matrix, and the optimized path transfer probability matrix is re-input into the ant colony path optimization mechanism for iteration until convergence to obtain a path structure optimization sequence.
[0023] The resource allocation sub-model performs distributed resource scheduling according to the query execution plan to form a final query execution plan; executing the final query execution plan, and recording the actual convergence time and resource occupancy of the path include:
[0024] The resource allocation sub-model collects processor performance indicators, memory capacity indicators, network bandwidth indicators and storage capacity indicators of each resource pool to construct a resource feature matrix; calculates resource pool scores according to the resource feature matrix, and performs initial resource allocation for the input query execution plan based on the resource pool scores;
[0025] The resource allocation submodel collects the real-time resource usage data of each node according to the initial resource allocation result, and calculates the node load status in combination with the resource type weight factor; when the node load status exceeds the preset load threshold, the optimal migration target node is determined based on the task migration cost and load balancing benefit, and an optimized query execution plan is generated;
[0026] The resource allocation sub-model performs distributed resource scheduling on the optimized query execution plan, allocates execution resources based on the real-time computing capability, storage capacity and network status of each node, and forms a final query execution plan.
[0027] The actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold are differentially analyzed to generate a performance optimization target; based on the performance optimization target, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm, and reconstructs and generates optimized global model parameters including:
[0028] Construct a path convergence time monitoring module and a resource occupancy monitoring module. The path convergence time monitoring module collects the actual path convergence time and compares it with the preset initial path convergence time threshold to generate a time dimension difference sequence. The resource occupancy monitoring module collects the resource occupancy data of each node and compares it with the preset resource utilization threshold to generate a resource dimension difference sequence.
[0029] Dynamic time warping is performed on the time dimension difference sequence, a time dimension cost matrix containing the deviation value of each time point is constructed, and the time difference score is obtained by solving the optimal matching path through a dynamic programming algorithm; the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate of the resource dimension difference sequence is calculated, and the resource difference score is obtained by combining the resource importance weight coefficient;
[0030] The time difference score and the resource difference score are normalized and fused to generate a performance difference index based on the harmonic mean method; a performance optimization objective function including a stability constraint is constructed according to the performance difference index, and a weight coefficient of the performance optimization objective function is adaptively adjusted according to the degree of fluctuation of the time difference score and the resource difference score;
[0031] The federated learning framework trains the local model in parallel on each node based on the performance optimization objective function, updates the local model parameters using the stochastic gradient descent method with momentum term, and dynamically adjusts the learning rate according to the time difference score during the training process; shards the updated local model parameters, and applies the homomorphic encryption algorithm to the parameter shards to generate encrypted parameter blocks;
[0032] Allocate a key shard related to the resource difference score to each node, combine the key shard with random mask information to generate a secure aggregation key; use the secure aggregation key to perform distributed aggregation operations on the encryption parameter blocks to generate a global parameter aggregation matrix; calculate the node weight coefficient according to the time difference score and the resource difference score of each node, perform weighted averaging on the global parameter aggregation matrix based on the node weight coefficient, perform sparse processing on the weighted averaged parameters, and reconstruct and generate optimized global model parameters.
[0033] Dynamic time warping is performed on the time dimension difference sequence, a time dimension cost matrix containing the deviation value of each time point is constructed, and the time difference score is obtained by solving the optimal matching path through the dynamic programming algorithm; the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate of the resource dimension difference sequence is calculated, and the resource difference score obtained by combining the resource importance weight coefficient includes:
[0034] Constructing a time dimension cost matrix according to the time dimension difference sequence, wherein each element of the time dimension cost matrix comprises a time point difference distance term and a time jump penalty term, wherein the time point difference distance term represents the degree of deviation of corresponding points in the time series, and the time jump penalty term constrains the nonlinear distortion of the time series through a preset penalty coefficient;
[0035] The time dimension cost matrix is subjected to cumulative cost calculation to obtain a cumulative cost matrix, and the cumulative cost matrix is used to find a minimum cumulative cost path through a dynamic programming algorithm; based on the minimum cumulative cost path, an optimal matching path is obtained by backtracking, and the cumulative cost on the optimal matching path is normalized to obtain a time dimension scoring sequence;
[0036] Collect the real-time resource occupancy data of each node in the system, including processor occupancy, memory occupancy and network bandwidth occupancy, calculate the difference between the real-time resource occupancy data and the preset resource utilization threshold to obtain the resource dimension difference sequence; use the hierarchical analysis method to construct a resource indicator judgment matrix, and obtain the importance weight coefficient corresponding to each resource indicator by calculating the eigenvector of the resource indicator judgment matrix;
[0037] The resource dimension difference sequence is sampled based on a sliding observation window of a preset size, and the statistical features of the resource dimension difference sequence within the sliding observation window are extracted; the resource dimension difference sequence and the importance weight coefficient are weighted and calculated, and the resource difference score is obtained using the weighted Euclidean distance metric.
[0038] Distributing the optimized global model parameters to the swarm intelligence collaborative network, dynamically adjusting the initial path convergence time threshold and the resource utilization threshold, and forming a new evaluation benchmark for guiding the next round of query optimization includes:
[0039] Constructing the topological structure of the swarm intelligence collaborative network, calculating the adjacency matrix based on the physical distance between nodes and the quality of network links; using the spectral clustering algorithm to divide the adjacency matrix into multiple sub-network groups, and selecting a regional coordinator in each of the sub-network groups based on the node computing power and network connection stability;
[0040] Collecting hardware configuration parameters and task load information of each node, performing differentiated processing on the optimized global model parameters based on the hardware configuration parameters and the task load information, and generating parameter deployment schemes adapted to different nodes; constructing a parameter integrity check matrix using a Bloom filter, and distributing the parameter deployment scheme to each node in the swarm intelligence collaborative network through the regional coordinator;
[0041] Deploy a performance collection agent in the swarm intelligence collaborative network to collect the path convergence time and resource utilization after the node executes the parameter deployment scheme; calculate the deviation value between the path convergence time and the initial path convergence time threshold, and the difference value between the resource utilization and the resource utilization threshold based on a sliding time window;
[0042] The deviation value and the difference value are input into an orthogonal test model to analyze the influence of the initial path convergence time threshold and the resource utilization threshold on system performance; a principal component analysis method is used to extract threshold combinations, and a support vector regression model is established to map the relationship between performance differences and threshold adjustment amounts;
[0043] Construct a dual Q-learning network, use the performance difference as the state input, and use the threshold adjustment direction and step size as the action output; design a reward function based on the performance improvement metric, train a threshold dynamic adjustment strategy, and adaptively adjust the exploration probability of the dual Q-learning network according to the degree of performance fluctuation;
[0044] Performing exponential moving average processing on the performance data after executing the threshold dynamic adjustment strategy, and generating a performance trend prediction value in combination with a time series prediction model; dynamically adjusting the initial path convergence time threshold and the resource utilization threshold based on the performance trend prediction value to form a new evaluation benchmark;
[0045] The performance data and the threshold dynamic adjustment strategy are stored in an optimization knowledge base, and the optimization rules are extracted to generate scenario templates; the scenario templates and the new evaluation benchmark are used as guidance for the next round of query optimization.
[0046] According to a second aspect of the embodiments of the present invention,
[0047] Provides an unstructured database federated learning collaborative system based on group intelligence, including:
[0048] The first unit is used to collect query statement feature vectors, query resource consumption indicators, query response time indicators, and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; the initial training sample set is input into a preset deep neural network, and a query pattern recognition sub-model, a path planning sub-model, and a resource allocation sub-model are obtained through a federated learning framework training, and corresponding initial model parameters are generated; a swarm intelligence collaborative network is established based on the initial model parameters, and an initial path convergence time threshold and a resource utilization threshold are set as a network evaluation benchmark;
[0049] The second unit is used to send the query request input by the user into the query pattern recognition submodel for semantic analysis, extract the query semantic features, and generate multiple candidate query paths through the path planning submodel according to the initial path convergence time threshold; input the candidate query path into the hybrid swarm intelligence algorithm integrating the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism, and generate a query execution plan based on the resource utilization threshold dynamic optimization; the resource allocation submodel performs distributed resource scheduling according to the query execution plan to form a final query execution plan; execute the final query execution plan, and record the actual convergence time and resource occupancy of the path;
[0050] The third unit is used to perform differential analysis on the actual convergence time of the path and the initial path convergence time threshold, and the resource occupancy and the resource utilization threshold, to generate a performance optimization target; based on the performance optimization target, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; the optimized global model parameters are distributed to the swarm intelligence collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization.
[0051] According to a third aspect of the embodiments of the present invention,
[0052] An electronic device is provided, comprising:
[0053] processor;
[0054] a memory for storing processor-executable instructions;
[0055] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0056] A fourth aspect of the embodiments of the present invention is:
[0057] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0058] The beneficial effects of this application are as follows:
[0059] 1. Improve query efficiency: Through the collaborative work of query pattern recognition, path planning and resource allocation sub-models, as well as the optimization of hybrid swarm intelligence algorithms, the optimal query execution plan can be dynamically generated, thereby reducing query response time and improving query efficiency.
[0060] 2. Optimize resource utilization: Dynamic resource scheduling based on resource utilization thresholds, and differentiated analysis and model parameter updates based on actual resource occupancy can effectively avoid resource waste and improve resource utilization.
[0061] 3. Adaptive learning and optimization: The combination of the federated learning framework and the swarm intelligence collaborative network enables the system to continuously learn and optimize model parameters, path convergence time thresholds, and resource utilization thresholds based on user query requests and system operating status, thereby achieving adaptive query optimization and improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 A flowchart of a collaborative method for federated learning of unstructured databases based on swarm intelligence according to an embodiment of the present invention;
[0063] Figure 2 It is a structural diagram of an unstructured database federated learning collaborative system based on swarm intelligence according to an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0066] Figure 1 FIG. 1 is a flow chart of a collaborative method for federated learning of unstructured databases based on swarm intelligence according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0067] S11. Collect query statement feature vectors, query resource consumption indicators, query response time indicators, and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; input the initial training sample set into a preset deep neural network, train the query pattern recognition sub-model, path planning sub-model, and resource allocation sub-model through a federated learning framework, and generate corresponding initial model parameters; establish a swarm intelligence collaborative network based on the initial model parameters, and set an initial path convergence time threshold and a resource utilization threshold as a network evaluation benchmark;
[0068] S12. Send the query request input by the user to the query pattern recognition sub-model for semantic analysis, extract the query semantic features, and generate a candidate query path through the path planning sub-model according to the initial path convergence time threshold; input the candidate query path into the hybrid swarm intelligence algorithm that integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism, and generate a query execution plan based on the resource utilization threshold dynamic optimization; the resource allocation sub-model performs distributed resource scheduling according to the query execution plan to form a final query execution plan; execute the final query execution plan, and record the actual convergence time and resource occupancy of the path;
[0069] S13. Perform differential analysis on the actual path convergence time and the initial path convergence time threshold, and the resource occupancy and the resource utilization threshold to generate a performance optimization target; based on the performance optimization target, synchronously update the local model parameters of each node, and fuse the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; distribute the optimized global model parameters to the swarm intelligent collaborative network, dynamically adjust the initial path convergence time threshold and the resource utilization threshold, and form a new evaluation benchmark to guide the next round of query optimization.
[0070] In an optional implementation, the candidate query path is input into a hybrid swarm intelligence algorithm that integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, and the query execution plan is generated based on the resource utilization threshold dynamic optimization, including:
[0071] Construct an initial candidate query path set, map the execution order of each path in the initial candidate query path set into a path code in the form of an integer sequence, map the resource allocation strategy of the query operation into a parameter code in the form of a real vector, and establish a two-layer path representation model of path structure and resource allocation; calculate the query cost ratio of each candidate path based on the two-layer path representation model to generate an initial path evaluation index;
[0072] The ratio of the initial path evaluation index to the preset benchmark threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligence optimization model, and the hybrid swarm intelligence optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; based on the path transfer probability matrix, a path structure optimization sequence is generated by an ant colony algorithm;
[0073] The evaluation score of the path structure optimization sequence is used as the pheromone increment, and the pheromone increment is converted into the fitness function of the particle swarm algorithm; based on the fitness function, the particle swarm algorithm is guided to iteratively optimize the resource allocation parameters to generate a parameter optimization vector; system resource utilization data is collected, and when the resource utilization data exceeds a preset benchmark threshold, the value space of the parameter optimization vector is adaptively adjusted according to the degree of excess; the path structure optimization sequence and the parameter optimization vector are combined to form a complete query execution plan.
[0074] A query execution plan generation method based on a hybrid swarm intelligence algorithm that integrates ant colony path optimization mechanism and particle swarm parameter optimization mechanism can dynamically optimize query paths and resource allocation strategies for database query operations to minimize query costs while satisfying resource utilization constraints.
[0075] First, construct an initial candidate query path set. For example, for a query task containing three query operations (A, B, C), the possible execution orders include ABC, ACB, BAC, BCA, CAB, and CBA, which constitute the initial candidate query path set. Map the execution order of each path to a path encoding in the form of an integer sequence, for example, ABC is encoded as [1, 2, 3], ACB is encoded as [1, 3, 2], and so on. At the same time, map the resource allocation strategy of the query operation (such as the allocation ratio of CPU, memory, IO and other resources) to a parameter encoding in the form of a real vector, for example, [0.5, 0.3, 0.2] means that 50% of CPU resources, 30% of memory resources, and 20% of IO resources are allocated to the current operation. In this way, a two-layer path representation model of path structure and resource allocation is established.
[0076] Next, the query cost ratio of each candidate path is calculated based on the two-layer path representation model to generate the initial path evaluation index. The query cost ratio can be defined as the weighted average of query execution time and resource consumption. For example, the execution time of path ABC is 10 seconds, the resource consumption is 5 units, and the weighted average cost ratio is 0.8; the execution time of path ACB is 8 seconds, the resource consumption is 6 units, and the weighted average cost ratio is 0.7, and so on.
[0077] Then, the ratio of the initial path evaluation index to the preset benchmark threshold is used as the path heuristic information. Assuming the preset benchmark threshold is 0.6, the heuristic information of path ABC is 0.8 / 0.6=1.33, the heuristic information of path ACB is 0.7 / 0.6=1.17, and so on. Combined with the initial pheromone concentration (for example, the initial pheromone concentration of all paths is 1), the path transition probability matrix is calculated. The path transition probability matrix represents the probability of the ant selecting the next node at the current node. The higher the probability value, the greater the possibility that the path will be selected.
[0078] The path transfer probability matrix is input into the hybrid swarm intelligence optimization model. The model combines the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism. Based on the path transfer probability matrix, the ant colony algorithm generates a path structure optimization sequence. For example, after several iterations, the ant colony algorithm may generate a path optimization sequence of ACB-BAC-ABC, indicating that the query efficiency of these three paths is high.
[0079] The evaluation score of the path structure optimization sequence is used as the pheromone increment. For example, ACB has the highest evaluation score, so its pheromone increment is the largest. The pheromone increment is converted into the fitness function of the particle swarm algorithm. Based on the fitness function, the particle swarm algorithm is guided to iteratively optimize the resource allocation parameters to generate a parameter optimization vector. For example, after optimization, the parameter optimization vector may be [0.6, 0.2, 0.2], which means that 60% of CPU resources, 20% of memory resources, and 20% of IO resources are allocated to the current operation.
[0080] Collect system resource utilization data. When the resource utilization data exceeds the preset benchmark threshold, the value space of the parameter optimization vector is adaptively adjusted according to the degree of excess. For example, when the CPU utilization exceeds 80%, the CPU resource allocation ratio is reduced, and the allocation ratio of other resources is adjusted accordingly.
[0081] Finally, the path structure optimization sequence and the parameter optimization vector are combined to form a complete query execution plan. For example, the final query execution plan may be: first execute the query according to the path order of ACB and adopt the resource allocation strategy of [0.6, 0.2, 0.2]; if the resource utilization exceeds the threshold, dynamically adjust the parameter optimization vector.
[0082] The solution of this application can:
[0083] Improve query efficiency: By optimizing query paths and resource allocation strategies, query execution time can be effectively reduced and query efficiency can be improved. Reduce resource consumption: By dynamically adjusting resource allocation strategies, excessive resource use can be avoided and resource consumption can be reduced. Enhance system stability: By monitoring resource utilization and dynamically adjusting parameters, system overload can be avoided and system stability can be enhanced.
[0084] In an optional implementation, the ratio of the initial path evaluation index to the preset reference threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligence optimization model, and the hybrid swarm intelligence optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; based on the path transfer probability matrix, generating a path structure optimization sequence through an ant colony algorithm includes:
[0085] Calculating a ratio between the initial path evaluation index and a preset reference threshold to obtain path heuristic information, wherein the initial path evaluation index includes a path length index, a path energy consumption index, and a path time index, and the preset reference threshold includes a reference path length threshold, a reference path energy consumption threshold, and a reference path time threshold;
[0086] The path heuristic information and the initial pheromone concentration are combined to obtain a path transfer probability matrix, wherein the initial pheromone concentration adopts a fixed initial value, and the calculation result of the path transfer probability matrix includes a pheromone concentration weight coefficient and a heuristic information weight coefficient;
[0087] Inputting the path transfer probability matrix into a hybrid swarm intelligent optimization model, wherein the hybrid swarm intelligent optimization model integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, wherein the ant colony path optimization mechanism is responsible for path structure search, and the particle swarm parameter optimization mechanism is responsible for optimizing the pheromone concentration weight coefficient and the heuristic information weight coefficient;
[0088] Generate candidate paths through an ant colony algorithm based on the path transfer probability matrix, calculate the path evaluation score corresponding to the candidate path, determine the pheromone release amount according to the path evaluation score, and update the path transfer probability matrix based on the pheromone release amount and a preset pheromone volatility coefficient;
[0089] The path evaluation score is used as the fitness value of the particle swarm algorithm, the individual optimal solution and the global optimal solution of the particle swarm are updated based on the fitness value, and the particle speed update amount and the position update amount are calculated according to the individual optimal solution and the global optimal solution;
[0090] The updated particle positions are used as weight coefficient configurations to optimize the path transfer probability matrix, and the optimized path transfer probability matrix is re-input into the ant colony path optimization mechanism for iteration until convergence to obtain a path structure optimization sequence.
[0091] First, initialize the path information. Construct a network structure of the path to be optimized, which contains multiple nodes and paths connecting the nodes. Set the initial path evaluation indicators for each path, including path length, path energy consumption, and path time. Set the preset benchmark thresholds, including the benchmark path length threshold, the benchmark path energy consumption threshold, and the benchmark path time threshold, for example, set them to 10 kilometers, 20 kWh, and 30 minutes respectively.
[0092] Next, calculate the path heuristic information. Compare the length, energy consumption, and time indicators of each path with the corresponding benchmark thresholds and calculate the ratio. For example, if the length of a path is 12 kilometers, the heuristic information of its length indicator is 12 / 10 = 1.2. Similarly, calculate the heuristic information of the energy consumption indicator and the time indicator.
[0093] Then, construct the initial path transfer probability matrix. Set the initial pheromone concentration, for example, to 1. Combine the heuristic information of each path with the initial pheromone concentration to obtain the path transfer probability matrix. The combination operation can be performed by multiplying the pheromone concentration and the heuristic information by the corresponding weight coefficients and then adding them. The initial weight coefficients can be set to an equal value, for example, both are 0.5.
[0094] After that, the path transfer probability matrix is input into the hybrid swarm intelligence optimization model. This model combines the ant colony algorithm and the particle swarm algorithm. The ant colony algorithm is responsible for path structure search. According to the path transfer probability matrix, it simulates the movement of ants in the path network and gradually constructs candidate paths. The particle swarm algorithm is responsible for optimizing the pheromone concentration weight coefficient and the heuristic information weight coefficient to improve the efficiency of path search.
[0095] Based on the path transition probability matrix, the candidate paths are generated using the ant colony algorithm. The path evaluation score of each candidate path is calculated, which can be the weighted sum of the path length, energy consumption, and time. For example, if a candidate path is 11 kilometers long, consumes 18 kWh of energy, and takes 28 minutes, its path evaluation score can be calculated as 0.5 * 11 / 10 + 0.25 *18 / 20 + 0.25 * 28 / 30 = 1.0083.
[0096] The pheromone release amount is determined based on the evaluation score of the candidate path. The higher the score, the more pheromone is released. For example, the score can be directly used as the pheromone release amount. Then, the path transition probability matrix is updated based on the pheromone release amount and the preset pheromone volatility coefficient. The pheromone volatility coefficient can be set to 0.1, which means that the pheromone concentration decreases by 10% after each iteration.
[0097] The path evaluation score is used as the fitness value of the particle swarm algorithm. According to the fitness value, the individual optimal solution and the global optimal solution of the particle swarm are updated. Then, according to the individual optimal solution and the global optimal solution, the particle velocity update amount and position update amount are calculated, and the particle position is updated, that is, the pheromone concentration weight coefficient and the heuristic information weight coefficient.
[0098] The updated particle positions are used as weight coefficients to optimize the path transfer probability matrix. The optimized path transfer probability matrix is re-input into the ant colony algorithm for iteration, and the above steps are repeated until the algorithm converges to obtain the final path structure optimization sequence.
[0099] The solution of this application can:
[0100] Improved path search efficiency. By integrating ant colony algorithm and particle swarm algorithm, combined with path heuristic information and pheromone concentration, it is possible to find high-quality path solutions more quickly. Comprehensive optimization of path evaluation indicators is achieved. By comprehensively considering multiple indicators such as path length, energy consumption and time, a path solution that better meets actual needs can be found. It has strong adaptability. This method can adjust the preset benchmark threshold, weight coefficient, pheromone volatility coefficient and other parameters according to different application scenarios and needs to obtain the best optimization effect.
[0101] In an optional implementation, the resource allocation submodel performs distributed resource scheduling according to the query execution plan to form a final query execution plan; executing the final query execution plan and recording the actual path convergence time and resource occupancy include:
[0102] The resource allocation sub-model collects processor performance indicators, memory capacity indicators, network bandwidth indicators and storage capacity indicators of each resource pool to construct a resource feature matrix; calculates resource pool scores according to the resource feature matrix, and performs initial resource allocation for the input query execution plan based on the resource pool scores;
[0103] The resource allocation submodel collects the real-time resource usage data of each node according to the initial resource allocation result, and calculates the node load status in combination with the resource type weight factor; when the node load status exceeds the preset load threshold, the optimal migration target node is determined based on the task migration cost and load balancing benefit, and an optimized query execution plan is generated;
[0104] The resource allocation sub-model performs distributed resource scheduling on the optimized query execution plan, allocates execution resources based on the real-time computing capability, storage capacity and network status of each node, and forms a final query execution plan.
[0105] First, construct a resource feature matrix. Collect processor performance indicators of each resource pool, such as CPU main frequency, number of cores, etc.; collect memory capacity indicators, such as total memory size, available memory size, etc.; collect network bandwidth indicators, such as upload bandwidth, download bandwidth, etc.; collect storage capacity indicators, such as total disk capacity, available disk capacity, etc. Construct these collected indicators into a resource feature matrix, where each row represents a resource pool and each column represents a resource indicator. For example, the CPU main frequency of resource pool 1 is 3GHz, the number of cores is 8, the memory capacity is 32GB, the network bandwidth is 10Gbps, and the storage capacity is 1TB. The corresponding feature vector of resource pool 1 is [3, 8, 32, 10, 1024].
[0106] Next, calculate the resource pool score. According to the resource feature matrix, calculate the comprehensive score of each resource pool. For example, the score can be calculated by weighted summation, and each resource indicator is assigned a different weight, such as CPU weight 0.4, memory weight 0.3, network bandwidth weight 0.2, and storage capacity weight 0.1. Then multiply each indicator of each resource pool by the corresponding weight and sum them to get the score of the resource pool. Assuming that the scores of each indicator of resource pool 1 are: CPU score 0.8, memory score 0.9, network bandwidth score 0.7, storage capacity score 0.6, then the comprehensive score of resource pool 1 is 0.4 *0.8 + 0.3 * 0.9 + 0.2 * 0.7 + 0.1 * 0.6 = 0.77.
[0107] Then, perform initial resource allocation. Perform initial resource allocation for the input query execution plan based on the calculated resource pool scores. For example, a query plan contains three subtasks, each of which requires different computing resources and storage resources. Based on the resource pool scores, these subtasks are assigned to different resource pools for execution, with resource pools with higher scores being given priority. Assuming that resource pool 1 has the highest score, the computationally intensive subtask 1 is assigned to resource pool 1 for execution.
[0108] After that, collect the real-time resource usage data of the nodes. After the initial resource allocation is completed, collect the real-time resource usage data of each node, such as CPU usage, memory usage, network bandwidth usage, and disk IO rate. For example, the CPU usage of resource pool 1 is 80%, the memory usage is 70%, the network bandwidth usage is 60%, and the disk IO rate is 50%.
[0109] Next, calculate the node load status. Combine the resource type weight factors, such as CPU weight 0.5, memory weight 0.3, network bandwidth weight 0.1, and disk IO weight 0.1, to calculate the node load status. For example, the load status of resource pool 1 is 0.5 * 80% + 0.3 * 70% + 0.1 * 60% + 0.1 * 50% = 62%.
[0110] Then, determine the optimal migration target node. When the node load status exceeds the preset load threshold, such as 70%, task migration is required. Based on the task migration cost, such as the amount of migrated data and the network transmission speed, and the load balancing benefit, such as the degree of reduction in the node load status after migration, determine the optimal migration target node. For example, migrating some tasks on resource pool 1 to resource pool 2 with a lower load can reduce the load status of resource pool 1 and improve the overall query execution efficiency.
[0111] After that, distributed resource scheduling is performed. Distributed resource scheduling is performed on the optimized query execution plan, and execution resources are allocated based on the real-time computing power, storage capacity and network status of each node to form the final query execution plan.
[0112] Finally, execute the final query execution plan and record relevant information. Execute the finalized query execution plan and record the actual path convergence time and resource usage, including CPU usage, memory usage, network bandwidth usage, etc. This information can be used for subsequent resource scheduling optimization.
[0113] The solution of this application can:
[0114] Improve resource utilization: By dynamically adjusting the resource allocation plan, you can make full use of the computing resources and storage resources of each resource pool, avoid resource waste, and improve overall resource utilization. Reduce query latency: By assigning tasks to nodes with lower loads, you can reduce task queue waiting time, reduce query latency, and improve query efficiency. Enhance system stability: By monitoring the node load status in real time and making dynamic adjustments, you can avoid system crashes caused by excessive load on a single node and enhance system stability.
[0115] In an optional implementation, the actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold are differentially analyzed to generate a performance optimization target; based on the performance optimization target, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm, and reconstructs and generates optimized global model parameters, including:
[0116] Construct a path convergence time monitoring module and a resource occupancy monitoring module. The path convergence time monitoring module collects the actual path convergence time and compares it with the preset initial path convergence time threshold to generate a time dimension difference sequence. The resource occupancy monitoring module collects the resource occupancy data of each node and compares it with the preset resource utilization threshold to generate a resource dimension difference sequence.
[0117] Dynamic time warping is performed on the time dimension difference sequence, a time dimension cost matrix containing the deviation value of each time point is constructed, and the time difference score is obtained by solving the optimal matching path through a dynamic programming algorithm; the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate of the resource dimension difference sequence is calculated, and the resource difference score is obtained by combining the resource importance weight coefficient;
[0118] The time difference score and the resource difference score are normalized and fused to generate a performance difference index based on the harmonic mean method; a performance optimization objective function including a stability constraint is constructed according to the performance difference index, and a weight coefficient of the performance optimization objective function is adaptively adjusted according to the degree of fluctuation of the time difference score and the resource difference score;
[0119] The federated learning framework trains the local model in parallel on each node based on the performance optimization objective function, updates the local model parameters using the stochastic gradient descent method with momentum term, and dynamically adjusts the learning rate according to the time difference score during the training process; shards the updated local model parameters, and applies the homomorphic encryption algorithm to the parameter shards to generate encrypted parameter blocks;
[0120] Allocate a key shard related to the resource difference score to each node, combine the key shard with random mask information to generate a secure aggregation key; use the secure aggregation key to perform distributed aggregation operations on the encryption parameter blocks to generate a global parameter aggregation matrix; calculate the node weight coefficient according to the time difference score and the resource difference score of each node, perform weighted averaging on the global parameter aggregation matrix based on the node weight coefficient, perform sparse processing on the weighted averaged parameters, and reconstruct and generate optimized global model parameters.
[0121] First, build a path convergence time monitoring module and a resource occupancy monitoring module. The path convergence time monitoring module is responsible for collecting the time taken for the path to actually converge during the model training process, and compares it with the preset initial path convergence time threshold to generate a sequence reflecting the time dimension difference. For example, if the preset threshold is 10 minutes, and the actual time is 9 minutes, 11 minutes, and 8 minutes, respectively, the generated time dimension difference sequence is -1, 1, and -2. The resource occupancy monitoring module is responsible for collecting the resource occupancy data of each node, including processor occupancy, memory occupancy, and network bandwidth occupancy, and compares the collected data with the preset resource utilization threshold to generate a sequence reflecting the resource dimension difference. For example, assuming that the processor occupancy threshold is 70%, and the actual occupancy is 65%, 75%, and 60%, respectively, the generated resource dimension difference sequence is -5, 5, and -10.
[0122] Next, the generated time dimension difference sequence and resource dimension difference sequence are analyzed and processed. The time dimension difference sequence is dynamically time-warped to construct a time dimension cost matrix containing the deviation values of each time point, and then the dynamic programming algorithm is used to solve the optimal matching path to obtain the time difference score. For example, the time difference score calculated by dynamic time warping and dynamic programming algorithms is 0.8. For the resource dimension difference sequence, the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate, and network bandwidth occupancy rate are calculated respectively, and the resource difference score is obtained by combining the resource importance weight coefficient (for example, the weight coefficients of the processor, memory, and network bandwidth are 0.5, 0.3, and 0.2, respectively). For example, the calculated resource difference score is 1.2.
[0123] Then, the time difference score and the resource difference score are normalized so that their values fall within the same numerical range, for example, between 0 and 1. After that, the two scores are fused based on the harmonic mean method to generate a performance difference index. For example, if the normalized time difference score is 0.7 and the resource difference score is 0.8, the performance difference index calculated according to the harmonic mean method is 0.75. A performance optimization objective function with stability constraints is constructed based on the performance difference index. The weight coefficient of the performance optimization objective function is adaptively adjusted according to the degree of fluctuation of the time difference score and the resource difference score. The greater the fluctuation, the smaller the weight coefficient.
[0124] The federated learning framework trains local models in parallel on each node according to the performance optimization objective function. The stochastic gradient descent method with momentum term is used to update the local model parameters, and the learning rate is dynamically adjusted according to the time difference score during the training process. The larger the time difference score, the smaller the learning rate. The updated local model parameters are sharded, and the homomorphic encryption algorithm is applied to the parameter shards to generate encrypted parameter blocks. For example, the model parameters are divided into 10 shards and encrypted separately.
[0125] Assign key shards related to the resource difference score to each node. The higher the resource difference score, the more important the assigned key shards. Combine the key shards with the random mask information to generate a secure aggregation key. Use the secure aggregation key to perform distributed aggregation operations on the encrypted parameter blocks to generate a global parameter aggregation matrix. Calculate the node weight coefficient based on the time difference score and resource difference score of each node. The lower the time difference score and resource difference score, the higher the node weight coefficient. Perform weighted averaging on the global parameter aggregation matrix based on the node weight coefficients, perform sparse processing on the weighted averaged parameters, remove the parameters with smaller weights, and reconstruct the optimized global model parameters.
[0126] The solution of this application can:
[0127] Improve model training efficiency: Accelerate model convergence and shorten training time by dynamically adjusting learning rate and optimizing model parameters. Optimize resource utilization: Improve resource utilization and reduce training costs by monitoring resource occupancy and adjusting resource allocation strategies. Enhance model robustness: Improve the generalization and robustness of the model and enhance the practical value of the model through the federated learning framework and secure aggregation algorithm.
[0128] In an optional implementation, the time dimension difference sequence is dynamically time-warped, a time dimension cost matrix containing deviation values at each time point is constructed, and the optimal matching path is solved by a dynamic programming algorithm to obtain a time difference score; the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate is calculated for the resource dimension difference sequence, and the resource difference score is obtained by combining the resource importance weight coefficient, including:
[0129] Constructing a time dimension cost matrix according to the time dimension difference sequence, wherein each element of the time dimension cost matrix comprises a time point difference distance term and a time jump penalty term, wherein the time point difference distance term represents the degree of deviation of corresponding points in the time series, and the time jump penalty term constrains the nonlinear distortion of the time series through a preset penalty coefficient;
[0130] The time dimension cost matrix is subjected to cumulative cost calculation to obtain a cumulative cost matrix, and the cumulative cost matrix is used to find a minimum cumulative cost path through a dynamic programming algorithm; based on the minimum cumulative cost path, an optimal matching path is obtained by backtracking, and the cumulative cost on the optimal matching path is normalized to obtain a time dimension scoring sequence;
[0131] Collect the real-time resource occupancy data of each node in the system, including processor occupancy, memory occupancy and network bandwidth occupancy, calculate the difference between the real-time resource occupancy data and the preset resource utilization threshold to obtain the resource dimension difference sequence; use the hierarchical analysis method to construct a resource indicator judgment matrix, and obtain the importance weight coefficient corresponding to each resource indicator by calculating the eigenvector of the resource indicator judgment matrix;
[0132] The resource dimension difference sequence is sampled based on a sliding observation window of a preset size, and the statistical features of the resource dimension difference sequence within the sliding observation window are extracted; the resource dimension difference sequence and the importance weight coefficient are weighted and calculated, and the resource difference score is obtained using the weighted Euclidean distance metric.
[0133] First, construct the time dimension cost matrix. Collect the time series data of the system to be analyzed, such as the response time series of a service. Assume that the collected time series is T1 = [10, 12, 11, 15, 13], and the reference time series is T2 = [9, 11, 12, 14, 16]. Calculate the degree of deviation of the corresponding time points in the time series T1 and T2. For example, the deviation of the first time point is |10-9|=1. At the same time, in order to constrain the nonlinear distortion of the time series, introduce a time jump penalty term. For example, if you jump from the first time point of T1 to the second time point of T2, introduce a penalty coefficient. Assuming that the penalty coefficient is 0.5, the jump penalty term is 0.5. Add the time point difference distance term and the time jump penalty term to form the elements of the time dimension cost matrix. For example, the cost of the first time point of T1 and the first time point of T2 is 1+0=1, and the cost of the first time point of T1 and the second time point of T2 is |10-11|+0.5=1.5. And so on, construct a complete time dimension cost matrix.
[0134] Next, the cumulative cost of the time dimension cost matrix is calculated. Starting from the upper left corner element of the cost matrix, the cumulative cost is calculated element by element. The cumulative cost is calculated as follows: the cost of the current element plus the minimum cumulative cost of the three adjacent elements to its left, above, and above the left. For example, the cumulative cost of the upper left corner element of the cost matrix is its own cost. The cumulative cost of the second element is its own cost plus the cumulative cost of the upper left corner element. And so on, the complete cumulative cost matrix is calculated.
[0135] Then, the optimal matching path is obtained by backtracking based on the cumulative cost matrix. Starting from the lower right corner element of the cumulative cost matrix, find the element with the smallest cumulative cost among the three adjacent elements on its left, above, and above the left, and record the position of the element. Repeat this process until you backtrack to the upper left corner element, and finally get a path from the upper left corner to the lower right corner, which is the optimal matching path.
[0136] After that, the accumulated cost on the optimal matching path is normalized to obtain the time dimension scoring sequence. The accumulated cost of each element on the optimal matching path is divided by the path length to obtain the normalized cost, which constitutes the time dimension scoring sequence.
[0137] Next, collect the real-time resource usage data of each node in the system, including processor usage, memory usage, and network bandwidth usage. Assume that the collected processor usage is 70%, memory usage is 80%, and network bandwidth usage is 60%. The preset resource utilization thresholds are 80%, 90%, and 70%, respectively. Then the resource dimension difference sequences are |70%-80%|=10%, |80%-90%|=10%, and |60%-70%|=10%, respectively.
[0138] Then, the hierarchical analysis method is used to construct the resource indicator judgment matrix, and the importance weight coefficient corresponding to each resource indicator is obtained by calculating the eigenvector of the resource indicator judgment matrix. Assume that the importance weight coefficients of processor, memory and network bandwidth are 0.4, 0.3 and 0.3 respectively.
[0139] Then, the resource dimension difference sequence is sampled based on a sliding observation window of a preset size. Assuming that the sliding observation window size is 5, the resource dimension difference sequence is sampled to extract statistical features within the window, such as the average value.
[0140] Finally, the resource dimension difference sequence and the importance weight coefficient are weighted and the resource difference score is obtained using the weighted Euclidean distance metric. Each element of the resource dimension difference sequence is multiplied by its corresponding weight coefficient, and then the three results are added together to obtain the final resource difference score. For example, the resource difference score is 0.4 * 10% + 0.3 *10% + 0.3 * 10% = 10%.
[0141] The solution of this application can:
[0142] Improve the accuracy of anomaly detection: Comprehensively consider the differences in time and resource dimensions to more accurately identify system anomalies. Reduce the false alarm rate: Through methods such as dynamic time warping and weighted Euclidean distance, the false alarm rate is effectively reduced. Enhance the robustness of anomaly detection: The use of technologies such as sliding observation windows and statistical feature extraction enhances the robustness of anomaly detection.
[0143] In an optional implementation, the optimized global model parameters are distributed to the swarm intelligence collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization, including:
[0144] Constructing the topological structure of the swarm intelligence collaborative network, calculating the adjacency matrix based on the physical distance between nodes and the quality of network links; using the spectral clustering algorithm to divide the adjacency matrix into multiple sub-network groups, and selecting a regional coordinator in each of the sub-network groups based on the node computing power and network connection stability;
[0145] Collecting hardware configuration parameters and task load information of each node, performing differentiated processing on the optimized global model parameters based on the hardware configuration parameters and the task load information, and generating parameter deployment schemes adapted to different nodes; constructing a parameter integrity check matrix using a Bloom filter, and distributing the parameter deployment scheme to each node in the swarm intelligence collaborative network through the regional coordinator;
[0146] Deploy a performance collection agent in the swarm intelligence collaborative network to collect the path convergence time and resource utilization after the node executes the parameter deployment scheme; calculate the deviation value between the path convergence time and the initial path convergence time threshold, and the difference value between the resource utilization and the resource utilization threshold based on a sliding time window;
[0147] The deviation value and the difference value are input into an orthogonal test model to analyze the influence of the initial path convergence time threshold and the resource utilization threshold on system performance; a principal component analysis method is used to extract threshold combinations, and a support vector regression model is established to map the relationship between performance differences and threshold adjustment amounts;
[0148] Construct a dual Q-learning network, use the performance difference as the state input, and use the threshold adjustment direction and step size as the action output; design a reward function based on the performance improvement metric, train a threshold dynamic adjustment strategy, and adaptively adjust the exploration probability of the dual Q-learning network according to the degree of performance fluctuation;
[0149] Performing exponential moving average processing on the performance data after executing the threshold dynamic adjustment strategy, and generating a performance trend prediction value in combination with a time series prediction model; dynamically adjusting the initial path convergence time threshold and the resource utilization threshold based on the performance trend prediction value to form a new evaluation benchmark;
[0150] The performance data and the threshold dynamic adjustment strategy are stored in an optimization knowledge base, and the optimization rules are extracted to generate scenario templates; the scenario templates and the new evaluation benchmark are used as guidance for the next round of query optimization.
[0151] The swarm intelligence collaborative query optimization method aims to improve the efficiency of query processing and resource utilization in a distributed environment. This method optimizes the query execution strategy by dynamically adjusting the path convergence time threshold and resource utilization threshold, and combining the swarm intelligence collaborative network.
[0152] First, construct a swarm intelligence collaborative network. Calculate the adjacency matrix based on the physical distance between nodes and the quality of the network link. For example, the physical distance between node A and node B is 10 kilometers, and the network link quality rating is 5 (the rating range is 1-5, 5 represents the best), then the element values corresponding to A and B in the adjacency matrix are 5 / 10=0.5. Then use the spectral clustering algorithm to divide the adjacency matrix into multiple sub-network groups. Assume that the spectral clustering algorithm divides the network into 3 groups. In each sub-network group, select the regional coordinator based on the computing power of the node and the stability of the network connection. For example, in group 1, node C has a computing power of 1000MIPS and a network connection stability of 99%, and is selected as the regional coordinator.
[0153] Next, collect the hardware configuration parameters and task load information of each node. For example, the hardware configuration parameters of node D include CPU main frequency 2.5GHz, memory 8GB, hard disk capacity 500GB, and current task load of 50%. Based on this information, the optimized global model parameters are differentiated to generate parameter deployment schemes adapted to different nodes. For example, for nodes with weaker computing power, a smaller subset of model parameters is allocated. The parameter integrity check matrix is constructed using the Bloom filter to ensure the consistency and integrity of parameter distribution. For example, the hash value of the model parameter is stored in the Bloom filter, and after receiving the parameter, the node verifies the integrity of the parameter by verifying the hash value. Finally, the parameter deployment scheme is distributed to each node in the swarm intelligence collaborative network through the regional coordinator.
[0154] Deploy a performance collection agent in the swarm intelligence collaborative network. The performance collection agent collects the path convergence time and resource utilization of the node after executing the parameter deployment scheme. For example, after node E executes the query task, the path convergence time is 10 milliseconds, the CPU utilization is 70%, and the memory utilization is 60%. Based on the sliding time window (for example, the window size is set to 10 seconds), calculate the deviation value of the path convergence time from the initial path convergence time threshold, and the difference value of the resource utilization from the resource utilization threshold. Assuming that the initial path convergence time threshold is 15 milliseconds and the resource utilization threshold is 80%, the path convergence time deviation value of node E is -5 milliseconds, and the resource utilization difference value is -10%.
[0155] The deviation value and the difference value are input into the orthogonal test model to analyze the impact of the initial path convergence time threshold and the resource utilization threshold on system performance. For example, through the orthogonal test, it is found that lowering the path convergence time threshold can significantly improve system performance, while increasing the resource utilization threshold has little impact on system performance. The principal component analysis method is used to extract the threshold combination, and a support vector regression model is established to map the relationship between performance differences and threshold adjustment amounts. For example, the support vector regression model can predict that lowering the path convergence time threshold by 2 milliseconds can improve system performance by 5%.
[0156] Construct a dual Q-learning network, use the performance difference as the state input, and the threshold adjustment direction and step size as the action output. For example, if the state is "the path convergence time deviation value is -5 milliseconds, and the resource utilization difference value is -10%", possible actions include "lower the path convergence time threshold by 2 milliseconds" and "increase the resource utilization threshold by 5%". Design a reward function based on the performance improvement metric. For example, the greater the performance improvement, the higher the reward value. Train a threshold dynamic adjustment strategy to adaptively adjust the exploration probability of the dual Q-learning network according to the degree of performance fluctuation. For example, when the performance fluctuation is large, increase the exploration probability to find a better threshold adjustment strategy.
[0157] Perform exponential moving average processing on the performance data after executing the threshold dynamic adjustment strategy, and generate performance trend prediction values in combination with the time series prediction model. For example, it is predicted that the system performance will increase by 10% in the next 5 minutes. Dynamically adjust the initial path convergence time threshold and resource utilization threshold based on the performance trend prediction value to form a new evaluation benchmark. For example, according to the prediction results, adjust the path convergence time threshold to 13 milliseconds and the resource utilization threshold to 75%.
[0158] Store performance data and threshold dynamic adjustment strategies in the optimization knowledge base, extract optimization rules and generate scenario-based templates. For example, generate corresponding threshold adjustment templates for high-concurrency query scenarios. Use scenario-based templates and new evaluation benchmarks as guidance for the next round of query optimization.
[0159] The solution of this application can:
[0160] Improve query efficiency: By dynamically adjusting thresholds and optimizing query execution strategies, query path convergence time can be shortened, thereby improving query efficiency. Optimize resource utilization: This method can perform differentiated parameter deployment based on the hardware configuration and task load of the node, and dynamically adjust the resource utilization threshold to optimize resource utilization and avoid resource waste. Enhance system adaptability: Through dual Q learning networks and performance trend prediction, this method can dynamically adjust parameters according to the system operating status, enhance the system's adaptability, and enable it to better cope with different query loads and environmental changes.
[0161] Figure 2 FIG. 1 is a schematic diagram of the structure of the unstructured database federated learning collaborative system based on swarm intelligence according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0162] The first unit is used to collect query statement feature vectors, query resource consumption indicators, query response time indicators, and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; the initial training sample set is input into a preset deep neural network, and a query pattern recognition sub-model, a path planning sub-model, and a resource allocation sub-model are obtained through a federated learning framework training, and corresponding initial model parameters are generated; a swarm intelligence collaborative network is established based on the initial model parameters, and an initial path convergence time threshold and a resource utilization threshold are set as a network evaluation benchmark;
[0163] The second unit is used to send the query request input by the user into the query pattern recognition submodel for semantic analysis, extract the query semantic features, and generate multiple candidate query paths through the path planning submodel according to the initial path convergence time threshold; input the candidate query path into the hybrid swarm intelligence algorithm integrating the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism, and generate a query execution plan based on the resource utilization threshold dynamic optimization; the resource allocation submodel performs distributed resource scheduling according to the query execution plan to form a final query execution plan; execute the final query execution plan, and record the actual convergence time and resource occupancy of the path;
[0164] The third unit is used to perform differential analysis on the actual convergence time of the path and the initial path convergence time threshold, and the resource occupancy and the resource utilization threshold, to generate a performance optimization target; based on the performance optimization target, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; the optimized global model parameters are distributed to the swarm intelligence collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization.
[0165] According to a third aspect of the embodiments of the present invention,
[0166] An electronic device is provided, comprising:
[0167] processor;
[0168] a memory for storing processor-executable instructions;
[0169] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0170] A fourth aspect of the embodiments of the present invention is:
[0171] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0172] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. Unstructured database federated learning collaborative method based on swarm intelligence, characterized by: include: Collect query statement feature vectors, query resource consumption indicators, query response time indicators, and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; input the initial training sample set into a preset deep neural network, train the query pattern recognition sub-model, path planning sub-model, and resource allocation sub-model through a federated learning framework, and generate corresponding initial model parameters; establish a swarm intelligence collaborative network based on the initial model parameters, and set an initial path convergence time threshold and a resource utilization threshold as a network evaluation benchmark; The query request input by the user is sent to the query pattern recognition sub-model for semantic analysis, query semantic features are extracted, and multiple candidate query paths are generated through the path planning sub-model according to the initial path convergence time threshold; the candidate query path is input into a hybrid swarm intelligence algorithm that integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism, and a query execution plan is generated based on the resource utilization threshold dynamic optimization; the resource allocation sub-model performs distributed resource scheduling according to the query execution plan to form a final query execution plan; the final query execution plan is executed, and the actual path convergence time and resource occupancy are recorded; Perform differential analysis on the actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold, and generate a performance optimization target; Based on the performance optimization goal, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; the optimized global model parameters are distributed to the swarm intelligence collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization.
2. The method according to claim 1, characterized in that: Inputting the candidate query path into a hybrid swarm intelligence algorithm that integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, and dynamically optimizing and generating a query execution plan based on the resource utilization threshold comprises: Construct an initial candidate query path set, map the execution order of each path in the initial candidate query path set into a path code in the form of an integer sequence, map the resource allocation strategy of the query operation into a parameter code in the form of a real vector, and establish a two-layer path representation model of path structure and resource allocation; calculate the query cost ratio of each candidate path based on the two-layer path representation model to generate an initial path evaluation index; The ratio of the initial path evaluation index to the preset benchmark threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligence optimization model, and the hybrid swarm intelligence optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; based on the path transfer probability matrix, a path structure optimization sequence is generated by an ant colony algorithm; The evaluation score of the path structure optimization sequence is used as the pheromone increment, and the pheromone increment is converted into the fitness function of the particle swarm algorithm; based on the fitness function, the particle swarm algorithm is guided to iteratively optimize the resource allocation parameters to generate a parameter optimization vector; system resource utilization data is collected, and when the resource utilization data exceeds a preset benchmark threshold, the value space of the parameter optimization vector is adaptively adjusted according to the degree of excess; the path structure optimization sequence and the parameter optimization vector are combined to form a complete query execution plan.
3. The method according to claim 2, characterized in that The ratio of the initial path evaluation index to the preset benchmark threshold is used as path heuristic information, and the path transfer probability matrix is calculated in combination with the initial pheromone concentration; the path transfer probability matrix is input into a hybrid swarm intelligent optimization model, and the hybrid swarm intelligent optimization model integrates the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism; Based on the path transition probability matrix, generating a path structure optimization sequence by using an ant colony algorithm includes: Calculating a ratio between the initial path evaluation index and a preset reference threshold to obtain path heuristic information, wherein the initial path evaluation index includes a path length index, a path energy consumption index, and a path time index, and the preset reference threshold includes a reference path length threshold, a reference path energy consumption threshold, and a reference path time threshold; The path heuristic information and the initial pheromone concentration are combined to obtain a path transfer probability matrix, wherein the initial pheromone concentration adopts a fixed initial value, and the calculation of the path transfer probability matrix includes a pheromone concentration weight coefficient and a heuristic information weight coefficient; Inputting the path transfer probability matrix into a hybrid swarm intelligent optimization model, wherein the hybrid swarm intelligent optimization model integrates an ant colony path optimization mechanism and a particle swarm parameter optimization mechanism, wherein the ant colony path optimization mechanism is responsible for path structure search, and the particle swarm parameter optimization mechanism is responsible for optimizing the pheromone concentration weight coefficient and the heuristic information weight coefficient; Generate candidate paths through an ant colony algorithm based on the path transfer probability matrix, calculate the path evaluation score corresponding to the candidate path, determine the pheromone release amount according to the path evaluation score, and update the path transfer probability matrix based on the pheromone release amount and a preset pheromone volatility coefficient; The path evaluation score is used as the fitness value of the particle swarm algorithm, the individual optimal solution and the global optimal solution of the particle swarm are updated based on the fitness value, and the particle speed update amount and the position update amount are calculated according to the individual optimal solution and the global optimal solution; The updated particle positions are used as weight coefficient configurations to optimize the path transfer probability matrix, and the optimized path transfer probability matrix is re-input into the ant colony path optimization mechanism for iteration until convergence to obtain a path structure optimization sequence.
4. The method according to claim 1, characterized in that: The resource allocation sub-model performs distributed resource scheduling according to the query execution plan to form a final query execution plan; executing the final query execution plan, and recording the actual convergence time and resource occupancy of the path include: Construct a resource allocation sub-model including a high-performance resource pool, a standard resource pool, and a basic resource pool, wherein the resource allocation sub-model collects processor performance indicators, memory capacity indicators, network bandwidth indicators, and storage capacity indicators of each resource pool to construct a resource feature matrix; calculate resource pool scores according to the resource feature matrix, and perform initial resource allocation for an input query execution plan based on the resource pool scores; The resource allocation submodel collects the real-time resource usage data of each node according to the initial resource allocation result, and calculates the node load status in combination with the resource type weight factor; when the node load status exceeds the preset load threshold, the optimal migration target node is determined based on the task migration cost and load balancing benefit, and an optimized query execution plan is generated; The resource allocation sub-model performs distributed resource scheduling on the optimized query execution plan, allocates execution resources based on the real-time computing capability, storage capacity and network status of each node, and forms a final query execution plan.
5. The method according to claim 1, characterized in that Perform differential analysis on the actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold, and generate a performance optimization target; Based on the performance optimization goal, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters including: Construct a path convergence time monitoring module and a resource occupancy monitoring module. The path convergence time monitoring module collects the actual path convergence time and compares it with the preset initial path convergence time threshold to generate a time dimension difference sequence. The resource occupancy monitoring module collects the resource occupancy data of each node and compares it with the preset resource utilization threshold to generate a resource dimension difference sequence. Dynamic time warping is performed on the time dimension difference sequence, a time dimension cost matrix containing the deviation value of each time point is constructed, and the time difference score is obtained by solving the optimal matching path through a dynamic programming algorithm; the weighted Euclidean distance of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate of the resource dimension difference sequence is calculated, and the resource difference score is obtained by combining the resource importance weight coefficient; The time difference score and the resource difference score are normalized and fused to generate a performance difference index based on the harmonic mean method; a performance optimization objective function including a stability constraint is constructed according to the performance difference index, and a weight coefficient of the performance optimization objective function is adaptively adjusted according to the degree of fluctuation of the time difference score and the resource difference score; The federated learning framework trains the local model in parallel on each node based on the performance optimization objective function, updates the local model parameters using the stochastic gradient descent method with momentum term, and dynamically adjusts the learning rate according to the time difference score during the training process; shards the updated local model parameters, and applies the homomorphic encryption algorithm to the parameter shards to generate encrypted parameter blocks; Allocate a key shard related to the resource difference score to each node, combine the key shard with random mask information to generate a secure aggregation key; use the secure aggregation key to perform distributed aggregation operations on the encryption parameter blocks to generate a global parameter aggregation matrix; calculate the node weight coefficient according to the time difference score and the resource difference score of each node, perform weighted averaging on the global parameter aggregation matrix based on the node weight coefficient, perform sparse processing on the weighted averaged parameters, and reconstruct and generate optimized global model parameters.
6. The method according to claim 5, characterized in that Dynamically time warping the time dimension difference sequence, constructing a time dimension cost matrix including the deviation value of each time point, and solving the optimal matching path through a dynamic programming algorithm to obtain a time difference score; The weighted Euclidean distances of the processor occupancy rate, memory occupancy rate and network bandwidth occupancy rate are calculated for the resource dimension difference sequence, and the resource difference score is obtained by combining the resource importance weight coefficient, including: Constructing a time dimension cost matrix according to the time dimension difference sequence, wherein each element of the time dimension cost matrix comprises a time point difference distance term and a time jump penalty term, wherein the time point difference distance term represents the degree of deviation of corresponding points in the time series, and the time jump penalty term constrains the nonlinear distortion of the time series through a preset penalty coefficient; The time dimension cost matrix is subjected to cumulative cost calculation to obtain a cumulative cost matrix, and the cumulative cost matrix is used to find a minimum cumulative cost path through a dynamic programming algorithm; based on the minimum cumulative cost path, an optimal matching path is obtained by backtracking, and the cumulative cost on the optimal matching path is normalized to obtain a time dimension scoring sequence; Collect the real-time resource occupancy data of each node in the system, including processor occupancy, memory occupancy and network bandwidth occupancy, calculate the difference between the real-time resource occupancy data and the preset resource utilization threshold to obtain the resource dimension difference sequence; use the hierarchical analysis method to construct a resource indicator judgment matrix, and obtain the importance weight coefficient corresponding to each resource indicator by calculating the eigenvector of the resource indicator judgment matrix; The resource dimension difference sequence is sampled based on a sliding observation window of a preset size, and the statistical features of the resource dimension difference sequence within the sliding observation window are extracted; the resource dimension difference sequence and the importance weight coefficient are weighted and calculated, and the resource difference score is obtained using the weighted Euclidean distance metric.
7. The method according to claim 1, characterized in that Distributing the optimized global model parameters to the swarm intelligence collaborative network, dynamically adjusting the initial path convergence time threshold and the resource utilization threshold, and forming a new evaluation benchmark for guiding the next round of query optimization includes: Constructing the topological structure of the swarm intelligence collaborative network, calculating the adjacency matrix based on the physical distance between nodes and the quality of network links; using the spectral clustering algorithm to divide the adjacency matrix into multiple sub-network groups, and selecting a regional coordinator in each of the sub-network groups based on the node computing power and network connection stability; Collecting hardware configuration parameters and task load information of each node, performing differentiated processing on the optimized global model parameters based on the hardware configuration parameters and the task load information, and generating parameter deployment schemes adapted to different nodes; constructing a parameter integrity check matrix using a Bloom filter, and distributing the parameter deployment scheme to each node in the swarm intelligence collaborative network through the regional coordinator; Deploy a performance collection agent in the swarm intelligence collaborative network to collect the path convergence time and resource utilization after the node executes the parameter deployment scheme; calculate the deviation value between the path convergence time and the initial path convergence time threshold, and the difference value between the resource utilization and the resource utilization threshold based on a sliding time window; The deviation value and the difference value are input into an orthogonal test model to analyze the influence of the initial path convergence time threshold and the resource utilization threshold on system performance; a principal component analysis method is used to extract key threshold combinations, and a support vector regression model is established to map the relationship between performance differences and threshold adjustment amounts; Construct a dual Q-learning network, use the performance difference as the state input, and use the threshold adjustment direction and step size as the action output; design a reward function based on the performance improvement metric, train a threshold dynamic adjustment strategy, and adaptively adjust the exploration probability of the dual Q-learning network according to the degree of performance fluctuation; Performing exponential moving average processing on the performance data after executing the threshold dynamic adjustment strategy, and generating a performance trend prediction value in combination with a time series prediction model; dynamically adjusting the initial path convergence time threshold and the resource utilization threshold based on the performance trend prediction value to form a new evaluation benchmark; The performance data and the threshold dynamic adjustment strategy are stored in an optimization knowledge base, and the optimization rules are extracted to generate scenario templates; the scenario templates and the new evaluation benchmark are used as guidance for the next round of query optimization.
8. An unstructured database federated learning collaborative system based on swarm intelligence, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect query statement feature vectors, query resource consumption indicators, query response time indicators and query result accuracy indicators of multiple unstructured database nodes, and construct the collected data into an initial training sample set; Input the initial training sample set into a preset deep neural network, train the query pattern recognition sub-model, the path planning sub-model and the resource allocation sub-model through a federated learning framework, and generate corresponding initial model parameters; establish a swarm intelligence collaborative network based on the initial model parameters, and set an initial path convergence time threshold and a resource utilization threshold as a network evaluation benchmark; The second unit is used to send the query request input by the user into the query pattern recognition submodel for semantic analysis, extract the query semantic features, and generate multiple candidate query paths through the path planning submodel according to the initial path convergence time threshold; input the candidate query path into the hybrid swarm intelligence algorithm integrating the ant colony path optimization mechanism and the particle swarm parameter optimization mechanism, and generate a query execution plan based on the resource utilization threshold dynamic optimization; the resource allocation submodel performs distributed resource scheduling according to the query execution plan to form a final query execution plan; execute the final query execution plan, and record the actual convergence time and resource occupancy of the path; The third unit is used to perform differential analysis on the actual path convergence time and the initial path convergence time threshold, the resource occupancy and the resource utilization threshold, and generate a performance optimization target; Based on the performance optimization goal, the federated learning framework synchronously updates the local model parameters of each node, and fuses the updated local model parameters through a secure aggregation algorithm to reconstruct and generate optimized global model parameters; the optimized global model parameters are distributed to the swarm intelligence collaborative network, and the initial path convergence time threshold and the resource utilization threshold are dynamically adjusted to form a new evaluation benchmark for guiding the next round of query optimization.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Database query optimization method based on ant colony genetic dynamic fusion algorithm
CN115391385A
Cost optimization method for efficient federated learning in Internet of Things scene
CN116193516A