Data distribution method, system and equipment for distributed computing and medium
By generating task-specific data flow identifiers in a distributed computing system and combining it with a multi-dimensional scoring mechanism, the singleness problem of task scheduling strategies in existing technologies is solved, and more efficient resource utilization and task execution are achieved.
Patent Information
- Application Number
- CN202510934261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, task scheduling strategies mostly focus on static resource allocation or single-dimensional optimization, and lack a mechanism to track the correlation between tasks, resulting in low resource utilization and task delays.
By generating task-specific data flow identifiers, combining resource status, cache status and data flow processing history information, and adopting a multi-dimensional comprehensive scoring mechanism, sub-query tasks are dynamically allocated to the most suitable computing nodes.
It improves the accuracy and execution efficiency of task allocation, optimizes the resource utilization and task response speed of distributed computing systems, and reduces the task failure rate.
Smart Images

Figure CN120803657A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data distribution, and in particular to a data distribution method, system, device and medium for distributed computing. BACKGROUND
[0002] With the rapid development of big data and cloud computing technology, distributed computing systems face challenges in task allocation efficiency and resource utilization when processing massive data. Traditional distributed computing frameworks (such as MapReduce, Spark) usually schedule tasks based on node resource status (such as CPU, memory, network bandwidth), and the core logic is to allocate computing tasks to nodes with sufficient resources to improve execution efficiency. However, such methods often ignore the impact of data storage characteristics (such as cache status, compression efficiency) and historical processing experience on task allocation, resulting in limited resource utilization and difficulty in adapting to dynamic needs in complex query scenarios.
[0003] In the prior art, task scheduling strategies focus on static resource allocation or single-dimensional optimization, such as reducing network transmission overhead through data locality, or avoiding node overload based on load balancing. However, in actual application scenarios, the cache management capabilities of computing nodes (such as compression data decompression efficiency, cache partition utilization) and historical task processing patterns (such as the execution success rate of specific type queries) have a significant impact on task allocation effectiveness. In addition, traditional methods lack tracking mechanisms for task correlation, making it difficult to achieve cross-node collaborative optimization through historical data flow identification. This limitation leads to problems such as resource waste and task delay when processing complex queries, and an intelligent scheduling scheme that integrates multi-dimensional state evaluation and historical experience analysis is urgently needed. SUMMARY
[0004] In view of the above problems, the present application is proposed.
[0005] Therefore, the technical problem solved by the present application is that in the prior art, task scheduling strategies focus on static resource allocation or single-dimensional optimization, and lack tracking mechanisms for task correlation.
[0006] To solve the above technical problems, the present application provides the following technical solutions: a data distribution method for distributed computing, comprising the following steps:
[0007] receiving a distributed query request, generating a task-specific data flow identifier for the distributed query request;
[0008] decomposing the distributed query request into multiple sub-query tasks;
[0009] obtaining node state information of multiple computing nodes in a distributed computing system, and calculating a comprehensive score for each computing node based on the node state information;
[0010] allocating a plurality of sub-query tasks to a target node in a plurality of computing nodes according to the comprehensive score;
[0011] executing the plurality of sub-query tasks on the target node and outputting a data stream identification.
[0012] As a preferred scheme of the data distribution method for distributed computing according to the application, wherein:
[0013] analyzing query characteristics of the distributed query request;
[0014] determining a specific computing task type based on the query characteristics;
[0015] allocating a unique identification code as the data stream identification for the specific computing task type.
[0016] The beneficial effect of the preferred technical scheme is that by analyzing query characteristics and generating task type identification, task classification management is realized, providing accurate basis for subsequent scheduling decisions, and improving the identification efficiency and execution accuracy of the system for complex tasks.
[0017] As a preferred scheme of the data distribution method for distributed computing according to the application, wherein:
[0018] determining query complexity of the distributed query request;
[0019] determining the number and type of sub-query tasks according to the query complexity, and inheriting the data stream identification for each sub-query task.
[0020] The beneficial effect of the preferred technical scheme is that the task is dynamically split according to the query complexity and the identification is inherited, which not only ensures the rationality of the task granularity, but also maintains the task correlation, and optimizes the parallelism and resource utilization of distributed processing.
[0021] As a preferred scheme of the data distribution method for distributed computing according to the application, wherein:
[0022] The resource state information includes processor speed, memory size, network delay and disk input / output performance.
[0023] The data stream processing history information includes historical data stream identification records processed by the computing node, processing time records corresponding to each historical data stream identification, and processing success rate statistics.
[0024] The cache state information comprises cache space partition information, cache space utilization information, data compression ratio information and cache hit rate information.
[0025] The cache space partition information comprises a first cache partition for storing compressed data and a second cache partition for storing instruction sets.
[0026] The compressed data is to-be-processed data stored in a compressed form.
[0027] As a preferred scheme of the data distribution method for distributed computing, the step of calculating the comprehensive score of each computing node comprises:
[0028] The resource score is calculated through the resource state information.
[0029] The cache efficiency score is calculated through the cache state information.
[0030] The data flow affinity score is calculated through the data flow processing history information and the data flow identifier.
[0031] The comprehensive score is calculated by weighting the resource score, the cache efficiency score and the data flow affinity score.
[0032] The calculation of the data flow affinity score comprises matching the data flow identifier with historical data flow identifier records of the computing nodes, and determining a processing experience value of the computing nodes for the specific computing task type according to the matching result.
[0033] The three-dimensional score weighting mechanism (resource, cache efficiency and data flow affinity) combines historical experience with real-time state, enhances the intelligent level of the scheduling strategy, reduces the task failure rate and improves the execution efficiency.
[0034] As a preferred scheme of the data distribution method for distributed computing, the step of distributing a plurality of subquery tasks to a target node in a plurality of computing nodes according to the comprehensive score comprises:
[0035] The plurality of computing nodes are sorted according to the comprehensive score.
[0036] The computing node with the highest comprehensive score is selected as the target node.
[0037] The current load status of the target node is checked, and when the current load of the target node exceeds a preset threshold, the computing node with the second highest comprehensive score is selected as the target node.
[0038] The beneficial effects of the preferred technical solutions are: based on the score ranking and the load threshold, the target node is dynamically selected, the single point overload is avoided, the system load is balanced, and the stability and task response speed in the high concurrency scenario are ensured.
[0039] As a preferred scheme of the data distribution method for distributed computing, wherein: the step of executing a plurality of subquery tasks on the target node comprises:
[0040] Loading the compressed data required by the subquery task in the first cache partition of the target node;
[0041] Decompressing the compressed data into uncompressed data;
[0042] Processing the uncompressed data according to the instruction set stored in the second cache partition, generating a subquery result and attaching the data stream identifier, and updating the data stream processing history information of the target node.
[0043] Another object of the present application is to provide a data distribution system for distributed computing.
[0044] To solve the above technical problems, the present application provides the following technical solutions: a data distribution system for distributed computing, comprising a data receiving and generating module, a task decomposition module, a computing module, and a task allocation module;
[0045] The data receiving and generating module is used for receiving a distributed query request and generating a task-specific data stream identifier for the request;
[0046] The task decomposition module is responsible for decomposing the received distributed query request into a plurality of subquery tasks;
[0047] The computing module obtains node state information of a plurality of computing nodes in the distributed computing system, and calculates a comprehensive score of each computing node based on the obtained node state information;
[0048] The task allocation module allocates the plurality of subquery tasks to a target node in the plurality of computing nodes according to the calculated comprehensive score.
[0049] The present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the data distribution method for distributed computing when executing the computer program.
[0050] The present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the data distribution method for distributed computing.
[0051] The beneficial effects of the present application: by introducing a multi-dimensional comprehensive scoring mechanism, the task allocation efficiency and resource utilization flexibility of the distributed computing system are improved. First, based on the three-dimensional scoring system of resource state, cache state and data flow processing history, the task allocation decision is upgraded from "resource availability" to "resource adaptability", so that the computing node can not only meet the basic resource demand of the task, but also play its cache management advantage and historical processing experience. Second, the generation and matching mechanism of data flow identification strengthens the traceability between tasks, optimizes the node selection strategy by associating historical processing records, so as to avoid repetitive resource consumption and improve the reliability of task execution. In addition, the dynamic management of cache partition (distinguishing between compressed data and instruction set storage) and the load threshold checking mechanism reduce the decompression overhead while ensuring system stability, especially suitable for complex scenarios with high concurrency and multiple types of queries. Finally, the scheme provides a more intelligent and efficient task scheduling path for distributed computing through the synergistic effect of multi-dimensional state perception and historical experience feedback. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0053] Fig. 1 The overall flowchart of a data distribution method for distributed computing provided by an embodiment of the present application.
[0054] Fig. 2 The structure diagram of a data distribution method for distributed computing provided by an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0056] Embodiment 1, refer to Figs. 1-2 For an embodiment of the present application, a data distribution method for distributed computing is provided, which includes the following steps S1-S4:
[0057] S1, receiving a distributed query request, generating a task-specific data flow identification for the distributed query request;
[0058] S2. Decompose the distributed query request into multiple sub-query tasks;
[0059] S3. Obtain node status information of multiple computing nodes in the distributed computing system, and calculate a comprehensive score for each computing node based on the node status information;
[0060] S4. Allocate multiple subquery tasks to target nodes among multiple computing nodes according to the comprehensive scores;
[0061] S5. Execute multiple subquery tasks on the target node and output the data flow identifier.
[0062] It should be noted that when distributed computing systems process large-scale query requests, traditional task allocation methods often only consider the basic resource status of computing nodes, such as CPU utilization and memory usage, and fail to fully utilize the node's cache status and historical processing experience. This leads to inefficient task allocation and limited overall performance. Furthermore, the lack of effective data flow identification and tracking mechanisms makes the data flow during task execution unclear, affecting manageability and troubleshooting capabilities.
[0063] Therefore, to address the above-mentioned task allocation and data flow management issues, an intelligent multi-dimensional scoring mechanism is constructed through steps S1-S5. A comprehensive evaluation is conducted based on resource status, cache efficiency, and data flow affinity to achieve the optimal task allocation strategy. At the same time, a complete data flow identification system is established to ensure the traceability of the entire task execution process and improve overall performance and management efficiency.
[0064] Example 2, reference Figs. 1-2 , is an embodiment of the present invention, based on the above embodiment, provides a data distribution method for distributed computing.
[0065] In the embodiment of the present application, in step S1, the steps of generating a task-specific data flow identifier include A1 to A3:
[0066] A1. Analyze the query characteristics of distributed query requests.
[0067] Analyzing the query features of distributed query requests includes: extracting keywords, operation types, number of data tables, connection relationships, aggregation function types, filter condition complexity, and estimated data processing volume in the query statement to form a query feature vector.
[0068] For example, query feature vector Feature query It can be expressed as:
[0069] Feature query =
[0070] {Type operation ,Count table ,Complexity join ,Count aggregate ,Level filter ,Size data};
[0071] In the formula, Type operation is an operation type identifier, such as SELECT, INSERT, UPDATE, etc.; Count table is the number of data tables involved; Complexity join is the table connection complexity level; Count aggregate is the number of aggregation functions; Level filter is the filtering condition complexity level; and Size data is the estimated data processing amount.
[0072] It should be noted that through multi-dimensional feature extraction, the computing characteristics of the query request can be comprehensively described, thereby providing accurate data basis for subsequent task type determination.
[0073] In an optional implementation, the analysis of the query features can also be performed by a SQL syntax parser for in-depth analysis to identify the execution plan complexity of the query, the index usage, and potential performance bottleneck points.
[0074] In another optional implementation, the analysis of the query features can also be combined with historical query statistical information, such as average execution time, resource consumption mode, etc., to enrich the description dimensions of the query features.
[0075] A2、based on the query feature to determine a specific computing task type.
[0076] Specifically, determining a specific computing task type based on the query features includes: performing pattern matching according to the query feature vector to classify the query into predefined computing task types such as data retrieval type, data aggregation type, data analysis type, and data update type.
[0077] It should be noted that accurate classification of the task type is a key basis for implementing data flow affinity scoring, and through a standardized task classification system, the historical processing experience of the nodes can be effectively matched.
[0078] In an optional implementation, determining the task type can also adopt a machine learning classification algorithm, such as a decision tree, a random forest, etc., to train a classification model based on historical query samples, thereby improving the accuracy of task type identification.
[0079] In another alternative embodiment, determining the task type can also introduce a fuzzy classification mechanism, allowing a query to have features of multiple task types simultaneously, and describing the mixed characteristics of the task through weight distribution.
[0080] A3. Assigning a unique identification code for a specific computing task type as a data stream identification.
[0081] Specifically, assigning a unique identification code for a specific computing task type includes generating a unique data stream identification code in combination with the task type identification, timestamp, system identification, and sequence number, to ensure global uniqueness and traceability of the identification.
[0082] For example, the generation rule of the data stream identification code FlowID is represented as:
[0083] FlowID =
[0084] TaskType code +Timestamp unix +SystemID hash +Sequence auto ;
[0085] Wherein, TaskType code is the task type code, such as DR (Data Retrieval), occupying 2 characters; Timestamp unix is the last 8 bits of the Unix timestamp, ensuring time uniqueness; SystemID hash is the last 4 bits of the hash value of the system identification, distinguishing different system sources; Sequence auto is a self-incrementing sequence number, occupying 4 digits, to handle concurrent requests at the same time.
[0086] It should be noted that the above identification code generation mechanism, through multi-level coding combination, ensures the global uniqueness of the identification in a distributed environment, while containing key information such as task type, time, and source, facilitating subsequent data stream tracking and management.
[0087] In an alternative embodiment, generating the identification code can also use the UUID algorithm in combination with a business feature prefix, to provide better readability and business relevance while ensuring uniqueness.
[0088] In another alternative embodiment, generating the identification code can also introduce a check bit mechanism, adding a check bit at the end of the identification code through the CRC check algorithm, to improve the reliability of the identification code in the transmission and storage process.
[0089] The step of decomposing the distributed query request into multiple sub-query tasks includes B1-B2:
[0090] B1, determining the query complexity of the distributed query request.
[0091] B2, determining the number and type of subquery tasks according to the query complexity, and inheriting the data flow identifier for each subquery task.
[0092] Specifically, the determination of the query complexity of the distributed query request in step B1 includes: analyzing the syntax structure of the query statement, counting the number of table connections, the number of nested query layers, the complexity of aggregation operation, the number of sorting operations and the data scanning range, and calculating the query complexity index by weighting.
[0093] Exemplarily, the query complexity calculation formula can be expressed as:
[0094] Complexity query =w1·Count join +w2·Level nested +w3·Score aggregate +
[0095] w4·Count sort +w5·Range scan ;
[0096] In the formula, Complexity query is the query complexity index; Count join is the number of table connections; Level nested is the number of nested query layers; Score aggregate is the aggregation operation complexity score; Count sort is the number of sorting operations; Range scan is the data scanning range coefficient; w1, w2, w3, w4 and w5 are the corresponding weight coefficients.
[0097] Exemplarily, the calculation method of each complexity parameter can be expressed as:
[0098]
[0099] Level nested =max(Depth subquery )+Count subquery ·0.5
[0100] Wherein, is the complexity weight function of different aggregation functions, such as COUNT = 1, SUM = 1.2, AVG = 1.5, MAX / MIN = 1.8; is the data volume influence factor; Index selectivity is the index selectivity coefficient.
[0101] It should be noted that by multi-dimensional complexity quantification, the computational load of the query can be accurately evaluated, providing a scientific basis for subsequent task decomposition and avoiding evaluation bias caused by simple counting.
[0102] In an optional embodiment, the query complexity determined in step B1 can also be analyzed by the execution plan of the query optimizer to extract the number of operation nodes, estimate the execution cost, and predict resource consumption, forming a more accurate complexity evaluation.
[0103] In another optional embodiment, the query complexity determined in step B1 can also be combined with historical execution statistical information, such as the average execution time of similar queries, CPU utilization, memory occupation, etc., to establish a complexity prediction model through regression analysis.
[0104] Specifically, the number and type of subquery tasks determined in step B2 according to the query complexity include: based on the query complexity index, the task decomposition strategy is selected, and the original query is decomposed into multiple independent executable subquery tasks according to data dependency, calculation logic and parallelism requirements.
[0105] For example, the subquery task number determination rule can be expressed as:
[0106]
[0107] Where Threshold simple = 10 is the simple query threshold; Threshold moderate = 5 is the medium complexity unit; Threshold complex = 50 is the complex query threshold; Max parallel = 8 is the maximum parallelism limit; Threshold unit = 8 is the task decomposition unit.
[0108] It should be noted that through a scientific task decomposition strategy, parallel processing of the query is realized, while maintaining hierarchical management of data flow identification, ensuring complete traceability and management convenience of the task execution process.
[0109] In an optional embodiment, the subquery task determined in step B2 can also use a graph-based query decomposition algorithm to model the query operation as a directed acyclic graph and determine the optimal task decomposition scheme through graph partitioning algorithm.
[0110] In another optional embodiment, the subquery task determined in step B2 can also combine a dynamic load balancing strategy to dynamically adjust the granularity and number of subtasks according to the current load and node performance, realizing adaptive task decomposition.
[0111] In yet another alternative embodiment, the inheriting of the data stream identifier for the sub-query task in step B2 can also establish a task dependency graph, record data dependencies and execution order among sub-tasks, and realize visual management of the task execution process through the association of the identification codes.
[0112] In the embodiment, the node state information includes resource state information, cache state information, and data stream processing history information; the resource state information includes processor speed, memory size, network delay, and disk input / output performance;
[0113] Specifically, the resource state information is acquired in the following manner: periodically collecting the hardware resource states of each computing node through a monitoring device. The processor speed is calculated by acquiring the CPU frequency and current utilization, reflecting the real-time computing capacity of the node; the memory size is determined by monitoring the available memory capacity, including the usage of physical memory and virtual memory; the network delay is obtained by inter-node communication testing, reflecting the real-time performance of data transmission; and the disk input / output performance is evaluated by monitoring the disk read / write speed and IOPS value, embodying the processing capacity of the storage subsystem.
[0114] For example, the acquisition frequency of the resource state information is set to be updated once every 30 seconds, ensuring the balance between real-time performance and overhead. When a significant change occurs in a resource indicator, such as a change in CPU utilization exceeding 20% or a decrease in available memory exceeding 1 GB, an instant update mechanism is triggered.
[0115] The data stream processing history information includes historical data stream identifier records processed by the computing node, processing time records corresponding to each historical data stream identifier, and processing success rate statistics;
[0116] Specifically, the management of the data stream processing history information includes: establishing a historical processing record database for each computing node to record all data stream identifiers processed by the node and their processing details. The historical data stream identifier records are stored in a time series manner, containing key information such as the type, source, and processing time of the data stream; the processing time records detail the entire process time from receiving to completing each data stream, including data loading time, computing processing time, and result output time; and the processing success rate statistics are calculated based on the processing records in the recent period of time, usually taking the last 100 processing or the processing in the last 24 hours as the statistical window.
[0117] It should be noted that the historical information is recorded using a sliding window mechanism, which automatically cleans up expired data to prevent unlimited growth of storage space. For processing failures, the failure reasons and error types are recorded to facilitate subsequent fault analysis and prevention.
[0118] The cache state information includes cache space partition information, cache space utilization information, data compression ratio information, and cache hit rate information; the cache space partition information includes a first cache partition for storing compressed data and a second cache partition for storing an instruction set; the compressed data is to-be-processed data stored in a compressed form.
[0119] Specifically, the monitoring manner of the cache state information includes real-time tracking of the running state and performance indicators of the cache system. The cache space partition information describes the physical layout of the cache, wherein the first cache partition is specially used for storing compressed to-be-processed data, and the second cache partition is used for storing an instruction set and program code required for query execution; the cache space utilization information reflects the current use of the cache, including used space, available space, and fragmentation degree; the data compression ratio information records the ratio of the data size before and after compression, reflecting the efficiency of the compression algorithm and the saving degree of the storage space; and the cache hit rate information statistics the success rate of cache access, which is a key indicator for evaluating the cache efficiency.
[0120] For example, the cache partition adopts a dynamic adjustment strategy to automatically adjust the size ratio of the first partition and the second partition according to the characteristics of the actual work load. When there are more data-intensive tasks, the space configuration of the first partition is increased; and when there are more computation-intensive tasks, the space of the second partition is appropriately increased.
[0121] Specifically, the processing mechanism of the compressed data includes: using an efficient compression algorithm to compress and store the to-be-processed data, common compression algorithms include LZ4, Snappy, and other fast compression algorithms, which ensure the compression efficiency while ensuring the decompression speed. Metadata information is added when the compressed data is stored, including the original data size, the compression algorithm type, the checksum, and the like, to ensure the integrity and recoverability of the data.
[0122] In an optional embodiment, the collection of resource state information can also include GPU utilization, network bandwidth usage, and environmental parameters such as temperature, to provide more comprehensive node evaluation basis for specific types of computing tasks.
[0123] In another optional embodiment, the management of cache state information can also introduce a predictive caching mechanism to predict the data that may be needed in the future based on historical access patterns, and pre-load the cache, further improving the cache hit rate.
[0124] It should be noted that through comprehensive collection and management of node state information, the real-time running state and historical performance of each computing node can be accurately mastered, providing reliable decision basis for intelligent task allocation, and effectively improving the overall performance and resource utilization efficiency of distributed computing.
[0125] In the embodiment, in step S3, the step of calculating the comprehensive score of each computing node includes C1-C4:
[0126] C1, calculating a resource score through resource status information;
[0127] The calculation of the resource score through the resource status information includes: comprehensively evaluating the processor performance, memory capacity, network communication capability and storage performance of the computing node to form a resource comprehensive score. The processor score is calculated based on the current available computing capability, considering the CPU frequency and the current load condition; the memory score is determined according to the available memory capacity and the memory use efficiency; the network score is evaluated through the inter-node communication delay and the bandwidth utilization; and the storage score is calculated based on the disk read-write performance and the available storage space.
[0128] For example, the resource score is calculated in a percentage system, and different weights are assigned to each resource indicator: the processor performance accounts for 40%, the memory capacity accounts for 30%, the network performance accounts for 20%, and the storage performance accounts for 10%. When the CPU utilization of the node is less than 60%, the processor score is high; when the available memory is more than 50% of the total memory, the memory score reaches a good level; and when the network delay is less than 10 milliseconds, the network score is in an excellent level.
[0129] It should be noted that the resource score adopts a dynamic evaluation mechanism, which adjusts the weight proportion of each indicator according to the resource demand characteristics of different types of tasks, to ensure that the score result can accurately reflect the adaptation degree of the node to a specific task.
[0130] C2, calculating a cache efficiency score through cache status information;
[0131] The calculation of the cache efficiency score through the cache status information includes: comprehensively considering the cache hit rate, storage space utilization efficiency, data compression effect and rationality of cache partition. The cache hit rate reflects the effectiveness of the cache system, and the higher the hit rate, the better the cache efficiency; the space utilization rate evaluates the use of cache capacity to avoid space waste or excessive occupation; the compression ratio evaluates the effect of data compression, and a high compression ratio means better storage efficiency; and the partition balance evaluates the use balance of the first cache partition and the second cache partition.
[0132] For example, the cache efficiency score is calculated as follows: when the cache hit rate is more than 80%, the basic score is 85 points; when the data compression ratio is more than 3:1, an additional 10 points are added; and when the utilization rate difference between the two cache partitions is less than 15%, the partition configuration is considered reasonable, and an additional 5 points are added. The final cache efficiency score is the weighted average of each score.
[0133] It should be noted that the cache efficiency score pays special attention to the adaptability of the cache system to the task to be executed. It will predict the cache performance based on the data access pattern of the task to improve the accuracy of the score.
[0134] C3. Calculate the data flow affinity score based on the data flow processing history information and data flow identifier;
[0135] Calculating a data flow affinity score based on historical data flow processing information and data flow identifiers involves analyzing the compute node's historical processing records, matching the task type identified by the current data flow, and evaluating the node's experience and success rate in processing tasks of that type. The matching process first extracts the task type information from the current data flow identifier and then searches the node's historical records for processing records of the same or similar types. Based on the matching results, the node's number of times it has processed tasks of that type, the average processing time, and the success rate are calculated. Based on these statistics, the node's experience and affinity score for the current task type are calculated.
[0136] Exemplarily, the calculation logic of the data flow affinity score is as follows: the weight of the record of the task type that fully matches is 1.0, the weight of the record of the similar task type is 0.7, and the weight of the record of the related task type is 0.3; the records with a processing success rate higher than 95% are additionally weighted; the records with an average processing time lower than the average level of similar nodes receive a time efficiency bonus; the records processed most recently receive a higher time decay weight.
[0137] C4: A comprehensive score is obtained by weighted calculation of the resource score, cache efficiency score, and data flow affinity score.
[0138] The calculation of the data flow affinity score includes matching the data flow identifier with the historical data flow identifier record of the computing node, and determining the processing experience value of the computing node for a specific computing task type based on the matching result.
[0139] The steps involved in processing experience points include: counting the total number of times a node processes matching task types as the experience base; calculating the average success rate of these tasks as a reliability indicator; analyzing the ratio of average processing time to standard time as an efficiency indicator; and analyzing the node's stability performance based on the time distribution of task processing. The final processing experience point comprehensively considers four dimensions: experience base, reliability, efficiency, and stability.
[0140] It should be noted that data stream affinity scoring introduces a learning mechanism. As the node processes more tasks, its historical records continue to be enriched, and the accuracy of the affinity score will gradually improve, achieving continuous performance optimization.
[0141] Specifically, the comprehensive score of the resource score, the cache efficiency score and the data stream affinity score in step C4 is calculated by weighting, including: setting the weight coefficients of the three scores, dynamically adjusting the weight distribution according to the current state and the task characteristics, and calculating the final comprehensive score by weighted summation. The weight distribution principle is: for a compute-intensive task, the resource score weight is higher; for a data-intensive task, the cache efficiency score weight is higher; for a high-repetition task, the data stream affinity score weight is higher.
[0142] For example, the weight distribution strategy is: by default, the resource score weight is 0.4, the cache efficiency score weight is 0.35, and the data stream affinity score weight is 0.25; when it is detected that the task is a new type of task for the first execution, the data stream affinity weight is reduced to 0.1, and the resource score and cache efficiency score weights are correspondingly increased; when the overall load is high, the resource score weight is increased to 0.5, ensuring that the task is assigned to a node with sufficient resources.
[0143] It should be noted that the comprehensive score calculation adopts normalization processing to ensure that the final score is within the range of 0-100, facilitating comparison and sorting between different nodes. At the same time, a score smoothing mechanism is introduced to avoid short-term fluctuations from having a large impact on the score.
[0144] In an optional embodiment, the resource score calculation in step C1 can also introduce historical load pattern analysis of the node to predict the resource availability of the node at different time periods, improving the forward-looking nature of the score.
[0145] In another optional embodiment, the data stream affinity score in step C3 can also consider the collaboration history between nodes, and when multiple nodes have successfully collaborated to process similar tasks, an additional collaboration bonus is given.
[0146] In yet another optional embodiment, the weight adjustment of the comprehensive score in step C4 can also be based on a machine learning algorithm to automatically optimize the weight parameters by analyzing the historical allocation results, achieving adaptive score optimization.
[0147] In the present embodiment, in step S4, the step of assigning the multiple subquery tasks to the target node in the multiple computing nodes according to the comprehensive score includes D1-D3:
[0148] D1, sorting the multiple computing nodes according to the comprehensive score;
[0149] Specifically, all available computing nodes in the system are arranged in descending order of comprehensive scores to form a candidate node priority list. The sorting process first obtains the latest comprehensive scores of all nodes, and then sorts efficiently using the quicksort or heapsort algorithm; for nodes with the same score, the secondary sorting standard is used, and the node with better historical stability is given priority; after sorting, the node candidate list is generated, which contains node identification, comprehensive score, sorting position and other information.
[0150] For example, the node sorting strategy is: when the difference between the comprehensive scores of two nodes is less than 2 points, the system compares the historical stability index of the nodes as the basis for sorting; for nodes newly added to the system, appropriate score weighting is given to avoid being ignored for a long time due to insufficient historical data; the sorting result is cached for a certain period of time to avoid system overhead caused by frequent rearrangement.
[0151] It should be noted that the sorting algorithm takes into account the dynamic change characteristics of node state in a distributed environment, and uses an incremental update mechanism. When the state of an individual node changes significantly, only the relevant part is reordered, improving the sorting efficiency.
[0152] D2, selecting the computing node with the highest comprehensive score as the target node;
[0153] Specifically, the first ranked node is selected from the sorted candidate node list as the preferred target node. The selection process needs to verify the current online state and availability of the node; confirm that the node has the basic ability and authority to execute the current task type; check the network connectivity between the node and the task initiator; verify the security authentication state and access permission of the node.
[0154] For example, the verification process of target node selection includes: first, node liveliness detection is performed to confirm the normal operation of the node through the heartbeat mechanism; then the task processing capability of the node is verified to ensure that the node supports the software environment and computing resources required by the current task; finally, the security state of the node is checked, including the compliance of identity authentication, access control and other security policies.
[0155] It should be noted that the target node selection adopts a fast failure mechanism. If the preferred node has problems in the verification process, the system will immediately switch to the suboptimal node, avoiding long-term blocking of the task allocation process.
[0156] D3, checking the current load status of the target node, and when the current load of the target node exceeds the preset threshold, selecting the computing node with the second highest comprehensive score as the target node.
[0157] Specifically, the current load condition of the target node is checked by monitoring the CPU usage, memory usage, network bandwidth usage and the number of tasks currently being executed in real time. The load check uses a multi-dimensional evaluation method, considering not only the usage of individual resources, but also the overall load level of the node; the preset threshold is dynamically set according to the node type and historical performance, and the CPU usage threshold is usually 80%, the memory usage threshold is 85%, and the concurrent task number threshold is determined according to the node specification.
[0158] For example, the load condition evaluation includes: instantaneous load check, obtaining the resource usage of the node at the current time; trend load analysis, predicting the load development direction of the node through the load change trend in the last 5 minutes; task queue check, counting the total number of tasks waiting to be executed and being executed. When any of the indicators exceeds the preset threshold, the system determines that the node load is too high.
[0159] The processing mechanism when the current load of the target node exceeds the preset threshold includes: the system automatically selects the node with the second highest comprehensive score from the candidate node list as the new target node; the load check process is repeated for the newly selected node; if multiple high-score nodes are continuously overloaded, the system will appropriately relax the threshold standard or start the load balancing mechanism; at the same time, the load overrun situation is recorded for subsequent capacity planning and system optimization.
[0160] For example, the alternative node selection strategy is: the system checks the first 5 high-score nodes at most, if none of them can meet the load requirement, the node with the lightest load is selected to execute the task; for urgent tasks, the load threshold can be temporarily increased to allow the node to execute the task under higher load; for tasks that can be delayed, the system will add the task to the waiting queue and wait for a suitable node to be available before assigning it.
[0161] It should be noted that the load check mechanism introduces predictive evaluation, not only considering the current load state, but also predicting the load change of the node during task execution according to the task execution time estimation and the historical load pattern of the node, to ensure that the task can be completed in a reasonable resource environment.
[0162] In an alternative embodiment, the node sorting in step D1 can also introduce geographical location factors, preferentially selecting nodes with closer network distance to reduce data transmission delay and network overhead.
[0163] In another alternative embodiment, the target node selection in step D2 can also consider the power state and heat dissipation of the node to avoid assigning important tasks to nodes in high temperature or unstable power state.
[0164] In yet another alternative embodiment, the load check in step D3 can also include resource status such as disk space usage and database connection number of the node, providing more comprehensive load evaluation.
[0165] In the present embodiment, the step of executing multiple subquery tasks on the target node in step S5 includes E1-E3:
[0166] E1, loading compressed data required by the subquery task in the first cache partition of the target node;
[0167] E2, decompressing the compressed data into uncompressed data;
[0168] E3, processing the uncompressed data according to the instruction set stored in the second cache partition, generating the subquery result and attaching the data stream identifier, and updating the data stream processing history information of the target node.
[0169] Specifically, in step E1, the data requirements of the subquery task are analyzed to determine the range and type of the data set that needs to be loaded; the available space of the first cache partition is checked to evaluate whether it is sufficient to accommodate the required compressed data; the compressed data is obtained from the distributed storage system or other nodes and loaded into the first cache partition in priority order. The data loading process uses a streaming loading method, which stores data while downloading to reduce memory usage; for data already existing in the cache, it is directly reused to avoid repeated loading; the network bandwidth and storage write speed are monitored in real time during the loading process to dynamically adjust the loading strategy.
[0170] For example, the data loading strategy includes: loading key data required at the initial stage of task execution in priority to ensure that the task can be started in time; for large data sets, using a block loading method to split the data into multiple compressed blocks for sequential loading; establishing a data loading queue to arrange the loading plan according to access frequency and time order. When the space of the first cache partition is insufficient, the system will clean up the least recently used data blocks according to the LRU algorithm to make room for new data.
[0171] It should be noted that a verification mechanism is introduced during the data loading process to check the integrity of the loaded compressed data, ensuring that the data is not damaged during transmission and storage, and avoiding errors in subsequent processing.
[0172] Specifically, in step E2, the compression algorithm type of the compressed data is identified and the corresponding decompression method is selected; the decompression operation is performed in sequence according to the data block order to restore the compressed data to the original format; the decompression process uses a pipeline method, which processes data while decompressing to reduce the storage demand of intermediate data; the CPU usage and memory occupation of the decompression process are monitored to avoid resource bottlenecks affecting performance. After decompression, the data is format-verified and integrity-checked to ensure that the data can be normally used for subsequent processing.
[0173] For example, the decompression optimization strategy includes: selecting the optimal decompression parameters according to the characteristics of the compression algorithm, such as buffer size, thread number, etc.; for multiple data blocks, using parallel decompression mode to fully utilize the computing power of multi-core processors; establishing a decompression cache mechanism to keep frequently accessed data in a decompressed state and reduce repeated decompression overhead. When an error occurs during the decompression process, the system will try to retrieve the compressed data or use backup data.
[0174] It should be noted that the decompression process uses an incremental processing method, which only decompresses the data actually needed in the current stage of the task, avoiding memory pressure and processing delay caused by decompressing a large amount of data at once.
[0175] Specifically, in step E3, processing non-compressed data according to the instruction set stored in the second cache partition includes: loading execution instructions and processing programs related to the current task type from the second cache partition; selecting appropriate data processing algorithms and calculation processes according to the specific requirements of the subquery task; performing data filtering, joining, aggregation, sorting, etc. operations to generate intermediate results and final results that meet the query requirements; monitoring the task execution status in real time during processing, recording key performance indicators and processing progress.
[0176] For example, the data processing execution strategy includes: selecting serial or parallel processing mode according to data size and complexity; establishing processing progress checkpoints, saving intermediate results regularly to prevent long tasks from losing calculation results due to unexpected interruptions; optimizing memory usage, using data paging technology to process large data sets to avoid memory overflow. For complex multi-table join operations, the system will select the optimal join algorithm according to the table size and join conditions.
[0177] Specifically, generating subquery results and attaching data stream identifiers includes: organizing the processed data into query result sets according to standard formats; embedding original data stream identifier information in the result data to ensure traceability of the results; adding result metadata, including processing time, data row count, processing node information, etc.; formatting and serializing the result data for easy transmission and storage.
[0178] Specifically, updating the data stream processing history information of the target node includes: recording the execution details of the current task, including start time, end time, processing data volume, resource consumption; updating the node's processing statistics information for this type of task, such as total processing times, average processing time, success rate, etc.; storing the processing records in the history database to provide data support for subsequent affinity scoring; analyzing the performance during task execution to identify potential optimization points and improvement directions.
[0179] For example, the historical information updating strategy includes: using an asynchronous updating mode to avoid the influence of historical record updating on task execution performance; establishing a hierarchical storage mechanism for historical data to save important statistical information for a long time and clean detailed execution logs periodically; and recording the failure cause and error information of a task that fails to be processed in detail for fault analysis and prevention.
[0180] It should be noted that the intelligent monitoring mechanism introduced in the task execution process can automatically detect abnormal conditions and take corresponding processing measures, such as automatic termination of a task that is overdue, degradation processing when resources are insufficient, retry mechanism when the network is abnormal, and the like, to ensure the stability and reliability of the system.
[0181] In an optional implementation, the data loading in step E1 can also use a predictive caching strategy to load data that can be needed in advance according to the task execution mode, thereby improving processing efficiency.
[0182] In another optional implementation, the decompression process in step E2 can also combine hardware acceleration technologies, such as GPU acceleration or a special decompression chip, to significantly improve the decompression speed.
[0183] In yet another optional implementation, the data processing in step E3 can also introduce a streaming computing framework to continuously process real-time data streams, supporting streaming query and real-time analysis scenarios.
[0184] In summary, by introducing a multi-dimensional comprehensive scoring mechanism, the task allocation efficiency and resource utilization flexibility of the distributed computing system are improved. First, based on the three-dimensional scoring system of resource status, cache status, and data stream processing history, the task allocation decision is upgraded from “resource availability” to “resource adaptability”, so that the computing node can not only meet the basic resource needs of the task, but also can exert its cache management advantages and historical processing experience. Second, the generation and matching mechanism of the data stream identifier strengthens the traceability between tasks, dynamically optimizes the node selection strategy by associating historical processing records, thereby avoiding repetitive resource consumption and improving task execution reliability. In addition, the dynamic management of cache partitioning (distinguishing between compressed data and instruction set storage) and the load threshold checking mechanism reduce the decompression overhead while ensuring system stability, which is particularly suitable for complex scenarios with high concurrency and multiple types of queries. Finally, the scheme provides a more intelligent and efficient task scheduling path for distributed computing through the synergistic effect of multi-dimensional state perception and historical experience feedback.
[0185] Embodiment 3, the above is a schematic solution of a data distribution method of distributed computing. It should be noted that the technical solution of the data distribution system of distributed computing and the technical solution of the data distribution method of distributed computing described above belong to the same concept. The technical details of the technical solution of the data distribution system of distributed computing in the present embodiment are not described in detail, and can be referred to the description of the technical solution of the data distribution method of distributed computing.
[0186] The present embodiment also provides a data distribution system of distributed computing, comprising a data receiving and generating module, a task decomposition module, a computing module, and a task allocation module.
[0187] The data receiving and generating module is used to receive a distributed query request and generate a task-specific data stream identifier for the request.
[0188] The task decomposition module is responsible for decomposing the received distributed query request into multiple sub-query tasks.
[0189] The computing module obtains node state information of multiple computing nodes in the distributed computing system, and calculates a comprehensive score of each computing node based on the obtained node state information.
[0190] The task allocation module allocates the multiple sub-query tasks to target nodes in the multiple computing nodes according to the calculated comprehensive scores.
[0191] The present embodiment also provides an electronic device suitable for the case of data distribution of distributed computing, comprising a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the data distribution method of distributed computing proposed in the above embodiments.
[0192] The present embodiment also provides a storage medium having a computer program stored thereon, which is executed by a processor to realize the data distribution method of distributed computing proposed in the above embodiments.
[0193] The storage medium proposed in the present embodiment and the data distribution method of distributed computing proposed in the above embodiments belong to the same inventive concept. The technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.
[0194] Those skilled in the art can clearly understand the present application by the above description of the embodiments, and the present application can be realized by software and necessary general hardware, and of course, can also be realized by hardware. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a FLASH, a hard disk, or an optical disc, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various embodiments of the present application.
[0195] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application, and although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and all should be covered in the scope of the claims of the present application.
Claims
1. A data distribution method for distributed computing, characterized in that: The following steps are involved: receiving a distributed query request, and generating a task-specific data flow identifier from the distributed query request; Decomposing the distributed query request into multiple sub-query tasks; Obtaining node status information of multiple computing nodes in a distributed computing system, and calculating a comprehensive score of each computing node based on the node status information; Allocating the plurality of subquery tasks to a target node among the plurality of computing nodes according to the comprehensive score; Execute multiple sub-query tasks on the target node and output a data flow identifier.
2. A data distribution method for distributed computing according to claim 1, characterized in that: The step of generating the task-specific data flow identifier includes: Analyzing query characteristics of the distributed query request; determining a specific computing task type based on the query characteristics; A unique identification code is allocated to the specific computing task type as the data flow identifier.
3. A data distribution method for distributed computing according to claim 2, characterized in that: The step of decomposing the distributed query request into multiple sub-query tasks includes: Determining the query complexity of the distributed query request; The number and type of sub-query tasks are determined according to the query complexity, and the data flow identifier is inherited for each sub-query task.
4. A data distribution method for distributed computing according to claim 3, characterized in that: The node status information includes resource status information, cache status information and data flow processing history information; The resource status information includes processor speed, memory size, network latency, and disk input and output performance; The data flow processing history information includes the historical data flow identification records processed by the computing node, the processing time records corresponding to each historical data flow identification, and the processing success rate statistics; The cache status information includes cache space partition information, cache space utilization information, data compression ratio information and cache hit rate information; The cache space partition information includes a first cache partition for storing compressed data and a second cache partition for storing an instruction set; The compressed data is data to be processed stored in a compressed form.
5. A distributed computing data distribution method according to claim 4, characterized in that: The steps to calculate the comprehensive score of each computing node include: Calculating a resource score based on the resource status information; Calculating a cache efficiency score based on the cache status information; Calculating a data flow affinity score using the data flow processing history information and the data flow identifier; The resource score, cache efficiency score, and data flow affinity score are weighted to obtain a comprehensive score. The calculation of the data flow affinity score includes matching the data flow identifier with the historical data flow identifier record of the computing node, and determining the processing experience value of the computing node for the specific computing task type according to the matching result.
6. A data distribution method for distributed computing according to claim 5, characterized in that: The step of allocating the multiple sub-query tasks to the target node among the multiple computing nodes according to the comprehensive score includes: sorting the plurality of computing nodes according to the comprehensive scores; Select the computing node with the highest comprehensive score as the target node; Check the current load status of the target node, and when the current load of the target node exceeds a preset threshold, select the computing node with the second highest comprehensive score as the target node.
7. A distributed computing data distribution method according to claim 6, characterized in that: The step of executing multiple sub-query tasks on the target node includes: Loading the compressed data required by the subquery task into the first cache partition of the target node; decompressing the compressed data into uncompressed data; The uncompressed data is processed according to the instruction set stored in the second cache partition, a sub-query result is generated and the data flow identifier is added, and the data flow processing history information of the target node is updated.
8. A distributed computing data distribution system, applying a distributed computing data distribution method according to any one of claims 1 to 7, characterized in that: It includes data receiving and generating module, task decomposition module, calculation module, and task allocation module; The data receiving and generating module is used to receive a distributed query request and generate a task-specific data flow identifier from the request; The task decomposition module is responsible for decomposing the received distributed query request into multiple sub-query tasks; The computing module obtains node status information of multiple computing nodes in the distributed computing system, and calculates a comprehensive score of each computing node based on the obtained node status information; The task assignment module assigns the plurality of sub-query tasks to a target node among the plurality of computing nodes according to the calculated comprehensive scores.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the data distribution method for distributed computing according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data distribution method for distributed computing according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Data checking method and device, storage medium and electronic device
CN121579580A