Parallel computing performance optimization method and device, equipment and medium

By using CostAwareSpliterator's custom cost calculation and incremental caching mechanism, parallel tasks are dynamically split, solving the problems of unbalanced load and full traversal in Spliterator, and improving the efficiency and adaptability of parallel computing.

CN120560866BActive Publication Date: 2025-11-11北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511064198.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing Spliterators fail to effectively consider the differences in element processing costs in parallel computing, resulting in unbalanced subtask loads. Furthermore, the computational cost of full traversal leads to inefficiency and a lack of flexibility, making them unsuitable for specific scenarios.

Method used

By introducing CostAwareSpliterator, a custom cost calculator and incremental calculation and caching mechanism are used to dynamically obtain the interval processing cost, realize dynamic task splitting and load balancing, execute subtasks in parallel, and execute them in a loop until resources are reclaimed.

Benefits of technology

It improves the parallel processing efficiency of large-scale uneven cost data, achieves load balancing, reduces computational overhead, is highly adaptable, has high resource utilization, and is suitable for processing large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560866B_ABST
    Figure CN120560866B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing and provides a method, apparatus, device, and medium for optimizing parallel computing performance. The method includes: initializing a CostAwareSpliterator, receiving a data list and a custom cost calculator, determining an initial processing interval and associating it with the custom cost calculator; within the initial processing interval, obtaining the element processing cost based on the custom cost calculator, and dynamically obtaining the interval processing cost through incremental calculation and caching mechanisms; dynamically splitting tasks based on the total interval cost to achieve subtask load balancing; executing multiple subtasks in parallel, processing elements within the interval corresponding to each subtask, and dynamically shrinking the interval based on the element processing status, updating the total interval cost after dynamic shrinkage; and iteratively executing the steps of dynamic splitting and parallel execution of subtasks based on the dynamically updated total interval cost, until the splitting conditions can no longer be met, automatically reclaiming all resources. This method improves the parallel processing efficiency of large-scale, uneven cost data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a method, apparatus, device, and medium for optimizing parallel computing performance. Background Technology

[0002] With the rapid development of big data technology, the efficient processing of massive amounts of data has become one of the core requirements in the field of information technology. Parallel computing, as a key means to improve data processing efficiency, has attracted much attention for performance optimization. To improve data processing efficiency in big data scenarios, the Spliterator interface can be used to split the dataset, dividing the original dataset into multiple subsets, thereby supporting multi-threaded parallel processing. This effectively improves the processing efficiency of large-scale data and has become an important technical support in the field of big data processing.

[0003] However, existing spliterator technologies typically split tasks based on quantity or index, leading to uneven load distribution among subtasks and impacting overall efficiency. Furthermore, the use of a fixed splitting strategy fails to consider cost differences, making load balancing difficult. Therefore, a novel technical solution is urgently needed to address at least one of the technical problems in these existing technologies. Summary of the Invention

[0004] This application addresses the technical problems existing in the prior art by providing a parallel computing performance optimization method, apparatus, device, and medium to solve at least one of the aforementioned technical problems.

[0005] In a first aspect, embodiments of this application provide a method for optimizing parallel computing performance, the method comprising:

[0006] Initialize the dynamic cost parallel task splitter (CostAwareSpliterator), receive the data list and custom cost calculator, determine the initial processing range and associate it with the custom cost calculator;

[0007] Within the initial processing range, the element processing cost is obtained based on a custom cost calculator, and the range processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal.

[0008] Dynamic task splitting is performed based on the total cost of the interval to achieve subtask load balancing and obtain multiple subtasks that support parallel execution;

[0009] Multiple subtasks are executed in parallel, each processing the elements within its corresponding interval. The intervals are then dynamically shrunk based on the processing status of the elements, and the total cost of the dynamically shrunk intervals is updated.

[0010] Based on the dynamically updated total cost of the interval, the process jumps to dynamically splitting the task according to the total cost of the interval, achieving load balancing of subtasks, and obtaining steps for multiple subtasks that support parallel execution. The process of dynamically splitting and parallel executing subtasks is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

[0011] Secondly, embodiments of this application provide a parallel computing performance optimization device, which includes at least the following units:

[0012] The initialization unit is configured to initialize the CostAwareSpliterator, receive a data list and a custom cost calculator, determine the initial processing range, and associate the custom cost calculator.

[0013] The cost unit is configured to obtain the element processing cost based on a custom cost calculator within the initial processing interval, and dynamically obtain the interval processing cost through incremental calculation and caching mechanisms to avoid full traversal.

[0014] The dynamic splitting unit is configured to dynamically split tasks based on the total cost of the interval, achieving subtask load balancing and obtaining multiple subtasks that support parallel execution. Multiple subtasks are executed in parallel, processing elements within their respective intervals, and the intervals are dynamically shrunk based on the element processing status, updating the total cost of the dynamically shrunk interval. Based on the dynamically updated total cost of the interval, the unit jumps to the step of dynamically splitting tasks based on the total cost of the interval, achieving subtask load balancing, and obtaining multiple subtasks that support parallel execution. This process of dynamically splitting and executing subtasks in parallel is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

[0015] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: at least one processor, a memory, and an input / output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the parallel computing performance optimization method of the first aspect.

[0016] Fourthly, a computer-readable storage medium is provided, comprising instructions that, when executed on a computer, cause the computer to perform the parallel computing performance optimization method of the first aspect.

[0017] The beneficial effects of this application are: it provides a method, apparatus, device, and medium for optimizing parallel computing performance. In this technical solution, the CostAwareSpliterator is first initialized, receiving a data list and a custom cost calculator, determining an initial processing interval, and associating it with the custom cost calculator. Then, within the initial processing interval, the element processing cost is obtained based on the custom cost calculator, and the interval processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal. Next, dynamic task splitting is performed based on the total interval cost to achieve subtask load balancing, resulting in multiple subtasks that support parallel execution. Then, multiple subtasks are executed in parallel, processing elements within their respective intervals, and the interval is dynamically shrunk based on the element processing status, updating the total interval cost after dynamic shrinkage. Finally, based on the dynamically updated total interval cost, the process jumps to the step of dynamically splitting tasks based on the total interval cost to achieve subtask load balancing and obtain multiple subtasks that support parallel execution. This process of dynamically splitting and parallel executing subtasks is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

[0018] This application's embodiments significantly improve the parallel processing efficiency of large-scale, uneven cost data by introducing dynamic cost calculation, incremental caching, and adaptive splitting mechanisms. Regarding load balancing, it breaks through the limitations of traditional splitting based on fixed data volume, dynamically dividing tasks based on the actual processing cost of elements. This keeps the load difference between subtasks within a preset threshold, avoiding thread idleness or overload and improving overall parallel efficiency. In terms of computational overhead, incremental calculation and caching mechanisms avoid full traversal of the dataset, making it particularly suitable for processing large datasets. Regarding scenario adaptability, it supports custom cost calculation logic, allowing flexible adjustment of evaluation rules based on element attributes, adapting to various business scenarios and significantly improving versatility. Regarding resource utilization, dynamic shrinking intervals and cyclic splitting mechanisms achieve on-demand resource allocation, reducing thread idle time, and combined with automatic resource reclamation, reducing memory consumption. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a parallel computing performance optimization method according to an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the structure of a parallel computing performance optimization device according to an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the structure of a medium according to an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] With the rapid development of big data technology, the efficient processing of massive amounts of data has become one of the core requirements in the field of information technology. Parallel computing, as a key means to improve data processing efficiency, has received much attention for performance optimization. For example, the Stream API introduced in Java 8 provides developers with convenient parallel computing tools. This API uses the Spliterator interface to split data sets, dividing the original dataset into multiple subsets, thereby supporting multi-threaded parallel processing, effectively improving the processing efficiency of large-scale data, and becoming an important technical support in the field of big data processing.

[0025] The applicant discovered that existing Spliterator implementations still suffer from numerous performance bottlenecks, limiting the full potential of parallel computing performance. These bottlenecks are specifically manifested as follows:

[0026] First, the applicant found that in handling uneven costs, data elements in real-world applications often have differentiated processing costs. Traditional Spliterator splitting methods are typically based on the quantity or index of data, failing to consider the actual processing cost differences for each element. This splitting logic easily leads to uneven load distribution among the split subtasks, with some threads potentially bearing an excessive computational burden while others are under light load, severely impacting the overall efficiency of parallel computing.

[0027] Secondly, the applicant found that in the total cost calculation stage, existing Spliterator implementations typically require a full traversal of the dataset to calculate the total cost, a process that presents significant efficiency issues when dealing with large datasets. For massive datasets, the initial cost calculation incurs a huge time overhead, resulting in a significant performance loss and hindering the efficiency of starting parallel computing.

[0028] Furthermore, the applicant found that most current Spliterator implementations use a fixed splitting strategy, such as splitting by interval size or number of data points. Such strategies do not take into account the differences in element processing costs, making it difficult for the splitting results to meet the requirements of load balancing, and further reducing the resource utilization of parallel computing.

[0029] Finally, existing Spliterator implementations lack sufficient flexibility, typically relying on default splitting strategies and failing to customize task splitting and load balancing strategies for specific application scenarios. This severely limits the optimization space for parallel computing performance when processing data with special cost characteristics.

[0030] Therefore, in order to solve at least one technical problem in the related art, embodiments of this application provide a parallel computing performance optimization method, apparatus, device, and medium.

[0031] In this technical solution, the CostAwareSpliterator is first initialized to receive a data list and a custom cost calculator, determining the initial processing interval and associating it with the custom cost calculator. Then, within the initial processing interval, the element processing cost is obtained based on the custom cost calculator. The interval processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal. Next, the total interval cost is dynamically split into tasks to achieve subtask load balancing, resulting in multiple subtasks that support parallel execution. Then, multiple subtasks are executed in parallel, processing elements within their respective intervals. The interval is dynamically shrunk based on the element processing status, and the total interval cost is updated after the dynamic shrinkage. Finally, based on the dynamically updated total interval cost, the process jumps back to the step of dynamically splitting tasks based on the total interval cost to achieve subtask load balancing and obtain multiple subtasks that support parallel execution. This process of dynamically splitting and parallel executing subtasks is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

[0032] Understandably, the initial dynamic cost parallel task splitter, by associating a custom cost calculator with initial interval locking, ensures precise matching between cost assessment logic and data range, providing a reliable foundation for subsequent splitting and controlling initial cost assessment errors. Incremental computation and caching mechanisms, when dynamically acquiring interval costs, avoid redundant calculations through cache reuse and incremental updates, shortening the response time for a single cost query and significantly improving efficiency in high-frequency splitting scenarios. Dynamic task splitting based on the total interval cost strategy improves subtask load balancing, resolving task backlog issues caused by traditional quantity-based splitting. Interval dynamic shrinkage and cost updates achieve interval shrinkage and cost calibration through atomic operations, ensuring data consistency in a multi-threaded environment. A dual-caching verification mechanism guarantees accurate total cost updates, avoiding splitting decision errors. Looping splitting and resource reclamation achieve fully automated load optimization through cyclical execution of splitting and parallel processing. The final resource reclamation mechanism reduces memory leak risks and improves system stability. In summary, by leveraging multi-dimensional technologies to create synergistic advantages in performance, adaptability, and resource efficiency, a highly efficient solution for big data parallel computing is provided.

[0033] The parallel computing performance optimization scheme provided in this application can also be executed by an electronic device, such as a server, server cluster, or cloud server. This electronic device can also be a terminal device such as a mobile phone, computer, tablet computer, wearable device, or dedicated device (e.g., a dedicated terminal device with a parallel computing performance optimization method). These electronic devices may also incorporate the chips or other hardware processing units described in the above embodiments. Alternatively, these electronic devices may also install a service program for executing the parallel computing performance optimization scheme.

[0034] Figure 1 This is a flowchart illustrating a parallel computing performance optimization method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0035] 101. Initialize CostAwareSpliterator, receive the data list and custom cost calculator, determine the initial processing range and associate it with the custom cost calculator;

[0036] 102. Within the initial processing interval, the element processing cost is obtained based on a custom cost calculator. The interval processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal.

[0037] 103. Based on the total cost of the interval, the task is dynamically split to achieve subtask load balancing and obtain multiple subtasks that support parallel execution.

[0038] 104. Execute multiple subtasks in parallel, process the elements in the interval corresponding to each subtask, and dynamically shrink the interval based on the element processing status, and update the total cost of the interval after dynamic shrinkage.

[0039] 105. Based on the dynamically updated total cost of the interval, jump to the step of dynamically splitting the task according to the total cost of the interval, realize the load balancing of the subtasks, and obtain the steps of multiple subtasks that support parallel execution. The steps of dynamically splitting and parallel execution of subtasks are executed in a loop until the splitting conditions can no longer be met, and all resources are automatically reclaimed.

[0040] In this embodiment, the Dynamic Cost Parallel Task Splitter (CostAwareSpliterator) is a high-efficiency parallel task splitter based on dynamic cost calculation. It implements the Spliterator interface and is primarily used to optimize parallel computing performance when processing elements with uneven costs. Its core logic involves accurately evaluating the processing cost of each element using a custom CostCalculator and implementing intelligent task partitioning in the trySplit method. It holds a reference to the original data list and uses start and end pointers to mark the current processing interval; initially, it represents the complete data list interval.

[0041] In practical applications, CostAwareSpliterator introduces a custom cost calculation method, allowing developers to define the calculation logic for element processing costs by implementing the `calculateCost` method of the `CostCalculator` interface. Leveraging incremental calculation and caching mechanisms, it caches interval processing costs using `cachedCost` and `totalCost`, avoiding full traversal and redundant calculations. Dynamic splitting based on the total cost of the interval ensures load balancing of the split subtasks, rather than splitting based on a fixed data volume. Through these mechanisms, CostAwareSpliterator can adapt to different business needs and data characteristics, effectively improving the efficiency of parallel computing in handling big data and complex cost calculation scenarios.

[0042] As an optional embodiment, in step 101, the CostAwareSpliterator is initialized to receive a data list and a custom cost calculator. This includes: receiving the data list to be processed (List) and the custom cost calculator (CostCalculator); initializing member variables, making data point to the List to be processed, setting the member variables start to 0 and end to the length of the List, and initially setting cachedCost and totalCost to empty to support subsequent incremental calculations. Specifically, cachedCost is used to cache the processing cost of the initial processing interval [start, end), and totalCost is used to cache the total processing cost of the initial processing interval [start, end). TotalCost is calculated using lazy loading. Furthermore, in step 101, the initial processing interval is determined and associated with the custom cost calculator. This includes: indirectly calling the calculateCost method through costCalculator to associate the passed-in cost calculator with the initial processing interval [start, end), providing a basis for subsequent dynamic splitting based on the total cost of the interval.

[0043] Specifically, in step 101, when initializing CostAwareSpliterator, it first receives the list of data to be processed (List) and the custom cost calculator (CostCalculator). During the initialization of member variables, data is set to point to the List to be processed, thereby obtaining the elements to be processed. The member variable start is set to 0, and end is set to the length of the List, thus defining the initial processing interval as [start, end), which covers the entire data list. At the same time, cachedCost and totalCost are initially empty, preparing for subsequent incremental calculations. Among them, cachedCost, as a variable used to cache the processing cost of the initial processing interval [start, end), will be cleared when the interval changes due to splitting or element processing to ensure the timeliness of cost data. When the cost needs to be calculated again, the estimatePartialCost method is used to perform incremental calculation on the current interval, and the calculation result will be stored in cachedCost for subsequent operations, reducing redundant calculations. `totalCost` is used to cache the total processing cost of the initial processing interval [start, end). It is calculated using a lazy loading method. When a task is split for the first time and `totalCost` has not been calculated, the `estimatePartialCost` method is used to traverse the elements of the current interval, accumulate the processing cost to obtain `totalCost`, and cache it to provide a basis for subsequent task splits.

[0044] In addition, step 101 also includes determining the initial processing interval and associating it with a custom cost calculator. Determining the initial processing interval is achieved by setting `start` to 0 and `end` to the length of the list. Associating the custom cost calculator is done by indirectly calling the `calculateCost` method through `costCalculator`. Since the passed-in custom cost calculator `CostCalculator` is associated with the member variable `costCalculator` during the initialization of `CostAwareSpliterator`, when it is necessary to associate a cost calculator with the initial processing interval [start, end), the `calculateCost` method implemented by `costCalculator` can be called. This associates the passed-in cost calculator with the initial processing interval, providing a reliable cost calculation basis for subsequent dynamic splitting based on the total cost of the interval, ensuring that subsequent task splitting can be based on accurate element processing costs.

[0045] Optionally, determining the initial processing interval and associating it with custom cost calculation logic is a core operation in the initialization phase of the dynamic cost parallel task splitter, CostAwareSpliterator. The specific implementation is as follows: During initialization, CostAwareSpliterator receives a list of data to be processed (List), sets its member variables `start` to 0 and `end` to the size of the data list, thus defining the initial processing interval as [0, data list size), covering the entire data set. This interval definition clearly defines the range of elements that the Spliterator needs to process in its initial state, providing a basic boundary for subsequent task splitting and element processing.

[0046] First, developers need to implement the `calculateCost` method of the `CostCalculator` interface. This method contains custom cost calculation logic based on element attributes, processing complexity, and other factors. When initializing the dynamic cost-parallel task splitter `CostAwareSpliterator`, the custom cost calculator implementing this interface is passed as a parameter, associating the `costCalculator` member variable of `Spliterator` with this custom calculator. Subsequently, when the dynamic cost-parallel task splitter `CostAwareSpliterator` needs to obtain the element processing cost, it can indirectly call the `calculateCost` method of `costCalculator` by calling the `getElementProcessingCost` method, thereby accessing and using the developer-defined custom cost calculation logic and providing a basis for subsequent cost-based dynamic splitting. Through this approach, the initial processing scope is clearly defined, and the custom cost calculation logic is successfully integrated, laying the foundation for optimizing the overall performance of the `Spliterator`.

[0047] 102. Within the initial processing interval, the element processing cost is obtained based on a custom cost calculator. The interval processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal.

[0048] As an optional embodiment, in step 102, within the initial processing interval, the element processing cost is obtained based on a custom cost calculator. This includes: defining custom cost calculation logic for each element's processing cost by implementing the `calculateCost` method of the `CostCalculator` interface; wherein the custom cost calculation logic is determined based on element attributes, processing complexity, and correlation with other elements. Furthermore, the custom cost calculation logic for each element is called through the `getElementProcessingCost` method to obtain the actual processing cost of each element, providing a basis for dynamic splitting.

[0049] Specifically, in step 102, when obtaining the element processing cost based on the custom cost calculator within the initial processing interval, the developer first needs to implement the `calculateCost` method of the `CostCalculator` interface to define custom cost calculation logic for each element. This custom cost calculation logic is not fixed or uniform, but is determined based on the element's attributes (such as file size, text length), processing complexity (such as whether it contains complex calculations, parsing difficulty), and its relationship with other elements (such as the impact of dependencies between elements on processing order and time consumption). For example, when processing documents with different formats, documents with many images and complex layouts have high processing complexity, and the cost calculation logic defined by the `calculateCost` method will assign them a higher processing cost; while plain text documents with simple structures will be assigned a lower processing cost. Then, the custom cost calculation logic for each element is called through the `getElementProcessingCost` method to obtain the actual processing cost of each element. In this process, the getElementProcessingCost method acts as an intermediary bridge, connecting CostAwareSpliterator with the custom cost calculation logic, so that the processing cost of each element can be accurately obtained, providing an important basis for subsequent dynamic splitting.

[0050] Regarding the dynamic acquisition of interval processing costs through incremental calculation and caching mechanisms, since the initial processing interval is [start, end), when the processing cost of this interval is needed, a full traversal and calculation of all elements within the interval is not performed. Instead, incremental calculation is used. When interval cost data is needed for the first time, the `estimatePartialCost` method is used to calculate and accumulate the cost of each element within the current interval. After obtaining the total interval cost, the result is stored in `cachedCost` for caching. In subsequent operations, if the interval remains unchanged, the interval cost is read directly from `cachedCost` when needed again, avoiding duplicate calculations. If the interval changes due to the processing or splitting of some elements, `cachedCost` is cleared. In this case, the `estimatePartialCost` method is used again to perform incremental calculation on the changed interval, obtain the new total interval cost, and cache it.

[0051] In this way, on the one hand, the custom cost calculation logic can accurately reflect the actual processing cost differences of different elements, providing a reliable foundation for subsequent cost-based dynamic splitting and ensuring that the splitting strategy is more in line with the actual processing situation. On the other hand, the incremental calculation and caching mechanism avoids a full traversal of the initial processing interval, greatly reducing unnecessary computational overhead. Especially when the initial processing interval contains a large number of elements, it can significantly improve the efficiency of obtaining the interval processing cost and save computing resources. At the same time, the use of the caching mechanism also reduces the time consumption caused by repeated calculations, laying the foundation for the efficient operation of the entire parallel computing process.

[0052] As an optional embodiment, in step 102, the interval processing cost is dynamically obtained through incremental calculation and caching mechanisms, including: when the initial processing interval changes due to splitting or element processing, the cachedCost corresponding to the initial processing interval is cleared to ensure the timeliness of cost data. When calculating the total interval cost, the estimatePartialCost method is called to perform incremental calculation for the current interval [start, end). The calculated total interval cost is stored in cachedCost and totalCost for direct use in subsequent operations to reduce redundant calculations.

[0053] Specifically, in step 102, when dynamically obtaining the interval processing cost through incremental calculation and caching mechanisms, if the initial processing interval changes due to splitting or element processing—for example, if some elements are processed causing the start value to increase, or if the interval is split into multiple sub-intervals—the cachedCost corresponding to the initial processing interval will be immediately cleared. This operation ensures that outdated cost data is not retained in the cache, guaranteeing the timeliness of subsequent cost calculations and preventing inaccurate interval total cost assessments due to the use of old data.

[0054] When calculating the total cost of an interval, the `estimatePartialCost` method is called to perform incremental calculations for the current interval [start, end). This method does not iterate through the entire original data list; instead, it only calculates the cost of each element within the current interval, starting from the `start` position, and then sums them up to obtain the total cost of the interval. For example, if the current interval is [2, 5), only the costs of elements with indices 2, 3, and 4 are calculated and summed, without involving other irrelevant elements. This incremental calculation method significantly reduces unnecessary computation.

[0055] After calculating the total cost for the interval, it is stored in both `cachedCost` and `totalCost`. `cachedCost` serves as an immediate cache of the current interval cost, readily available for subsequent cost queries within a short period. `totalCost`, on the other hand, caches the total cost using lazy loading, serving as the basis for calculating the target split cost during subsequent task splitting. When the cost data for that interval is needed again, if the interval has not changed, it can be read directly from either `cachedCost` or `totalCost` without needing to recalculate using the `estimatePartialCost` method, thus reducing redundant calculations.

[0056] Therefore, clearing the cachedCost of the changed intervals fundamentally avoids the interference of outdated data on cost assessment, ensuring the accuracy of each cost calculation and providing a reliable data foundation for cost-based dynamic splitting. The incremental calculation mode skips processed or irrelevant elements, significantly reducing computational overhead, especially in scenarios with large datasets and frequent interval changes, greatly improving computational efficiency. Storing the total cost in cachedCost and totalCost enables data reuse, reducing the number of repeated calculations and further saving computational resources. This makes the entire parallel computing process more efficient in the cost assessment stage, indirectly improving overall parallel processing performance.

[0057] Understandably, in methods for optimizing Spliterator performance, `cachedCost` and `totalCost` are key member variables of the dynamic cost parallel task splitter `CostAwareSpliterator` for cost calculation and caching. Their specific definitions and functions are as follows: `cachedCost` is a variable used to cache the processing cost of the current interval [start, end). When the interval changes due to splitting or element processing, `cachedCost` is cleared to ensure the timeliness of cost data. When the cost needs to be calculated again, the `estimatePartialCost` method performs incremental calculation on the current interval, and the result is stored in `cachedCost` for direct use in subsequent operations, reducing redundant calculations. This is closely related to the incremental calculation and caching mechanism. `totalCost` is used to cache the total processing cost of the current interval, calculated using lazy loading. When the task is first split (by calling the `trySplit` method) and `totalCost` has not yet been calculated, the `estimatePartialCost` method iterates through the elements of the current interval, accumulates the processing cost to obtain `totalCost`, and caches it. In subsequent splitting, totalCost serves as the basis for calculating the target splitting cost; that is, the splitting point is determined based on half of totalCost to achieve load balancing of subtasks. When the interval changes, totalCost is updated accordingly to the total cost of the remaining interval, ensuring the accuracy of dynamic splitting. In summary, cachedCost and totalCost, by caching interval cost data, support the incremental calculation mechanism and dynamic splitting logic, and are important variables for reducing the overhead of full traversal and improving the efficiency of parallel computing.

[0058] 103. Based on the total cost of the interval, the task is dynamically split to achieve subtask load balancing and obtain multiple subtasks that support parallel execution.

[0059] As an optional implementation, in step 103, the `trySplit` method is called; if `totalCost` is not cached, the total cost of the current interval is obtained through the `estimatePartialCost` method. The target split cost is half of the total interval cost. Starting from the `start` position, the element cost is accumulated to determine the split point `mid`, ensuring load balancing of the subtasks after splitting. The first dynamic split interval cost of the left half interval [start, mid) is obtained. An instance of a parallel task splitter `CostAwareSpliterator` is created to process the left half interval [start, mid) to split it into the first subtask, and `totalCost` is updated to the accumulated element processing cost of all elements within the left half interval [start, mid). The processing interval of the original `Spliterator` is updated to the right half interval [mid, end), and the second dynamic split interval cost of the right half interval [mid, end) is obtained. The corresponding second subtask is obtained through the instance, and `totalCost` is updated to the accumulated element processing cost of all elements within the right half interval [mid, end) as the total interval cost of the right half interval [mid, end). `cachedCost` is simultaneously cleared.

[0060] Specifically, in step 103, the `trySplit` method is called to dynamically split the task based on the total cost of the interval, achieving load balancing of subtasks and obtaining multiple subtasks that support parallel execution. This process begins by obtaining the total cost of the current interval. If `totalCost` is not yet cached, the `estimatePartialCost` method is called to calculate the total cost of the current interval. The `estimatePartialCost` method iterates through the elements in the interval, calculates the processing cost element by element using a custom `CostCalculator` interface, and accumulates the result, which is the total cost of the interval. This process reflects the characteristics of incremental calculation and avoids a full traversal of the entire dataset.

[0061] After obtaining the total cost of the interval, half of it is used as the target split cost. Starting from the start position of the interval, the processing cost of each element is accumulated until the accumulated value reaches or exceeds the target split cost. This position is then determined as the split point mid. This splitting logic no longer relies on a fixed number of elements, but is based on the actual processing cost of the elements, ensuring that the left and right intervals after splitting have a relatively balanced processing load. For example, when processing data containing orders with different complexities, if an interval contains high-cost cross-store orders and low-cost ordinary orders, this cost-based cumulative splitting can make the total costs of the left and right intervals closer, avoiding the situation where a subtask concentrates on processing high-cost elements due to simple quantity splitting. After determining the split point mid, a CostAwareSpliterator instance is created to process the left half of the interval [start, mid). This instance serves as the first subtask, and its totalCost is set to the accumulated value of the processing costs of all elements in the left half of the interval, i.e., the cost of the first dynamically split interval. At the same time, the processing range of the original Spliterator is updated to the right half range [mid, end), and its totalCost is also updated accordingly to the total cost of the right half range, that is, the cost of the second dynamically split range. Meanwhile, cachedCost will be cleared synchronously because the change of the range makes the previous cached cost data no longer applicable. This operation is in line with the design of incremental calculation and caching mechanism, ensuring that subsequent cost calculations are based on the latest range.

[0062] In this way, by dynamically splitting based on the total cost of the interval, the limitations of traditional fixed-quantity splitting strategies are overcome, and the advantages of custom cost calculation methods are fully utilized. This results in a more balanced load on the split subtasks, reducing thread idleness or overload, and improving the overall efficiency of parallel computing. The application of the estimatePartialCost method enables incremental cost calculation, avoiding the performance loss caused by full traversal, which is particularly effective when processing large datasets. Timely clearing of cachedCost ensures the accuracy of cost data, providing a reliable basis for possible subsequent splitting. Furthermore, newly created subtasks execute in parallel with the original task, independently managing their intervals with their respective CostAwareSpliterator instances, further enhancing the flexibility and efficiency of parallel processing. This aligns with the core objective of this invention: optimizing load balancing and improving parallel computing performance through dynamic splitting.

[0063] Alternatively, CostAwareSpliterator can be upgraded to a distributed version. The core of this upgrade is to build a parallel computing framework that is compatible with large-scale cluster environments such as Spark and Flink. By using shard-level cost awareness and a layered splitting architecture, the load balancing problem in distributed scenarios can be solved, allowing incremental computing capabilities to be extended efficiently in the cluster environment.

[0064] The shard-level cost-aware mechanism uses data shards as the basic unit of partitioning, rather than focusing on individual elements. The cost of each shard is estimated by the average processing time of its internal elements: randomly selected sample elements within the shard are processed, the processing time is recorded, and the average is calculated to estimate the processing cost of the entire shard. This approach retains the lightweight nature of incremental computation while adapting to data organization methods in a distributed environment. To achieve dynamic load balancing across nodes, a heartbeat mechanism is introduced. Each node reports its load status to the cluster at fixed intervals (e.g., 100ms), including the number of shards currently being processed, remaining computing resources, network I / O pressure, and other data. After the central coordinator aggregates this information, it can monitor the load differences between nodes in real time, providing a basis for shard allocation and migration decisions.

[0065] The layered splitting architecture employs a two-tier system: coarse-grained splitting at the cluster level and fine-grained splitting at the node level, minimizing cross-node communication overhead. Coarse-grained splitting occurs at the cluster level, dividing the entire dataset into several large initial shards. The shard size is typically matched to the storage and computing capabilities of the cluster nodes (e.g., each shard corresponds to 1GB of data). During the splitting process, a consistent hashing algorithm with virtual nodes is used: multiple virtual nodes are mapped to each physical node. A hash value is calculated based on the shard's characteristic values ​​(e.g., data hash, business identifier), and the shard is allocated to the physical node associated with the corresponding virtual node. This design ensures that shards with the same characteristics (such as the behavioral data of the same user) consistently reside on the same node, reducing cross-node data dependencies. Fine-grained splitting is executed within each node. After receiving the coarse-grained shard allocated by the cluster, each node, following the logic of the single-machine version of CostAwareSpliterator, splits the shard into smaller sub-shards based on element cost, utilizing local multi-threaded parallel processing to fully release the node's CPU and memory resources.

[0066] In actual operation, coarse-grained partitioning first completes the initial allocation of all data: the central coordinator maps shards to each node based on the initial node load using consistent hashing; after receiving the data, the nodes initiate fine-grained partitioning, processing sub-shards through incremental calculation and updating shard costs in real time. When a node experiences excessive load due to a backlog of high-cost shards, a heartbeat mechanism synchronizes its status to the coordinator, triggering dynamic load adjustment: the coordinator selects low-dependency shards that can be migrated from the node, recalculates the target node using consistent hashing, and migrates the shards to nodes with lower loads. During the migration process, leveraging the caching mechanism of incremental calculation, only unprocessed sub-shard data is transmitted, avoiding full data copying.

[0067] The core value of this solution lies in extending single-machine parallel optimization logic to distributed scenarios for the first time: shard-level cost awareness enables more accurate cross-node load assessment; the layered splitting architecture reduces network overhead through "cluster-level control of distribution and node-level efficiency improvement"; and consistent hashing with virtual nodes balances the stability and uniformity of data distribution. Ultimately, the system can control the load difference between nodes within 20% when processing petabyte-scale data, solving the technical problem of "uneven sharding distribution leading to node busy / idle imbalance" in big data processing, and enabling incremental computing to maintain high-efficiency parallel processing capabilities in a distributed environment.

[0068] 104. Execute multiple subtasks in parallel, process the elements in the interval corresponding to each subtask, and dynamically shrink the interval based on the element processing status, and update the total cost of the interval after dynamic shrinkage.

[0069] As an optional embodiment, in 104, multiple subtasks are started in parallel in independent threads. Each subtask locks its respective responsible interval by holding an independent CostAwareSpliterator instance. The intervals corresponding to the subtasks are calibrated by subtask numbers, and the intervals corresponding to each subtask do not overlap and cover the original data set. Furthermore, when any subtask executes element processing, it repeatedly calls the tryAdvance method. Through atomic operations, it determines whether the start value of the current interval is less than the end value. If the splitting condition is met, it processes the current element through action.accept (data.get (start)). During the processing, it records the actual processing time of the current element and compares it in real time with the output value of the custom cost calculation logic. If the deviation exceeds the preset threshold, it triggers the dynamic calibration mechanism of the cost calculator. After the current element is processed, it atomically increments the start value by 1 in a thread-safe manner, dynamically shrinking the interval corresponding to the current unprocessed elements from [start, end) to [start + 1, end). At the same time, it generates an interval change log, recording the interval ranges and timestamps before and after the shrinkage. If the interval undergoes a shrinkage change, it triggers an atomic clearing operation of cachedCost to ensure that any subsequent cost calculation operations cannot read expired cost data. Among them, any subsequent cost calculation operations at least include: internal resplitting within a subtask or cross-task load monitoring. When a subtask detects an interval shrinkage, if the current remaining number of elements is greater than the preset splitting threshold, it automatically triggers a cost recalculation process: when calling the estimatePartialCost method, it first verifies the boundary validity of the current interval [start + 1, end) to ensure that start + 1 < end and the element index does not exceed the boundary; then it repeatedly calls the calibrated cost calculator logic based on the latest start value for each element, and accumulates to obtain the new total cost of the interval. Then, it writes the recalculated total cost into the totalCost variable through a double-buffer mechanism: first write it into the temporary buffer and perform verification to ensure that the total cost of the interval is non-negative and has a reasonable proportional relationship with the number of elements; after passing the verification, it atomically replaces the original totalCost value, and at the same time updates the cost calculation timestamp, providing a reference basis in the time dimension for subsequent splitting strategies. Finally, during the parallel execution of multiple subtasks, they synchronize their respective interval shrinkage situations and the latest total costs in real time through the interval status table in shared memory, for the system-level load balancer to perform global monitoring. When it detects that the remaining total cost of any subtask deviates from that of other subtasks by more than the set threshold, it triggers cross-subtask dynamic load scheduling to ensure the overall parallel efficiency.

[0070] It can be understood that in the above step 104, multiple subtasks are started in parallel in independent threads. Each subtask holds an independent CostAwareSpliterator instance and locks the interval it is responsible for. These intervals are calibrated by subtask numbers, do not overlap with each other, and jointly cover the original data set. Taking the processing of text data with different complexities as an example, each subtask corresponds to an interval of text data, and the elements within the interval are managed through their respective CostAwareSpliterator instances.

[0071] When any subtask executes element processing, it will repeatedly call the tryAdvance method. This method uses atomic operations to determine whether the start value of the current interval is less than the end value. If the condition is met, it will process the current element through action.accept(data.get(start)). During the processing, the actual processing time of the element will be recorded and compared in real time with the cost value obtained from the custom cost calculation logic. If the deviation between the two exceeds the preset threshold, the dynamic calibration mechanism of the cost calculator will be triggered to ensure the accuracy of cost evaluation. After the current element is processed, the start value is atomically incremented by 1 in a thread-safe manner, so that the interval of unprocessed elements dynamically shrinks from [start, end) to [start + 1, end). At the same time, an interval change log is generated to record the interval range and timestamp before and after the shrinkage, providing a basis for subsequent tracking.

[0072] Due to the contractive change of the interval, cachedCost will be atomically cleared to ensure that any subsequent operation involving cost calculation, whether it is the re-splitting within the subtask or the load monitoring across tasks, cannot read expired cost data. When a subtask detects the interval shrinkage, if the current remaining number of elements is still greater than the preset splitting threshold, it will automatically trigger the cost recalculation process. When calling the estimatePartialCost method, first verify the boundary validity of the current interval [start + 1, end) to ensure that start + 1 < end and the element index does not exceed the boundary. Then, based on the latest start value, call the calibrated cost calculator logic for each element and accumulate to obtain the total cost of the new interval. This process reflects the characteristics of incremental calculation and avoids full traversal.

[0073] The recalculated total cost is written to the `totalCost` variable using a dual-caching mechanism. First, it's written to a temporary cache for verification, ensuring the total cost is non-negative and proportional to the number of elements. Once verified, the original `totalCost` value is atomically replaced, and the cost calculation timestamp is updated to provide a time-based reference for subsequent splitting strategies. When multiple subtasks execute in parallel, their respective interval shrinkage and latest total cost are synchronized in real-time through an interval status table in shared memory, allowing the system-level load balancer to monitor globally. When the remaining total cost of a subtask deviates from a set threshold from other subtasks, dynamic load scheduling across subtasks is triggered to ensure overall parallel efficiency.

[0074] Thus, by processing subtasks in parallel through independent threads, combined with atomic operations and thread safety mechanisms, the accuracy and stability of data processing in a multi-threaded environment are ensured. The dynamic calibration mechanism improves the evaluation accuracy of the custom cost calculator, making cost calculations more closely reflect actual processing time and providing a reliable basis for subsequent dynamic splitting. Clearing cachedCost and triggering incremental recalculation during interval shrinkage avoids the impact of duplicate calculations and expired data, conforming to the design of incremental calculation and caching mechanisms and improving performance. Dual cache verification and boundary validity verification further guarantee the reliability of cost data, while cross-task load scheduling optimizes overall load balancing, reduces thread idle time and resource waste, improves the efficiency of parallel computing, and also reduces memory consumption, fully leveraging the advantages of this invention in custom cost calculation, incremental calculation and caching, and dynamic splitting.

[0075] It can be explained that the real-time feedback adaptive mechanism is an innovative extension of the dynamic cost perception solution. Its core lies in building a closed-loop feedback system, which dynamically adjusts the splitting strategy according to the real-time status of task execution. Through the synergistic effect of dynamic cost correction and elastic scaling strategy, a learning-optimization closed loop is achieved, and the splitting strategy evolves dynamically with the load.

[0076] The dynamic cost adjustment mechanism focuses on improving the accuracy of cost prediction. During task execution, the system periodically samples the actual processing time of elements and compares the sampled results with the cost predicted by the CostCalculator model. When the actual processing time of an element exceeds the prediction by 20%, a local rebalancing is triggered. In this process, the Online Gradient Descent (OGDA) algorithm plays a crucial role, adjusting the parameters of the cost prediction model in real time, such as the weighting coefficients of element processing time. Through continuous iterative optimization, the model's prediction results become closer to the actual processing situation. Simultaneously, reinforcement learning (Q-Learning) is also incorporated. It defines a state space encompassing indicators reflecting the current system state, such as queue backlog and thread activity; and an action space including operations such as adjusting the splitting granularity by ±10%. The reward function uses task completion time and resource utilization as evaluation criteria. Through continuous learning via Q-Learning, the system can explore optimal splitting strategies, such as appropriate target cost ratios and reasonable rebalancing trigger thresholds, making cost prediction and task splitting more consistent with the system's actual load.

[0077] The elastic scaling strategy primarily focuses on optimizing task splitting granularity and thread resource allocation. The system closely monitors queue backlog. When queue backlog is severe, it automatically increases the target cost ratio for splitting tasks, reducing the number of small tasks and preventing excessive task consumption of system resources. Conversely, when the queue is relatively idle, it decreases the target cost ratio for splitting tasks, increasing task parallelism. Simultaneously, it dynamically starts and stops worker threads based on thread pool status, such as the number of active threads and queue length. Adaptive threshold detection (CUSUM algorithm) is used to detect sudden changes in processing time; if the processing time for a certain type of element suddenly slows down, the system can promptly detect and respond. A queuing theory model is used to optimize thread pool size and task allocation efficiency. By analyzing the thread pool load, it determines the optimal number of threads and task allocation method, ensuring efficient utilization of thread resources.

[0078] By combining dynamic cost adjustment and elastic scaling strategies, the real-time feedback adaptive mechanism forms a complete learning-optimization closed loop. Dynamic cost adjustment provides accurate cost data support for the elastic scaling strategy, which in turn adjusts the partitioning and resource allocation based on the real-time system status. This continuous cyclical interaction allows the partitioning strategy to dynamically adjust with changes in system load, continuously optimizing the performance and resource utilization of parallel computing. Whether facing fluctuations in element processing costs or changes in system load, this mechanism can respond quickly and make reasonable adjustments, significantly improving the system's adaptability and efficiency.

[0079] 105. Based on the dynamically updated total cost of the interval, jump to the step of dynamically splitting the task according to the total cost of the interval, realize the load balancing of the subtasks, and obtain the steps of multiple subtasks that support parallel execution. The steps of dynamically splitting and parallel execution of subtasks are executed in a loop until the splitting conditions can no longer be met, and all resources are automatically reclaimed.

[0080] Specifically, in step 105, based on the dynamically updated total cost of the interval, the system jumps to the step of dynamically splitting the task according to the total cost of the interval, thereby achieving load balancing of subtasks and obtaining multiple subtasks that can be executed in parallel. Afterwards, the system will repeatedly execute the process of dynamic splitting and parallel execution of subtasks until the splitting conditions can no longer be met, and finally automatically reclaim all resources. During this loop, when processing elements, the subtask calls the tryAdvance method. If the start position of the current interval is less than the end position, the corresponding element is processed, and then the start position is incremented, causing the interval to dynamically shrink. Due to the interval change, the cached interval cost is cleared, ensuring that subsequent cost calculations are based on the latest interval. Based on the dynamically updated total cost of the adjusted interval, the subtask will trigger the dynamic splitting logic again. The newly created subtask is also a dynamic cost parallel task splitter instance, responsible for processing the left half of the interval, while the current subtask is updated to the right half of the interval and the corresponding cost data is adjusted. The newly generated subtask and the current subtask are executed in parallel, continuously adjusting the interval and updating the cost during element processing, and repeating the dynamic splitting operation based on the updated total cost. This loop continues until a subtask meets the condition that it cannot be split: when the number of elements in the interval is so small that splitting is unnecessary, or when the total cost of the interval is zero, rendering splitting meaningless, the dynamic splitting method returns a null value, terminating the splitting process for that subtask. When all subtasks have terminated splitting and completed processing of all elements—that is, when the start and end positions of the intervals in each subtask coincide—all instances of the dynamic cost parallel task splitter will no longer be used. At this point, the system will automatically reclaim the memory resources occupied by these instances, including data references, interval pointers, cost caches, etc. Thus, the entire process is complete, achieving efficient parallel processing of large-scale data.

[0081] This process leverages the advantages of custom cost calculation and incremental caching mechanisms through the linkage of element processing and interval adjustment, the iterative optimization of dynamic splitting, and the automatic recycling of resources, ensuring the load balance and resource utilization efficiency of parallel computing.

[0082] For example, in step 105, based on the dynamically updated total cost of the interval, the system jumps to the dynamic task splitting stage, repeatedly executing the splitting and parallel execution steps until it can no longer be split and resources are reclaimed. Taking the processing of a batch of text data containing different complexities as an example, the initial data list contains 10 articles, the initial interval of CostAwareSpliterator is [0, 10), and the custom cost calculator assigns different costs to each article according to the number of words and sentence complexity, such as a cost of 1 for a simple short article and a cost of 5 for a complex long article.

[0083] After the initial split, a left half [0, 5) (total cost 12) and a right half [5, 10) (total cost 13) are obtained. The subtask in the left half begins processing elements. Assuming that after processing the short article at index 0, the interval shrinks to [1, 5), and the total cost is updated to 11. Based on this updated total cost, the subtask triggers a second split. The target split cost is calculated as 5.5 using the trySplit method. The element cost is accumulated starting from index 1. When the accumulated cost reaches index 3, the split point is determined to be 3. The new subtask is responsible for [1, 3), while the current subtask is adjusted to [3, 5). The new subtask and the current subtask execute in parallel. When the new subtask processes articles at indices 1 and 2, the interval shrinks to [3, 3) and the split terminates. When the current subtask processes articles at indices 3 and 4, the interval becomes [5, 5) and the split stops. The processing of the right half interval [5, 10) is similar. After multiple splits and parallel executions, all articles are finally processed. The start and end of each subtask are equal, and the memory resources occupied by the CostAwareSpliterator instance are automatically reclaimed. For example, the list references storing article data and the pointers to the marked intervals are released.

[0084] Further optionally, the splitting condition includes at least one of the following: the number of elements in the interval is greater than 1, and the total cost of the interval is greater than 0. Further optionally, the automatic reclamation of all resources when the splitting condition cannot be met includes: returning null using the trySplit method to terminate the dynamic data interval splitting process of the data list.

[0085] For example, if the number of elements in the interval is less than or equal to 1, there is no need to split it further. Or, if the total cost of the interval is 0, splitting it is meaningless. In this case, the trySplit() method returns null, terminating the splitting process for this subtask.

[0086] In one optional embodiment of this application, CostAwareSpliterator is a generic class that implements the Spliterator interface. Its core function is to achieve efficient parallel task splitting based on dynamic cost calculation. Its overall implementation revolves around custom cost calculation, incremental caching, and dynamic splitting.

[0087] First, a custom cost calculation interface, `CostCalculator`, is defined. This interface only defines the `calculateCost` method, used to calculate the processing cost of a single element. The specific logic is written by the implementation class according to the business scenario. For example, calculation rules can be set based on element size, complexity, etc., providing accurate cost basis for subsequent task splitting. The member variables of `CostAwareSpliterator` include `Listdata` storing elements, `start` and `end` marking the interval, `cachedCost` caching the current interval cost, `totalCost` caching the total processing cost, and `CostCalculator` for cost calculation. The constructor receives a data list and a cost calculator. During initialization, `start` is set to 0 and `end` is set to the size of the data list. At this point, the processing interval is the entire dataset, and the custom cost calculator is associated with it.

[0088] For example, the core method `trySplit` is used for task splitting: First, it checks whether `totalCost` and `cachedCost` are cached. If not, it calls the `estimatePartialCost` method to calculate the total cost of the current interval. This helper method iterates through the elements in the interval, calling `getElementProcessingCost` (which internally calls `costCalculator.calculateCost`) to accumulate the cost element by element. If the total cost is 0 or the number of elements is ≤1, splitting is not possible. Otherwise, it aims for half of the total cost, starts from `start`, accumulates the element cost to determine the split point `mid`, creates a new instance to process the left interval [start, mid) (its `totalCost` is the cost of the left interval), updates the original instance to the right interval [mid, end) and updates `totalCost`, and clears `cachedCost` to ensure that the subtasks are load-balanced after splitting.

[0089] For example, the tryAdvance method is used to process a single element: if start < end, the current element is processed through action.accept, start is incremented by 1 to shrink the interval, and cachedCost is cleared to ensure the timeliness of cost data. If the processing is successful, true is returned; otherwise, false is returned. The estimateSize method directly returns end - start to provide real-time feedback on the remaining number of elements. In the overall implementation process, the complete processing interval is determined during initialization and associated with the cost calculator; during the first split, the total cost is calculated through lazy loading, and the interval and cost cache are updated based on the cost after splitting; when processing elements, the interval is dynamically shrunk and the cache is cleared to prepare accurate data for the next split. Through incremental calculation (only processing the current interval), cache reuse (reducing duplicate calculations), and cost-based dynamic splitting, efficient parallel computing is finally achieved in scenarios with significant differences in element processing costs, improving the overall processing efficiency and load balancing.

[0090] In the technical solution of this application, by introducing a dynamic cost calculation, incremental caching, and adaptive splitting mechanism, the parallel processing efficiency of large-scale uneven cost data is significantly improved. In terms of load balancing, it breaks through the limitation of traditional fixed data volume splitting, dynamically divides tasks based on the actual processing cost of elements, and controls the load difference of each subtask within a preset threshold to avoid thread idleness or overload, thus improving the overall parallel efficiency. In terms of computational overhead, the incremental calculation and caching mechanism are used to avoid full traversal of the data set, which is especially suitable for processing large data sets. In terms of scenario adaptability, it supports custom cost calculation logic, can flexibly adjust the evaluation rules according to element attributes, and adapts to multiple business scenarios, significantly improving the generality. In terms of resource utilization, through the dynamic shrinking interval and cyclic splitting mechanism, resources are allocated on demand, reducing thread idle time, and combined with automatic resource recycling to reduce memory occupancy.

[0091] In another embodiment of this application, a device for optimizing parallel computing performance is also provided. Refer to Figure 2 as described, the device includes the following units:

[0092] An initialization unit, configured to initialize CostAwareSpliterator, receive a data list and a custom cost calculator, determine the initial processing interval, and associate the custom cost calculator;

[0093] A cost unit, configured to obtain the element processing cost based on the custom cost calculator within the initial processing interval, and dynamically obtain the interval processing cost through the incremental calculation and caching mechanism to avoid full traversal;

[0094] The dynamic splitting unit is configured to dynamically split tasks based on the total cost of the interval, achieving subtask load balancing and obtaining multiple subtasks that support parallel execution. Multiple subtasks are executed in parallel, processing elements within their respective intervals, and the intervals are dynamically shrunk based on the element processing status, updating the total cost of the dynamically shrunk interval. Based on the dynamically updated total cost of the interval, the unit jumps to the step of dynamically splitting tasks based on the total cost of the interval, achieving subtask load balancing, and obtaining multiple subtasks that support parallel execution. This process of dynamically splitting and executing subtasks in parallel is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

[0095] The above-described apparatus can implement the various steps in the above method embodiments, which will not be elaborated here.

[0096] Please see Figure 3 , Figure 3 A schematic diagram illustrating an embodiment of the electronic device provided in this application. For example... Figure 3 As shown, this application provides an electronic device 500, including a memory 510, a processor 520, and a software program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the software program 511, it implements the parallel computing performance optimization method described in the above embodiments.

[0097] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided in this application. For example... Figure 4 As shown, this embodiment provides a computer-readable storage medium 600 on which a computer program 611 is stored. When the computer program 611 is executed by a processor, it implements the parallel computing performance optimization method described in the above embodiment.

[0098] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. Although preferred embodiments of this application have been described, those skilled in the art, once they understand the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application. Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A method for optimizing the performance of parallel computing, characterized in that, The method includes: Initialize the dynamic cost parallel task splitter CostAwareSpliterator, receive a data list and a custom cost calculator, determine the initial processing range and associate it with the custom cost calculator; Within the initial processing range, the element processing cost is obtained based on a custom cost calculator, and the range processing cost is dynamically obtained through incremental calculation and caching mechanisms to avoid full traversal. Dynamic task splitting is performed based on the total cost of the interval to achieve subtask load balancing and obtain multiple subtasks that support parallel execution; Multiple subtasks are executed in parallel, each processing the elements within its corresponding interval. The intervals are then dynamically shrunk based on the processing status of the elements, and the total cost of the dynamically shrunk intervals is updated. Based on the dynamically updated total cost of the interval, the process jumps to dynamically splitting the task according to the total cost of the interval, achieving load balancing of subtasks, and obtaining steps for multiple subtasks that support parallel execution. The process of dynamically splitting and parallel executing subtasks is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

2. The parallel computing performance optimization method according to claim 1, characterized in that, The initialization dynamic cost parallel task splitter, CostAwareSpliterator, receives a data list and a custom cost calculator, including: Receives a list of data to be processed (List) and a custom cost calculator (CostCalculator); Initialize member variables, making data point to the List to be processed, set the member variable start in the List to 0, and set end to the length of the List. The cost cache cachedCost and the interval total cost cache totalCost are initially empty to support subsequent incremental calculations. Among them, cachedCost is used to cache the variable of processing cost for the initial processing interval [start, end), and totalCost is used to cache the total processing cost for the initial processing interval [start, end). totalCost is calculated using a lazy loading method. The process of determining the initial processing interval and associating it with the custom cost calculator includes: The `calculateCost` method is indirectly called through `costCalculator` to associate the input cost calculator with the initial processing interval [start, end), providing a basis for subsequent dynamic splitting based on the total cost of the interval.

3. The parallel computing performance optimization method according to claim 2, characterized in that, The step of obtaining the element processing cost based on a custom cost calculator within the initial processing interval includes: By implementing the calculateCost method of the CostCalculator interface, a custom cost calculation logic for the element processing cost is defined for each element; the custom cost calculation logic is determined based on the element attributes, processing complexity, and correlation with other elements. The getElementProcessingCost method calls the custom cost calculation logic for each element to obtain the actual processing cost of each element, providing a basis for dynamic splitting.

4. The parallel computing performance optimization method according to claim 2, characterized in that, The method of dynamically obtaining the interval processing cost through incremental calculation and caching mechanism includes: When the initial processing range changes due to splitting or element processing, the cachedCost corresponding to the initial processing range is cleared to ensure the timeliness of cost data. When calculating the total cost of an interval, the estimatePartialCost method is called to perform incremental calculations for the current interval [start, end). The calculated total cost for the interval is stored in cachedCost and totalCost for direct use in subsequent operations to reduce redundant calculations.

5. The parallel computing performance optimization method according to claim 2, characterized in that, The dynamic task splitting based on the total cost of the interval achieves subtask load balancing, resulting in multiple subtasks that support parallel execution, including: Call the trySplit method; if totalCost is not cached, obtain the total cost of the current interval using the estimatePartialCost method; The cost is split with half of the total cost of the interval as the target. Starting from the start position, the cost of each element is accumulated to determine the split point mid, ensuring that the subtasks are load-balanced after splitting. Get the cost of the first dynamic split interval of the left half interval [start, mid), create an instance of CostAwareSpliterator to process the left half interval [start, mid) to split it into the first subtask, and update totalCost to the sum of the element processing costs of all elements in the left half interval [start, mid). Update the processing range of the original Spliterator to the right half range [mid, end), and obtain the second dynamic split range cost of the right half range [mid, end). Obtain the corresponding second subtask by splitting the instance. Update the totalCost to the sum of the element processing costs of all elements in the right half range [mid, end) as the total range cost of the right half range [mid, end). Simultaneously clear the cachedCost.

6. The parallel computing performance optimization method according to claim 5, characterized in that, The parallel execution of multiple subtasks, each processing elements within its corresponding interval, and dynamically shrinking the interval based on the element processing status, and updating the total cost of the dynamically shrunken interval, includes: Multiple subtasks are launched in parallel in independent threads. Each subtask holds an independent CostAwareSpliterator instance and locks its own responsible interval. The interval corresponding to the subtask is marked by the subtask number, and the intervals corresponding to each subtask do not overlap and cover the original data set. When any subtask is processing an element, the tryAdvance method is called repeatedly. The atomic operation determines whether the start value of the current interval is less than the end value. If the splitting condition is met, the current element is processed by action.accept(data.get(start)). During the processing, the actual processing time of the current element is recorded and compared in real time with the output value of the custom cost calculation logic. If the deviation exceeds the preset threshold, the dynamic calibration mechanism of the cost calculator is triggered. After the current element is processed, atomically increment the start value by 1 in a thread-safe manner, so that the interval corresponding to the currently unprocessed elements dynamically shrinks from [start, end) to [start + 1, end). At the same time, generate an interval change log to record the interval ranges and timestamps before and after the shrinkage; If a contractive change occurs in the interval, trigger an atomic clearing operation of cachedCost to ensure that any subsequent cost calculation operations cannot read expired cost data. Among them, any subsequent cost calculation operations at least include: further splitting within a subtask or cross-task load monitoring; When a subtask detects an interval shrinkage, if the current remaining number of elements is greater than the preset splitting threshold, automatically trigger a cost recalculation process: when calling the estimatePartialCost method, first verify the boundary validity of the current interval [start + 1, end) to ensure that start + 1 < end and the element index does not exceed the boundary; then call the calibrated cost calculator logic element by element based on the latest start value, and accumulate to obtain the new total cost of the interval; Write the recalculated total cost into the totalCost variable through a double-buffer mechanism: first write it into the temporary buffer and perform verification to ensure that the total cost of the interval is non-negative and has a reasonable proportional relationship with the number of elements; after passing the verification, atomically replace the original totalCost value, and at the same time update the cost calculation timestamp to provide a reference basis for the subsequent splitting strategy in the time dimension; During the parallel execution of multiple subtasks, the interval shrinkage situations and the latest total costs of each subtask are synchronized in real time through the interval status table in the shared memory, for the system-level load balancer to perform global monitoring. When it is detected that the remaining total cost of any subtask deviates from that of other subtasks by more than the set threshold, trigger cross-subtask dynamic load scheduling to ensure the overall parallel efficiency.

7. The parallel computing performance optimization method according to claim 1, characterized in that, The splitting conditions at least include one of the following: the number of elements in the interval is greater than 1, the total cost of the interval is greater than 0; When the splitting conditions cannot be met, automatically recycle all resources, including: Return null using the trySplit method to terminate the dynamic splitting process of the data interval in the data list.

8. A parallel computing performance optimization device, characterized in that, The device executes the parallel computing performance optimization method according to any one of claims 1 to 7. The device includes the following units, where, An initialization unit, configured to initialize a CostAwareSpliterator, receive a data list and a custom cost calculator, determine an initial processing interval, and associate the custom cost calculator; A cost unit, configured to obtain the element processing cost based on the custom cost calculator within the initial processing interval, and dynamically obtain the interval processing cost through incremental calculation and caching mechanism to avoid full traversal; The dynamic splitting unit is configured to dynamically split tasks based on the total cost of the interval, achieving subtask load balancing and obtaining multiple subtasks that support parallel execution. Multiple subtasks are executed in parallel, processing elements within their respective intervals, and the intervals are dynamically shrunk based on the element processing status, updating the total cost of the dynamically shrunk interval. Based on the dynamically updated total cost of the interval, the unit jumps to the step of dynamically splitting tasks based on the total cost of the interval, achieving subtask load balancing, and obtaining multiple subtasks that support parallel execution. This process of dynamically splitting and executing subtasks in parallel is repeated until the splitting conditions can no longer be met, at which point all resources are automatically reclaimed.

9. An electronic device, characterized in that, The electronic device includes: at least one processor, a memory, and an input / output unit; The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the parallel computing performance optimization method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the parallel computing performance optimization method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Financial data batch processing method and device and medium

    CN118096408A

  • Intelligent data splitting and optimizing method and system based on multiple dimensions

    CN120045613A