Data parallel processing method and device, computer equipment and readable storage medium
In the data parallel processing method, operators and aggregate functions are obtained using the programming interface, array matching relationship tables are constructed and split, and sub-relational arrays are processed in parallel, the problem of low computing efficiency in the existing technology is solved and efficient data processing is achieved.
Patent Information
- Application Number
- CN202510328531.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology has low computing efficiency and is difficult to efficiently process big data while ensuring data security.
Through a data parallel processing method, the operators and aggregation functions entered by the preset programming interface are obtained by using the preset programming interface, the array matching relationship table is constructed, and the table is split based on the preset array range and the number of target subtasks. Finally, the sub-relational array is processed in parallel to determine the target aggregation result of the private input data.
Through the sub-task division strategy, this method improves the parallel processing capability of computer equipment and significantly improves the computing efficiency.
Smart Images

Figure CN120145453A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of secure computing technology, and particularly to a data parallel processing method, apparatus, computer device, and readable storage medium. Background Art
[0002] In the current rapid development of information technology, big data has become an important driving force for promoting innovation in various industries. However, while the wide application of data brings convenience, it is also accompanied by severe privacy and security challenges. Therefore, how to perform efficient computing while ensuring data security has become an urgent problem to be solved.
[0003] In the prior art, to solve these problems, security operations are mainly performed on private input data based on Secret Sharing (SS). Secret Sharing requires splitting data into multiple shares and performing calculations and communications among multiple participants.
[0004] However, the above method has the problem of low computing efficiency. Summary of the Invention
[0005] Based on this, it is necessary to provide a data parallel processing method, apparatus, computer device, and readable storage medium that can improve computing efficiency for the above technical problems.
[0006] In a first aspect, this application provides a data parallel processing method, including:
[0007] Obtaining an operator and an aggregation function input by a user based on a pre-set programming interface;
[0008] Performing construction processing on a first input array and a second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0009] Performing splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0010] Based on a pre-configured target vector length and the aggregation function, processing at least two sub-relationship arrays in parallel to determine a target aggregation result corresponding to the private input data.
[0011] In one embodiment, the determining process of the above target vector length includes:
[0012] Obtaining performance metric parameters input by a user based on the programming interface; the performance metric parameters at least include the maximum number of concurrent subtasks and the memory constraint amount;
[0013] Determine the maximum vector length based on the memory constraint amount;
[0014] Split the array matching relationship table based on the array range and the maximum number of concurrent subtasks to obtain at least two maximum sub-relationship arrays;
[0015] Aggregate at least two maximum sub-relationship arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table;
[0016] Determine the target vector length according to the task type.
[0017] In one embodiment, the above-mentioned aggregating at least two maximum sub-relationship arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table includes:
[0018] Perform a single aggregation process on at least two maximum sub-relationship arrays using the maximum vector length and the aggregation function, and obtain the bandwidth utilization rate during the aggregation process;
[0019] Determine the task type based on the bandwidth utilization rate.
[0020] In one embodiment, determine the target duration when the above-mentioned determined bandwidth utilization rate exceeds the preset utilization rate threshold;
[0021] When the target duration is greater than or equal to the preset duration, determine the task type as a communication bottleneck task;
[0022] When the target duration is less than the preset duration, determine the task type as a computing bottleneck task.
[0023] In one embodiment, when the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0024] In one embodiment, when the task type is a computing bottleneck task, expand the pre-set minimum vector length by a preset multiple to determine multiple intermediate vector lengths;
[0025] Respectively determine the local computing time corresponding to each intermediate vector length;
[0026] Determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0027] In one embodiment, determine the task establishment time corresponding to the target aggregation result according to the pre-set first relational expression; the first relational expression is used to represent the corresponding relationship between the task establishment time, the shortest local computing time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length;
[0028] Determine the number of target subtasks according to a pre-set second relational expression; the second relational expression is used to represent the corresponding relationship between the number of target subtasks, the task establishment time, and the hardware parameters.
[0029] In a second aspect, the present application further provides a data parallel processing device, including:
[0030] An acquisition module, configured to acquire an operator and an aggregation function input by a user based on a pre-set programming interface;
[0031] A construction module, configured to perform construction processing on a first input array and a second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0032] A splitting module, configured to split the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0033] A processing module, configured to process at least two sub-relationship arrays in parallel based on a pre-configured target vector length and an aggregation function to determine a target aggregation result corresponding to the private input data.
[0034] In a third aspect, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0035] Acquire an operator and an aggregation function input by a user based on a pre-set programming interface;
[0036] Perform construction processing on a first input array and a second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0037] Split the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0038] Process at least two sub-relationship arrays in parallel based on a pre-configured target vector length and an aggregation function to determine a target aggregation result corresponding to the private input data.
[0039] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0040] Acquire an operator and an aggregation function input by a user based on a pre-set programming interface;
[0041] Perform construction processing on the first input array and the second input array based on an operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0042] Perform splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0043] Based on a pre-configured target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine the target aggregation result corresponding to the private input data.
[0044] In a fifth aspect, the present application further provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0045] Obtain the operator and the aggregation function input by the user based on a pre-set programming interface;
[0046] Perform construction processing on the first input array and the second input array based on an operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0047] Perform splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0048] Based on a pre-configured target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine the target aggregation result corresponding to the private input data.
[0049] The above data parallel processing method, device, computer device and readable storage medium obtain the operator and the aggregation function input by the user based on a pre-set programming interface; perform construction processing on the first input array and the second input array based on an operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance; perform splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays; based on a pre-configured target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine the target aggregation result corresponding to the private input data. In this method, a subtask division strategy is adopted to split the array matching relationship table into at least two sub-relationship arrays, so that the computing tasks can be executed simultaneously by multiple processing units, reducing the processing time of a single computing task. This task splitting method improves the parallel processing ability of the computer device and thus improves the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0051] Figure 1 Internal structure diagram of a computer device in one embodiment;
[0052] Figure 2 Flow schematic diagram of a data parallel processing method in one embodiment;
[0053] Figure 3 Flow schematic diagram of a data parallel processing method in another embodiment;
[0054] Figure 4 Flow schematic diagram of a data parallel processing method in another embodiment;
[0055] Figure 5 Flow schematic diagram of a data parallel processing method in another embodiment;
[0056] Figure 6 Flow schematic diagram of a data parallel processing method in another embodiment;
[0057] Figure 7 Flow schematic diagram of a data parallel processing method in another embodiment;
[0058] Figure 8 Structure block diagram of a data parallel processing device in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 1As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data during the data parallel processing process. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a data parallel processing method.
[0061] Those skilled in the art can understand that Figure 1 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0062] In an exemplary embodiment, as Figure 2 shown, a data parallel processing method is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps 201 to 204. Among them:
[0063] Step 201, obtain the operator and aggregation function input by the user based on a pre-set programming interface.
[0064] Among them, the programming interface is a set of pre-designed rules and protocols in the computer device, which provides a way for external users or programs to interact with the computer device. Through this programming interface, users can pass the operators and aggregation functions defined by themselves to the computer device, so that the computer device can perform data processing according to the specific needs of users. The programming interface can have various forms. For example, a command-line interface, where the user enters the code or relevant parameters of the function in the command line; a graphical interface input box, where the user fills in the function information in the input box through a visual interface.
[0065] The operator is a key factor in constructing the array matching relationship table, which determines how two array elements are operated and associated. Users can customize operators according to specific business requirements to meet different calculation requirements. The operator can be addition, multiplication, difference, etc.
[0066] An aggregation function is a function that performs summary calculations on a set of data. The aggregation function is used to perform aggregation operations on the elements in the array matching relationship table to obtain the final target aggregation result. The aggregation function can be sum, average, maximum, etc.
[0067] In the embodiments of the present application, the computer device obtains the operator and the aggregation function input by the user through a pre-set programming interface. The user selects an appropriate operator and aggregation function according to specific requirements. These input items can be set manually by the user or automatically selected through a preset default configuration, aiming to provide operation instructions for subsequent data processing.
[0068] Step 202, perform construction processing on the first input array and the second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance.
[0069] The array matching relationship table is a two-dimensional table obtained by performing arithmetic operations on the elements in the first input array and the second input array through the operator. The rows of the table correspond to the elements (or sub-arrays) in the first input array, the columns correspond to the elements in the second input array, and each cell stores the arithmetic result of the corresponding element combination. The array matching relationship table shows the association and calculation results between the elements of the two input arrays, providing a data basis for subsequent aggregation processing.
[0070] The protected private input data refers to data containing sensitive information, such as users' personal information, business secrets, medical records, etc. These data need to be strictly protected to prevent leakage and abuse.
[0071] In the embodiments of the present application, the computer device performs abstract representation on the protected private input data. The computer device converts it into the first input array and the second input array according to the attributes or characteristics of the data in the first input array and the second input array.
[0072] After the data abstraction is completed, the computer device processes the first input array and the second input array based on the operator provided by the user to construct an array matching relationship table. Specifically, according to the obtained operator, the computer device will perform corresponding mathematical operations, such as addition, subtraction, multiplication or other custom operations, to combine or transform the elements in the two arrays, and finally generate a matching relationship table.
[0073] It should be noted that in the embodiments of the present application, both the first input array and the second input array are arrays related to medical data, that is, the first input array and the second input array are the first medical input array and the second medical input array respectively.
[0074] Step 203: Based on a preset array range and a preconfigured number of target subtasks, split the array matching relation table to obtain at least two sub-relation arrays.
[0075] In the embodiment of the present application, after obtaining the array matching relation table, the computer device splits the matching relation table according to a preset array range and a preconfigured number of target subtasks to form at least two sub-relation arrays. During the splitting process, the computer device segments the matching relation table according to a preset partitioning strategy, enabling each sub-relation array to be calculated independently. The splitting process involves reconstructing data indexes to ensure that the split sub-relation arrays can still maintain the matching characteristics of the original data and meet the requirements of subsequent calculations.
[0076] Step 204: Based on a preconfigured target vector length and an aggregation function, process at least two sub-relation arrays in parallel to determine the target aggregation result corresponding to the private input data.
[0077] Among them, the target aggregation result is the final result obtained by aggregating the elements in the array matching relation table through the aggregation function. It is the product of analyzing and summarizing the private input data, reflecting a certain comprehensive relationship between the two input arrays.
[0078] In the embodiment of the present application, after completing the array splitting, the computer device processes multiple sub-relation arrays in parallel based on a preconfigured target vector length and the obtained aggregation function. The computer device batches the data of the sub-relation arrays and loads them into the computing unit according to the set target vector length, and performs operations according to the requirements of the aggregation function. During the parallel computing process, the computer device schedules computing resources for the computing tasks of different sub-relation arrays and processes the data according to the vectorized computing method to accelerate the computing efficiency and ensure the consistency of data calculation.
[0079] After completing the parallel computing of the sub-relation arrays, the computer device performs a reduction process on the calculation results of each sub-relation array to determine the final target aggregation result. During the reduction process, the computer device sequentially combines the calculation outputs of each sub-relation array according to the definition of the aggregation function to finally form a complete calculation result. After the reduction is completed, the computer device outputs the target aggregation result.
[0080] In the above data parallel processing method, the operator and aggregation function input by the user are obtained based on a preset programming interface; the first input array and the second input array are constructed and processed based on the operator to obtain an array matching relationship table; wherein the first input array and the second input array are determined by abstractly representing the protected private input data in advance; the array matching relationship table is split and processed based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relation arrays; based on the pre-configured target vector length and aggregation function, at least two sub-relation arrays are processed in parallel to determine the target aggregation result corresponding to the private input data. In this method, a sub-task division strategy is adopted to split the array matching relationship table into at least two sub-relation arrays, so that the computing task can be executed simultaneously by multiple processing units, reducing the processing time of a single computing task. This task splitting method improves the parallel processing capability of the computer device, thereby improving the computing efficiency.
[0081] In an exemplary embodiment, Figure 3 As shown, the process of determining the target vector length includes:
[0082] Step 301: Obtain performance measurement parameters input by a user based on a programming interface. The performance measurement parameters at least include a maximum concurrent subtask quantity and a memory constraint quantity.
[0083] In the embodiment of the present application, the process of determining the target vector length first obtains the performance measurement parameters input by the user by the computer device. The process is based on a pre-set programming interface, allowing the user to input key parameters related to computing performance, including at least the maximum number of concurrent subtasks and memory constraints. After receiving these parameters, the computer device will store them and use them as an important basis for the execution of the computing task to ensure that the subsequent steps can be completed within the established computing resource range.
[0084] Step 302: determine the maximum vector length based on the memory constraint.
[0085] In an embodiment of the present application, after obtaining the performance measurement parameters, the computer device needs to determine the maximum vector length based on the memory constraint. This process involves analyzing factors such as the total capacity of the system memory, the current available memory space, the data storage structure, and the complexity of the computing task. The computer device determines the maximum vector length that can be supported in the current computing environment by calculation or table lookup to ensure that the computing task does not exceed the memory resource limit during execution, while improving the data processing capability as much as possible. After determining the maximum vector length, the device will store it and use it as a basis for data segmentation and processing in subsequent computing processes.
[0086] Step 303: Split the array matching relationship table based on the array range and the maximum number of concurrent subtasks to obtain at least two maximum sub-relationship arrays.
[0087] In the embodiment of the present application, the computer device splits the array matching relationship table based on the array range and the maximum number of concurrent subtasks, thereby obtaining at least two maximum sub-relationship arrays. To achieve efficient data parallel computing, the device evenly splits the original array matching relationship table according to the size of the maximum number of concurrent subtasks, so that the scale of each sub-relationship array is within the range allowed by the computing resources, and at the same time ensures that the computing tasks can be evenly distributed to multiple subtasks for processing. During the splitting process, the device will select an appropriate splitting strategy according to different task characteristics and data structures, such as splitting by rows, splitting by columns, or performing block processing based on specific rules, to improve the balance of data distribution and reduce the uneven computing load between different tasks.
[0088] Step 304: Aggregate at least two maximum sub-relationship arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table.
[0089] In the embodiment of the present application, after the splitting is completed, the computer device aggregates at least two of the maximum sub-relationship arrays based on the maximum vector length and the aggregation function. First, the device slices each sub-relationship array according to the maximum vector length to ensure that the size of the data block meets the requirements of the maximum vector length, so as to be able to efficiently perform vectorized computing in the subsequent computing process. Subsequently, for each slice, the computer device calculates the elements in the slice based on the aggregation function and generates a preliminary aggregation result corresponding to the sub-relationship array. In this process, the computer will call appropriate mathematical operations or statistical methods, such as summation, mean calculation, maximum value calculation, etc., to meet the aggregation requirements specified by the user.
[0090] Step 305: Determine the target vector length according to the task type.
[0091] In the embodiment of the present application, after the preliminary aggregation process is completed, the computer device needs to further determine the target vector length based on the task type corresponding to the array matching relationship table. Specifically, the device will select a suitable target vector length by combining the computing characteristics of the task, the data structure, and the previously determined maximum vector length, so that the computing process can not only make full use of the available resources, but also ensure the stability and efficiency of data processing. In this process, the computer device may adopt a dynamic adjustment strategy to optimize the target vector length in real time according to the execution situation of the current task and the resource occupancy situation to adapt to the changes in the computing environment.
[0092] Finally, after determining the target vector length, the computer device stores it and applies it to subsequent computing tasks. At this time, the execution strategy of the computing task is adjusted according to the target vector length, so that each stage of data processing can be carried out within the optimized vector length range, maximizing the computing efficiency while ensuring the reasonable utilization of computing resources.
[0093] In an exemplary embodiment, as Figure 4 shown, the above "aggregating at least two maximum sub-relation arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relation table" includes:
[0094] Step 401, performing a first aggregation process on at least two maximum sub-relation arrays using the maximum vector length and the aggregation function, and obtaining the bandwidth utilization rate during the aggregation process.
[0095] In the embodiment of the present application, the computer device first performs a first aggregation process on at least two of the maximum sub-relation arrays using the maximum vector length and the aggregation function. During this process, the computer device slices each sub-relation array according to the maximum vector length and calculates the elements in each vector slice based on the aggregation function to obtain the corresponding aggregation result. When specifically executing, the computer will select an appropriate computing mode, such as parallel computing, to ensure the efficient execution of the computing task and minimize the waste of computing resources.
[0096] During the execution of the aggregation process, the computer device monitors the usage of the bandwidth in real time and obtains the bandwidth utilization rate during the aggregation process. The calculation of the bandwidth utilization rate usually includes factors such as the data transfer rate, the total amount of data during the execution of the computing task, and the memory access pattern. The computer device measures the data transfer rate between the computing unit and the memory, as well as the data interaction between different computing modules, through system-level performance monitoring tools or based on internal data flow tracing methods. The computer device records the bandwidth utilization rate and evaluates the overall efficiency of the aggregation process in combination with possible memory bottlenecks or data contention situations during the task execution process.
[0097] Step 402, determining the task type based on the bandwidth utilization rate.
[0098] In an embodiment of the present application, after obtaining the bandwidth utilization rate, the computer device needs to determine the task type based on the bandwidth utilization rate. First, the device will compare the currently measured bandwidth utilization rate with a preset threshold range to determine whether the computing task is a high-bandwidth occupancy task, a compute-intensive task, or an I / O-bound task. If the bandwidth utilization rate is close to the maximum available bandwidth, it indicates that the task may be a data-intensive task, mainly limited by the data transmission speed; if the bandwidth utilization rate is low while the computing resource occupancy is high, it is a compute-intensive task, mainly limited by the computing power demand.
[0099] After determining the task type, the computer device may further analyze the specific mode of task execution, such as the concurrency of the task, the locality of data access, and the complexity of the computing operation. The device may combine historical operation data or, based on existing task classification criteria, make further optimization decisions to adopt more suitable scheduling and resource allocation strategies during the subsequent execution of computing tasks. In addition, the computer device stores the task type information so that the determined computing mode can be quickly reused when similar tasks are executed in the future, improving the computing efficiency.
[0100] In an exemplary embodiment, the task types include communication bottleneck tasks and compute bottleneck tasks. On this basis, as Figure 5 shown, the above “determine the task type based on the bandwidth utilization rate” includes:
[0101] Step 501, determine the target duration for which the bandwidth utilization rate exceeds the preset utilization threshold.
[0102] Among them, the preset utilization threshold can be 90%.
[0103] In an embodiment of the present application, the computer device first monitors the bandwidth utilization rate during the aggregation process in real time and records the change of the bandwidth utilization rate over time. The computer device will statistically analyze the bandwidth occupancy in each time window based on the sampling period and form a continuous data stream of bandwidth utilization rates. For each sampling point, the computer device will compare its current bandwidth utilization rate with the preset utilization threshold to determine whether it exceeds the threshold.
[0104] In the case where it is determined that the bandwidth utilization rate exceeds the preset utilization threshold, the computer device needs to calculate the duration of this state, that is, the target duration. The calculation methods of the target duration usually include the continuous time window statistical method, that is, statistically analyzing the time period during which the bandwidth utilization rate continuously exceeds the threshold, or the sliding window method to evaluate the fluctuation of the bandwidth utilization rate within a certain time range.
[0105] Step 502, in the case where the target duration is greater than or equal to the preset duration, determine that the task type is a communication bottleneck task.
[0106] In an embodiment of the present application, after calculating the target duration, the computer device needs to compare it with a preset duration. If the target duration is greater than or equal to the preset duration, it indicates that the computer device has been in a high bandwidth occupancy state for a long time, suggesting that the performance of the computing task is limited by the data transmission capacity rather than the computing capacity. Therefore, the computer device determines that the task type is a communication bottleneck task, that is, the execution efficiency of this task is mainly restricted by the data transmission rate, and the optimization direction should focus on methods such as data prefetching, compression, and communication strategy adjustment.
[0107] Step 503, in the case where the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0108] In an embodiment of the present application, if the target duration is less than the preset duration, it indicates that the state of high bandwidth utilization only lasts for a short time, and the computing resources of the computer device may still be the main performance bottleneck. In this case, the computer device determines that the task type is a computing bottleneck task, that is, the execution of this task is mainly restricted by the computing capacity, and the optimization direction should focus on strategies such as algorithm optimization, parallel computing acceleration, and the use of hardware accelerators.
[0109] After the task type is determined, the computer device stores the task type information in the task scheduling module or the task management system, so that when subsequent similar tasks are executed, the optimal computing resource configuration scheme can be selected based on the characteristics of the task type to improve the overall computing efficiency. In addition, the device may input the task type information as feedback to the computing optimization module to further adjust parameters such as bandwidth allocation and data scheduling methods to meet the requirements of different types of tasks.
[0110] In an exemplary embodiment, determining the target vector length according to the task type includes:
[0111] In the case where the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0112] In an embodiment of the present application, the computer device first determines that the current task belongs to a communication bottleneck task based on the determination result of the task type. This determination usually comes from the analysis process of bandwidth utilization, that is, during the aggregation process, if the bandwidth utilization continuously exceeds the preset threshold and the exceeded time reaches or exceeds the preset duration, this task will be classified as a communication bottleneck task.
[0113] After the task type is determined, the computer device needs to evaluate whether the current maximum vector length is suitable for communication bottleneck tasks. The computer device analyzes the impact of different vector lengths on data transmission efficiency based on historical execution data, bandwidth utilization trends, and computing resource utilization. If the current maximum vector length leads to excessive bandwidth utilization and affects task execution efficiency, adjustments are required.
[0114] Next, the computer device will adopt a vector length optimization strategy to adjust the vector length. In the case of communication bottleneck tasks, the usual optimization direction is to reduce the vector length to reduce the load of a single data transmission, thereby alleviating the bandwidth bottleneck. The computer device will calculate the appropriate target vector length by combining bandwidth occupancy analysis, data throughput measurement, and task execution time evaluation.
[0115] Specifically, the computer device can use empirical formulas, historical data modeling, or dynamic adjustment methods to iteratively find the optimal target vector length within the set vector length range. Usually, the computer device will first reduce the vector length to an initial adjustment value and monitor the change in bandwidth utilization. If the bandwidth utilization drops to a reasonable level after the reduction and the computing efficiency is improved, then this vector length can be used as the final target vector length. If there is still a communication bottleneck after the adjustment, further optimization and adjustment are carried out until the optimal solution is found.
[0116] Finally, the computer device stores the calculated target vector length and uses it for subsequent calculation steps of the current task. At the same time, the computer device updates the vector length optimization strategy so that it can quickly determine the appropriate vector length when performing similar tasks in the future, thereby improving the overall computing efficiency.
[0117] In an exemplary embodiment, as Figure 6 shown, the above "determining the target vector length according to the task type" includes:
[0118] Step 601, in the case where the task type is a computing bottleneck task, expand the pre-set minimum vector length by a preset multiple to determine multiple intermediate vector lengths.
[0119] In the embodiments of the present application, after the computer device determines that the task type is a computing bottleneck task, it first needs to obtain the pre-set minimum vector length and expand it based on the preset multiple to generate multiple intermediate vector lengths. The selection of the preset multiple is usually determined by system configuration or historical calculation data to ensure that the processing capacity is fully utilized within the range allowed by the computing resources and improve the computing efficiency.
[0120] Step 602, respectively determine the local computing time corresponding to each intermediate vector length.
[0121] In an embodiment of the present application, after obtaining multiple intermediate vector lengths, the computer device sequentially uses these vector lengths for calculation processing and measures the local calculation time corresponding to each vector length. The measurement method of the local calculation time can be based on actual execution time statistics, calculation resource occupancy evaluation, or performance analysis tools for measurement. The computer device ensures that the calculation tasks of each intermediate vector length are carried out under the same conditions to ensure the accuracy and comparability of the measurement results.
[0122] Step 603: Determine the intermediate vector length corresponding to the shortest local calculation time as the target vector length.
[0123] In an embodiment of the present application, next, the computer device compares the measured local calculation times to determine the vector length with the shortest duration. The calculation device usually adopts a minimum value search algorithm to find the optimal solution among the calculation times corresponding to all intermediate vector lengths. If there are multiple calculation times corresponding to the same vector length, factors such as bandwidth utilization rate and memory occupancy can be further combined for refined selection to ensure that the finally determined vector length can reduce the impact of calculation bottlenecks without introducing new performance limitations.
[0124] Finally, the computer device determines the optimal intermediate vector length obtained by calculation as the target vector length and stores the result for subsequent steps of the current calculation task.
[0125] In an exemplary embodiment, as Figure 7 shown, the above "configuration method for the number of target subtasks" includes:
[0126] Step 701: Determine the task establishment time corresponding to the target aggregation result according to a preset first relationship. The first relationship is used to characterize the corresponding relationship between the task establishment time, the shortest local calculation time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length.
[0127] In an embodiment of the present application, the computer device first determines the task establishment time based on a preset first relationship. For this purpose, the device needs to obtain relevant calculation parameters, including the shortest local calculation time, the array length of the first input array, the array length of the second input array, and the target vector length. Then, these parameters are input into the first relationship for calculation to determine the task establishment time of the current task. The first relationship is used to describe the mapping relationship between the task establishment time and the above parameters to ensure that the computer device can accurately evaluate the establishment overhead of the task.
[0128] Step 702: Determine the number of target subtasks according to a preset second relational expression. The second relational expression is used to represent the corresponding relationship between the number of target subtasks, the task establishment time, and the hardware parameters.
[0129] In an embodiment of the present application, after determining the task establishment time, the computer device calculates the number of target subtasks based on a preset second relational expression. At this time, the computer device needs to obtain hardware parameters related to the task establishment time, such as the number of computing cores, memory bandwidth, and scheduling overhead. Then, the task establishment time and the hardware parameters are input into the second relational expression to calculate the corresponding number of target subtasks. The second relational expression is used to describe the relationship between the number of target subtasks, the task establishment time, and the hardware performance, so as to ensure that the determination of the number of subtasks can achieve optimal task scheduling under the constraint of computing resources.
[0130] Finally, the computer device uses the calculated number of target subtasks for subsequent task scheduling and execution, and may store this result to optimize the subtask partitioning strategy for future tasks, improving computing efficiency and resource utilization.
[0131] In some embodiments, the preset first relational expression may be as shown in formula (1):
[0132]
[0133] In formula (1), t 1 represents the task establishment time, represents the shortest local computing time, n represents the length of the first input array, m represents the length of the second input array, and B 1 represents the target vector length.
[0134] In some embodiments, the preset second relational expression may be as shown in formula (2):
[0135]
[0136] In formula (2), c represents the number of target subtasks, t 1 represents the task establishment time, and a represents the hardware parameter.
[0137] In an exemplary embodiment, the above method further includes:
[0138] Step 1: Obtain performance metric parameters input by the user based on a programming interface. The performance metric parameters at least include the maximum concurrent subtask volume and the memory constraint volume.
[0139] Step 2: Determine the maximum vector length based on the memory constraint volume.
[0140] Step 3, split the array matching relationship table based on the array range and the maximum number of concurrent subtasks to obtain at least two maximum sub-relationship arrays.
[0141] Step 4, perform an aggregation process on at least two maximum sub-relationship arrays using the maximum vector length and an aggregation function, and obtain the bandwidth utilization rate during the aggregation process.
[0142] Step 5, determine the target duration when the bandwidth utilization rate exceeds the preset utilization rate threshold.
[0143] Step 6, when the target duration is greater than or equal to the preset duration, determine that the task type is a communication bottleneck task. When the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0144] Step 7, when the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0145] Step 8, when the task type is a computing bottleneck task, expand the preset minimum vector length by a preset multiple to determine multiple intermediate vector lengths.
[0146] Step 9, respectively determine the local computing time corresponding to each intermediate vector length.
[0147] Step 10, determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0148] Step 11, determine the task establishment time corresponding to the target aggregation result according to the preset first relationship. The first relationship is used to represent the corresponding relationship between the task establishment time, the shortest local computing time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length.
[0149] Step 12, determine the target number of subtasks according to the preset second relationship. The second relationship is used to represent the corresponding relationship between the target number of subtasks, the task establishment time, and the hardware parameters.
[0150] Step 13, obtain the operator and aggregation function input by the user based on the preset programming interface.
[0151] Step 14, perform a construction process on the first input array and the second input array based on the operator to obtain an array matching relationship table. Among them, the first input array and the second input array are determined by abstractly representing the protected private input data in advance.
[0152] Step 15, split the array matching relationship table based on the preset array range and the preset target number of subtasks to obtain at least two sub-relationship arrays.
[0153] Step 16: Based on a pre-configured target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine a target aggregation result corresponding to the private input data.
[0154] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0155] Based on the same inventive concept, an embodiment of the present application further provides a data parallel processing device for implementing the data parallel processing method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data parallel processing device provided below can refer to the limitations on the data parallel processing method in the above text, and will not be repeated here.
[0156] In an exemplary embodiment, as Figure 8 shown, a data parallel processing device is provided, including: an acquisition module 801, a construction module 802, a splitting module 803, and a processing module 804, where:
[0157] The acquisition module 801 is configured to acquire an operator and an aggregation function input by a user based on a pre-set programming interface;
[0158] The construction module 802 is configured to perform construction processing on a first input array and a second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0159] The splitting module 803 is configured to perform splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0160] The processing module 804 is configured to process at least two sub-relationship arrays in parallel based on a pre-configured target vector length and an aggregation function to determine a target aggregation result corresponding to the private input data.
[0161] In an exemplary embodiment, the above-mentioned device may further include:
[0162] An input module, configured to obtain performance metric parameters input by a user based on a programming interface; the performance metric parameters at least include the maximum number of concurrent subtasks and the memory constraint amount;
[0163] A first determination module, configured to determine the maximum vector length based on the memory constraint amount;
[0164] A second determination module, configured to perform a splitting process on an array matching relationship table based on an array range and the maximum number of concurrent subtasks to obtain at least two maximum sub-relationship arrays;
[0165] An aggregation module, configured to perform an aggregation process on at least two maximum sub-relationship arrays based on the maximum vector length and an aggregation function to determine the task type corresponding to the array matching relationship table;
[0166] A third determination module, configured to determine a target vector length according to the task type.
[0167] In an exemplary embodiment, the above-mentioned aggregation module is specifically configured to perform an aggregation process on at least two maximum sub-relationship arrays by using the maximum vector length and an aggregation function, and obtain the bandwidth utilization rate during the aggregation process;
[0168] Determine the task type based on the bandwidth utilization rate.
[0169] In an exemplary embodiment, the above-mentioned aggregation module is specifically configured to determine a target duration during which the bandwidth utilization rate exceeds a preset utilization threshold;
[0170] In the case where the target duration is greater than or equal to a preset duration, determine that the task type is a communication bottleneck task;
[0171] In the case where the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0172] In an exemplary embodiment, the above-mentioned third determination module is specifically configured to, in the case where the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0173] In an exemplary embodiment, the above-mentioned third determination module is specifically configured to, in the case where the task type is a computing bottleneck task, perform an expansion process on a preset minimum vector length at a preset magnification to determine a plurality of intermediate vector lengths;
[0174] Respectively determine the local computing time corresponding to each intermediate vector length;
[0175] Determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0176] In an exemplary embodiment, the above device may further include:
[0177] A fourth determination module, configured to determine a task establishment time corresponding to a target aggregation result according to a preset first relational expression; the first relational expression is used to characterize the corresponding relationship between the task establishment time, the local calculation time with the shortest duration, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length;
[0178] A fifth determination module, configured to determine a target subtask quantity according to a preset second relational expression; the second relational expression is used to characterize the corresponding relationship between the target subtask quantity, the task establishment time, and the hardware parameters.
[0179] Each module in the above data parallel processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0180] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0181] Obtain an operator and an aggregation function input by a user based on a preset programming interface;
[0182] Perform construction processing on a first input array and a second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly characterizing the protected private input data in advance;
[0183] Perform splitting processing on the array matching relationship table based on a preset array range and a preset target subtask quantity to obtain at least two sub-relationship arrays;
[0184] Based on a preset target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine a target aggregation result corresponding to the private input data.
[0185] In an embodiment, when the processor executes the computer program, the following steps are further implemented:
[0186] Obtain performance metric parameters input by a user based on the programming interface; the performance metric parameters at least include the maximum concurrent subtask quantity and the memory constraint quantity;
[0187] Determine a maximum vector length based on the memory constraint quantity;
[0188] Split the array matching relationship table based on the array range and the maximum number of concurrent subtasks to obtain at least two maximum sub-relationship arrays;
[0189] Based on the maximum vector length and the aggregation function, perform an aggregation process on at least two maximum sub-relationship arrays to determine the task type corresponding to the array matching relationship table;
[0190] Determine the target vector length according to the task type.
[0191] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0192] Perform an aggregation process on at least two maximum sub-relationship arrays using the maximum vector length and the aggregation function, and obtain the bandwidth utilization rate during the aggregation process;
[0193] Determine the task type based on the bandwidth utilization rate.
[0194] In one embodiment, the task type includes communication bottleneck tasks and computing bottleneck tasks. When the processor executes the computer program, the following steps are further implemented:
[0195] Determine the target duration when the bandwidth utilization rate exceeds the preset utilization rate threshold;
[0196] When the target duration is greater than or equal to the preset duration, determine that the task type is a communication bottleneck task;
[0197] When the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0198] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0199] When the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0200] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0201] When the task type is a computing bottleneck task, expand the pre-set minimum vector length by a preset multiple to determine multiple intermediate vector lengths;
[0202] Respectively determine the local computing time corresponding to each intermediate vector length;
[0203] Determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0204] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0205] Determine the task establishment time corresponding to the target aggregation result according to a preset first relational expression; the first relational expression is used to characterize the corresponding relationship between the task establishment time, the local calculation time with the shortest duration, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length.
[0206] Determine the number of target subtasks according to a preset second relational expression; the second relational expression is used to characterize the corresponding relationship between the number of target subtasks, the task establishment time, and the hardware parameters.
[0207] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0208] Obtain the operator and aggregation function input by the user based on a preset programming interface;
[0209] Perform construction processing on the first input array and the second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly characterizing the protected private input data in advance.
[0210] Perform splitting processing on the array matching relationship table based on a preset array range and a preset number of target subtasks to obtain at least two sub-relationship arrays;
[0211] Based on the preset target vector length and the aggregation function, process at least two sub-relationship arrays in parallel to determine the target aggregation result corresponding to the private input data.
[0212] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0213] Obtain the performance metric parameters input by the user based on the programming interface; the performance metric parameters at least include the maximum concurrent subtask volume and the memory constraint volume;
[0214] Determine the maximum vector length based on the memory constraint volume;
[0215] Perform splitting processing on the array matching relationship table based on the array range and the maximum concurrent subtask volume to obtain at least two maximum sub-relationship arrays;
[0216] Perform aggregation processing on at least two maximum sub-relationship arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table;
[0217] Determine the target vector length according to the task type.
[0218] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0219] Perform an aggregation process on at least two maximum sub-relation arrays using the maximum vector length and an aggregation function, and obtain the bandwidth utilization rate during the aggregation process;
[0220] Determine the task type based on the bandwidth utilization rate.
[0221] In one embodiment, the task type includes communication bottleneck tasks and computing bottleneck tasks. When the computer program is executed by a processor, the following steps are further implemented:
[0222] Determine the target duration when the bandwidth utilization rate exceeds a preset utilization rate threshold;
[0223] When the target duration is greater than or equal to the preset duration, determine that the task type is a communication bottleneck task;
[0224] When the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0225] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0226] When the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0227] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0228] When the task type is a computing bottleneck task, perform an expansion process on a preset minimum vector length at a preset magnification factor to determine multiple intermediate vector lengths;
[0229] Respectively determine the local computing time corresponding to each intermediate vector length;
[0230] Determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0231] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0232] According to a preset first relationship formula, determine the task establishment time corresponding to the target aggregation result; the first relationship formula is used to represent the corresponding relationship between the task establishment time, the shortest local computing time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length;
[0233] According to a preset second relationship formula, determine the target sub-task quantity; the second relationship formula is used to represent the corresponding relationship between the target sub-task quantity, the task establishment time, and the hardware parameters.
[0234] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:
[0235] Obtain the operator and aggregation function input by the user based on a pre-set programming interface;
[0236] Perform construction processing on the first input array and the second input array based on the operator to obtain an array matching relationship table; wherein, the first input array and the second input array are determined by abstractly representing the protected private input data in advance;
[0237] Perform splitting processing on the array matching relationship table based on a preset array range and a pre-configured number of target subtasks to obtain at least two sub-relationship arrays;
[0238] Based on a pre-configured target vector length and an aggregation function, process at least two sub-relationship arrays in parallel to determine the target aggregation result corresponding to the private input data.
[0239] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented:
[0240] Obtain the performance metric parameters input by the user based on the programming interface; the performance metric parameters at least include the maximum concurrent subtask amount and the memory constraint amount;
[0241] Determine the maximum vector length based on the memory constraint amount;
[0242] Perform splitting processing on the array matching relationship table based on the array range and the maximum concurrent subtask amount to obtain at least two maximum sub-relationship arrays;
[0243] Perform aggregation processing on at least two maximum sub-relationship arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table;
[0244] Determine the target vector length according to the task type.
[0245] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented:
[0246] Perform an aggregation process on at least two maximum sub-relationship arrays using the maximum vector length and the aggregation function, and obtain the bandwidth utilization rate during the aggregation process;
[0247] Determine the task type based on the bandwidth utilization rate.
[0248] In one embodiment, the task type includes a communication bottleneck task and a computing bottleneck task. When the computer program is executed by the processor, the following steps are also implemented:
[0249] Determine the target duration when the bandwidth utilization rate exceeds a preset utilization rate threshold;
[0250] When the target duration is greater than or equal to the preset duration, determine that the task type is a communication bottleneck task;
[0251] When the target duration is less than the preset duration, determine that the task type is a computing bottleneck task.
[0252] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0253] When the task type is a communication bottleneck task, determine the maximum vector length as the target vector length.
[0254] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0255] When the task type is a computing bottleneck task, perform an expansion process on the preset minimum vector length at a preset magnification to determine multiple intermediate vector lengths;
[0256] Respectively determine the local computing time corresponding to each intermediate vector length;
[0257] Determine the intermediate vector length corresponding to the shortest local computing time as the target vector length.
[0258] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0259] According to a preset first relationship, determine the task establishment time corresponding to the target aggregation result; the first relationship is used to characterize the corresponding relationship between the task establishment time, the shortest local computing time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length;
[0260] According to a preset second relationship, determine the number of target subtasks; the second relationship is used to characterize the corresponding relationship between the number of target subtasks, the task establishment time, and the hardware parameters.
[0261] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0262] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0263] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.
[0264] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A data parallel processing method, characterized in that: The method comprises: Operators and aggregation functions that obtain user input based on pre-set programming interfaces; Based on the operator, the first input array and the second input array are constructed and processed to obtain an array matching relationship table; wherein the first input array and the second input array are determined in advance by abstractly representing the protected private input data; Splitting the array matching relationship table based on a preset array range and a pre-configured target subtask quantity to obtain at least two sub-relation arrays; Based on a preconfigured target vector length and the aggregation function, at least two of the sub-relation arrays are processed in parallel to determine a target aggregation result corresponding to the private input data.
2. The method according to claim 1, characterized in that The process of determining the target vector length includes: Acquiring performance measurement parameters input by a user based on the programming interface; the performance measurement parameters at least include a maximum concurrent subtask amount and a memory constraint amount; Based on the memory constraint, determining a maximum vector length; Splitting the array matching relationship table based on the array range and the maximum concurrent subtask amount to obtain at least two maximum sub-relation arrays; Based on the maximum vector length and the aggregation function, performing aggregation processing on at least two of the maximum sub-relation arrays to determine the task type corresponding to the array matching relationship table; The target vector length is determined according to the task type.
3. The method according to claim 2, characterized in that The step of performing aggregation processing on at least two of the maximum sub-relation arrays based on the maximum vector length and the aggregation function to determine the task type corresponding to the array matching relationship table includes: Performing an aggregation process on at least two of the maximum sub-relation arrays by using the maximum vector length and the aggregation function, and obtaining bandwidth utilization during the aggregation process; The task type is determined based on the bandwidth utilization.
4. The method according to claim 3, characterized in that: The task type includes a communication bottleneck task and a computing bottleneck task, and the determining of the task type based on the bandwidth utilization includes: Determine a target duration for which the bandwidth utilization exceeds a preset utilization threshold; When the target duration is greater than or equal to the preset duration, determining the task type as the communication bottleneck task; When the target duration is less than the preset duration, the task type is determined to be the computing bottleneck task.
5. The method according to claim 4, characterized in that The determining the target vector length according to the task type includes: When the task type is the communication bottleneck task, the maximum vector length is determined as the target vector length.
6. The method according to claim 4, characterized in that The determining the target vector length according to the task type includes: In the case where the task type is the computing bottleneck task, the preset minimum vector length is enlarged by a preset multiple to determine a plurality of intermediate vector lengths; respectively determining the local computing time corresponding to the length of each intermediate vector; The intermediate vector length corresponding to the local calculation time with the shortest duration is determined as the target vector length.
7. The method according to claim 5, characterized in that The method for configuring the target subtask quantity includes: Determine the task creation time corresponding to the target aggregation result according to a preset first relationship; the first relationship is used to characterize the correspondence between the task creation time and the shortest local computing time, the array length corresponding to the first input array, the array length corresponding to the second input array, and the target vector length; The target subtask quantity is determined according to a preset second relational expression; the second relational expression is used to characterize the corresponding relationship between the target subtask quantity and the task establishment time and hardware parameters.
8. A data parallel processing device, characterized in that: The device comprises: An acquisition module, used to acquire operators and aggregation functions input by users based on a preset programming interface; A construction module, used for constructing and processing the first input array and the second input array based on the operator to obtain an array matching relationship table; wherein the first input array and the second input array are determined in advance by abstractly representing the protected private input data; A splitting module, used for splitting the array matching relationship table based on a preset array range and a pre-configured target subtask quantity to obtain at least two sub-relation arrays; A processing module is used to process at least two of the sub-relation arrays in parallel based on a pre-configured target vector length and the aggregation function to determine a target aggregation result corresponding to the private input data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.