Data processing method and device, electronic equipment and storage medium

When the to-processed operator meets the preset conditions, it is split into K subtasks according to its split ratio and allocates them to K cores of the second processor for processing, the problem of low allocation efficiency of computing tasks in the prior art is solved, and the resource utilization rate and computing efficiency of the processor are improved.

CN120179402APending Publication Date: 2025-06-20MOORE THREADS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510315521.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively allocate computing tasks to computing cores of different specifications and combinations, resulting in low processor performance and resource utilization.

Method used

When the to-processed operator meets the preset splitting conditions, it obtains its first split ratio, and divides the operator into K subtasks based on the ratio, allocates it to K cores of the second processor for processing, and finally determines the second result of the operator based on the subtask results.

Benefits of technology

It realizes that the pending operator is split into sub-tasks of different sizes as needed and allocates them to processor cores that match their size, improving the resource utilization and computing efficiency of the second processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179402A_ABST
    Figure CN120179402A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first splitting proportion used for indicating the scale proportion among K cores in a second processor under the condition that a to-be-processed operator meets a preset splitting condition, the operator to be processed can be split into K sub-tasks according to the obtained first splitting proportion, the K sub-tasks are sent to K cores of a second processor to be processed to obtain K first results, and a second result of the operator to be processed is determined according to the K first results from the K cores and the category of the operator to be processed. According to the embodiment of the invention, the to-be-processed operator can be divided into the sub-tasks of different scales according to needs, the sub-tasks are allocated to the processor cores matched with the scales of the sub-tasks, the scales and performance proportions of different processor cores are supported, the tasks are allowed to be reasonably allocated to the asymmetric processor cores, and the utilization efficiency of the processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] An image processor (Graphics Processing Unit, GPU) or other computing devices usually have multiple computing cores. These computing cores may have different core specifications or different combined scales of cores. Therefore, in order to improve the processor performance and the utilization rate of computing resources, it becomes increasingly important how to appropriately allocate computing tasks to different computing cores. Summary of the Invention

[0003] The present disclosure proposes a technical solution for data processing.

[0004] According to an aspect of the present disclosure, there is provided a data processing method, which is applied to a first processor. The method includes: when a to-be-processed operator meets a preset splitting condition, obtaining a first splitting ratio of the to-be-processed operator, where the first splitting ratio is used to indicate a scale ratio among K cores in a second processor, and K is an integer greater than or equal to 2; splitting the to-be-processed operator into K subtasks according to the first splitting ratio; sending the K subtasks to the K cores of the second processor, where the cores are used to process the received subtasks to obtain first results of the subtasks; receiving the K first results from the K cores, and determining a second result of the to-be-processed operator according to the K first results and the category of the to-be-processed operator.

[0005] In a possible implementation manner, the splitting the to-be-processed operator into K subtasks according to the first splitting ratio includes: when the to-be-processed operator conforms to a first splitting manner, splitting the to-be-processed operator into K subtasks according to the first splitting ratio and the first splitting manner, where a splitting dimension of the first splitting manner is a batch dimension of the to-be-processed operator; or when the to-be-processed operator conforms to a second splitting manner, splitting the to-be-processed operator into K subtasks according to the first splitting ratio and the second splitting manner, where a splitting dimension of the second splitting manner is the largest dimension in data dimensions of the to-be-processed operator.

[0006] In a possible implementation, when the operator to be processed conforms to the first splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the first splitting method includes: determining a second splitting ratio and a first splitting non-uniformity according to the batch dimension of the operator to be processed and the first splitting ratio; splitting the batch dimension into K first values according to the second splitting ratio; and when the minimum value among the K first values is greater than a first preset threshold and the first splitting non-uniformity is less than a second preset threshold, splitting the operator to be processed into K subtasks according to the K first values.

[0007] In a possible implementation, when the operator to be processed conforms to the second splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the second splitting method includes: determining a third splitting ratio and a second splitting non-uniformity according to the maximum dimension in the data dimension of the operator to be processed and the first splitting ratio; splitting the maximum dimension in the data dimension into K second values according to the third splitting ratio; and when the minimum value among the K second values is greater than a third preset threshold and the second splitting non-uniformity is less than a fourth preset threshold, splitting the operator to be processed into K subtasks according to the K second values.

[0008] In a possible implementation, the method further includes: when the operator to be processed does not conform to the first splitting method and does not conform to the second splitting method, allocating the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0009] In a possible implementation, determining a second result of the operator to be processed according to the K first results and the category of the operator to be processed includes: when the category of the operator to be processed belongs to the target category, reprocessing the K first results according to the category of the operator to be processed to obtain the second result of the operator to be processed; and when the category of the operator to be processed belongs to other than the target category, merging the K first results to obtain the second result of the operator to be processed.

[0010] In a possible implementation, when the operator to be processed does not meet the splitting conditions, allocating the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0011] According to one aspect of the present disclosure, a data processing device is provided. The device is applied to a first processor and includes: an acquisition module configured to acquire a first splitting ratio of the operator to be processed when the operator to be processed meets a preset splitting condition, where the first splitting ratio is used to indicate the scale ratio among K cores in a second processor, and K is an integer greater than or equal to 2; a splitting module configured to split the operator to be processed into K subtasks according to the first splitting ratio; a sending module configured to send the K subtasks to the K cores of the second processor, where the cores are configured to process the received subtasks to obtain first results of the subtasks; and a receiving module configured to receive the K first results from the K cores and determine a second result of the operator to be processed according to the K first results and the category of the operator to be processed.

[0012] In a possible implementation, the splitting module is configured to: when the operator to be processed conforms to a first splitting mode, split the operator to be processed into K subtasks according to the first splitting ratio and the first splitting mode, where the splitting dimension of the first splitting mode is the batch dimension of the operator to be processed; or, when the operator to be processed conforms to a second splitting mode, split the operator to be processed into K subtasks according to the first splitting ratio and the second splitting mode, where the splitting dimension of the second splitting mode is the maximum dimension in the data dimension of the operator to be processed.

[0013] In a possible implementation, the step of splitting the operator to be processed into K subtasks according to the first splitting ratio and the first splitting mode when the operator to be processed conforms to the first splitting mode includes: determining a second splitting ratio and a first splitting non-uniformity degree according to the batch dimension of the operator to be processed and the first splitting ratio; splitting the batch dimension into K first values according to the second splitting ratio; and when the minimum value among the K first values is greater than a first preset threshold and the first splitting non-uniformity degree is less than a second preset threshold, splitting the operator to be processed into K subtasks according to the K first values.

[0014] In a possible implementation, when the operator to be processed conforms to the second splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the second splitting method includes: determining a third splitting ratio and a second splitting non-uniformity according to the maximum dimension in the data dimension of the operator to be processed and the first splitting ratio; splitting the maximum dimension in the data dimension into K second values according to the third splitting ratio; and when the minimum value in the K second values is greater than a third preset threshold and the second splitting non-uniformity is less than a fourth preset threshold, splitting the operator to be processed into K subtasks according to the K second values.

[0015] In a possible implementation, the apparatus further includes an allocation module, configured to: when the operator to be processed does not conform to the first splitting method and does not conform to the second splitting method, allocate the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0016] In a possible implementation, the receiving module is configured to: when the category of the operator to be processed belongs to the target category, reprocess the K first results according to the category of the operator to be processed to obtain a second result of the operator to be processed; and when the category of the operator to be processed belongs to a category other than the target category, merge the K first results to obtain a second result of the operator to be processed.

[0017] In a possible implementation, the allocation module is further configured to: when the operator to be processed does not meet the splitting condition, allocate the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0018] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.

[0019] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0020] In an embodiment of the present disclosure, when a to-be-processed operator meets a preset splitting condition, a first splitting ratio of the to-be-processed operator is obtained. The first splitting ratio is used to indicate the scale ratio among K cores in a second processor, where K is an integer greater than or equal to 2. The to-be-processed operator is split into K subtasks according to the first splitting ratio. The K subtasks are sent to the K cores of the second processor, where each core is used to process the received subtask to obtain a first result of the subtask. K first results from the K cores are received, and a second result of the to-be-processed operator is determined according to the K first results and the category of the to-be-processed operator.

[0021] In this way, the to-be-processed operator can be split into subtasks of different scales as needed and allocated to processor cores that match their scales, supporting the scale and performance ratios of different processor cores, allowing tasks to be reasonably allocated on asymmetric processor cores, making the calculation time required by each core in the second processor basically the same, and being beneficial to improving the utilization efficiency of the second processor. For example, in the related art, deep learning operators are not further split but are treated as a whole with the smallest granularity for processing, resulting in the calculation being completed on only one computing core and having low efficiency. The embodiments of the present disclosure can combine deep learning operators as to-be-processed operators with an asymmetric, multi-core scenario, and the to-be-processed operators can be split to the second processor with asymmetric, multi-core, improving the utilization efficiency of the second processor.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.

[0024] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure.

[0025] Figure 2 A schematic diagram showing a data processing method according to an embodiment of the present disclosure.

[0026] Figure 3 A schematic diagram showing a first splitting method according to an embodiment of the present disclosure.

[0027] Figure 4 A schematic diagram showing a second splitting method according to an embodiment of the present disclosure.

[0028] Figure 5 A schematic diagram showing a reprocessing method according to an embodiment of the present disclosure.

[0029] Figure 6 A schematic diagram showing another data processing method according to an embodiment of the present disclosure.

[0030] Figure 7 A block diagram showing a data processing apparatus according to an embodiment of the present disclosure.

[0031] Figure 8 A block diagram showing an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0032] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0033] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0034] The term "and / or" in this article merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.

[0035] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0036] In the related art, a task is only placed on one (or a group of) cores, for example, on the core (group) with the largest scale, which causes the remaining cores to be idle and results in insufficient utilization of the overall resources of the processor. Alternatively, the computing task is evenly divided into subtasks of the same size and evenly distributed to each core, which makes the cores with small scale undertake the same tasks as the large cores. Obviously, this will cause the large cores to complete the tasks before the small cores and wait for the small cores. This also leads to a decrease in the overall resource utilization of the processor. Among them, the scale of a core represents the number of integrated electronic components or the number of logic gates. The number of integrated electronic components or the number of logic gates of a large core is much larger than that of a small core.

[0037] In view of this, in order to improve the resource utilization rate of a heterogeneous-core processor, an embodiment of the present disclosure provides a data processing method. Figure 1 The flowchart showing the data processing method according to the embodiment of the present disclosure is as Figure 1 shown. The data processing method includes:

[0038] In step S11, when the operator to be processed meets the preset splitting condition, obtain a first splitting ratio of the operator to be processed, where the first splitting ratio is used to indicate the scale ratio among K cores in a second processor, and K is an integer greater than or equal to 2;

[0039] In step S12, split the operator to be processed into K subtasks according to the first splitting ratio;

[0040] In step S13, send the K subtasks to the K cores of the second processor, where the cores are used to process the received subtasks to obtain first results of the subtasks;

[0041] In step S14, receive K first results from the K cores, and determine a second result of the operator to be processed according to the K first results and the category of the operator to be processed.

[0042] In this way, the operator to be processed can be split into subtasks of different scales as needed and allocated to the processor cores that match their scales, supporting the scale and performance ratio of different processor cores, allowing tasks to be reasonably allocated on asymmetric processor cores, making the calculation time required by each core in the second processor basically the same, which is beneficial to improving the utilization efficiency of the second processor. For example, in the related art, the deep learning operator is not further split but is treated as an integral whole with the smallest granularity for processing, resulting in the calculation being completed on only one computing core and having low efficiency. The embodiments of the present disclosure can combine the deep learning operator, which is the operator to be processed, with the asymmetric and multi-core scenario, and can split the operator to be processed to the second processor with asymmetric and multi-core, improving the utilization efficiency of the second processor.

[0043] In a possible implementation manner, the data processing method can be applied to electronic devices such as terminal devices or servers. The electronic device may include a first processor serving as a main control processor and a second processor serving as a coprocessor. The first processor can call the computer-readable instructions stored in the memory to implement the data processing method of the embodiments of the present disclosure and allocate tasks to the second processor. The electronic device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.

[0044] Among them, the first processor may include but is not limited to: a central processing unit (CPU), a digital signal processing unit (DSP), an application specific integrated circuit (ASIC), a tensor processing unit (TPU), a field programmable gate array (FPGA), etc. The second processor may include but is not limited to a graphics processing unit (GPU) with multiple cores, a general-purpose computing on graphics processing units (GPGPU), etc. The embodiments of the present disclosure do not limit the types of the first processor and the second processor.

[0045] In the example, it is allowed to divide the computing cores on a second processor into multiple cores (e.g., two), and each core can have a different scale. These cores can share the same video memory and other resources to achieve a more flexible computing mode. Among them, the scale of a core represents the number of integrated electronic components or the number of logic gates it integrates.

[0046] In a possible implementation manner, the data processing method can be executed by a controller. The controller can be software or program code running in a first processor, and can be implemented through a hardware description language, assembly language, high-level language (such as C, C++ etc.), or script language; the controller can also be a logic circuit embedded in the first processor. The embodiments of the present disclosure do not limit the form of the controller.

[0047] In a possible implementation manner, the data processing method can be applied to a distributed training scenario, where multiple second processors (e.g., multi-core GPUs, GPUs with multiple stream processor units) and multiple first processors (e.g., CPUs) can communicate with each other. The data processing method of the embodiments of the present disclosure can effectively allocate tasks to the cores of the second processor (such as GPU), thereby improving the speed and efficiency of distributed training.

[0048] In a possible implementation manner, in step S11, when the operator to be processed meets a preset splitting condition, the controller can obtain a first splitting ratio of the operator to be processed, and the first splitting ratio is used to indicate the scale ratio among K cores in the second processor, where K is an integer greater than or equal to 2;

[0049] In a possible implementation manner, the operator to be processed can be an operator used in deep learning, such as including matrix calculation operators, sliding window calculation operators, broadcast calculation operators, element-wise calculation operators, reduction calculation operators, non-computing task operators, etc. The embodiments of the present disclosure do not specifically limit the operator to be processed.

[0050] In a possible implementation manner, in deep learning, an operator or a network layer is a building block of a deep neural network, which defines the structure and operation process of the deep neural network. These operators can be used to perform various mathematical operations and operations, can accept tensor or scalar inputs, and produce tensor or scalar outputs.

[0051] Among them, the deep neural network may include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Generative Adversarial Networks (GANs), Attention-based Networks, Graph Neural Networks (GNNs), etc. Embodiments of the present disclosure do not specifically limit the deep neural network.

[0052] In a possible implementation, the obtained operator to be processed can be used to process the data to be processed in a deep learning task, and the data to be processed can be any one of image data, voice data, and text data. Embodiments of this application do not limit the data types processed by the operator to be processed.

[0053] For example, in the scenario of using a deep neural network for face recognition of a target object, the operator to be processed can be used to process the image data of the target object to obtain the image processing result required by the deep neural network. Another example, in the scenario of using a deep neural network for speech recognition of a target object, the operator to be processed can be used to process the voice data of the target object to obtain the voice processing result required by the deep neural network. Another example, in the scenario of using a deep neural network for text recognition of a target document, the operator to be processed can be used to process the text data of the target object to obtain the text processing result required by the deep neural network.

[0054] In a possible implementation, the splitting condition is a preset rule or threshold, which can be used to determine whether to split the operator to be processed. By setting the splitting condition in advance, the splitting mode of the operator to be processed can be managed.

[0055] Optionally, the black and white list can be used as the splitting condition. The black and white list can include a black list and a white list. Any operator in the operator list recorded in the black list refuses to perform the splitting operation, and any operator in the operator list recorded in the white list allows the splitting operation. When the controller obtains the operator to be processed, the controller can check the black and white list. If the operator to be processed is in the black list, the operator to be processed does not meet the splitting condition; if the operator to be processed is in the white list, the operator to be processed meets the splitting condition.

[0056] Optionally, the configuration list input by the user can be used as the splitting condition. The configuration list records the operators specified by the user that do not perform the splitting operation. When the controller obtains the operator to be processed, the controller can check the configuration list. If the operator to be processed is in the configuration list, the operator to be processed does not meet the splitting condition; if the operator to be processed is not in the configuration list, the operator to be processed meets the splitting condition.

[0057] Optionally, a preset threshold can be used as the splitting condition. When the controller obtains the operator to be processed, the controller can check whether the amount of data to be processed by the operator to be processed is greater than the preset threshold. If the amount of data to be processed by the operator to be processed is greater than or equal to the preset threshold, the operator to be processed meets the splitting condition; if the amount of data processed by the operator to be processed is less than the preset threshold, the operator to be processed does not meet the splitting condition.

[0058] It should be understood that the splitting condition is a preset rule or threshold for determining whether to split the operator to be processed. The embodiments of the present disclosure do not specifically limit the splitting condition, and it can be set according to the actual application scenario.

[0059] In a possible implementation manner, if the controller determines that the operator to be processed meets the preset splitting condition, the controller can obtain the first splitting ratio of the operator to be processed by reading the environment variable or the configuration file. The embodiments of the present disclosure do not specifically limit the specific manner of obtaining the first splitting ratio.

[0060] Among them, the first splitting ratio is used to indicate the scale ratio among K cores in the second processor. For example, the second processor may have K cores, and the scale ratio of the K cores can be expressed as P1:P2:……:P K , where, P1~P K These K values can be the same or different, and the present disclosure does not limit this.

[0061] After obtaining the first splitting ratio of the operator to be processed in step S11, the controller can split the operator to be processed into K subtasks according to the first splitting ratio in step S12, and send the K subtasks to the K cores of the second processor in step S13. The K cores of the second processor process the received subtasks to obtain K first results.

[0062] For example, assume that the second processor may have K cores, that is: core 1~core K, and the first splitting ratio is P1:P2:……:P K .

[0063] Optionally, the controller can follow the first splitting ratio P1:P2:……:P K, the operator to be processed is split into K sub-tasks, namely: sub-task 1 to sub-task K, and sub-task 1 is assigned to core 1 of the second processor for processing to obtain the first result 1, sub-task 2 is assigned to core 2 of the second processor for processing to obtain the first result 2, and so on, sub-task K is assigned to core K of the second processor for processing to obtain the first result K. Among them, the ratio of the data processing amounts of sub-task 1 to sub-task K is the first splitting ratio P1:P2:……:P K .

[0064] Optionally, the controller can appropriately transform the first splitting ratio P1:P2:……:P K to obtain the second splitting ratio Q1:Q2:……:Q K , and the controller can split the operator to be processed into K sub-tasks according to the second splitting ratio Q1:Q2:……:Q K , namely: sub-task 1 to sub-task K, and sub-task 1 is assigned to core 1 of the second processor for processing to obtain the first result 1, sub-task 2 is assigned to core 2 of the second processor for processing to obtain the first result 2, and so on, sub-task K is assigned to core K of the second processor for processing to obtain the first result K. Among them, the ratio of the data processing amounts of sub-task 1 to sub-task K is the second splitting ratio Q1:Q2:……:Q K .

[0065] In this way, according to the first splitting ratio, the operator to be processed can be decomposed into two or more sub-tasks, realizing symmetric or asymmetric partitioning of the operator to be processed. It should be understood that the above decomposition is a logical split, and in actual operation, it can be achieved by methods such as pointer offset, non-continuous tensor access, partial copy, etc., and its specific logic can be determined according to the situation, and the embodiments of the present disclosure do not limit this.

[0066] In step S13, the second processor processes the K sub-tasks to obtain K first results. In step S14, the controller can receive the K first results from the K cores, and according to the K first results and the type of the operator to be processed, choose to directly merge the K first results from the K cores, and use the merged result as the second result of the operator to be processed; or choose to reprocess the K first results from the K cores, and use the reprocessed result as the second result of the operator to be processed.

[0067] For example, if the type of the operator to be processed is an element-wise calculation operator, the K first results from the K cores can be directly merged to obtain the second result of the operator to be processed.

[0068] For example, if the type of the operator to be processed is a mean operator, the K first results from K cores can be reprocessed, and the K first results can be weighted averaged according to the splitting ratio to obtain the second result of the operator to be processed.

[0069] It should be understood that in the embodiments of the present disclosure, a second processor with asymmetric cores and its related support components are provided, which support dependencies such as storage alignment, and are used to implement the logic of distributing computing tasks to specified cores (for example, the controller splits K subtasks into K cores of the second processor), and synchronizing data after the calculation is completed. The support components may include tools and interfaces provided at the hardware, system, and driver levels. The embodiments of the present disclosure do not limit the implementation form of the support components.

[0070] Through steps S11 to S14, when the operator to be processed meets the preset splitting conditions, a first splitting ratio indicating the scale ratio between K cores in the second processor can be obtained. The operator to be processed can be split into K subtasks according to the obtained first splitting ratio, and the K subtasks can be assigned to the K cores of the second processor for processing to obtain K first results. The second result of the operator to be processed can be determined according to the K first results from the K cores and the category of the operator to be processed. In this way, the operator to be processed can be split into subtasks of different scales as needed and assigned to the processor cores that match their scales, supporting the scale and performance ratio of different processor cores, allowing tasks to be reasonably allocated on asymmetric processor cores, making the calculation time required by each core in the second processor basically the same, which is beneficial to improving the utilization efficiency of the second processor.

[0071] Figure 2 A schematic diagram showing a data processing method according to an embodiment of the present disclosure Figure 2 Taking the second processor having two cores (K = 2) as an example, an exemplary description of the data processing method of the embodiments of the present disclosure is given.

[0072] In the example, the controller can be software or program code running in the first processor (such as a CPU), or can be a logic circuit embedded in the first processor (such as a CPU). The embodiments of the present disclosure do not limit the form of the controller. The controller can obtain information such as environment variables, black and white lists, and control parameters, and determine whether the operator to be processed currently running needs to be split, as well as the number of split subtasks, the task scale ratio of each subtask, etc., in combination with the input data scale and task complexity of the operator to be processed.

[0073] Among them, the environment variables are used to indicate system information such as the scale ratio between multiple cores in the second processor, the asymmetric core partitioning method, the number of partitions, the computing power, and the computing power ratio of the second processor. The black and white list may include a black list and a white list. Any operator in the operator list recorded in the black list refuses to execute the splitting operation, and any operator in the operator list recorded in the white list is allowed to execute the splitting operation. The control parameter may include the running state information for controlling whether the splitting mode is started or stopped. The environment variables, the black and white list, and the control parameters can be used to determine the mode of the controller, and the environment variables, the black and white list, and the control parameters can be obtained through user input or automatic acquisition, etc. The embodiments of the present disclosure do not limit the acquisition methods of the environment variables, the black and white list, and the control parameters.

[0074] In the example, the deep learning operator is a unit of the computing task and can be split through the mode of the controller.

[0075] Such as Figure 2 As shown, the controller can obtain information from the black and white list or the control parameters configured by the user to determine whether to split the operator to be processed. And in the case of determining to split the operator to be processed, according to the first splitting ratio read from the environment variable, the operator to be processed is split to obtain two split operators. The split operators can be sent to different computing streams 1 and 2 and issued as different computing tasks (such as subtask 1 and subtask 2) for calculation in two asymmetric cores of the same second processor. For example, subtask 1 is calculated in core 1 of the second processor, and subtask 2 is calculated in core 2 of the second processor.

[0076] Among them, the scale of the split operator can be divided into a size that matches the core scale ratio corresponding to the second processor according to the function of the operator to be processed and the size of the input data of the operator to be processed. For example, when the core scale ratio indicated by the first splitting ratio is 3:1, subtask 1 can account for three-quarters of the computing amount of the operator to be processed, and task 2 can account for one-quarter of the computing amount of the operator to be processed, so that the task amounts borne by the large and small cores of the second processor match their scales.

[0077] In this way, the scale of the operator after splitting conforms to the splitting ratio specified by the controller, thereby allowing tasks to be reasonably allocated on the asymmetric computing cores, making the computing time required by each core basically the same to improve utilization. And it supports the second processor with asymmetric cores to provide the function of specifying computing tasks on different cores from the driver and hardware levels and provide multiple computing streams. The split operator will issue relevant subtasks to the computing streams, and each computing stream corresponds to a different processor core, thereby realizing the reasonable splitting of tasks and the efficiency of the second processor with cores of different scales.

[0078] It should be understood that Figure 2As an example of a data processing method, the operator to be processed can be various deep learning operators. The splitting method is not limited to splitting into 2, and can be split into multiple. The splitting scale can vary according to the actual situation, and the embodiments of the present disclosure do not limit this.

[0079] The data processing method of the embodiments of the present disclosure will be described in detail below.

[0080] In a possible implementation, in step S11, it can be first determined whether the operator to be processed meets a preset splitting condition. When the operator to be processed does not meet the splitting condition, the operator to be processed can be allocated to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0081] Alternatively, when the operator to be processed meets the preset splitting condition, a first splitting ratio of the operator to be processed can be obtained, and in step S12, the operator to be processed can be split into K subtasks according to the first splitting ratio.

[0082] By setting the splitting condition, when the operator to be processed does not meet the splitting condition, it is directly allocated to the largest core of the second processor for processing. When the operator to be processed meets the splitting condition, it is split into K subtasks according to the first splitting ratio. This not only dynamically determines whether to split the operator to be processed, increasing the flexibility and scalability of the system (such as a system composed of a first processor and a second processor), but also is conducive to the optimal utilization of the resources of the second processor and improves the efficiency of the second processor.

[0083] In a possible implementation, step S12 may include: when the operator to be processed conforms to the first splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the first splitting method, where the splitting dimension of the first splitting method is the batch dimension of the operator to be processed; or, when the operator to be processed conforms to the second splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the second splitting method, where the splitting dimension of the second splitting method is the largest dimension among the data dimensions of the operator to be processed.

[0084] Optionally, it can be first determined whether the operator to be processed conforms to the first splitting method. When the operator to be processed does not conform to the first splitting method, it is then determined whether the operator to be processed conforms to the second splitting method; or, it can also be first determined whether the operator to be processed conforms to the second splitting method. When the operator to be processed does not conform to the second splitting method, it is then determined whether the operator to be processed conforms to the first splitting method. The embodiments of the present disclosure do not limit this.

[0085] Exemplarily, the first splitting method is a method of splitting in the batch dimension of the operator to be processed. The input data of the operator to be processed and the operator itself can be split in the batch dimension to obtain K sub-tasks after splitting. Among them, the batch can represent the amount of data input into the deep learning model each time for the deep learning model to calculate, and can be a combination of some data of the same scale. For example, 16 pictures can form a batch tensor with a size of 16, and the batch dimension of this tensor is 16. It should be understood that the embodiments of the present disclosure only use 16 as an example for the batch dimension, and the embodiments of the present disclosure do not limit the size of the batch dimension.

[0086] Among them, the first splitting method is applicable to most operators, such as matrix calculation operators, sliding window calculation operators, broadcast calculation operators, element-wise calculation operators, reduction operators, non-computation task operators, etc. The embodiments of the present disclosure do not make specific limitations on this.

[0087] Exemplarily, the second splitting method is a method of splitting in the largest dimension among the data dimensions of the operator to be processed. First, the largest dimension of the input data of the operator to be processed can be obtained, and the input data of the operator to be processed is split in this dimension to obtain K sub-tasks after splitting. For example, assuming that the size of the input data of the operator to be processed is W×H×C, and its data dimensions include the width dimension W, the height dimension H, and the channel dimension W, where W represents the number of elements in the width dimension, H represents the number of elements in the height dimension, and C represents the number of elements in the channel dimension, the dimension where the maximum value among the width dimension W, the height dimension H, and the channel dimension W is located can be used as the largest dimension.

[0088] Among them, the second splitting method is applicable to matrix calculation operators, element-wise calculation operators, sliding window calculation operators, broadcast calculation operators, etc. The embodiments of the present disclosure do not make specific limitations on this.

[0089] In this way, the splitting method can be selected according to the self-information of the operator to be processed (such as data dimension, batch dimension), which can more flexibly adapt to different scenarios and help improve the generality and adaptability of the splitting process.

[0090] In a possible implementation, the splitting dimension of the first splitting method is the batch dimension of the operator to be processed. When the operator to be processed conforms to the first splitting method, according to the first splitting ratio and the first splitting method, the operator to be processed is split into K subtasks, including: determining a second splitting ratio and a first splitting non-uniformity according to the batch dimension of the operator to be processed and the first splitting ratio; splitting the batch dimension into K first values according to the second splitting ratio; and when the minimum value among the K first values is greater than a first preset threshold and the first splitting non-uniformity is greater than a second preset threshold, splitting the operator to be processed into K subtasks according to the K first values.

[0091] Among them, the first splitting non-uniformity is used to describe the degree of balance in the quantity of input sub-data when the input data of the operator to be processed is split into several parts of input sub-data in the batch dimension.

[0092] Exemplarily, assume that the batch dimension of the operator to be processed is N, and the first splitting ratio is P1:P2:……:P K , where P1≥P2≥……≥P K , P1, P2,……P K are relatively prime. According to the batch dimension N of the operator to be processed and the first splitting ratio P1:P2:……:P K , the second splitting ratio Q1:Q2:……:Q K can be determined, where and so on. Int() represents the rounding function.

[0093] And, according to the batch dimension N of the operator to be processed and the first splitting ratio P1:P2:……:P K , the first splitting non-uniformity

[0094] According to the second splitting ratio Q1:Q2:……:Q K , the batch dimension N can be split into K first values, that is: the first value DQ1 to the first value DQ K , where the first value DQ K is the minimum number among the K first values DQ1 to DQ K . For example, assume that the batch dimension N is 10, the number of splitting parts K is 2, and the second splitting ratio Q1:Q2 is 3:2. Then, the split first value DQ1 is 6 and the first value DQ2 is 4. It should be understood that the embodiments of the present disclosure are for the batch dimension N, the number of splitting parts K, and the split first values DQ1 to DQ KThe value of

[0095] If the first value DQ K is greater than the first preset threshold (e.g., 1), and the first splitting non-uniformity H1 is less than the second preset threshold (e.g., 0.1), the operator to be processed can be split into K subtasks in the batch dimension N of the operator to be processed according to the K first values DQ1 to the first value DQ K .

[0096] If the first value DQ K is less than or equal to the first preset threshold (e.g., 1), or the first splitting non-uniformity H1 is greater than or equal to the second preset threshold (e.g., 0.1), it indicates that the operator to be processed does not conform to the first splitting method. It can continue to determine whether the operator to be processed conforms to the second splitting method. If the operator to be processed conforms to the second splitting method, the second splitting method can be used; if the operator to be processed does not conform to the second splitting method, the operator to be processed can be directly calculated without performing the splitting operation.

[0097] Among them, the first preset threshold and the second preset threshold can be set according to the actual application scenario, and the embodiments of the present disclosure do not limit the specific values of the first preset threshold and the second preset threshold.

[0098] Figure 3 FIG. shows a schematic diagram of the first splitting method according to an embodiment of the present disclosure, as Figure 3 shown, the input data 31 of the operator to be processed 34 can be decomposed into two input sub-data 32 and input sub-data 33 with different sizes in the batch dimension. The operator to be processed 34 can also be split into two splitting operators 35 and splitting operator 36 in the batch dimension. The two input sub-data 32 and input sub-data 33 with different sizes can share the same weight. Among them, the operation processing of the splitting operator 35 on the input sub-data 32 and the weight can be used as a subtask to obtain the first result 37 corresponding to the output of this subtask; the operation processing of the splitting operator 36 on the input sub-data 33 and the weight can be used as another subtask to obtain the first result 38 corresponding to the output of this subtask. The first results 37 and first result 37 of the two (K = 2) subtasks can be merged to obtain the second result 39 of the operator to be processed 34. According to Figure 3 it can be seen that before and after the calculation, the input and output of the operator to be processed can be decomposed and merged.

[0099] It should be understood that Figure 3 the decomposition and merging in

[0100] Through the first splitting method, the operator to be processed can be decomposed into two or more subtasks to achieve asymmetric partitioning, so as to be applicable to a second processor with an asymmetric core, which is beneficial to improving the efficiency of the second processor.

[0101] In a possible implementation, the splitting dimension of the second splitting method is the maximum dimension among the data dimensions of the operator to be processed. When the operator to be processed conforms to the second splitting method, according to the first splitting ratio and the second splitting method, the operator to be processed is split into K subtasks, including: determining a third splitting ratio and a second splitting non-uniformity according to the maximum dimension among the data dimensions of the operator to be processed and the first splitting ratio; splitting the maximum dimension among the data dimensions into K second values according to the third splitting ratio; and when the minimum value among the K second values is greater than a third preset threshold and the second splitting non-uniformity is less than a fourth preset threshold, splitting the operator to be processed into K subtasks according to the K second values.

[0102] Among them, the second splitting non-uniformity is used to describe the degree of balance in the quantity of input sub-data when the input data of the operator to be processed is split into several parts of input sub-data in the data dimension. The data dimension of the operator to be processed may include a width dimension W, a height dimension H, and a channel dimension C, and the maximum dimension M among the data dimensions of the operator to be processed is M = max(W, H, C), where max() is a function for taking the maximum value.

[0103] Exemplarily, assume that the maximum dimension of the operator to be processed is M, and the first splitting ratio is P1:P2:……:P K , where P1≥P2≥……≥P K , P1, P2,……P K are relatively prime. According to the batch dimension M of the operator to be processed and the first splitting ratio P1:P2:……:P K , a third splitting ratio R1:R2:……:R K can be determined, where And so on, Int() represents the integer function.

[0104] And, according to the batch dimension M of the operator to be processed and the first splitting ratio P1:P2:……:P K , the second splitting non-uniformity

[0105] According to the third splitting ratio R1:R2:……:R K , the batch dimension M can be split into K second values, that is: the second value DR1~the second value DR K, where the second value DR K is the smallest number among K second values DR1 to DR K .

[0106] If the second value DR K is greater than the third preset threshold (e.g., 1), and the second splitting non-uniformity H2 is less than the fourth preset threshold (e.g., 0.1), the operator to be processed can be split into K subtasks in the batch dimension M of the operator to be processed according to K second values DR1 to DR K .

[0107] If the second value DR K is less than or equal to the third preset threshold (e.g., 1), or the second splitting non-uniformity H2 is greater than or equal to the fourth preset threshold (e.g., 0.1), it means that the operator to be processed does not conform to the second splitting method. It can continue to judge whether the operator to be processed conforms to the first splitting method. If the operator to be processed conforms to the first splitting method, the first splitting method can be used; if the operator to be processed does not conform to the first splitting method, the operator to be processed can be directly calculated without performing the splitting operation.

[0108] Among them, the third preset threshold and the fourth preset threshold can be set according to the actual application scenario. The embodiments of the present disclosure do not limit the specific values of the third preset threshold and the fourth preset threshold.

[0109] Figure 4 shows a schematic diagram of the second splitting method according to an embodiment of the present disclosure, as Figure 4 shown. Assuming that the operator to be processed is a matrix multiplication operator, in the example of matrix multiplication, the input data 41 can be decomposed into input sub-data 42 and input sub-data 43 by using a matrix decomposition method. The matrix multiplication operation of the input sub-data 42 and the weight can be used as one subtask to obtain the first result 44 corresponding to the output of this subtask; the matrix multiplication operation of the input sub-data 43 and the weight can be used as another subtask to obtain the first result 45 corresponding to the output of this subtask. The first results 44 and 45 of the two (K = 2) subtasks can be merged to obtain the second result 46 of the matrix multiplication operator. According to Figure 4 , it can be known that before and after the calculation, the input and output of the operator to be processed can be decomposed and merged.

[0110] It should be understood that in the second splitting method, the division of the input data of the operator to be processed is not strictly non-overlapping. For some operators to be processed (such as convolution operators), in order to ensure the correctness of the data, a part of the input data may be used by two subtasks at the same time.

[0111] Through the second splitting method, the operator to be processed can be decomposed into two or more subtasks to achieve asymmetric partitioning, so as to be applicable to the second processor with an asymmetric core, which is beneficial to improving the efficiency of the second processor.

[0112] In a possible implementation manner, the method further includes: when the operator to be processed does not conform to the first splitting method and does not conform to the second splitting method, allocating the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0113] Exemplarily, if the first numerical value DQ K is less than or equal to the first preset threshold, or the first splitting non-uniformity H1 is greater than or equal to the second preset threshold, it indicates that the operator to be processed does not conform to the first splitting method. If the second numerical value DR K is less than or equal to the third preset threshold, or the second splitting non-uniformity H2 is greater than or equal to the fourth preset threshold, it indicates that the operator to be processed does not conform to the second splitting method.

[0114] When the operator to be processed does not conform to the first splitting method and also does not conform to the second splitting method, no splitting operation is performed on the operator to be processed, and the operator to be processed can be directly calculated, and the operator to be processed is allocated to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0115] In this way, when the operator to be processed does not conform to the first splitting method and the second splitting method, it is directly allocated to the largest core of the second processor for processing, which is beneficial to reducing the deviation or calculation error caused by inappropriate splitting, and can improve the accuracy of the second processor calculation.

[0116] If the operator to be processed conforms to the first splitting method or the second splitting method, in step S12, the operator to be processed can be split into K subtasks according to the first splitting method or the second splitting method, and in step S13, the controller allocates the K subtasks to the K cores of the second processor, and the K cores of the second processor can process the K subtasks to obtain K first results. In step S14, according to the K first results from the K cores and the category of the operator to be processed, the second result of the operator to be processed is determined.

[0117] In a possible implementation manner, step S14 may include: when the category of the operator to be processed belongs to the target category, reprocessing the K first results according to the category of the operator to be processed to obtain the second result of the operator to be processed; when the category of the operator to be processed belongs to other than the target category, merging the K first results to obtain the second result of the operator to be processed.

[0118] In the example, the target category is a reduction computing operator, that is, an operator for which directly merging K first results will cause an error in the second result of the operator to be processed. For example, it includes mean operators, maximum value operators, minimum value operators, summation operators, product operators, etc. Embodiments of the present disclosure do not make specific limitations on the target category.

[0119] For the operator to be processed in the target category, the K first results of the operator to be processed can be reprocessed. After the reduction calculation of the target category, a vector or scalar calculation, or a reduction calculation of the same type can be supplemented. For example, for the calculation of mean operators, summation operators, and product operators, mathematical calculations of vectors or scalars can be added (for example, a weighted average calculation is added to the mean operator, an addition calculation is added to the summation operator, and a multiplication calculation is added to the product operator), and a reduction calculation of the same type is added to the maximum and minimum values.

[0120] Figure 5 A schematic diagram showing a reprocessing method according to an embodiment of the present disclosure is as follows Figure 5 As shown, assume that the operator to be processed in the target category is a reduction operator 51. The input data 52 of the reduction operator 51 can be split into input sub-data 53 and input sub-data 54. The reduction results (for example, their respective means) of the input sub-data 53 and the input sub-data 54 can be calculated respectively to obtain an intermediate result 55 and an intermediate result 56. The intermediate results can be recomputed by merging to obtain the second result 57 of the reduction operator 51. For example, the intermediate result 55 and the intermediate result 56 can be weighted averaged according to a second splitting ratio or a third splitting ratio, that is, a combination of vector calculations of a set of multiplication, addition, and division, to obtain the second result 57 of the reduction operator 51.

[0121] In this way, a merging recomputation method is designed to correctly merge the calculation results of subtasks. It can be determined whether to directly merge the K first results or reprocess the K first results according to the category of the operator to be processed. It can be applied to operators with reduction logic, expanding the application scenario of splitting processing and improving the accuracy of the results of the operator to be processed.

[0122] Figure 6 A schematic diagram showing another data processing method according to an embodiment of the present disclosure is as follows Figure 6 As shown, referring to step S11, it can be determined whether the operator to be processed meets the splitting condition. If it does not meet the splitting condition, it is directly calculated. A certain core (for example, the largest core) of the second processor can process the operator to be processed from the first processor to obtain the third result of the operator to be processed, and the second processor can return this result to the first processor. If the operator to be processed meets the splitting condition, the first splitting ratio can be obtained, and the first splitting ratio is used to indicate the scale ratio of K cores of the second processor.

[0123] Referring to steps S12 to S13, it is possible to sequentially determine whether the operator to be processed can use the first splitting method (corresponding to the batch splitting strategy) and the second splitting method (corresponding to the data splitting strategy). If the first splitting condition and the second splitting condition are not satisfied, direct calculation is performed. A certain core (such as the largest core) of the second processor can process the operator to be processed from the first processor to obtain a third result of the operator to be processed, and the second processor can return this result to the first processor.

[0124] If there is an available splitting strategy, for example, if the operator to be processed satisfies the first splitting method, the input data of the operator to be processed can be decomposed using the first splitting method and the calculation can be completed. Or, if the operator to be processed satisfies the second splitting method, the input data of the operator to be processed can be decomposed using the second splitting method and the calculation can be completed.

[0125] Referring to step S14, after the calculation is completed, it is possible to determine whether the operator to be processed is a target category (such as a reduction operator). If it is a target category operator, data reprocessing and data merging are appended to obtain a second result of the operator to be processed; if it is not a target category operator, data merging is directly performed to obtain a second result of the operator to be processed. The second processor can return this result to the first processor.

[0126] In summary, the data processing method of the embodiments of the present disclosure can split the operator to be processed into subtasks of different scales as needed and allocate them to processor cores that match their scales, support the scale and performance ratio of different processor cores, allow tasks to be reasonably allocated on asymmetric processor cores, so that the calculation times required by each core in the second processor are basically the same, which is beneficial to improving the utilization efficiency of the second processor.

[0127] Compared with the scheduling optimization in the compilation stage in the related art, and retrieving the possible optimization solution space to optimize the performance on the processor device. Among them, the compiler optimization scheme depends on the modification and adjustment of the compiler. At the same time, this modification is limited by the data that the compiler can obtain, and it cannot obtain the complete data characteristics and scale of the current calculation task, nor can it freely split or adjust it. The data processing method of the embodiments of the present disclosure does not make adjustments to the compilation period, does not depend on the underlying hardware components or the compiler, and does not require modification or adjustment of the compiler, processor device, etc. It involves the deep learning scenario and has good versatility and can be applied to various asymmetric multi-core devices. At the same time, the optimization strategy of the data processing method of the embodiments of the present disclosure is established at a higher perspective of the operator granularity, and each optimization can be obtained through mathematical calculations based on the existing parameters without searching the solution space.

[0128] Similarly, compared with the related art that modifies the hardware computing power by physically adjusting the hardware voltage and power consumption to achieve an optimization effect, the data processing method of the embodiments of the present disclosure does not involve requirements for hardware, and the optimization scope is limited to the software level. Moreover, compared with the related art that optimizes methods such as dynamic scheduling through the detection and real-time analysis of the running state at the processor level, the optimization scheme of the processor is limited by the monitoring of the processor state, making it difficult to know the characteristics of the computing tasks and unable to optimize according to specific computing tasks and operator classifications. It is difficult to ensure the matching between task scheduling and the actual processor core scale. The data processing method of the embodiments of the present disclosure is an optimization strategy of pre-computation and pre-allocation, which can achieve on-demand division of operators in the deep learning scenario without the need for real-time detection or real-time analysis of the processor.

[0129] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0130] In addition, the present disclosure also provides a data processing device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any data processing method provided by the present disclosure. For the corresponding technical solutions and descriptions, please refer to the corresponding records in the method part and will not be elaborated here.

[0131] Figure 7 The block diagram of the data processing device according to the embodiments of the present disclosure is shown as Figure 7 shown, the device is applied to a first processor, and the device includes:

[0132] An obtaining module 71, configured to obtain a first splitting ratio of the to-be-processed operator when the to-be-processed operator meets a preset splitting condition, where the first splitting ratio is used to indicate the scale ratio among K cores in a second processor, and K is an integer greater than or equal to 2;

[0133] A splitting module 72, configured to split the to-be-processed operator into K sub-tasks according to the first splitting ratio;

[0134] A sending module 73, configured to send the K sub-tasks to the K cores of the second processor, where the cores are configured to process the received sub-tasks to obtain first results of the sub-tasks;

[0135] A receiving module 74, configured to receive K first results from the K cores, and determine a second result of the to-be-processed operator according to the K first results and the category of the to-be-processed operator.

[0136] In a possible implementation, the splitting module 72 is configured to: when the operator to be processed conforms to the first splitting method, split the operator to be processed into K subtasks according to the first splitting ratio and the first splitting method, where the splitting dimension of the first splitting method is the batch dimension of the operator to be processed; or, when the operator to be processed conforms to the second splitting method, split the operator to be processed into K subtasks according to the first splitting ratio and the second splitting method, where the splitting dimension of the second splitting method is the maximum dimension among the data dimensions of the operator to be processed.

[0137] In a possible implementation, the step of splitting the operator to be processed into K subtasks according to the first splitting ratio and the first splitting method when the operator to be processed conforms to the first splitting method includes: determining a second splitting ratio and a first splitting non-uniformity degree according to the batch dimension of the operator to be processed and the first splitting ratio; splitting the batch dimension into K first values according to the second splitting ratio; and when the minimum value among the K first values is greater than a first preset threshold and the first splitting non-uniformity degree is less than a second preset threshold, splitting the operator to be processed into K subtasks according to the K first values.

[0138] In a possible implementation, the step of splitting the operator to be processed into K subtasks according to the first splitting ratio and the second splitting method when the operator to be processed conforms to the second splitting method includes: determining a third splitting ratio and a second splitting non-uniformity degree according to the maximum dimension among the data dimensions of the operator to be processed and the first splitting ratio; splitting the maximum dimension in the data dimension into K second values according to the third splitting ratio; and when the minimum value among the K second values is greater than a third preset threshold and the second splitting non-uniformity degree is less than a fourth preset threshold, splitting the operator to be processed into K subtasks according to the K second values.

[0139] In a possible implementation, the apparatus further includes an allocation module, configured to: when the operator to be processed does not conform to the first splitting method and does not conform to the second splitting method, allocate the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0140] In a possible implementation, the receiving module 74 is configured to: when the category of the operator to be processed belongs to the target category, reprocess the K first results according to the category of the operator to be processed to obtain a second result of the operator to be processed; when the category of the operator to be processed belongs to a category other than the target category, merge the K first results to obtain a second result of the operator to be processed.

[0141] In a possible implementation, the distribution module is further configured to: when the operator to be processed does not meet the splitting condition, allocate the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

[0142] This method has a specific technical association with the internal structure of a computer system and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing the data storage volume, reducing the data transmission volume, and increasing the hardware processing speed, etc.), thereby obtaining a technical effect of improving the internal performance of the computer system that conforms to the laws of nature.

[0143] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0144] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0145] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.

[0146] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.

[0147] The electronic device can be provided as a terminal, a server, or other forms of devices.

[0148] Figure 8 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Refer to Figure 8, the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0149] The electronic device 1900 may also include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.

[0150] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0151] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0152] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0153] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0154] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0155] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0156] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner. Thus, the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0157] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0158] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0159] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), and so on.

[0160] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they will not be elaborated herein.

[0161] Those skilled in the art can understand that in the above methods of the specific implementation manners, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0162] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0163] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A data processing method, characterized in that: The method is applied to a first processor, and the method includes: When the operator to be processed meets a preset splitting condition, obtaining a first splitting ratio of the operator to be processed, where the first splitting ratio is used to indicate a scale ratio between K cores in the second processor, where K is an integer greater than or equal to 2; Splitting the operator to be processed into K subtasks according to the first splitting ratio; Sending K subtasks to K cores of the second processor, wherein the cores are used to process the received subtasks to obtain first results of the subtasks; The K first results are received from the K cores, and a second result of the operator to be processed is determined according to the K first results and the category of the operator to be processed.

2. The method according to claim 1, characterized in that The step of splitting the operator to be processed into K subtasks according to the first split ratio includes: When the operator to be processed conforms to the first splitting method, the operator to be processed is split into K subtasks according to the first splitting ratio and the first splitting method, wherein the splitting dimension of the first splitting method is the batch dimension of the operator to be processed; Alternatively, when the operator to be processed conforms to the second splitting method, the operator to be processed is split into K subtasks according to the first splitting ratio and the second splitting method, wherein the splitting dimension of the second splitting method is the maximum dimension among the data dimensions of the operator to be processed.

3. The method according to claim 2, characterized in that When the operator to be processed conforms to the first splitting method, splitting the operator to be processed into K subtasks according to the first splitting ratio and the first splitting method includes: Determine a second split ratio and a first split unevenness according to the batch dimension of the operator to be processed and the first split ratio; Splitting the batch dimension into K first values ​​according to the second splitting ratio; When the minimum value among the K first values ​​is greater than a first preset threshold and the first splitting unevenness is less than a second preset threshold, the operator to be processed is split into K subtasks according to the K first values.

4. The method according to claim 2, characterized in that: When the operator to be processed conforms to the second splitting method, the operator to be processed is split into K subtasks according to the first splitting ratio and the second splitting method, including: Determine a third splitting ratio and a second splitting unevenness according to the maximum dimension of the operator data to be processed and the first splitting ratio; Splitting the largest dimension among the data dimensions into K second values ​​according to the third splitting ratio; When the minimum value among the K second values ​​is greater than the third preset threshold and the second splitting unevenness is less than the fourth preset threshold, the operator to be processed is split into K subtasks according to the K second values.

5. The method according to any one of claims 2 to 4, characterized in that: The method further includes: when the operator to be processed does not conform to the first splitting method and the operator to be processed does not conform to the second splitting method, allocating the operator to be processed to the largest core of the second processor for processing to obtain a third result of the operator to be processed.

6. The method according to claim 1, characterized in that The determining, according to the K first results and the category of the operator to be processed, the second result of the operator to be processed includes: When the category of the operator to be processed belongs to the target category, reprocessing the K first results according to the category of the operator to be processed to obtain a second result of the operator to be processed; When the category of the operator to be processed belongs to a category other than the target category, the K first results are merged to obtain a second result of the operator to be processed.

7. The method according to claim 1, characterized in that The method further includes: when the operator to be processed does not meet the splitting condition, allocating the operator to be processed to the largest core of the second processor for processing, to obtain a third result of the operator to be processed.

8. A data processing device, characterized in that: The device is applied to a first processor, and the device includes: an acquisition module, configured to acquire a first split ratio of the operator to be processed when the operator to be processed satisfies a preset split condition, wherein the first split ratio is used to indicate a scale ratio between K cores in the second processor, where K is an integer greater than or equal to 2; A splitting module, used for splitting the operator to be processed into K subtasks according to the first splitting ratio; A sending module, used for sending K subtasks to K cores of the second processor, wherein the cores are used for processing the received subtasks to obtain first results of the subtasks; The receiving module is used to receive the K first results from the K cores, and determine the second result of the operator to be processed according to the K first results and the category of the operator to be processed.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Data processing device and method, electronic equipment and storage medium

    CN121807558A

  • Data processing apparatus and method, electronic device, storage medium

    CN121807558B