Task allocation method, instruction processor, device, medium and program product

By storing tasks and data in adjacent processing cores and memory in the processor, and using the association relationship between task identifiers and core serial numbers for task allocation, the problem of handling core communication bandwidth bottlenecks is solved and task processing efficiency is improved.

CN120540707APending Publication Date: 2025-08-26SHANGHAI BIREN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510572306.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the existing processor technology, the task allocation process is carried out independently from the data storage process, resulting in the communication bandwidth between the processing cores becoming a bottleneck and affecting the task processing efficiency.

Method used

When allocating tasks, tasks and data are pre-stored in adjacent processing cores and memory, avoiding communication processes across processing cores, and using the association relationship between task identifiers and core serial numbers for task allocation.

Benefits of technology

Improve the overall efficiency of task processing, avoid communication bottlenecks between processing cores, and improve the speed of task allocation and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540707A_ABST
    Figure CN120540707A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of processors, and discloses a task allocation method, an instruction processor, equipment, a medium and a program product, and the method comprises the steps: receiving a to-be-allocated target task, and allocating the target task to a target processing core in a plurality of preset processing cores; wherein target data required for executing the target task is stored in a memory adjacent to the target processing core in advance, and when the target processing core obtains the target data from the memory, a cross-processing-core communication process does not need to be executed. According to the technical scheme provided by one or more embodiments, the overall efficiency of task processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of processor technology, and in particular to a task allocation method, instruction processor, device, medium, and program product. Background Art

[0002] In the current field of processor technology, tasks are usually distributed through a command processor (CP). Specifically, the command processor can distribute tasks to be executed among multiple processing cores, so that the tasks can be processed by a streaming processor cluster (SPC) in the processing cores.

[0003] Currently, task assignment and data storage are performed independently. This can cause a processing core to communicate across multiple processing cores to retrieve the required data. This can easily lead to inter-core communication bottlenecks, resulting in lower processing efficiency.

[0004] In view of this, a more efficient task allocation method is currently needed. Summary of the Invention

[0005] The present application provides a task allocation method, instruction processor, device, medium and program product, which can improve the overall efficiency of task processing.

[0006] The first aspect of the present application provides a task allocation method, which is applied to an instruction processor, and the method includes: receiving a target task to be assigned, and assigning the target task to a target processing core among a preset plurality of processing cores; wherein, the target data required to execute the target task is pre-stored in a memory adjacent to the target processing core, and when the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

[0007] In this embodiment, the data required to execute a task is pre-stored in memory. When assigning a target task, the instruction processor can determine the memory location of the target data required to execute the target task. In an integrated chip, the memory and processing cores can be determined to be adjacent based on their respective locations. If a processing core can directly read data from the memory, then the processing core and the memory are considered to be adjacent. However, if a processing core attempts to read data from the memory and needs to obtain the required data through another processing core using cross-processing core communication, then the processing core and the memory are considered to be non-adjacent. In this embodiment, when assigning a target task, the instruction processor can assign the target task to a target processing core that is adjacent to the memory where the target data is located. This way, when the target processing core subsequently executes the target task, it can directly retrieve the target data from the adjacent memory. This task assignment method avoids cross-processing core communication and is not limited by the communication bandwidth bottleneck between processing cores, thereby achieving higher task processing efficiency.

[0008] In one embodiment, allocating the target task to a target processing core among a preset plurality of processing cores includes: identifying the task identifier of the target task and querying the core serial number associated with the task identifier; and allocating the target task to the target processing core having the core serial number.

[0009] In this embodiment, when assigning a target task, the instruction processor can quickly determine the core number associated with the target task's task identifier based on the association between the task identifier and the core number, and then assign the target task to the processing core with that core number. This assignment method can quickly complete task assignment and improve the overall efficiency of subsequent task processing.

[0010] In one embodiment, querying the core sequence number associated with the task identifier includes: reading pre-generated association information, the association information is used to represent the association relationship between the task identifier and the core sequence number; and querying the core sequence number associated with the task identifier of the target task in the association information.

[0011] In this embodiment, the association relationship between the task identifier and the core serial number can be represented by pre-generated association information. Task allocation according to the association information can ensure that the processing core where the task is located and the memory where the data is located have a proximity relationship, thereby improving the efficiency of task allocation and processing.

[0012] In one embodiment, the association information is pre-generated in the following manner: obtaining a data processing comparison table in the target application, the data processing comparison table being used to characterize the matching relationship between tasks and data during the execution of the target application; extracting the task identifier of any set of matching tasks and data in the data processing comparison table, and determining the memory for storing the data; identifying the processing core adjacent to the memory, and obtaining the core serial number of the processing core; constructing an association relationship between the task identifier and the core serial number, and generating association information based on each set of constructed association relationships.

[0013] In this embodiment, when the target application is written, the data required for the task to be executed can be determined, and the data can be limited to which memory or memories the data is stored in. In this way, after the storage location of the data is clarified, it is only necessary to limit the subsequent allocation of tasks to be executed to the processing core with a neighboring relationship, so as to ensure that the processing core reads the required data directly from the memory with a neighboring relationship when processing the task. In other words, when the target application is running, the matching relationship between the task and the data has been clarified. For such a matching relationship, the memory where the data is located can be determined first, and then the processing core with a neighboring relationship with the memory can be determined. By associating the task identifier of the task with the core serial number of the processing core, the association information can be constructed. Subsequently, when tasks are assigned according to this association information, it can be ensured that the processing core where the task is located and the memory where the data is located have a neighboring relationship, thereby improving the efficiency of task allocation and processing.

[0014] In one embodiment, after obtaining the core serial number of the processing core, the method further includes: initiating a memory request carrying the core serial number for the data, so that the memory management module responds to the memory request and allocates memory for storing the data in the memory that is adjacent to the core serial number.

[0015] In this embodiment, when writing data to the memory, it is necessary to request the corresponding memory space from the memory. Specifically, the memory request initiated for the data can include a core number. When the memory management module receives the memory request, it can identify the core number and allocate the corresponding memory in the memory that is adjacent to the core number. This ensures that data and tasks are placed in adjacent memory and processing cores, respectively, thereby improving the overall processing efficiency of the task.

[0016] In one embodiment, allocating the target task to a target processing core among a preset plurality of processing cores includes: determining a memory where the target data required to execute the target task is currently located; identifying a target processing core that has an adjacent relationship with the memory, and allocating the target task to the target processing core.

[0017] In this embodiment, when assigning a target task, the instruction processor can assign the target task to a target processing core that is adjacent to the memory where the target data resides. This allows the target processing core to directly retrieve the target data from the adjacent memory when subsequently executing the target task. This task assignment method avoids inter-processing core communication and is not limited by the communication bandwidth bottleneck between processing cores, thereby achieving higher task processing efficiency.

[0018] In one embodiment, the preset multiple processing cores include a first processing core and a second processing core, a first memory has an adjacent relationship with the first processing core, and a second memory has an adjacent relationship with the second processing core; wherein, data required for tasks with even-numbered task identifiers are stored in the first memory, and data required for tasks with odd-numbered task identifiers are stored in the second memory; allocating the target task to the target processing core among the preset multiple processing cores includes: identifying the task identifier of the target task, and when the task identifier is an even-numbered task identifier, allocating the target task to the first processing core; when the task identifier is an odd-numbered task identifier, allocating the target task to the second processing core.

[0019] In this embodiment, tasks can be assigned based on the parity of the task identifiers, where even-numbered task identifiers can be assigned to the first processing core, and odd-numbered task identifiers can be assigned to the second processing core. To ensure that task assignment in this manner can avoid communication across processing cores, when storing each task, the data required for tasks with even-numbered task identifiers can be stored in the first memory, and the data required for tasks with odd-numbered task identifiers can be stored in the second memory. In this way, when the processing core subsequently executes a task, it can directly read the required data from the adjacent memory, thereby improving the overall efficiency of task processing.

[0020] In one embodiment, after receiving the target task to be assigned, the method further includes: determining whether a preset optimization allocation function is enabled, and if the optimization allocation function is enabled, assigning the target task to the target processing core.

[0021] In this embodiment, an optimized allocation function can be pre-set. Only when this function is enabled will the designated processor allocate the target task to the target processing core according to the logic of avoiding cross-processing core communication. If this function is not enabled, task allocation can be carried out according to the conventional task allocation method. This measure improves the overall compatibility of the system, is compatible with existing task allocation mechanisms, and ensures the stability of the task processing process.

[0022] On the other hand, the present application also provides an instruction processor, which includes: a receiving unit for receiving a target task to be assigned; an assigning unit for assigning the target task to a target processing core among a preset plurality of processing cores; wherein, the target data required to execute the target task is pre-stored in a memory adjacent to the target processing core, and when the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

[0023] On the other hand, the present application also provides an electronic device, including: a memory storing computer instructions; and at least one processor configured to execute the computer instructions in the memory to perform the task allocation method in the above embodiment.

[0024] On the other hand, the present application further provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the task allocation method in the above embodiment.

[0025] On the other hand, the present application further provides a computer program product, including computer instructions, which, when executed by a processor, enable the processor to execute the task allocation method in the above embodiment. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0027] Figure 1 Assign a schematic diagram to tasks in related technologies;

[0028] Figure 2 A schematic diagram of a communication method across processing cores in related technologies;

[0029] Figure 3 A schematic diagram of the steps of a task allocation method provided in one embodiment of the present application;

[0030] Figure 4 A schematic diagram of steps for generating association information provided in one embodiment of the present application;

[0031] Figure 5 A diagram of task allocation for a specific application example;

[0032] Figure 6 A schematic diagram of the functional modules of an instruction processor provided in one embodiment of the present application;

[0033] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0035] In addition, the descriptions of "first", "second", etc. in this application are for descriptive purposes only and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values ​​can be based on additional conditions or values ​​beyond the described values ​​in practice.

[0036] See also Figure 1 In related technologies, the instruction processor usually uses a round-robin method to allocate the tasks to be assigned among multiple preset processing cores. Specifically, each processing core can contain several stream processor clusters. When allocating tasks, the instruction processor can refer to the confidence of each stream processor cluster and then give priority to allocating tasks to the processing core where the stream processor cluster with higher confidence is located. For example, Figure 1 In the example, the instruction processor can assign tasks to the processing core on the left. However, in actual applications, the data required for the task to be executed is likely to be located in the memory on the right. In this case, the execution unit (CU) in the processing core on the left needs to be allocated according to Figure 2The communication method across processing cores shown reads the required data from the memory on the right through the processing core on the right.

[0037] In one practical application example, the memory positioned relative to the processing core can be High Bandwidth Memory (HBM). Of course, with technological advancements and process differences, the memory can also be other types of memory, such as quantum memory or High Bandwidth Flash (HBF). For ease of explanation, the technical solution will be explained using HBM as an example.

[0038] When multiple kernel functions are executing within a processing core, the above-mentioned cross-processing core communication method and the transmission bandwidth between the processing cores will become the bottleneck of the entire system, thereby reducing the system throughput, resulting in low task processing efficiency, and further increasing the average memory access latency.

[0039] An embodiment of the present application provides a task allocation method, which can place tasks and data in processing cores and memories with adjacent relationships respectively. When the processing core executes the task, it can directly read the required data from the memory, thereby avoiding communication across processing cores.

[0040] The task allocation method provided in this application can be applied to an artificial intelligence processor with a task allocation function, specifically any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose Graphics Processing Unit).

[0041] See also Figure 3 , a task allocation method provided in one embodiment of the present application can be applied to an instruction processor, and the method can include the following steps.

[0042] S1: Receive the target task to be assigned.

[0043] S3: Allocate the target task to a target processing core among a plurality of preset processing cores; wherein, the target data required for executing the target task is pre-stored in a memory adjacent to the target processing core, and when the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

[0044] In this embodiment, the tasks executed within the kernel function can be issued to the instruction processor, which then distributes them among a plurality of pre-defined processing cores. To ensure that tasks and required data are allocated to adjacent processing cores and memories, the matching relationship between tasks and data can be predefined when the target application is written. The target application can be any application that requires data processing on a processing core. During runtime, the application can execute the various tasks within the application by calling one or more kernel functions.

[0045] In this embodiment, before calling the kernel function, the data required for each task can be written into the memory in advance according to certain rules. Specifically, taking the target task to be assigned as an example, the target data required to execute the target task can be written into a memory in advance. On the integrated chip, the memory and the processing core can be determined to be adjacent based on their respective locations. If the processing core can read data directly from the memory, then the processing core and the memory can be considered to have an adjacent relationship; and when the processing core tries to read data from the memory, if it needs to obtain the required data through another processing core in accordance with the communication method across the processing cores, then the processing core and the memory can be considered to have no adjacent relationship. For example, in Figure 2 In the diagram, the memory on the left and the processing core have a neighboring relationship, and the memory on the right and the processing core also have a neighboring relationship, but the memory on the left and the processing core on the right do not have a neighboring relationship. Similarly, the memory on the right and the processing core on the left do not have a neighboring relationship.

[0046] In this embodiment, after the target data is written to the memory, the memory location of the target data can be recorded, and the target processing core that is adjacent to the memory can be determined. Subsequently, when the instruction server assigns a target task, it can assign the target task to the target processing core. In this way, when executing the target task, the target processing core can directly read the target data from the adjacent memory, thus avoiding cross-processing core communication.

[0047] As can be seen, in this embodiment, the data required to execute the task is pre-stored in the memory. When assigning a target task, the instruction processor can determine the memory where the target data required to execute the target task is currently located. When assigning a target task, the instruction processor can assign the target task to a target processing core that is adjacent to the memory where the target data is located. In this way, when the target processing core subsequently executes the target task, it can directly obtain the target data from the adjacent memory. This task allocation method can avoid communication across processing cores and will not be limited by the bottleneck of communication bandwidth between processing cores, thereby achieving higher task processing efficiency.

[0048] In one embodiment, when assigning tasks, the instruction processor can complete the assignment process by querying the relationship between the task identifier and the core sequence number. Specifically, taking the target task as an example, if the target data required by the target task is stored in a certain memory, the core sequence number of the processing core that is adjacent to the memory can be determined. To avoid cross-processing core communication, when assigning the target task, the target task should be assigned to the processing core with the core sequence number. Therefore, the task identifier of the target task and the core sequence number can be pre-associated. When the target application is written, it is already clear which tasks to be executed in the target application and which memory or memories the data required by each task can be stored in. Therefore, for each task to be executed, an association between the task identifier and the core sequence number can be established in the aforementioned manner. This association can be stored in the instruction processor as a key-value pair. Subsequently, when the instruction processor receives the target task to be assigned, it can identify the task identifier of the target task and query the core sequence number associated with the task identifier, ultimately assigning the target task to the target processing core with the core sequence number.

[0049] In this embodiment, when assigning a target task, the instruction processor can quickly determine the core number associated with the target task's task identifier based on the association between the task identifier and the core number, and then assign the target task to the processing core with that core number. This assignment method can quickly complete task assignment and improve the overall efficiency of subsequent task processing.

[0050] In one embodiment, this association relationship between the task identifier and the core serial number can be summarized as association information, which can be, for example, the information stored in the form of key-value pairs as described above. By traversing the various tasks to be executed in the target application, association relationships between multiple groups of task identifiers and core serial numbers can be established, and then such association relationships can be summarized in the form of formatted data to form association information. This association information can be stored in the instruction processor, and when the instruction processor queries the core serial number associated with the task identifier, it can read the pre-generated association information and query the core serial number associated with the task identifier of the target task in the association information.

[0051] In this embodiment, the association relationship between the task identifier and the core serial number can be represented by pre-generated association information. Task allocation according to the association information can ensure that the processing core where the task is located and the memory where the data is located have a proximity relationship, thereby improving the efficiency of task allocation and processing.

[0052] See also Figure 4 In one embodiment, the above-mentioned association information can be as follows: Figure 4 The following method is pre-generated as shown.

[0053] S41: Obtain a data processing comparison table in the target application, where the data processing comparison table is used to represent the matching relationship between tasks and data during the execution of the target application.

[0054] In this embodiment, when the target application is written, the matching relationship between tasks and data can be predetermined. This matching relationship can be used to represent the data required to execute the task. The matching relationship between tasks and data can be represented using a data processing comparison table. When generating association information, this data processing comparison table can be obtained first to clarify the matching relationship between tasks and data.

[0055] S43: For any set of matching tasks and data in the data processing comparison table, extract the task identifier of the task and determine a memory for storing the data.

[0056] S45: Identify a processing core adjacent to the memory, and obtain a core serial number of the processing core.

[0057] S47: Constructing an association relationship between the task identifier and the core serial number, and generating association information based on each set of constructed association relationships.

[0058] In this embodiment, for any set of matching tasks and data, the task identifier of the task can be first extracted, and then the memory used to store the matching data can be determined. To establish an association between the task identifier and the core number, a processing core that is adjacent to the memory can be identified and the core number of the processing core can be obtained. Then, an association between the task identifier and the core number can be established. For each set of matching tasks and data, a corresponding association relationship can be established. Ultimately, each set of association relationships established can constitute association information.

[0059] In this embodiment, when the target application is written, the data required for the task to be executed can be determined, and the data can be limited to which memory or memories the data is stored in. In this way, after the storage location of the data is clarified, it is only necessary to limit the subsequent allocation of tasks to be executed to the processing core with a neighboring relationship, so as to ensure that the processing core reads the required data directly from the memory with a neighboring relationship when processing the task. In other words, when the target application is running, the matching relationship between the task and the data has been clarified. For such a matching relationship, the memory where the data is located can be determined first, and then the processing core with a neighboring relationship with the memory can be determined. By associating the task identifier of the task with the core serial number of the processing core, the association information can be constructed. Subsequently, when tasks are assigned according to this association information, it can be ensured that the processing core where the task is located and the memory where the data is located have a neighboring relationship, thereby improving the efficiency of task allocation and processing.

[0060] In one embodiment, considering that when writing data to a memory, it is necessary to request corresponding memory space from the memory, in order to ensure that the task to be assigned and the required data are located in adjacent processing cores and memory, respectively, after obtaining the core sequence number of the processing core in step S45, a memory request carrying the core sequence number can be initiated for the data, so that the memory management module responds to the memory request and allocates memory for storing the data in the memory adjacent to the core sequence number.

[0061] In this embodiment, a memory request for data can include a core number. Upon receiving the memory request, the memory management module can identify the core number and allocate the corresponding memory in a memory adjacent to the core number. This ensures that data and tasks are placed in adjacent memory and processing cores, respectively, thereby improving overall task processing efficiency.

[0062] In one embodiment, when assigning a target task to a target processing core among a plurality of preset processing cores, the memory in which target data required for executing the target task is currently located can be determined. Then, a target processing core that is adjacent to the memory can be identified, and the target task can be assigned to the target processing core.

[0063] In this embodiment, when assigning a target task, the instruction processor can assign the target task to a target processing core that is adjacent to the memory where the target data resides. This allows the target processing core to directly retrieve the target data from the adjacent memory when subsequently executing the target task. This task assignment method avoids inter-processing core communication and is not limited by the communication bandwidth bottleneck between processing cores, thereby achieving higher task processing efficiency.

[0064] See also Figure 5 In one specific application example, the plurality of pre-defined processing cores include a first processing core and a second processing core. A first memory is adjacent to the first processing core, and a second memory is adjacent to the second processing core. Data required for tasks with even-numbered task identifiers is stored in the first memory, and data required for tasks with odd-numbered task identifiers is stored in the second memory. This allows subsequent instruction processors to allocate tasks based on the parity of the task identifiers, ensuring that the processing cores can directly read the required data from the adjacent memory when executing tasks.

[0065] In this application example, when the target task is assigned to a target processing core among a preset plurality of processing cores, the task identifier of the target task can be identified. When the task identifier is an even-numbered task identifier, the target task can be assigned to the first processing core; when the task identifier is an odd-numbered task identifier, the target task can be assigned to the second processing core.

[0066] In this embodiment, tasks can be assigned based on the parity of the task identifiers, where even-numbered task identifiers can be assigned to the first processing core, and odd-numbered task identifiers can be assigned to the second processing core. To ensure that task assignment in this manner can avoid communication across processing cores, when storing each task, the data required for tasks with even-numbered task identifiers can be stored in the first memory, and the data required for tasks with odd-numbered task identifiers can be stored in the second memory. In this way, when the processing core subsequently executes a task, it can directly read the required data from the adjacent memory, thereby improving the overall efficiency of task processing.

[0067] In one embodiment, to be compatible with existing task allocation mechanisms, an optimized allocation function can be pre-set. When this optimized allocation function is enabled, task allocation is performed according to the logic of avoiding cross-processing core communication. If this optimized allocation function is not enabled, task allocation is performed according to the existing task allocation mechanism.

[0068] Specifically, after receiving the target task to be assigned, the instruction processor may determine whether a preset optimization allocation function is enabled. Only when the optimization allocation function is enabled will the subsequent step of allocating the target task to the target processing core be executed.

[0069] In practice, the optimized allocation feature can be enabled and disabled using a pre-set function field. For example, this field could be a select field. When the select field is assigned a value of 1, the optimized allocation feature is enabled, and the instruction processor will allocate tasks according to logic that avoids cross-processing core communication. If the select field is assigned a value of 0, the optimized allocation feature is disabled, and tasks are allocated according to the existing task allocation mechanism.

[0070] In this embodiment, an optimized allocation function can be pre-set. Only when this function is enabled will the designated processor allocate the target task to the target processing core according to the logic of avoiding cross-processing core communication. If this function is not enabled, task allocation can be carried out according to the conventional task allocation method. This measure improves the overall compatibility of the system, is compatible with existing task allocation mechanisms, and ensures the stability of the task processing process.

[0071] See also Figure 6 On the other hand, the present application further provides an instruction processor, the instruction processor comprising:

[0072] A receiving unit 100 is configured to receive a target task to be assigned;

[0073] The allocation unit 200 is used to allocate the target task to a target processing core among a plurality of preset processing cores; wherein, the target data required for executing the target task is pre-stored in a memory adjacent to the target processing core, and when the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

[0074] On the other hand, the present application also provides an electronic device, including: a memory storing computer instructions; and at least one processor configured to execute the computer instructions in the memory to perform the task allocation method in the above embodiment.

[0075] On the other hand, the present application further provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the task allocation method in the above embodiment.

[0076] On the other hand, the present application further provides a computer program product, including computer instructions, which, when executed by a processor, enable the processor to execute the task allocation method in the above embodiment.

[0077] See also Figure 7 , Figure 7 is a structural diagram of an electronic device provided by an optional embodiment of the present invention, such as Figure 7 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.

[0078] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0079] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0080] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0081] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0082] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.

[0083] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0084] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0085] The present application is described with reference to the flowcharts and / or block diagrams of the methods and systems according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0086] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0088] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0089] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0090] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

[0091] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.

Claims

1. A task allocation method, characterized in that: The method is applied to an instruction processor, and the method includes: receiving a target task to be assigned, and assigning the target task to a target processing core among a plurality of preset processing cores; The target data required to execute the target task is pre-stored in a memory adjacent to the target processing core. When the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

2. The method according to claim 1, characterized in that Allocating the target task to a target processing core among the preset multiple processing cores includes: Identify the task identifier of the target task and query the core sequence number associated with the task identifier; Allocate the target task to the target processing core with the core sequence number.

3. The method according to claim 2, characterized in that Querying the core sequence number associated with the task identifier includes: Reading pre-generated association information, wherein the association information is used to represent the association relationship between the task identifier and the core sequence number; The core sequence number associated with the task identifier of the target task is searched in the association information.

4. The method according to claim 3, characterized in that The association information is pre-generated in the following manner: Obtaining a data processing comparison table in the target application, wherein the data processing comparison table is used to represent the matching relationship between tasks and data during the execution of the target application; For any set of matching tasks and data in the data processing comparison table, extract the task identifier of the task and determine a memory for storing the data; Identifying a processing core adjacent to the memory and obtaining a core serial number of the processing core; An association relationship between the task identifier and the core serial number is constructed, and association information is generated based on each set of constructed association relationships.

5. The method according to claim 4, characterized in that After obtaining the core serial number of the processing core, the method further includes: A memory application carrying the core sequence number is initiated for the data, so that the memory management module responds to the memory application and applies for memory for storing the data in the memory that is adjacent to the core sequence number.

6. The method according to claim 1, characterized in that Allocating the target task to a target processing core among the preset multiple processing cores includes: Determine the memory where the target data required to perform the target task is currently located; A target processing core that is adjacent to the memory is identified, and the target task is allocated to the target processing core.

7. The method according to claim 1, characterized in that The plurality of preset processing cores include a first processing core and a second processing core, the first memory and the first processing core are in a neighboring relationship, and the second memory and the second processing core are in a neighboring relationship; wherein data required for tasks with even-numbered task identifiers is stored in the first memory, and data required for tasks with odd-numbered task identifiers is stored in the second memory; Allocating the target task to a target processing core among the preset multiple processing cores includes: Identify the task identifier of the target task, and when the task identifier is an even-numbered task identifier, allocate the target task to the first processing core; when the task identifier is an odd-numbered task identifier, allocate the target task to the second processing core.

8. The method according to any one of claims 1 to 7, characterized in that After receiving the target task to be assigned, the method further includes: It is determined whether a preset optimization allocation function is enabled, and if the optimization allocation function is enabled, the target task is allocated to the target processing core.

9. An instruction processor, characterized in that The instruction processor includes: A receiving unit, configured to receive target tasks to be assigned; an allocating unit, configured to allocate the target task to a target processing core among a plurality of preset processing cores; The target data required to execute the target task is pre-stored in a memory adjacent to the target processing core. When the target processing core obtains the target data from the memory, there is no need to execute a communication process across processing cores.

10. An electronic device, characterized in that: include: a memory storing computer instructions; At least one processor is configured to execute the computer instructions in the memory to perform the method according to any one of claims 1-8.

11. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.

12. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method and device

    CN117519957A

  • Task allocation method and device, electronic equipment and storage medium

    CN118312311A

  • Cache system, cache processing method, electronic equipment and storage medium

    CN119271574A

  • Apparatus and method for allocating a task

    US20130024868A1

  • A processing unit that enables asyncronous task dispatch

    WO2011028986A2