Network-on-chip mapping method and device, network processor and readable storage medium
By combining static mapping and dynamic mapping, an accurate static mapping solution is built and dynamic selection is solved, the problem of inability to take into account both response speed and accuracy in the existing technology is solved, and the mapping effect is achieved with fast response and high-precision, which improves the performance of network processors.
Patent Information
- Application Number
- CN202311491392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-13
AI Technical Summary
The existing on-chip network mapping methods cannot take into account both mapping response speed and mapping accuracy, and cannot meet the requirements of network processors that are both fast response speed and high accuracy.
By combining static mapping with dynamic mapping, an accurate static mapping scheme is built using static mapping to ensure mapping accuracy, and then dynamically select the mapping area through dynamic mapping to achieve rapid response.
It realizes the effect of fast response and maintaining high mapping accuracy in network processors, effectively reducing network congestion and energy consumption of on-chip networks and improving the performance of network processors.
Smart Images

Figure CN119988304A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network communications, and in particular to an on-chip network mapping method, device, network processor and readable storage medium. Background Art
[0002] With the rapid development of the Internet, users have higher and higher requirements for the quality of network services, which also puts higher requirements on the performance of network processors. As a multi-processor system-on-chip (MPSoC), network processors usually use network-on-chip (NoC) to achieve high-speed communication between multiple processing cores.
[0003] For network processors, reducing network congestion and energy consumption of on-chip networks can effectively improve the performance of network processors. The key to reducing network congestion and energy consumption of on-chip networks lies in the design of mapping schemes between application tasks and processing cores. An efficient mapping scheme can effectively reduce communication costs and avoid resource contention for data transmission by reasonably allocating application tasks to processing units, thereby reducing network congestion and energy consumption. At present, the mapping technologies of on-chip networks mainly include dynamic mapping and static mapping. In static mapping, by constructing an on-chip network congestion assessment model and an energy consumption model, the model is integrated into the algorithm for solving to find a mapping scheme with good energy saving and network performance; in dynamic mapping, the method of neighbor allocation of communication tasks is often considered to reduce the communication cost as much as possible.
[0004] However, most of the existing research on on-chip network mapping problems only focuses on static mapping or dynamic mapping. Static mapping has high solution accuracy but long solution time, while dynamic mapping often has fast calculation time but low solution accuracy. With the rapid development of the Internet, for real-time systems such as network processors, the existing on-chip network mapping methods cannot meet the requirements of network processors for both fast mapping response speed and high mapping accuracy. Summary of the invention
[0005] The present application provides a network-on-chip mapping method, device, network processor and readable storage medium to solve the technical problem that the existing network-on-chip mapping method cannot take into account both mapping response speed and mapping accuracy.
[0006] According to a first aspect disclosed in the present application, the present application provides an on-chip network mapping method, comprising:
[0007] Obtaining a target mapping request; wherein the target mapping request carries an application task of a target application;
[0008] Acquire a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes a plurality of static mapping schemes, the static mapping schemes include a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to the target application in the static mapping library;
[0009] Performing a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies the target mapping scheme, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor;
[0010] If a continuous mapping area that satisfies the target mapping scheme is found, the target mapping scheme is matched and mapped with the continuous mapping area.
[0011] In a feasible implementation manner, obtaining a target mapping request includes:
[0012] receiving an application mapping request; wherein the application mapping request carries an application task of a program application;
[0013] Adding the application mapping requests to a mapping request queue in the order in which they are received;
[0014] The application mapping request at the head of the mapping request queue is extracted to obtain the target mapping request.
[0015] In a feasible implementation manner, before performing a region search on a processing core in the network processor, the method further includes:
[0016] Acquire a first number of processing cores that the target application needs to occupy based on the target mapping scheme;
[0017] Obtaining a second number of idle processing cores in the network processor;
[0018] If the first number is less than the second number, performing a step of performing a region search on the processing cores in the network processor;
[0019] If the first number is not less than the second number, then after a preset period, the process jumps to the step of executing the step of obtaining the second number of idle processing cores in the network processor.
[0020] In a feasible implementation manner, the method further includes:
[0021] If a continuous mapping area satisfying the static mapping scheme cannot be found, obtaining a third number of idle central processing cores in the network processor and a fourth number of central processing cores in the network processor;
[0022] If the ratio of the third number to the fourth number is less than a preset threshold, obtaining a running application currently running on the network processor;
[0023] According to the preset migration priority of each running application, the running applications are sorted to obtain a migration queue;
[0024] Acquire an application to be migrated; wherein the initial application to be migrated is a running application at the head of the migration queue;
[0025] Migrating the application to be migrated to a migration area, and performing a step of performing a regional search on the processing cores in the network processor; wherein the migration area includes a plurality of processing cores in the network processor, and the migration area is obtained by a preset optimization search algorithm;
[0026] If a continuous mapping area that satisfies the target mapping scheme is found, matching and mapping the target mapping scheme with the continuous mapping area;
[0027] If a continuous mapping area satisfying the static mapping solution cannot be found, the application to be migrated is updated according to the running application located after the migration application in the migration queue, and the process proceeds to the step of obtaining the application to be migrated.
[0028] In a feasible implementation, if the ratio of the third number to the fourth number is not less than a preset threshold, then after an application running on the network processor is completed, a step of performing a region search on a processing core in the network processor is performed.
[0029] In a feasible implementation manner, after matching and mapping the target mapping scheme with the continuous mapping area, the method further includes:
[0030] updating the state of the processing core in the network processor located in the continuous mapping area to an occupied state;
[0031] Moving the target mapping request from the mapping request queue to the application running queue, and executing the application task of the target application;
[0032] After the application task of the target application is executed, the target application request is removed from the run queue, and the state of the processing core in the network processor located in the continuous mapping area is updated to an idle state.
[0033] In a feasible implementation manner, the method further includes:
[0034] Acquire a program application set; wherein the program application set includes a plurality of program applications running on the network processor;
[0035] For each program application, a static mapping scheme corresponding to the program application is obtained based on a preset static mapping algorithm, and the static mapping scheme is added to the static mapping library.
[0036] According to a second aspect disclosed in the present application, the present application provides an on-chip network mapping device, comprising:
[0037] A request acquisition module, used to acquire a target mapping request; wherein the target mapping request carries an application task of a target application;
[0038] A scheme query module, used to obtain a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes a plurality of static mapping schemes, the static mapping schemes include a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to a target application in the static mapping library;
[0039] A region search module, used to perform a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies the target mapping scheme, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor;
[0040] The matching and mapping module is used to match and map the target mapping scheme with the continuous mapping area if a continuous mapping area satisfying the target mapping scheme is searched.
[0041] According to a third aspect disclosed in the present application, there is provided a network processor, comprising a processor, and a memory communicatively connected to the processor;
[0042] The memory stores computer-executable instructions;
[0043] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspects.
[0044] According to a fourth aspect disclosed in the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement any one of the methods described in the first aspect.
[0045] According to a fifth aspect disclosed in the present application, a computer program product is provided, including a computer program, wherein the computer program is used to implement any one of the methods in the first aspect when executed by a processor.
[0046] Compared with the prior art, this application has the following beneficial effects:
[0047] The present application provides an on-chip network mapping method, device, network processor and readable storage medium. By combining static mapping with dynamic mapping, a precise static mapping scheme is constructed using a static mapping method to ensure mapping accuracy, and then a mapping area is dynamically selected through a dynamic mapping method to achieve fast response, thereby meeting the on-chip network of the network processor's requirements for both fast mapping response speed and high mapping accuracy, effectively reducing network congestion and energy consumption of the on-chip network, and improving the performance of the network processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0049] Figure 1 A flowchart of an on-chip network mapping method provided in an embodiment of the present application;
[0050] Figure 2 A flowchart of another on-chip network mapping method provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of the structure of an on-chip network mapping device provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of the structure of a network processor provided in an embodiment of the present application.
[0053] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0054] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0055] With the rapid development of the Internet, users have higher and higher requirements for the quality of network services, which also puts higher requirements on the performance of network processors. As a multi-processor system-on-chip (MPSoC), network processors usually use network-on-chip (NoC) to achieve high-speed communication between multiple processing cores.
[0056] Among them, a multi-processor system-on-chip (MPSoC) is a system-on-chip that integrates multiple processing cores and other IP cores to achieve a high-performance, low-latency, and low-power multi-processor architecture. For the interaction between multiple cores in a multi-processor system-on-chip, the traditional communication method uses a bus interconnection structure. However, with the increasing demand for chip functions and the continuous advancement of semiconductor process technology, the complexity of multi-processor systems-on-chip is increasing. Multi-processor systems-on-chip based on bus interconnection structures face severe challenges in multi-core processor communication issues, mainly in terms of bandwidth, throughput limitations, scalability, reusability, signal integrity, clock synchronization, and power consumption.
[0057] Due to the increase in bus load and the decrease in network communication performance of the traditional bus interconnection structure, the bus structure is no longer suitable for interaction between multiple cores. Therefore, the Network-on-Chip (NoC) is proposed as a new technology for high-speed communication between multiple cores, drawing on the idea of computer networks to meet the growing communication needs of multi-processor systems on chip. The Network-on-Chip is a communication architecture for multi-processor systems on chip. It connects processing units, storage units and other IP cores based on the Network-on-Chip structure to support high-performance, low-latency and low-power communication.
[0058] For the on-chip network, reducing the on-chip network congestion and energy consumption is the focus of the on-chip network design. An efficient mapping scheme can effectively reduce the communication cost and avoid resource contention for data transmission by reasonably allocating application tasks to processing units, thereby reducing network congestion and energy consumption. At present, the mapping technology of on-chip networks mainly includes dynamic mapping and static mapping. In static mapping, by constructing an on-chip network congestion assessment model and an energy consumption model, the model is integrated into the algorithm for solving to find a mapping scheme with good energy saving and network performance; in dynamic mapping, the method of neighbor allocation of communication tasks is often considered to reduce the communication cost as much as possible.
[0059] Among them, static mapping can be subdivided into precise mapping and search / heuristic mapping. The algorithm constructs an energy consumption model and performs task sorting, mapping and voltage distribution in an integrated manner, explicitly reducing communication competition and communication energy consumption. Module optimality does not guarantee that each module combined is the optimal solution for the whole, so many researchers focus on swarm intelligence algorithms to solve combinatorial optimization problems such as mapping. Swarm intelligence algorithms are meta-heuristic algorithms, which are bionic intelligent algorithms based on group random search. They have the advantages of high efficiency, simple execution, high self-organization, and independence from problem characteristics. Compared with traditional methods, it can obtain near-optimal calculation results in a faster time and has achieved better results in solving NP (Nondeterministic Polynomially) problems. In addition, the update method of particle velocity is optimized by using the roulette selection method and adding perturbation particles.
[0060] Among them, dynamic mapping can map the incoming application tasks continuously or discontinuously. Continuous mapping stipulates that tasks must be assigned to physically connected processing cores in the on-chip network space. If there is no such mapping area, no mapping is performed, while discontinuous area mapping has no such mandatory constraints. Continuous mapping can reduce the communication cost of application tasks and reduce network congestion within the application, while ensuring that tasks between applications are spatially collocated, isolating their communication traffic, and reducing network congestion outside the application. Discontinuous mapping optimizes system utilization at the expense of communication. It does not need to wait for the system to have a continuous idle core area that meets the required number, but only requires a sufficient number of idle processing cores, reducing the time for tasks to wait for mapping. However, due to the increase in inter-task communication overhead, discontinuous mapping will affect the performance of the entire system execution process. When executing multiple application workloads, discontinuous allocation will not only reduce the execution efficiency of new incoming application tasks, but also reduce the execution efficiency of mapped tasks due to on-chip network congestion.
[0061] Common discontinuous mapping algorithms include NN (Neural Network) algorithm and BN (Batch Normalization) algorithm. NN algorithm first sorts application tasks from large to small according to the communication volume, and also sorts processing cores from large to small according to the number of their idle neighbor nodes. Then, the first task is assigned to the first processing core node, and the nearest neighbor nodes around the first node of the processing core are found and assigned to the remaining task nodes in turn. By giving priority to assigning tasks with large communication volume to neighbors, energy consumption and communication delay are reduced. The BN algorithm is similar to the NN algorithm, except that the link congestion is also considered when sorting processing core nodes. In addition to ensuring neighbor allocation, it also considers whether the neighbor allocation will be congested. Real-time program applications are mapped continuously, while non-real-time program applications are allocated discontinuously. The CASQA method adopts a discontinuous mapping heuristic method. Initially, a continuous mapping attempt is made. If there is not enough continuous unallocated processing core area, the idle processing cores near the remaining task allocation are mapped discontinuously.
[0062] However, most of the existing research on on-chip network mapping problems only focuses on static mapping or dynamic mapping. Static mapping has high solution accuracy but long solution time, while dynamic mapping often has fast calculation time but low solution accuracy. With the rapid development of the Internet, for real-time systems such as network processors, the existing on-chip network mapping methods cannot meet the requirements of network processors for both fast mapping response speed and high mapping accuracy.
[0063] In response to the above technical problems, the present application proposes an on-chip network mapping method, which combines static mapping with dynamic mapping, uses static mapping to construct an accurate static mapping scheme to ensure mapping accuracy, and then uses dynamic mapping to dynamically select the mapping area to achieve fast response, thereby meeting the requirements of the network processor for both fast mapping response speed and high mapping accuracy.
[0064] The technical solution of the on-chip network mapping method provided by the present application is described in detail below through specific embodiments. It should be noted that the following embodiments may exist independently or in combination with each other, and the same or similar contents may not be described repeatedly in different embodiments.
[0065] It should be noted that the execution subject of the on-chip network mapping method provided in the embodiment of the present application is a network processor, and accordingly, the on-chip network mapping device is also arranged in the network processor. Specifically, the network processor is a network processor based on a 2D Mesh (two-dimensional mesh structure) on-chip network.
[0066] Figure 1 A flowchart of a network-on-chip mapping method provided in an embodiment of the present application is shown in FIG. Figure 1In some embodiments, the process of the on-chip network mapping method includes the following steps:
[0067] S101, obtaining a target mapping request; wherein the target mapping request carries an application task of a target application.
[0068] The network processor receives a target mapping request from a target application, and based on the target mapping request, the network processor performs mapping matching of a processing core for an application task of the target application.
[0069] S102, obtaining a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes multiple static mapping schemes, the static mapping scheme includes a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to the target application in the static mapping library.
[0070] After receiving the target mapping request, the network processor first searches for a static mapping method corresponding to the target application from a pre-stored static mapping library according to the target mapping request.
[0071] Specifically, the static mapping library is a static mapping scheme corresponding to each program application obtained in advance through a preset static mapping algorithm according to the program application that needs to be run on the network processor. By building a static mapping library in advance, the mapping accuracy can be improved without affecting the subsequent response speed of the network processor.
[0072] S103, performing a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies a target mapping solution, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor.
[0073] After the static mapping solution is obtained, the processing cores in the network processor are searched for regions to find continuous mapping regions that meet the target mapping solution.
[0074] Preferably, in order to improve the response speed of the on-chip network mapping, the region search algorithm starts searching from the region with the largest number of idle processing cores in the network processor, and adopts a first-fit method. The first-fit method means that the search stops after a continuous mapping region that meets the conditions is found, so as to improve the search efficiency.
[0075] Preferably, in order to avoid the mapped target applications being concentrated in the network processor, which would cause the local temperature of the chip to be too high and lead to a decrease in the performance of the network processor, the regional search algorithm allocates processing cores in an overall manner from the periphery of the chip to the center of the chip, which is beneficial to the chip heat dissipation of the network processor.
[0076] S104: If a continuous mapping area that satisfies the target mapping solution is found, the target mapping solution is matched with the continuous mapping area.
[0077] Among them, if a continuous mapping area that meets the target mapping scheme is searched through the area search operation, then after obtaining the target mapping scheme and the continuous mapping area, the target mapping scheme and the continuous mapping area are matched and mapped. The purpose is to match the mapping relationship between the application task and the processing core in the target mapping scheme to the processing core of the continuous mapping area in the network processor.
[0078] Specifically, the continuous mapping area is a plurality of interconnected processing cores, which can reduce the communication cost of application tasks and reduce the network congestion within the application, while ensuring that the tasks between applications are spatially collocated, isolating their communication traffic and reducing the network congestion outside the application.
[0079] In this embodiment, static mapping is combined with dynamic mapping, an accurate static mapping scheme is constructed using the static mapping method to ensure mapping accuracy, and then the mapping area is dynamically selected through the dynamic mapping method to achieve fast response, thereby meeting the requirements of the network processor's on-chip network for both fast mapping response speed and high mapping accuracy, effectively reducing network congestion and energy consumption of the on-chip network, and improving the performance of the network processor.
[0080] exist Figure 1 Based on the embodiment shown, the following Figure 2 , the technical solution of the above-mentioned on-chip network mapping method is further introduced.
[0081] Figure 2 A flowchart of another on-chip network mapping method provided in an embodiment of the present application is shown in FIG. Figure 2 In some embodiments, the process of the on-chip network mapping method includes the following steps:
[0082] S201, obtaining a program application set; wherein the program application set includes a plurality of program applications running on a network processor.
[0083] Among them, because the program applications that the network processor will run are known, the program applications that will run on the network processor can be collected in advance.
[0084] S202: for each program application, obtain a static mapping solution corresponding to the program application based on a preset static mapping algorithm, and add the static mapping solution to a static mapping library.
[0085] Among them, through the preset static mapping algorithm, the program application can be mapped offline, and the static mapping library can be prepared in advance. In addition, if the network processor subsequently adds a new program application to be run, a static mapping scheme for the new program application can be constructed based on this method before the program application is run, and the static mapping scheme for the new program application can also be added to the static mapping library.
[0086] Specifically, for the static mapping algorithm, an existing static mapping algorithm based on a swarm intelligence algorithm may be used, wherein the swarm intelligence algorithm may be a bat optimization algorithm, a whale optimization algorithm, and the like.
[0087] S203, receiving an application mapping request; wherein the application mapping request carries an application task of the program application.
[0088] S204: Add the application mapping requests to a mapping request queue in the order in which they are received.
[0089] S205, extracting the application mapping request at the head of the mapping request queue to obtain a target mapping request.
[0090] The received application mapping requests are processed sequentially based on the order of precedence to ensure the timeliness of processing each application mapping request.
[0091] S206, obtaining a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes multiple static mapping schemes, the static mapping scheme includes a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to the target application in the static mapping library.
[0092] S207: Obtain a first number of processing cores that the target application needs to occupy based on the target mapping solution.
[0093] Among them, since there is a mapping relationship between the application tasks of the program application and the processing cores in the static mapping scheme, the number of processing cores required to run the target application can be known based on the target mapping scheme.
[0094] S208: Obtain a second number of idle processing cores in the network processor.
[0095] Wherein, according to the running status of the processing cores in the network processor, the number of the processing cores in the network processor that are in an idle state is obtained.
[0096] S209, determining whether the first quantity is smaller than the second quantity.
[0097] S210: If the first number is not less than the second number, then after a preset period, jump to the step of executing obtaining the second number of idle processing cores in the network processor.
[0098] Among them, if the first number is not less than the second number, it means that the number of idle processing cores in the network processor does not meet the basic operating requirements of the target application. At this time, the second number is updated by periodically obtaining the number of idle processing cores in the network processor, and after the second number is updated, the relationship between the number of idle processing cores in the network processor and the number of processing cores required by the target application continues to be judged to monitor whether the network processor meets the basic requirements for the operation of the target application.
[0099] S211, if the first number is smaller than the second number, executing a step of performing a region search on the processing cores in the network processor.
[0100] If the first number is smaller than the second number, it means that the number of idle processing cores in the network processor is larger than the number of processing cores required by the target application. At this time, the network processor meets the basic operating requirements for the target application and can perform subsequent area search operations.
[0101] S212: Determine whether a continuous mapping area satisfying the static mapping solution is found.
[0102] S213, if a continuous mapping area satisfying the static mapping solution cannot be searched, obtaining a third number of idle central processing cores in the network processor and a fourth number of central processing cores in the network processor;
[0103] Among them, the running of program applications on the network processor will cause the fragmentation of network processing core resources, resulting in a slow mapping response speed. For this reason, a defragmentation algorithm based on threshold monitoring is designed, and the specific defragmentation method is implemented by application migration. Since continuous mapping response will cause the processing core to be idle but unusable, the existing method considers a combination of continuous mapping and discontinuous mapping. However, this method applies discontinuous mapping to non-real-time applications, and for network processors, most of them are real-time applications, so this method is not very applicable. In addition, it ignores the impact of discontinuous mapping on completed continuous mapping, so it is considered to defragment the processing core resources by borrowing the memory defragmentation method to optimize continuous mapping.
[0104] Specifically, since the network processor belongs to a multi-processor system on chip, in a multi-processor system on chip (MPSoC), the central processing unit (CPU) refers to the main core in the multi-core processor, which is mainly responsible for major computing tasks such as the operating system kernel, system-level tasks and applications.
[0105] S214, determining whether the ratio of the third quantity to the fourth quantity is less than a preset threshold.
[0106] Whether the ratio of the third number to the fourth number is less than a preset threshold is a trigger condition for determining whether to perform defragmentation, and its purpose is to determine whether the central processing core of the network processor is too fragmented.
[0107] S215, if the ratio of the third number to the fourth number is not less than the preset threshold, then after the application running on the network processor is completed, a step of performing a region search on the processing cores in the network processor is performed.
[0108] If the ratio of the third number to the fourth number is not less than a preset threshold, it indicates that the processing core resources of the network processor are not too fragmented and there is no need to trigger a defragmentation process.
[0109] S216: If the ratio of the third number to the fourth number is less than a preset threshold, obtain the running application currently running on the network processor.
[0110] If the ratio of the third number to the fourth number is less than a preset threshold, it indicates that the processing core resources of the network processor are too fragmented, and a defragmentation process needs to be triggered to integrate the scattered processing core resources together to meet the conditions of the regional search.
[0111] S217: Sort the running applications according to the preset migration priority of each running application to obtain a migration queue.
[0112] The priority of application migration is obtained by presetting, and its purpose is to reduce the cost of processing core defragmentation by prioritizing different running applications.
[0113] S218, obtaining an application to be migrated; wherein the initial application to be migrated is a running application at the head of the migration queue.
[0114] The applications to be migrated in the migration queue are migrated in sequence according to the migration priorities of the applications to be migrated.
[0115] S219, migrate the application to be migrated to the migration area, and perform a step of performing a regional search on the processing cores in the network processor; wherein the migration area includes multiple processing cores in the network processor, and the migration area is obtained by a preset optimization search algorithm.
[0116] The migration area of the application to be migrated is obtained by an optimization search algorithm. Specifically, the optimization search algorithm determines the migration location of the application to be migrated by searching in the minimum migration direction cost area. In addition, the migration direction of the application is set to the reverse process of the mapping, that is, migrating from the center of the processing core of the network processor to the surroundings, so that the processing core fragments are moved to the central area of the network processor.
[0117] S220: Determine whether a continuous mapping area satisfying the static mapping solution is found.
[0118] S221, if a continuous mapping area satisfying the static mapping solution cannot be found, the application to be migrated is updated according to the running application located after the migration application in the migration queue, and the process proceeds to the step of obtaining the application to be migrated.
[0119] S222: If a continuous mapping area that satisfies the target mapping solution is found, the target mapping solution is matched with the continuous mapping area.
[0120] S223, updating the status of the processing cores in the network processor located in the continuous mapping area to an occupied status.
[0121] After the target mapping method matches and maps the continuous mapping area, the processing core of the continuous mapping area will subsequently execute the application task of the target application, so its state is updated and changed to an occupied state.
[0122] S224, moving the target mapping request from the mapping request queue to the application running queue, and executing the application task of the target application.
[0123] After the matching mapping between the target mapping scheme and the continuous mapping area is completed, the target mapping request is added to the queue of application running, and the target application is run.
[0124] S225, when the application task of the target application is executed, the target application request is removed from the running queue, and the state of the processing core located in the continuous mapping area of the network processor is updated to an idle state.
[0125] After the application task of the target application is executed, the processing core resources are released, and the status of the processing cores in the occupied continuous mapping area is updated to an idle state.
[0126] In this embodiment, by adding a defragmentation process, the present application can integrate the scattered processing core resources together to meet the needs of continuous mapping.
[0127] Figure 3 is a schematic diagram of a structure of an on-chip network mapping device provided in an embodiment of the present application, see Figure 3 The on-chip network mapping device includes various functional modules for implementing the above-mentioned on-chip network mapping method, and any functional module can be implemented by software and / or hardware.
[0128] In some embodiments, the on-chip network mapping device 300 includes a request acquisition module 301, a solution query module 302, a region search module 303 and a matching mapping module 304. Among them:
[0129] The request acquisition module 301 is used to acquire a target mapping request; wherein the target mapping request carries an application task of a target application;
[0130] The solution query module 302 is used to obtain a target mapping solution from a pre-stored static mapping library; wherein the static mapping library includes multiple static mapping solutions, the static mapping solution includes a mapping relationship between application tasks of a program application and a processing core, and the target mapping solution is a static mapping solution corresponding to the target application in the static mapping library;
[0131] The region search module 303 is used to perform a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies the target mapping scheme, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor;
[0132] The matching and mapping module 304 is used to match and map the target mapping scheme with the continuous mapping area if a continuous mapping area that satisfies the target mapping scheme is found.
[0133] In some embodiments, the request acquisition module 301 is specifically used to:
[0134] receiving an application mapping request; wherein the application mapping request carries an application task of a program application;
[0135] Add application mapping requests to the mapping request queue in the order in which they are received;
[0136] The application mapping request at the head of the mapping request queue is extracted to obtain the target mapping request.
[0137] In some embodiments, the area search module 303 is further used to:
[0138] Acquire a first number of processing cores that the target application needs to occupy based on the target mapping scheme;
[0139] Obtaining a second number of idle processing cores in the network processor;
[0140] If the first number is less than the second number, performing a step of performing a region search on the processing cores in the network processor;
[0141] If the first number is not less than the second number, then after a preset period, the process jumps to the step of executing the step of obtaining the second number of idle processing cores in the network processor.
[0142] In some embodiments, the device further includes a defragmentation module 305, and the defragmentation module 305 is specifically used to:
[0143] If a continuous mapping area satisfying the static mapping scheme cannot be searched, obtaining a third number of idle central processing cores in the network processor and a fourth number of central processing cores in the network processor;
[0144] If the ratio of the third number to the fourth number is less than a preset threshold, obtaining a running application currently running on the network processor;
[0145] According to the preset migration priority of each running application, the running applications are sorted to obtain a migration queue;
[0146] Obtaining the application to be migrated; wherein the initial application to be migrated is the running application at the head of the migration queue;
[0147] Migrating the application to be migrated to a migration area, and performing a step of performing a regional search on the processing cores in the network processor; wherein the migration area includes a plurality of processing cores in the network processor, and the migration area is obtained by a preset optimization search algorithm;
[0148] If a continuous mapping area that satisfies the target mapping scheme is found, the target mapping scheme is matched with the continuous mapping area;
[0149] If a continuous mapping area satisfying the static mapping solution cannot be found, the application to be migrated is updated according to the running application located after the migration application in the migration queue, and the process proceeds to the step of obtaining the application to be migrated.
[0150] In some embodiments, the defragmentation module 305 is further configured to:
[0151] If the ratio of the third number to the fourth number is not less than the preset threshold, then after the application running on the network processor is completed, a step of performing a region search on the processing core in the network processor is performed.
[0152] In some embodiments, the device further includes a task execution module 306, and the task execution module 306 is specifically used to:
[0153] updating the status of the processing cores in the network processor located in the continuous mapping area to an occupied state;
[0154] Move the target mapping request from the mapping request queue to the application run queue, and execute the application task of the target application;
[0155] After the application task of the target application is executed, the target application request is removed from the running queue, and the state of the processing core located in the continuous mapping area of the network processor is updated to an idle state.
[0156] In some embodiments, the apparatus further includes a mapping library construction module 307, and the mapping library construction module 307 is specifically used to:
[0157] Acquire a program application set; wherein the program application set includes a plurality of program applications running on a network processor;
[0158] For each program application, a static mapping solution corresponding to the program application is obtained based on a preset static mapping algorithm, and the static mapping solution is added to a static mapping library.
[0159] The matching and mapping device 300 provided in the embodiment of the present application is used to execute the technical solution provided in the aforementioned matching and mapping method embodiment. Its implementation principle and technical effect are similar to those in the aforementioned method embodiment, and will not be repeated here.
[0160] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, all or part of them can be integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements, or all in the form of hardware, or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the request acquisition module can be a separately established processing element, or it can be integrated in a certain chip of the above device for implementation. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.
[0161] Figure 4 A schematic diagram of the structure of a network processor provided in an embodiment of the present application is shown in FIG. Figure 4 , the network processor 400 includes: a processor 401, and a memory 402 communicatively connected to the processor 401;
[0162] Memory 402 stores computer executable instructions;
[0163] The processor 401 executes the computer-executable instructions stored in the memory 402 to implement the technical solution of the aforementioned matching mapping method.
[0164] In the above network processor 400, the memory 402 and the processor 401 are electrically connected directly or indirectly to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines, such as through a bus connection. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus. The memory 402 stores computer execution instructions for implementing the above matching mapping method, including at least one software function module that can be stored in the memory 402 in the form of software or firmware. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402.
[0165] The memory 402 includes at least one type of readable storage medium, including but not limited to random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable read-only memory (EEPROM). The memory 402 is used to store programs, and the processor 401 executes the programs after receiving the execution instruction. Furthermore, the software programs and modules in the above-mentioned memory 402 may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide an operating environment for other software components.
[0166] The processor 401 may be an integrated circuit chip having the ability to process signals. The processor 401 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc. The methods, steps, and logic diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor, or the processor 401 may be any conventional processor, etc.
[0167] The network processor 400 is used to execute the technical solution provided by the aforementioned matching mapping method embodiment. Its implementation principle and technical effect are similar to those in the aforementioned method embodiment and will not be repeated here.
[0168] The embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the technical solution of the aforementioned matching mapping method is implemented.
[0169] The computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general or special purpose computer.
[0170] An exemplary readable storage medium is coupled to the processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the control device of the network processor.
[0171] An embodiment of the present application also provides a computer program product, including a computer program, which is used to implement the technical solution of the aforementioned matching mapping method when executed by a processor.
[0172] In the above embodiments, it will be appreciated by those skilled in the art that the above-mentioned various method embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer program instructions on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless network, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).
[0173] In the above embodiments, the description of each embodiment has its own emphasis. For the part not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0174] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0175] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A network-on-chip mapping method, characterized in that: include: Obtaining a target mapping request; wherein the target mapping request carries an application task of a target application; Acquire a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes a plurality of static mapping schemes, the static mapping schemes include a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to the target application in the static mapping library; Performing a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies the target mapping scheme, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor; If a continuous mapping area that satisfies the target mapping scheme is found, the target mapping scheme is matched and mapped with the continuous mapping area.
2. The method according to claim 1, characterized in that Get target mapping request, including: receiving an application mapping request; wherein the application mapping request carries an application task of a program application; Adding the application mapping requests to a mapping request queue in the order in which they are received; The application mapping request at the head of the mapping request queue is extracted to obtain the target mapping request.
3. The method according to claim 1, characterized in that Before performing a region search on the processing cores in the network processor, the method further includes: Acquire a first number of processing cores that the target application needs to occupy based on the target mapping scheme; Obtaining a second number of idle processing cores in the network processor; If the first number is less than the second number, performing a step of performing a region search on the processing cores in the network processor; If the first number is not less than the second number, then after a preset period, the process jumps to the step of executing the step of obtaining the second number of idle processing cores in the network processor.
4. The method according to claim 1, characterized in that: The method further comprises: If a continuous mapping area satisfying the static mapping scheme cannot be found, obtaining a third number of idle central processing cores in the network processor and a fourth number of central processing cores in the network processor; If the ratio of the third number to the fourth number is less than a preset threshold, obtaining a running application currently running on the network processor; According to the preset migration priority of each running application, the running applications are sorted to obtain a migration queue; Acquire an application to be migrated; wherein the initial application to be migrated is a running application at the head of the migration queue; Migrating the application to be migrated to a migration area, and performing a step of performing a regional search on the processing cores in the network processor; wherein the migration area includes a plurality of processing cores in the network processor, and the migration area is obtained by a preset optimization search algorithm; If a continuous mapping area that satisfies the target mapping scheme is found, matching and mapping the target mapping scheme with the continuous mapping area; If a continuous mapping area satisfying the static mapping solution cannot be found, the application to be migrated is updated according to the running application located after the migration application in the migration queue, and the process proceeds to the step of obtaining the application to be migrated.
5. The method according to claim 4, characterized in that The method further comprises: If the ratio of the third number to the fourth number is not less than a preset threshold, then after the application running on the network processor is completed, a step of performing a region search on the processing cores in the network processor is performed.
6. The method according to claim 1, characterized in that After matching and mapping the target mapping scheme with the continuous mapping area, the method further includes: updating the state of the processing core in the network processor located in the continuous mapping area to an occupied state; Moving the target mapping request from the mapping request queue to the application running queue, and executing the application task of the target application; After the application task of the target application is executed, the target application request is removed from the run queue, and the state of the processing core in the network processor located in the continuous mapping area is updated to an idle state.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Acquire a program application set; wherein the program application set includes a plurality of program applications running on the network processor; For each program application, a static mapping scheme corresponding to the program application is obtained based on a preset static mapping algorithm, and the static mapping scheme is added to the static mapping library.
8. A network-on-chip mapping device, characterized in that: include: A request acquisition module, used to acquire a target mapping request; wherein the target mapping request carries an application task of a target application; A scheme query module, used to obtain a target mapping scheme from a pre-stored static mapping library; wherein the static mapping library includes a plurality of static mapping schemes, the static mapping schemes include a mapping relationship between application tasks of a program application and a processing core, and the target mapping scheme is a static mapping scheme corresponding to a target application in the static mapping library; A region search module, used to perform a region search on the processing cores in the network processor; wherein the region search is used to search for a continuous mapping region that satisfies the target mapping scheme, and the continuous mapping region includes a plurality of idle and interconnected processing cores in the network processor; The matching and mapping module is used to match and map the target mapping scheme with the continuous mapping area if a continuous mapping area satisfying the target mapping scheme is searched.
9. A network processor, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.