Instruction processing apparatus, system, and instruction processing method
By setting up an arbitration module in the processor and allocating computing resources in the order of kernel priority, the problem of excessive hardware resource consumption in multi-core scenarios is solved, and the sharing of computing resources and the reduction of hardware costs is achieved.
Patent Information
- Application Number
- CN202510510289.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In multi-core scenarios, a large amount of hardware resources will be consumed during the V expansion process, and the area of large computing units even exceeds the kernel itself, resulting in a large overhead cost.
By setting up an arbitration module, authorization signals and kernel identifications are generated in sequence according to the kernel priority order, so that each processor core can access the resources of the computing processing unit in sequence, realizing computing resource sharing in multi-core mode.
There is no need to configure independent computing units for each core, which reduces hardware costs and improves resource utilization.
Smart Images

Figure CN120045230A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of processors, and specifically, to an instruction processing device, system, and instruction processing method. Background Art
[0002] With the rapid development of modern computing systems, vector processing capabilities are crucial for applications in fields such as high-performance computing (HPC), machine learning (ML), and cryptography. As an open-source instruction set architecture (ISA), RISC-V is used to support standard instruction set extensions for vector operations and has received wide attention due to its variable-length vectors, rich instruction set, modular design, and low-power consumption characteristics. The RISC-V architecture enhances its vector processing capabilities by defining the "V" extension, enabling the processor to efficiently execute complex vector operations.
[0003] Currently, although the "V" extension has powerful vector processing capabilities in related technologies, in a multi-core scenario, a large amount of hardware resources are consumed during the execution of the V extension. The area of large arithmetic units is relatively large, and in many cases, the area of large arithmetic units even exceeds that of the core itself. If each core is equipped with an independent arithmetic unit, it will increase the area of the processor core, thereby resulting in a relatively large overhead cost. Summary of the Invention
[0004] An instruction processing device, system, and instruction processing method are provided in an embodiment of this application.
[0005] In a first aspect of the embodiment of this application, an instruction processing device is provided. The instruction processing device includes: a plurality of processor cores, an instruction arithmetic module, an arbitration module, and a first multiplexer; the arbitration module is respectively connected to each processor core and the instruction arithmetic module, and the first multiplexer respectively establishes communication connections with each processor core and the instruction arithmetic module; Each of the processor cores is configured to: send an operation request signal to the arbitration module; the operation request signal is used to request access to the resources of the instruction arithmetic module; The arbitration module is configured to: receive the operation request signals of each of the processor cores, generate an authorization signal and a core identifier in the order of core priorities, and send them to the target core; the target core is the processor core that is granted the access permission to the resources of the instruction arithmetic module; The instruction arithmetic module is configured to: in response to a target operation request signal, perform an operation to obtain an operation result, and send the operation result and the core identifier to the first multiplexer; the target operation request signal is the operation request signal sent by the target core; The first multiplexer is configured to: send the operation result to the target core according to the core identifier.
[0006] In an optional embodiment of the present application, the arbitration module is specifically configured to: According to the kernel priority order and the polling mechanism, starting from the operation request signal with the highest priority among the operation request signals, poll each operation request signal in turn to determine the target operation request signal until all operation request signals have been traversed; Generate an authorization signal and a corresponding kernel identifier according to the target operation request signal and send them to the target kernel.
[0007] In an optional embodiment of the present application, the arbitration module includes: a counter and a second multiplexer; the counter is connected to the second multiplexer; The counter is configured to: count according to the operation request signal in clock cycles, generate a current output result and send it to the second multiplexer; the current output result is used to indicate the number of the operation request signal with the highest priority currently; the change of the current output result is used to implement the rotation of the kernel priority order; The second multiplexer is configured to: select the target operation request signal from all operation request signals according to the current output result; generate an authorization signal and a kernel identifier according to the target operation request signal and send them to the target kernel.
[0008] In an optional embodiment of the present application, the counter is further configured to: According to the operation request signal, starting from a preset initial value, update the count value according to the counting rule every time a clock cycle passes to obtain the corresponding current output result; and when the count value reaches the maximum value, return to the initial value and start counting again.
[0009] In an optional embodiment of the present application, the counter is further configured to: For the current clock cycle, use the operation request signal corresponding to the kernel with the highest priority in the current clock cycle among the processor kernels as the target operation request signal, and number the target operation request signal to obtain the current output result; When the next clock cycle arrives, the counter moves to the next operation request signal, increments the count value by one, and determines the next operation request signal with the highest priority in the next clock cycle according to the kernel priority order; the next operation request signal is the operation request signal whose priority in the current clock cycle is after the highest priority.
[0010] In an optional embodiment of the present application, the counter is further configured to: Generate a busy signal and send it to other cores, where the other cores are the remaining cores among all the processor cores except the target core, and the busy signal is used to indicate that the resources of the instruction operation module are currently occupied.
[0011] In an optional embodiment of the present application, the first multiplexer is further configured to: decode all cores according to the core identifier to determine the target core, and send the operation result to the target core.
[0012] In an optional embodiment of the present application, the bit width of the arbitration module is determined according to the number of cores, and the counter is represented in binary coding or Gray coding form.
[0013] In a second aspect of the embodiments of the present application, an instruction processing system is provided, including the instruction processing device provided in the above embodiment.
[0014] In a third aspect of the embodiments of the present application, an instruction processing method is provided, and the method includes: Receiving operation request signals sent by each of the processor cores; the operation request signals are used to request access to the resources of the instruction operation module; Generating an authorization signal and a core identifier in the order of core priorities and sending them to the target core; the target core is the processor core granted the access right to the resources of the instruction operation module; Responding to the target operation request signal, performing an operation to obtain an operation result; the target operation request signal is the operation request signal sent by the target core; Sending the operation result to the target core according to the core identifier.
[0015] In an embodiment of the present application, an instruction processing device, system, and instruction processing method are provided. The device includes: a plurality of processor cores, an instruction operation module, an arbitration module, and a first multiplexer. The arbitration module is respectively connected to each processor core and the instruction operation module, and the first multiplexer respectively establishes communication connections with each processor core and the instruction operation module. Each processor core is configured to: send an operation request signal to the arbitration module; the arbitration module is configured to: receive the operation request signals of each processor core, generate an authorization signal and a core identifier in the order of core priorities and send them to the target core; the instruction operation module is configured to: in response to the target operation request signal, execute an operation to obtain an operation result and send the operation result and the core identifier to the first multiplexer; the first multiplexer is configured to: according to the core identifier, send the operation result to the target core. Compared with the prior art, since the arbitration module is provided in the technical solution of the present application, the authorization signal and the core identifier can be generated in sequence according to the priority order to the corresponding target core, so that each processor core can access the resources of the operation processing unit in sequence to execute the operation, without configuring an independent operation unit for each core, realizing the sharing of computing resources in the multi-core mode, and then sending the operation result to the target core through the first multiplexer, greatly reducing the hardware cost. Description of the Drawings
[0016] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings: Figure 1 It is a schematic structural diagram of a computer device provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of an instruction processing device provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an instruction processing device provided by another embodiment of the present application; Figure 4 It is a schematic flowchart of an instruction processing method provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of an instruction processing device provided by another embodiment of the present application.
[0017] Description of the Reference Numerals: Processor core - 10; Instruction operation module - 20; Arbitration module - 30; Counter - 31; Second multiplexer - 32; First multiplexer - 40. Detailed Embodiments
[0018] In the process of implementing this application, the inventors found that traditional instruction processing requires each core to be equipped with an independent arithmetic unit, which increases the area of the processor core and thus leads to a relatively high overhead cost.
[0019] To make the technical solutions and advantages in the embodiments of this application clearer and more understandable, the following further details the exemplary embodiments of this application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0020] Based on the above-mentioned defects, this application provides an instruction processing device. Compared with the related art, since the arbitration module is provided in the technical solution of this application, authorization signals and core identifiers can be sequentially generated to the corresponding target cores in the order of priority, enabling each processor core to access the resources of the arithmetic processing unit in sequence to perform arithmetic operations, without the need to configure an independent arithmetic unit for each core, realizing the sharing of computing resources in the multi-core mode, and then sending the arithmetic result to the target core through the first multiplexer, greatly reducing the hardware cost.
[0021] Please refer to Figure 1 , a schematic structural diagram of a computer device provided by an embodiment of this application. As Figure 1 shown, the computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium can be, for example, a disk. Files (which can be files to be processed or processed files), an operating system, and computer programs are stored in the non-volatile storage medium. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements an instruction processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0022] Please refer to Figure 2 shown, Figure 2 is a schematic structural diagram of the instruction processing device provided by an embodiment of this application. Please refer to Figure 2As shown in the figure, the instruction processing device includes: a plurality of processor cores 10, an instruction operation module 20, an arbitration module 30, and a first multiplexer 40; the arbitration module 30 is respectively connected to each processor core 10 and the instruction operation module 20, and the first multiplexer 40 respectively establishes communication connections with each processor core 10 and the instruction operation module 20.
[0023] Each processor core 10 is configured to: send an operation request signal to the arbitration module; the operation request signal is used to request access to the resources of the instruction operation module; the arbitration module 30 is configured to: receive the operation request signals of each processor core, generate an authorization signal and a core identifier in the order of core priorities and send them to the target core; the target core is the processor core that is granted the access permission to the resources of the instruction operation module; the instruction operation module 20 is configured to: in response to the target operation request signal, perform an operation to obtain an operation result and send the operation result and the core identifier to the first multiplexer 40; the target operation request signal is the operation request signal sent by the target core; the first multiplexer 40 is configured to: according to the core identifier, send the operation result to the target core.
[0024] It should be noted that the above-mentioned processor cores 10 can be one, two or more. The multiple processor cores 10 are respectively, for example, core 0, core 1,..., core n, core (n + 1), and are the main bodies that initiate operation access requests to the instruction operation module. When a specific operation needs to be performed, an operation request signal will be sent to the instruction operation module. Among them, the above-mentioned operation request signal is used to indicate the execution of an operation, and the operation request signal may include operation data and an operation method.
[0025] The above-mentioned instruction operation module 20 can be a VPU operation unit, which refers to the "V" instruction operation module and provides the operation ability for the entire device, including operation resources that are competitively used by multiple processor cores. The above-mentioned instruction operation module is configured to respond to the operation request signal, perform a corresponding operation to generate an operation result and send it to the first multiplexer.
[0026] The above-mentioned arbitration module 30 can be an ARBIT module, which is responsible for coordinating the access requests of multiple processor cores to the operation module, ensuring that each processor core has the opportunity to preempt the computing resources of the instruction operation module, and avoiding conflicts caused by multiple processor cores accessing the instruction operation module simultaneously.
[0027] The above-mentioned first multiplexer 40 is configured to receive the operation result and the core identifier sent by the instruction operation module, and according to the core identifier, send the operation result to the target core. In the process of determining the target core, the target core is any one of all the processor cores.
[0028] Specifically, when multiple processor cores need to perform arithmetic operations, the multiple processor cores simultaneously compete for the access right to the computing resources in the instruction arithmetic module, and they send their respective arithmetic request signals to the arbitration module. After receiving each arithmetic request signal, the arbitration module counts through a fixed clock cycle according to the core priority order, determines the arithmetic request number with the highest priority currently, and selects the arithmetic request signal corresponding to the arithmetic request number from the multiple arithmetic request signals as the target arithmetic request signal, and takes the processor core corresponding to the target arithmetic request signal as the target core. An authorization signal ("grant" signal) is generated according to the target arithmetic request signal to allow the processor core corresponding to the target arithmetic request signal to access the instruction arithmetic module. When a processor core successfully preempts the computing resources of the instruction arithmetic module, the arbitration module will block the arithmetic request signals of other processor cores until the target core has used up the resources of the instruction arithmetic module. After that, the counter in the arbitration module continues to count, and passes the priority to the next processor core in turn, and so on, to ensure that each processor core can obtain the access opportunity to the computing resources in the instruction arithmetic module within a certain time.
[0029] An instruction processing device is provided in an embodiment of the present application. The instruction processing device includes: multiple processor cores, an instruction arithmetic module, an arbitration module, and a first multiplexer. The arbitration module is respectively connected to each processor core and the instruction arithmetic module, and the first multiplexer respectively establishes communication connections with each processor core and the instruction arithmetic module. Each processor core is used to: send an arithmetic request signal to the arbitration module; the arbitration module is used to: receive the arithmetic request signals of each processor core, generate an authorization signal and a core identifier according to the core priority order and send them to the target core; the instruction arithmetic module is used to: in response to the target arithmetic request signal, perform an arithmetic operation to obtain an arithmetic result and send the arithmetic result and the core identifier to the first multiplexer; the first multiplexer is used to: according to the core identifier, send the arithmetic result to the target core. Compared with the prior art, since the arbitration module is provided in the technical solution of the present application, the authorization signal and the core identifier can be generated in sequence according to the priority order to the corresponding target core, so that each processor core can access the resources of the arithmetic processing unit in sequence to perform arithmetic operations, without configuring an independent arithmetic unit for each core, realizing the sharing of computing resources in the multi-core mode, and then sending the arithmetic result to the target core through the first multiplexer, which greatly reduces the hardware cost.
[0030] In an optional embodiment of the present application, the arbitration module is specifically used to: According to the kernel priority order and the polling mechanism, starting from the operation request signal with the highest priority among the operation request signals, each operation request signal is polled in turn to determine the target operation request signal until all operation request signals have been traversed; an authorization signal and a corresponding kernel identifier are generated according to the target operation request signal and sent to the target kernel.
[0031] It can be understood that the above kernel priority order can be custom - set according to actual needs. The polling mechanism means polling each operation request signal in turn, so that each processor core can access the computing resources in the instruction operation module, that is, the corresponding operation request signals are all executed. The purpose of the polling mechanism is to fairly and orderly allocate the usage rights of the instruction operation module when multiple processor cores simultaneously initiate access requests to the instruction operation module, ensuring that each processor core can be served within a reasonable time.
[0032] The above authorization signal is also called the "grant" signal, which is used to permit a specific processor core (target kernel) to access the computing resources of the instruction operation module and is a signal indicating resource allocation authorization.
[0033] Exemplarily, the above - mentioned device is configured with a kernel priority order and usually adopts a polling method, that is, processing the requests of each kernel in turn in a loop. Starting from the operation request signal with the highest priority, each processor core's operation request signal is checked one by one according to the kernel priority order. During the polling process, when the operation request signal of a certain kernel is valid (usually high - level), then this signal is determined as the target operation request signal, and continuous polling processing is carried out until all operation request signals have been traversed.
[0034] After the target operation request signal is determined, the arbitration module generates an authorization signal (grant signal), indicating that the target kernel corresponding to the target operation request signal is allowed to access the instruction operation module. At the same time, the arbitration module generates a corresponding kernel identifier (ID) for the target operation request signal, which is used to track and identify the source of the request during subsequent operations. Then the generated authorization signal and kernel identifier are sent to the target kernel, notifying it that it can start using the arithmetic unit for arithmetic operations.
[0035] In the embodiment of the present application, by setting the arbitration module, starting from the operation request signal with the highest priority among the operation request signals according to the kernel priority order and the polling mechanism, polling each operation request signal in turn can avoid conflicts caused by multiple cores accessing simultaneously, ensure that each processor core has the opportunity to preempt the computing resources of the instruction operation module, and realize the orderly sharing of resources.
[0036] In an optional embodiment of the present application, please refer to Figure 3As shown in the figure, the above arbitration module 30 includes: a counter 31 and a second multiplexer 32; the counter 31 is connected to the second multiplexer 32.
[0037] The counter 31 is used for: counting according to the operation request signal in clock cycles, generating the current output result and sending it to the second multiplexer; the current output result is used to indicate the number of the operation request signal with the highest priority currently; the change of the current output result is used to implement the rotation of the kernel priority order.
[0038] The second multiplexer 32 is used for: selecting a target operation request signal from all operation request signals according to the current output result; generating an authorization signal and a kernel identifier according to the target operation request signal and sending them to the target kernel.
[0039] It should be noted that the above counter 31 is the core part of the arbitration module, which counts at the rising edge or falling edge of the clock, and sequentially outputs different states, representing the number of the operation request signal with the highest priority currently.
[0040] The second multiplexer 32 (MUX) is used to select the corresponding request signal according to the current output value of the counter. If the operation request signal is valid (high level), the second multiplexer MUX will select this request signal as the target operation request signal, and output to the authorization signal ("grant" signal) according to the target operation request signal, indicating that the processor kernel (target kernel) corresponding to the target operation request signal is allowed to access the computing resources of the instruction operation module. The above kernel identifier is used to indicate which processor kernel this operation request signal comes from.
[0041] Taking the kernel identifier as the kernel ID and the instruction operation module as the VPU operation unit as an example, in the process of sending the authorization signal ("grant" signal) and the kernel ID to the target kernel by the second multiplexer, the target kernel is notified through the "grant" signal that it can access the computing resources of the VPU operation unit, and the kernel ID enables the VPU operation unit and the subsequent result write-back path to clarify the source and target of the operation.
[0042] After receiving the "grant" signal and the kernel ID, the VPU operation unit performs initialization and preparation work according to the "grant" signal, and then starts to receive the specific operation data and instructions sent by the authorized processor kernel, so as to officially start the operation, obtain the operation result and send the operation result to the first multiplexer, so as to return the operation result to the target kernel through the first multiplexer.
[0043] In the embodiments of the present application, by setting a counter and a second multiplexer, fair resource allocation can be achieved, and the right to use the computing resources in the instruction operation module can be dynamically allocated according to the operation request signal, improving resource utilization.
[0044] In an alternative embodiment of the present application, the counter is further configured to: According to the operation request signal, starting from a preset initial value, every time a clock cycle passes, update the count value according to the counting rule to obtain the corresponding current output result; and when the count value reaches the maximum value, return to the initial value and start counting again.
[0045] It should be noted that the above preset initial value is custom-set according to actual needs. For example, it is 0. When updating the count value according to the counting rule, the current count value can be incremented by one to obtain the corresponding count value.
[0046] During the counting process by the counter, starting from the initial value, every time a clock cycle passes, the value of the counter is counted according to the counting rule. When the count reaches the maximum value, it returns to the initial value and starts counting again. For example, a 4-bit binary counter with an initial value of 0000, every time a clock pulse comes, the value will increase sequentially to 0001, 0010, 0011... 1111, and then return to 0000 to continue cycling.
[0047] In this embodiment, by counting according to a fixed clock cycle, the count value of the counter corresponds to the priority order of different processor cores. The change of the count value realizes the rotation of the priority of the processor cores, ensuring that each core has the opportunity to obtain the access right to the arithmetic unit. And the counting state of the counter determines the selection timing of the multiplexer. Each time the count is updated, it triggers the multiplexer to select the corresponding signal from multiple operation request signals according to the new count value, so as to realize polling each operation request signal in turn.
[0048] In the embodiments of the present application, by starting from a preset initial value, every time a clock cycle passes, updating the count value according to the counting rule to obtain the corresponding current output result, and when the count value reaches the maximum value, returning to the initial value and starting counting again, it is possible to ensure that only one core obtains the access right each time, avoiding conflicts and competitions caused by multiple cores accessing the arithmetic unit simultaneously, ensuring the normal operation of the system. At the same time, it reduces the complexity of the system, the hardware cost, and the design difficulty.
[0049] In an alternative embodiment of the present application, the counter is further configured to: For the current clock cycle, use the operation request signal corresponding to the core with the highest priority in the current clock cycle among the processor cores as the target operation request signal, and number the target operation request signal to obtain the current output result.
[0050] When the next clock cycle arrives, the counter moves to the next operation request signal, increments the count value by one, and determines the next operation request signal with the highest priority in the next clock cycle according to the core priority order; the next operation request signal is the operation request signal whose priority in the current clock cycle is after the highest priority.
[0051] Exemplarily, assume that there are 4 processor cores, namely CORE1, CORE2, CORE3, and CORE4, which simultaneously send operation request signals to the arbitration module, and the corresponding operation request signals are req1, req2, req3, and req4 respectively. When the system starts and all operation request signals are not activated (i.e., req1 = 0, req2 = 0, req3 = 0, req4 = 0), for the first clock cycle, the count value in the counter is 0 and the current output result is output to the second multiplexer (MUX). Assume that the count value 0 corresponds to req1. Assume that in this clock cycle, req1 becomes high level (i.e., req1 = 1), indicating that CORE1 initiates a resource access request. Since the counter output corresponds to req1 and req1 is valid (high level), the second multiplexer (MUX) will select req1 as the target operation request signal and generate an authorization grant signal and a page ID to the target core CORE1, indicating that CORE1 is allowed to access the shared resource.
[0052] For the second clock cycle, increment the counter value by 1 to become 1. Assume that 1 corresponds to req2, and req2 is high level (req2 = 1) in this clock cycle, indicating that requester 2 (CORE2) initiates a resource access request. Although the counter output corresponds to req2 and req2 is valid, since the "busy" signal is high level (busy = 1), indicating that the resource is being used, the second multiplexer MUX will not generate a new authorization grant signal, and the grant authorization signal remains 0. When req1 is executed and completed by the instruction operation module, then according to the core priority order, the next operation request signal with the highest priority at this time is req2, and a corresponding new authorization grant signal and core identifier are generated for req2 and sent to the target core CORE2, indicating that CORE2 is allowed to access the shared resource.
[0053] In the embodiment of the present application, for the current clock cycle, the operation request signal corresponding to the core with the highest priority in the processor core during the current clock cycle is used as the target operation request signal, and the current output result is obtained. When the next clock cycle arrives, the counter moves to the next operation request signal, increments the count value by one, and determines the next operation request signal with the highest priority in the next clock cycle according to the core priority order, enabling the counter to change the priority order by continuous counting and providing a basis for selection for the multiplexer, thereby ensuring that each processor core has the opportunity to obtain access to the arithmetic unit.
[0054] In an alternative embodiment of the present application, the counter is further configured to: generate a busy signal and send it to other cores, where the other cores are the remaining cores among all the processor cores except the target core, and the busy signal is used to indicate that the resources of the instruction operation module are currently occupied.
[0055] It should be noted that the above-mentioned other cores may include multiple processor cores or may include one processor core. The busy signal is also referred to as the "busy" signal or the "occupation signal". This busy signal is mainly used to feedback the occupation status of the computing resources in the instruction operation module, indicating that the computing resources in the instruction operation module are being used and are in a busy state, and are temporarily unable to respond to other requests. When the "busy" signal is set to a high level (busy = 1), it indicates that the requester starts to use the resources at this time.
[0056] Among them, when the "busy" signal indicates that the instruction operation module is busy, even if the counter continues to count, the multiplexer will not generate a new authorization signal until the "busy" signal becomes low level, indicating that the arithmetic unit is idle, and then the normal selection and authorization operations will continue.
[0057] For example, when the operation request signals corresponding to 4 processor cores (CORE1, CORE2, CORE3, and CORE4) are req1, req2, req3, and req4 respectively, when it is determined that req1 is the target operation request signal, an authorization signal is generated and sent to the target core CORE1 corresponding to req1 at this time, and the busy signal is set to a high level and sent to other cores (CORE2, CORE3, and CORE4) to inform other cores that the current instruction operation module is in a busy state and is temporarily unable to respond to the operation request signals of other cores. After receiving the busy signal, other cores will suspend the request for this resource.
[0058] In the embodiment of the present application, by generating a busy signal and sending it to other cores, it is possible to prevent multiple cores from accessing the same resource simultaneously, avoid data conflicts or errors, and contribute to the orderly allocation and access of resources.
[0059] In an optional embodiment of the present application, the first multiplexer is further used to: decode all cores according to the core identification, determine the target core, and send the operation result to the target core.
[0060] It is understandable that when the instruction operation module completes the operation and is ready to write the result back, it needs to send the kernel identifier to the first multiplexer. The first multiplexer performs a decoding operation based on the kernel identifier, selects the correct one from multiple possible write-back paths, and accurately sends the operation result to the processor core that initiated the request. In other words, at this stage, the kernel identifier plays a routing role, which guides the result data to be accurately returned to the request source.
[0061] Exemplarily, in the process where the first multiplexer selects the corresponding operation result to be output to the processor core according to the received core identifier, for example, there are multiple processor cores (such as core 1, core 2, core 3, etc.) in the system, and each core has a corresponding operation result. When the first multiplexer MUX receives the identifier of core 2, it decodes the identifier, then selects the operation result corresponding to core 2 from all the operation results, and sends it to core 2, thereby realizing the processing and response of the specific core operation request.
[0062] In the embodiment of the present application, the calculation result can be accurately sent to the corresponding processor core according to the core identification, ensuring that the calculation request signal of each processor core can be accurately identified and processed, avoiding signal confusion and erroneous selection.
[0063] In an optional embodiment of the present application, the bit width of the arbitration module is determined according to the number of cores, and the counter is represented according to binary coding or Gray coding.
[0064] It should be noted that the characteristic of the above Gray code is that there is only one binary digit different between any two adjacent codes. For example, a typical 4-bit Gray code starts with 0000, the next 0001, only the last digit is different; the next one is 0011, which is different from the previous 0001 only by the second to last digit, and so on. This characteristic gives Gray code a unique advantage in signal transmission and digital systems.
[0065] The arbitration module can adapt to different numbers of processor cores through the number of bits in the counter. If there are n processor cores, an m-bit binary counter is usually required, where m satisfies 2m≥n. In this way, the count value of the counter can cover all core numbers, thereby realizing polling of all core requests.
[0066] In the embodiment of the present application, the counter is represented in the form of binary coding or Gray coding, which can number all the cores, so as to realize polling of the operation request signals of all the cores and avoid the situation of multiple bits changing simultaneously during adjacent code conversion, thereby reducing the competition phenomenon.
[0067] On the other hand, the embodiment of the present application also provides an instruction processing method. Please refer to Figure 4 As shown, the instruction processing method includes the following steps 201-step 204: Step 201, receive the operation request signals sent by each processor core; the operation request signals are used to request access to the resources of the instruction operation module.
[0068] Step 202, generate an authorization signal and a core identifier according to the core priority order and send them to the target core; the target core is the processor core authorized to access the resources of the instruction operation module.
[0069] Step 203, in response to the target operation request signal, perform an operation to obtain an operation result; the target operation request signal is the operation request signal sent by the target core.
[0070] Step 204, send the operation result to the target core according to the core identifier.
[0071] It should be noted that the content included in each operation request signal may vary according to the system design and the functional requirements of the processor core. It may include: request identifier, operation type, data source information, priority information, control parameters, etc. Among them, the request identifier is used to uniquely identify the operation request for tracking and management in the system. For example, a specific number or identifier enables the arbitration module and other components to distinguish different requests. The operation type specifies the type of operation requested, such as addition, multiplication, logical operation, etc. Different operation types may require the operation unit to adopt different processing methods and algorithms. For example, for a processor core requesting a floating-point multiplication operation, the operation request signal will clearly indicate that this is a floating-point multiplication operation type.
[0072] Data source information refers to that if a specific data input is required for an operation, the operation request signal may contain information about the data source, such as which memory address or register the data is stored in. In this way, the operation unit can accurately obtain the required data for the operation. For example, when requesting a data processing operation, the signal will indicate that the data to be processed is stored in a specific area of the memory. Optionally, the above operation request signal may include priority information indicating the urgency or importance of the request. The arbiter can decide which requests to process first based on the priority information. For example, requests with a higher priority may be processed prior to those with a lower priority. For example, some operation requests with high real-time requirements are assigned a higher priority.
[0073] Control parameters refer to the parameters that control the operation process, such as the operation precision requirement, the number of iterations (for some complex iterative algorithms), etc. These parameters can help the operation unit execute the operation task more precisely. For example, for a numerical calculation request, the precision requirement of the calculation result may be specified.
[0074] In a system where multiple processor cores compete to use the operation unit, a polling arbitration module is adopted to manage resource allocation. The entire logical process is as follows: Please refer to Figure 5 As shown, taking the instruction operation module as the VPU operation module, the first multiplexer as the first MUX, the arbitration module as the ARBRT module, and multiple processor cores as CORE0, CORE1, CORE2,..., CORE(n + 1) as examples. In the initial state, when the system starts, the operation request signals of all processor cores are in an inactive state, and the "busy" signal is at a low level indicating that the instruction operation module is idle. The counter in the ARBRT module starts counting from 0. Multiple processor cores (CORE0 - CORE(n + 1)) send operation request signals (such as req1, req2, etc.) to the ARBRT module according to their needs, requesting to use the computing resources in the VPU operation module, and the requesters set their respective operation request signals to a high level.
[0075] The counter in the ARBRT module increments by 1 according to a fixed clock cycle, changing the priority order. The second multiplexer (MUX) selects the corresponding operation request signal as the target operation request signal according to the current output value of the counter. If the operation request signal is valid (high level), the second multiplexer MUX in the ARBRT module selects this operation request signal from all operation request signals, generates a "grant" signal (authorization signal), and determines the corresponding core identifier (ID) at the same time. Then, the "grant" signal and the core identifier are sent to the corresponding target core, and the target core obtains the access right to the operation unit and starts using the VPU operation unit, and the "busy" signal is set to a high level.
[0076] When the "busy" signal is high, even if the counter continues to count, the second multiplexer MUX will not generate a new "grant" signal, and the operation requests of other cores are blocked, and they cannot obtain access to the VPU operation unit until the "busy" signal becomes low, that is, the current core has used up the VPU operation unit and released the computing resources. In the next clock cycle, the counter moves to the next requester (the count value increases by 1), and the above process is repeated, starting from the highest priority operation request signal, polling each operation request signal in turn, determining the new target operation request signal, and generating a new "grant" signal and core identifier and sending it to the new target core, so that all requesters (all cores) take turns to obtain the right to use the computing resources in the VPU operation unit in a predetermined order.
[0077] After the VPU operation module completes the operation, it writes the operation result back to the corresponding processor core that initiated the request through the first MUX according to the previously recorded core identifier.
[0078] Resource sharing is achieved in this embodiment. In a multi-core scenario, there are often many large computing units whose area even exceeds the core itself. If this sharing mechanism is not used, the area of the processor core will increase exponentially. This application is a solution that does not require multiple computing units, and achieves computing resource sharing in multi-core mode, which greatly reduces hardware costs.
[0079] Compared with the prior art, the instruction processing method in this embodiment, due to the provision of an arbitration module, can generate authorization signals and core identifiers to the corresponding target cores in order of priority, so that each processor core can access the resources of the operation processing unit in sequence to perform operation operations, without the need to configure an independent operation unit for each core, thereby realizing computing resource sharing in multi-core mode, and then sending the operation results to the target core through the first multiplexer, thereby greatly reducing hardware costs.
[0080] On the other hand, an embodiment of the present application further provides an instruction processing system, which includes the instruction processing device provided by the above embodiment.
[0081] Compared with the prior art, the processing system of the present application, due to the provision of an arbitration module, can generate authorization signals and core identifiers to the corresponding target cores in order of priority, so that each processor core can access the resources of the operation processing unit in sequence to perform operation operations, without the need to configure an independent operation unit for each core, thereby realizing computing resource sharing in multi-core mode, and then sending the operation results to the target core through the first multiplexer, thereby greatly reducing hardware costs.
[0082] It should be understood that although the steps in the flowchart are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0083] In one embodiment, a computer device is provided. The internal structure diagram of the computer device can be as Figure 1 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an instruction processing method as described above. It includes: including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements any step in the above instruction processing method.
[0084] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it can implement any step in the above instruction processing method.
[0085] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0089] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0090] Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. An instruction processing device, characterized in that: The instruction processing device comprises: a plurality of processor cores, an instruction operation module, an arbitration module, and a first multiplexer; the arbitration module is respectively connected to each processor core and the instruction operation module, and the first multiplexer is respectively connected to each processor core and the instruction operation module; Each of the processor cores is used to: send a calculation request signal to the arbitration module; the calculation request signal is used to request access to the resources of the instruction calculation module; The arbitration module is used to: receive the operation request signal of each processor core, generate the authorization signal and the core identification according to the core priority order, and send them to the target core; the target core is the processor core that is granted the instruction operation module resource access permission; The instruction operation module is used to: in response to a target operation request signal, perform an operation operation to obtain an operation result and send the operation result and the core identifier to the first multiplexer; the target operation request signal is an operation request signal sent by the target core; The first multiplexer is used to send the operation result to the target core according to the core identifier.
2. The device according to claim 1, characterized in that The arbitration module is specifically used for: According to the kernel priority order and the polling mechanism, starting from the highest priority operation request signal among the operation request signals, each operation request signal is polled in turn to determine the target operation request signal until all operation request signals are traversed; An authorization signal and a corresponding kernel identifier are generated according to the target operation request signal and sent to the target kernel.
3. The device according to claim 1, characterized in that The arbitration module includes: a counter and a second multiplexer; the counter is connected to the second multiplexer; The counter is used to: count according to the operation request signal according to the clock cycle, generate a current output result and send it to the second multiplexer; the current output result is used to indicate the number of the operation request signal currently at the highest priority; the change of the current output result is used to realize the rotation of the kernel priority order; The second multiplexer is used to: select the target operation request signal from all operation request signals according to the current output result; generate an authorization signal and a core identifier according to the target operation request signal and send them to the target core.
4. The device according to claim 3, characterized in that The counter is also used for: According to the operation request signal, starting from the preset initial value, each clock cycle, the count value is updated according to the counting rule to obtain the corresponding current output result; and when the count value reaches the maximum value, it returns to the initial value and restarts counting.
5. The device according to claim 3, characterized in that The counter is also used for: For the current clock cycle, the operation request signal corresponding to the core with the highest priority in the current clock cycle among the processor cores is used as the target operation request signal, and the target operation request signal is numbered to obtain the current output result; When the next clock cycle arrives, the counter moves to the next operation request signal, increases the count value by one, and determines the next operation request signal with the highest priority in the next clock cycle according to the kernel priority order; The next operation request signal is an operation request signal whose priority in the current clock cycle is subsequent to the highest priority.
6. The device according to claim 3, characterized in that The counter is also used to: A busy signal is generated and sent to other cores, where the other cores are the remaining cores in all processor cores except the target core, and the busy signal is used to indicate that the resources of the instruction operation module are currently in an occupied state.
7. The device according to claim 1, characterized in that The first multiplexer is further configured to: perform decoding processing on all cores according to the core identifier, determine a target core, and send the operation result to the target core.
8. The device according to claim 6, characterized in that The bit width of the arbitration module is determined according to the number of the cores, and the counter is represented according to binary coding or Gray coding.
9. A processing system, characterized in that: include: An instruction processing device as claimed in any one of claims 1 to 8.
10. A method for processing an instruction, characterized in that: Applied to the instruction processing device according to any one of claims 1 to 8, the method comprises: Receiving operation request signals sent by each of the processor cores; the operation request signals are used to request access to resources of the instruction operation module; Generate an authorization signal and a core identification according to the core priority order and send them to a target core; the target core is a processor core that is granted access rights to instruction operation module resources; In response to a target operation request signal, executing an operation to obtain an operation result; the target operation request signal is an operation request signal sent by the target core; The operation result is sent to the target kernel according to the kernel identifier.
Citation Information
Patent Citations
Digital signal processing device and method for electric energy metering chip
CN112834819A
Method and apparatus for bus arbitration capable of effectively altering a priority order
US20020133654A1