Reconfigurable processor for distributed energy transaction, control method, device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-11
AI Technical Summary
然而,由于传统处理器中的硬件核仅固化单一算法,当需要处理包含不同运算类型的交易时,则必须依次调用不同的处理器,或者在单个处理器内集成多个独立的算法专用核心,其中在单个处理器内集成多个独立的算法专用核心的设计会导致处理器面积冗余严重,在某个核心工作时,处理器内的其他核心会处于闲置状态,从而降低了硬件利用率,增加了对分布式能源交易的数据进行处理的成本
[0059] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array-shared configuration bus that is physically isolated from the on-chip interactive cache. Since the configuration bus and data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
Smart Images

Figure CN122547747A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of semiconductor technology, and specifically to a reconfigurable processor, control method, apparatus, device, and medium for distributed energy trading. Background Technology
[0002] In recent years, with the widespread application of distributed photovoltaic and energy storage terminals, the frequency of distributed energy transactions per unit time has increased dramatically. Distributed energy transactions typically require processing massive amounts of concurrent data. Each transaction needs to complete different types of operations, such as signing, verification, hash packaging, and data encryption. These operations place high demands on the throughput, latency, algorithm adaptability, and power efficiency of the relevant computing units.
[0003] In related technologies, traditional processors can be used to process distributed energy transaction data. However, because the hardware cores of traditional processors only have a single algorithm embedded in them, when processing transactions involving different computational types, different processors must be called sequentially, or multiple independent algorithm-specific cores must be integrated within a single processor. The design of integrating multiple independent algorithm-specific cores within a single processor leads to severe processor area redundancy. When one core is working, other cores within the processor are idle, thus reducing hardware utilization and increasing the cost of processing distributed energy transaction data. Furthermore, frequent hardware module switching introduces additional bus arbitration and context saving and recovery overhead, thereby reducing the efficiency of processing distributed energy transaction data. Summary of the Invention
[0004] To address the problems in related technologies, embodiments of this disclosure provide a reconfigurable processor, control method, apparatus, device, and medium for distributed energy trading.
[0005] In a first aspect, this disclosure provides a reconfigurable processor for distributed energy trading, including an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0006] The array scheduling controller is used to obtain the current task queue of distributed energy trading, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0007] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0008] The configuration register group is used to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions.
[0009] The multiplexer array is used to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group. The cross-connection switch matrix is used to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group. The pipeline stage enable switch is used to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0010] In one embodiment of this disclosure, the reconfigurable computing core further includes a handshake interaction interface.
[0011] The array scheduling controller is also used to allocate independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache.
[0012] The reconfigurable computing core is also used to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, and the multiple computing macro units perform computing processing on the task data based on the computing pipeline. The reconfigurable computing core writes the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
[0013] In one embodiment of this disclosure, the corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the computation steps in the target task corresponding to the reconfigurable computing core.
[0014] In one embodiment of this disclosure, the plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including:
[0015] The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm.
[0016] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm.
[0017] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm.
[0018] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
[0019] In one embodiment of this disclosure, the reconfigurable computing core further includes a key buffer.
[0020] The key buffer is used to store the public key, session key, and hash initialization vector.
[0021] In one embodiment of this disclosure, the arithmetic macrounit includes a modular multiplication macrounit, a modular exponentiation macrounit, a substitution box macrounit, a linear shift XOR macrounit, a hash compression iteration macrounit, a register file macrounit, and a pipelined latch macrounit.
[0022] In one embodiment of this disclosure, sending the corresponding register configuration instruction to the corresponding reconfigurable computing core via the array shared configuration bus includes: The array shares a configuration bus to broadcast register configuration instructions to multiple reconfigurable computing cores.
[0023] Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
[0024] In one embodiment of this disclosure, the array scheduling controller is further configured to: Obtain the current working status of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores.
[0025] If the overall computing load rate is greater than or equal to the first computing load rate threshold, a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores.
[0026] If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold.
[0027] The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: At least one reconfigurable computing core is identified as being in the startup state among the plurality of reconfigurable computing cores.
[0028] The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
[0029] Secondly, this disclosure provides a control method for a reconfigurable processor for distributed energy trading. The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0030] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0031] The control method includes: The array scheduling controller obtains the current task queue of distributed energy transactions, determines the operation type of each task in the current task queue, generates corresponding register configuration instructions according to the operation type of each task, and sends the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0032] The configuration register group is controlled to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions.
[0033] The multiplexer array is controlled to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group; the cross-connection switch matrix is controlled to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group; and the pipeline stage enable switch is controlled to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0034] In one embodiment of this disclosure, the reconfigurable computing core further includes a handshake interaction interface.
[0035] The control method further includes: The array scheduling controller allocates independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache.
[0036] The reconfigurable computing core is controlled to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, control the multiple computing macrounits to perform computing processing on the task data based on the computing pipeline, and control the reconfigurable computing core to write the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
[0037] In one embodiment of this disclosure, the corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the computation steps in the target task corresponding to the reconfigurable computing core.
[0038] In one embodiment of this disclosure, the plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm.
[0039] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm.
[0040] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm.
[0041] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
[0042] In one embodiment of this disclosure, the reconfigurable computing core further includes a key buffer.
[0043] The method further includes: The key buffer is controlled to store the public key, session key, and hash initialization vector.
[0044] In one embodiment of this disclosure, the arithmetic macrounit includes a modular multiplication macrounit, a modular exponentiation macrounit, a substitution box macrounit, a linear shift XOR macrounit, a hash compression iteration macrounit, a register file macrounit, and a pipelined latch macrounit.
[0045] In one embodiment of this disclosure, sending the corresponding register configuration instruction to the corresponding reconfigurable computing core via the array shared configuration bus includes: The array shares a configuration bus to broadcast register configuration instructions to multiple reconfigurable computing cores.
[0046] Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
[0047] In one embodiment of this disclosure, the method further includes: The array scheduling controller is instructed to perform the following steps: Obtain the current working status of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores.
[0048] If the overall computing load rate is greater than or equal to the first computing load rate threshold, a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores.
[0049] If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold.
[0050] The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: At least one reconfigurable computing core is identified as being in the startup state among the plurality of reconfigurable computing cores.
[0051] The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
[0052] Thirdly, this disclosure provides a control device for a reconfigurable processor for distributed energy trading. The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0053] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0054] The control device includes: The instruction acquisition module is configured to control the array scheduling controller to acquire the current task queue of distributed energy transactions, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0055] The level output module is configured to control the configuration register group to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions.
[0056] The unit reconfiguration module is configured to control the multiplexer array to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group, control the cross-connection switch matrix to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group, and control the pipeline stage enable switch to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0057] Fourthly, embodiments of this disclosure provide an electronic device including a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method as described in any one of the first aspects.
[0058] Fifthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method as described in any one of the first aspects.
[0059] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array-shared configuration bus that is physically isolated from the on-chip interactive cache. Since the configuration bus and data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
[0060] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0061] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings: Figure 1 A schematic structural diagram of a reconfigurable processor for distributed energy trading according to an embodiment of the present disclosure is shown.
[0062] Figure 2 A schematic structural diagram of a reconfigurable computing core according to an embodiment of the present disclosure is shown.
[0063] Figure 3 A flowchart illustrating a control method for a reconfigurable processor for distributed energy trading according to an embodiment of the present disclosure is shown.
[0064] Figure 4 A structural block diagram of a control device for a reconfigurable processor for distributed energy trading according to an embodiment of the present disclosure is shown.
[0065] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0066] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown. Detailed Implementation
[0067] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of exemplary embodiments have been omitted from the drawings.
[0068] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.
[0069] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0070] In this disclosure, any operation involving the acquisition of user information or user data, or the display of user information or user data to others, is an operation authorized or confirmed by the user, or actively selected by the user.
[0071] In related technologies, traditional processor architectures are typically used to process data from distributed energy transactions. However, the hardware cores in traditional processors employ a hard-wired, fixed design, with each core dedicated to only one specific cryptographic algorithm. For example, it may only be able to perform SM2 signature verification or only SM3 hash operations. When a distributed energy transaction requires sequentially performing various types of operations such as signing, hash packaging, and data encryption, traditional processors must achieve this through two approaches: one is to sequentially call multiple external, independent dedicated processors, each responsible for one algorithm; the other is to integrate multiple independent algorithm-specific cores within a single processor, each core corresponding to a fixed algorithm.
[0072] However, the first approach requires frequent switching between different processors, each involving bus arbitration, data transfer, and core state saving and recovery, significantly increasing transaction processing latency. The second approach, while integrating all algorithm cores onto a single chip, results in each core being physically independent. The chip area increases linearly with the number of integrated algorithms. In actual operation, only the currently invoked core is active, while the rest remain idle. This leads to significant underutilization of chip space, severe waste of hardware resources, and directly increases the hardware cost of distributed energy transaction data processing.
[0073] Furthermore, regardless of the method used, the system must perform a hardware module switching operation whenever the type of computation involved in a transaction changes. Each switch requires pausing the current pipeline, saving the intermediate state of the current core (such as register values and computation progress), obtaining access to the target core through bus arbitration, and restoring the target core's context before starting a new computation. This process introduces additional bus arbitration overhead and context saving and restoration overhead, and the switching frequency increases dramatically with the increase in transaction concurrency, severely reducing overall processing efficiency and making it difficult to meet the high-throughput, low-latency real-time processing requirements of distributed energy trading scenarios.
[0074] To address the aforementioned issues, this disclosure provides a reconfigurable processor, control method, apparatus, device, and medium for distributed energy trading.
[0075] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array shared configuration bus that is physically isolated from the on-chip interactive cache. Since the array shared configuration bus and the data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
[0076] Figure 1 A schematic structural diagram of a reconfigurable processor for distributed energy trading according to embodiments of the present disclosure is shown. Figure 1 As shown, the reconfigurable processor 100 for distributed energy trading includes an array scheduling controller 101, multiple independent reconfigurable computing cores 102, an array shared configuration bus 103, and an on-chip interactive cache 104. The data transmission buses of the array shared configuration bus 103 and the on-chip interactive cache 104 are physically isolated. The array scheduling controller 101 is connected to each of the multiple reconfigurable computing cores 102 through the array shared configuration bus 103, and each reconfigurable computing core 102 is connected to the on-chip interactive cache 104.
[0077] The array scheduling controller 101 is used to obtain the current task queue of distributed energy trading, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core 102 through the array shared configuration bus 103.
[0078] Figure 2 A schematic structural diagram of a reconfigurable computing core according to an embodiment of the present disclosure is shown. Figure 2 As shown, the reconfigurable computing core 102 includes a configuration register group 105, a topology reconstruction module 106, and a computational macrocell pool 107. The computational macrocell pool 107 includes multiple computational macrocells 117. The configuration register group 105 is connected to the topology reconstruction module 106. The topology reconstruction module 106 includes a multiplexer array 108, a cross-connection switch matrix 109, and a pipeline stage enable switch 110. The multiplexer array 108 is connected to the input port of each of the multiple computational macrocells 117, and the cross-connection switch matrix 109 is connected to the output port of each of the multiple computational macrocells 117.
[0079] The configuration register group 105 is used to receive register configuration instructions corresponding to the reconfigurable computing core 102 through the array shared configuration bus 103, latch the register configuration instructions, and output a level to the topology reconfiguration module 106 according to the register configuration instructions.
[0080] The multiplexer array 108 is used to configure the connection relationship of the input ports of each arithmetic macro unit 117 according to the level output by the configuration register group 105. The cross-connection switch matrix 109 is used to configure the connection relationship of the output ports of each arithmetic macro unit 117 according to the level output by the configuration register group 105. The pipeline stage enable switch 110 is used to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output by the configuration register group 105, so that the plurality of arithmetic macro units 117 form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0081] In one implementation of this disclosure, distributed energy trading can be understood as business activities such as buying and selling surplus electricity, energy storage settlement, green electricity certificate trading, or electricity data storage, primarily using distributed energy devices such as distributed photovoltaics, energy storage terminals, and electric vehicles, conducted in a peer-to-peer network. During distributed energy trading, each transaction can be decomposed into multiple cryptographic steps, and one or more of these steps together constitute a task in the distributed energy trading task queue. For example, a distributed energy trading task typically includes performing multiple steps of different types of cryptographic operations, such as signing, verifying signatures, hashing, and data encryption, to ensure transaction security. The original input data corresponding to each distributed energy trading task, such as the transaction digest to be signed, the message block to be hashed, the plaintext of the transaction to be encrypted, and related keys and initialization vectors, constitutes the task data for that task.
[0082] In one implementation of this disclosure, the array shared configuration bus and the data transmission bus of the on-chip interactive cache are physically isolated. This can be understood as follows: in the physical layout of the processor, the array shared configuration bus and the data transmission bus connecting the on-chip interactive cache use independent wiring channels, and there are no direct metal connections or electrical couplings between them. The array shared configuration bus is used to transmit register configuration instructions issued by the array scheduler, while the data transmission bus connecting the on-chip interactive cache is used for task data read / write between the reconfigurable computing core and the on-chip interactive cache. Because the two buses are physically separated, the transition of the configuration signal carrying the register configuration instructions will not interfere with the data signal carrying the task data, and the data flow on the data transmission bus will not affect the timing integrity of the register configuration instructions. Furthermore, the array shared configuration bus does not need to arbitrate or time-multiplex with the data transmission bus, and the scheduler can independently issue register configuration instructions at any time.
[0083] In one implementation of this disclosure, the operation type of the tasks in the current task queue of distributed energy trading can be understood as the type of cryptographic operation required by the tasks in the distributed energy trading scenario. The cryptographic operation types can include SM2 asymmetric encryption, SM3 hash iteration, SM4 block cipher, and SHA-256 hash. These four cryptographic operation types correspond to different hardware data path topologies and algorithmic logics.
[0084] In one implementation of this disclosure, corresponding register configuration instructions are generated based on the operation type of each task. By parsing each task, the required cryptographic algorithm type can be identified. By determining the operation type of each task, the array scheduler can select a suitable reconfigurable computing core for each task and generate register configuration instructions matching that operation type, thereby driving the corresponding reconfigurable computing core to be reconfigured into the required algorithm hardware path.
[0085] In one implementation of this disclosure, the array scheduler can also identify the operation type of each task, as well as analyze the task priority and computational load of each task. Specifically, the array scheduler can determine the task allocation order based on priority, with high-priority tasks preferentially allocated to idle or soon-to-be-idle reconfigurable computing cores to ensure low-latency processing of critical transactions. The array scheduler can also dynamically select a task allocation strategy by combining the load size with the current busy / idle status of each reconfigurable computing core. For low-load tasks, they can be merged and allocated to the same core for batch processing; for high-load tasks, they are allocated to cores with lighter current loads. For example, when a high-priority transaction has multiple computational steps and a heavy computational load, the array scheduler can adopt a single-transaction pipeline operation mode, splitting the transaction into multiple sub-tasks and allocating them to different reconfigurable computing cores for serial collaborative completion, thereby reducing the end-to-end latency of the transaction; conversely, for a large number of low-priority, low-load transactions, a multi-transaction parallel operation mode is adopted, allowing each core to independently process unrelated transactions to improve overall throughput.
[0086] In one implementation of this disclosure, the array scheduler controller can internally store hardware configuration word libraries corresponding to various algorithms. This hardware configuration word library refers to a set of binary configuration word tables, pre-fixed during the chip design phase and corresponding one-to-one with the underlying reconfigurable hardware topology. Each configuration word has a fixed bit width (e.g., 32 bits), where each bit is directly mapped to the control pins of the multiplexers, cross-connection switch matrices, and pipeline enable switches within the reconfigurable computing core. After determining the operation type of any task, the array scheduler controller can retrieve the corresponding hardware configuration word from the hardware configuration word library based on the operation type of the task and encapsulate the configuration word into a register configuration instruction.
[0087] In one implementation of this disclosure, the array scheduling controller may further include a global synchronization timing unit, which outputs a unified global clock signal and a global reset signal to all reconfigurable computing cores. In high-concurrency scenarios of distributed energy trading, multiple reconfigurable computing cores may execute different transactions in parallel or collaboratively complete the pipeline computation of the same transaction. Data interaction across reconfigurable computing cores requires strict alignment of the timing boundaries of each reconfigurable computing core. By outputting a unified global clock signal and a global reset signal to all reconfigurable computing cores through the global synchronization timing unit, it can be ensured that the same clock edge arrives simultaneously at the registers and pipeline latches of each core, eliminating the risk of data misalignment caused by clock skew.
[0088] In one implementation of this disclosure, the configuration register group latches register configuration instructions and outputs a level to the topology reconfiguration module according to the register configuration instructions. This can be understood as the configuration register group within each reconfigurable computing core first receiving register configuration instructions from the array scheduler controller via the array shared configuration bus. The configuration register group latches the received register configuration instructions into internal registers to prevent other data or noise on the array shared configuration bus from overwriting the current configuration. After latching, each output bit of the configuration register group is hardwired to the control pin of the topology reconfiguration module (i.e., the multiplexer array, cross-connection switch matrix, and pipeline inter-stage enable switches), and outputs a corresponding DC level according to the latched binary bits. This DC level directly controls the on / off state and selection of each switch without any intermediate logic conversion, thereby instantly changing the connection topology and pipeline clock enable state between different computational macrounits. Since the level is continuously and stably output once latched until the next configuration update, the reconfigurable computing core maintains a fixed hardware path during computation, ensuring the determinism of the data flow.
[0089] In one implementation of this disclosure, a multiplexer array is used to configure the connection relationship of the input ports of each operational macrounit according to the level output of the configuration register group. This can be understood as follows: within the reconfigurable computing core, a multiplexer is provided before the input port of each basic operational macrounit. After the configuration register group latch register is configured, its output DC level serves as the control input of the multiplexer. Based on the received level value, the multiplexer selects one of multiple possible data sources to connect to the input of the current macrounit. These multiple possible data sources include task data output from the on-chip interactive buffer, the key output from the key buffer, the output of the previous-level macrounit, feedback of intermediate iteration results, or constant initialization vectors, etc. By configuring different level combinations, the data source of each macrounit can be dynamically changed, thereby constructing differentiated data flow topologies under different algorithm modes.
[0090] In one implementation of this disclosure, a cross-connection switch matrix is used to configure the connection relationship of the output ports of each computational macrounit according to the level output of the configuration register group. This can be understood as follows: within the reconfigurable computing core, the output ports of all computational macrounits are connected to the input side of the cross-connection switch matrix. The cross-connection switch matrix has multiple interleaved switch nodes, and whether each switch node is on or off is directly controlled by the level signal output by the configuration register group. When the configuration register latches the register configuration instruction for the corresponding algorithm, the level output by the configuration register causes a specific set of switches in the cross-connection switch matrix to close, thereby directing the output ports of certain computational macrounits to specific output terminals of the cross-connection switch matrix. These output terminals return to the multiplexer array before the input terminals of each computational macrounit, thus the cross-connection switch matrix determines which computational macrounit's input port the computation result of each computational macrounit will be sent to next.
[0091] In one implementation of this disclosure, an inter-pipeline enable switch is used to configure the clock of the pipeline latches in the computation pipeline according to the level output of the configuration register group. This can be understood as follows: within the reconfigurable computing core, the computation pipeline consists of multiple pipeline latches connected in series. Each pipeline latch is used to temporarily store intermediate calculation results, and the clock input of each pipeline latch is connected to a corresponding clock gating circuit. The control terminal of this clock gating is controlled by a certain bit output level of the configuration register group. When the configuration register group latches the register configuration instruction corresponding to the current algorithm, the level signal output by the configuration register group is sent to the clock gating enable terminal of each pipeline latch. For pipeline segments that need to participate in computation, the level signal output by the configuration register group enables the clock gating of that stage, ensuring that the clock signal can reach the pipeline latch normally, allowing data to be temporarily stored and transferred at that stage. For pipeline segments that do not require computation, the level signal output by the configuration register group disables the clock of the corresponding stage latch, keeping that stage's pipeline latch in a static state—that is, it does not toggle, does not consume dynamic power, and data is directly passed through or bypassed. Since different algorithms require different numbers of pipeline stages, the actual effective pipeline depth can be dynamically adjusted according to the current algorithm through pipeline stage enable switches, thus adapting to the different iteration round requirements of various algorithms without changing the fixed hardware structure.
[0092] In one implementation of this disclosure, multiple computational macrounits form a computational pipeline that matches the corresponding computational logic. This can be understood as follows: within the reconfigurable computing core, the computational macrounit pool contains various basic cryptographic computational macrounits, each with a fixed function and independent of the others. When the configuration register set latches the register configuration instruction corresponding to a specific algorithm, the level signal output by the configuration register set drives the multiplexer array, the cross-connection switch matrix, and the pipeline-level enable switches, dynamically combining these computational macrounits into a complete computational pipeline corresponding to the algorithm according to the required order, connection method, and data flow direction. The computational pipeline starts from the data input end, sequentially passes through different selected computational macrounits, completes all the iterations and transformations required by the algorithm, and finally outputs the computation result.
[0093] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array shared configuration bus that is physically isolated from the on-chip interactive cache. Since the array shared configuration bus and the data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
[0094] In one embodiment of this disclosure, the reconfigurable computing core further includes a handshake interaction interface.
[0095] The array scheduling controller is also used to allocate independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache.
[0096] The reconfigurable computing core is also used to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, and the multiple computing macro units perform computing processing on the task data based on the computing pipeline. The reconfigurable computing core writes the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
[0097] In one implementation of this disclosure, the array scheduler allocates independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache. This can be understood as the array scheduler assigning a task to a reconfigurable computing core while simultaneously specifying a dedicated storage area for that core from the on-chip interactive cache. The read address indicates from which starting position of the on-chip interactive cache the reconfigurable computing core reads the task data to be processed, and the write address indicates from which starting position of the on-chip interactive cache the reconfigurable computing core writes the processing result back to. Because the reconfigurable computing cores are allocated independent read and write addresses—that is, different reconfigurable computing cores are allocated non-overlapping read and write address ranges—each reconfigurable computing core performs read and write operations only within its exclusive address range. This address allocation strategy, from a hardware perspective, forces the on-chip interactive cache to be divided into multiple logically independent, physically coexisting read and write partitions. Even if different reconfigurable computing cores simultaneously read task data from the on-chip interactive cache according to their respective read addresses, or simultaneously write the computation results to the on-chip interactive cache according to their respective write addresses, the cache controller can respond in parallel without conflict because the read and write addresses they use belong to different partitions. Therefore, there is no need to set up additional mutex locks or arbitration logic inside the cache; multi-partition isolation of the cache can be naturally achieved simply by uniformly allocating independent addresses through the scheduling controller, supporting parallel access by multiple reconfigurable computing cores.
[0098] In one implementation of this disclosure, the handshake interface can be understood as a hardware module within each reconfigurable computing core responsible for data exchange with the on-chip interactive cache and other reconfigurable computing cores. The handshake interface communicates with the on-chip interactive cache via a set of standard synchronous handshake signals. Once an independent read address has been allocated in the configuration register set by the array scheduler, the handshake interface automatically sends out the read address and initiates a read request. After the on-chip interactive cache returns a read confirmation, it receives the task data and stores it in its internal input buffer. This process is repeated until all data is retrieved. Subsequently, the task data is sent in parallel to the reconfigured computation macrocell pool according to the bit width required by the computation pipeline. After the computation macrocell completes the cryptographic operation and generates the processing result, the handshake interface obtains an independent write address and length from the configuration register set and writes the processing result back to the specified location in the on-chip interactive cache via a write request and write confirmation timing sequence.
[0099] According to the technical solution provided in this disclosure, by setting a handshake interaction interface in the reconfigurable computing core and cooperating with the independent read and write addresses allocated by the array scheduler in the on-chip interactive cache, each reconfigurable computing core can actively initiate a synchronous handshake read transaction through this handshake interaction interface, obtain its own task data from the dedicated cache space of the on-chip interactive cache, and write the result back to the dedicated cache space of the on-chip interactive cache through a handshake write transaction after the operation is completed. Since the scheduler allocates non-overlapping read / write addresses to different reconfigurable computing cores, combined with the output enable control of the handshake interface, multiple reconfigurable computing cores can simultaneously exchange data with the on-chip interactive cache through the data bus without write conflicts or electrical contention, thereby improving the reliability of data transmission between the reconfigurable computing core and the on-chip interactive cache and reducing transmission latency.
[0100] In one embodiment of this disclosure, the corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the computation steps in the target task corresponding to the reconfigurable computing core.
[0101] In one implementation of this disclosure, the task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue. This can be understood as follows: when the array scheduler operates in multi-transaction parallel mode, the array scheduler can allocate a complete, independent energy transaction from the current task queue to each reconfigurable computing core as the task currently being processed by each reconfigurable computing core. At this time, the task data to be processed by each reconfigurable computing core is all the original input data carried in the task currently being processed by that reconfigurable computing core, such as the transaction digest to be signed, the message block to be hashed and packaged, the plaintext to be encrypted, or the signature value to be verified. That is, each reconfigurable computing core is independently responsible for all the cryptographic operations of a transaction, and the data read by the reconfigurable computing core directly corresponds to the complete input of the entire transaction.
[0102] In one implementation of this disclosure, the task data includes task data for the computational steps corresponding to the reconfigurable computing cores in the target task. This can be understood as follows: when the array scheduler operates in single-transaction pipeline mode, a complete transaction containing multiple cryptographic operations is broken down into multiple sequentially executed computational steps. For example, a complete transaction can be broken down into three steps: SM2 signing, SM3 hashing, and SM4 encryption. In this case, each reconfigurable computing core is only responsible for one computational step, not the entire transaction. Therefore, the task data read by a single reconfigurable computing core is no longer the complete input of the entire transaction, but rather specific task data required only for the corresponding computational step undertaken by that reconfigurable computing core. For example, the task data required for the signing step includes the private key and message digest; the task data required for the hashing step includes the message block; and the task data required for the encryption step includes the transaction plaintext and the key. Through this splitting, multiple reconfigurable computing cores can serially and collaboratively complete different computational steps of the same transaction, achieving pipelined accelerated processing.
[0103] According to the technical solution provided in the embodiments of this disclosure, when the array scheduler is working in multi-task parallel mode, each reconfigurable computing core can directly obtain the complete input data of a task, thereby independently and in parallel processing multiple unrelated tasks, significantly improving task processing per unit time; when the array scheduler is working in single-task pipeline mode, each reconfigurable computing core only obtains the specific input data of a certain operation step it is responsible for, thereby enabling it to serially and collaboratively complete different operation processing steps of the same task with other cores, effectively reducing task processing latency.
[0104] In one embodiment of this disclosure, the plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm.
[0105] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm.
[0106] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm.
[0107] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
[0108] According to the technical solution provided in this disclosure, by configuring the level output of the register group to drive the multiplexer array, the cross-connection switch matrix, and the pipeline stage enable switch, multiple computational macrocells in the same set of computational macrocell pools can be dynamically reconstructed into dedicated hardware paths required by the corresponding algorithms under different configuration instructions. When processing SM2 signature verification tasks, multiple computational macrocells are reconstructed into an asymmetric computational pipeline of modular exponentiation and modular multiplication in series; when processing SM3 or SHA-256 hash tasks, multiple computational macrocells are reconstructed into a closed-loop iterative pipeline of hash compression iteration units; when processing SM4 encryption tasks, multiple computational macrocells are reconstructed into a group round function pipeline of S-box permutation and linear shift XOR units in series. Through the above scheme, a single reconfigurable computing core can switch flexibly between four cryptographic algorithms without integrating multiple independent fixed algorithm circuits, thereby significantly reducing chip area redundancy and improving hardware reuse rate. At the same time, since the reconstruction process is directly controlled by the register level, algorithm switching can be completed in a short time, reducing the latency caused by algorithm switching and improving the processing efficiency of hybrid algorithm tasks in distributed energy trading scenarios.
[0109] In one embodiment of this disclosure, the reconfigurable computing core further includes a key buffer.
[0110] The key buffer is used to store the public key, session key, and hash initialization vector.
[0111] According to the technical solution provided in this disclosure, by setting an independent key buffer within each reconfigurable computing core to store the public key, session key, and hash initialization vector, the reconfigurable computing core can directly read the required key data from the local key buffer when performing operations corresponding to the SM2, SM3, SM4, or SHA-256 algorithms. This eliminates the need for frequent access to on-chip interactive caches or remote storage, significantly reducing memory access latency and bus occupancy. Furthermore, storing keys and task data in different caches avoids resource conflicts caused by mixed access.
[0112] In one embodiment of this disclosure, the arithmetic macrounit includes a modular multiplication macrounit, a modular exponentiation macrounit, a substitution box macrounit, a linear shift XOR macrounit, a hash compression iteration macrounit, a register file macrounit, and a pipelined latch macrounit.
[0113] In one implementation of this disclosure, the modular multiplication macrounit is used to perform large number modular multiplication operations and is the core operation unit of the SM2 asymmetric encryption algorithm. The modular exponentiation macrounit is used to implement modular exponentiation operations based on modular multiplication iterations, and is used to perform encryption, decryption, signing, and verification operations of the SM2 asymmetric encryption algorithm. The substitution box macrounit, also known as the S-box permutation macrounit, is used to perform nonlinear permutations from 8-bit input to 8-bit output, and is a core component for implementing obfuscation characteristics in the SM4 block cipher algorithm. The linear shift XOR macrounit supports linear shift and XOR operations, and can be used for linear transformations in the SM4 round function, as well as for mixed operations in the hash compression process. The hash compression iteration macrounit has built-in multi-round addition, shift, and Boolean function circuits to adapt to the compression function iteration requirements of both SM3 and SHA-256 hash algorithms. The register file macrounit is used to uniformly store intermediate data during the operation process, reducing access to external caches. The pipeline latch macrounit is responsible for data timing and synchronization between pipeline stages, ensuring that data is transmitted between macrounits according to the correct clock cycle.
[0114] Among them, the input and output pins of the modular multiplication macrounit, modular exponentiation macrounit, substitution box macrounit, linear shift XOR macrounit, hash compression iterative macrounit, register file macrounit, and pipeline latch macrounit all adopt standardized design, with unified bit width and handshake timing, facilitating free interconnection through multiplexer arrays and cross-connection switch matrices. Within the arithmetic macrounit pool, these macrounits can be combined into dedicated arithmetic pipelines required for SM2, SM3, SM4, or SHA-256 algorithms by controlling the multiplexer array and cross-connection switch matrix through register configuration instructions issued by the configuration register, thereby achieving efficient hardware reuse of a single core for multiple algorithms.
[0115] According to the technical solution provided in this disclosure, by solidifying the computational macrounits in the computational macrounit pool into modular multiplication macrounits, modular exponentiation macrounits, substitution box macrounits, linear shift XOR macrounits, hash compression iteration macrounits, register file macrounits, and pipeline latch macrounits, and dynamically connecting them in the computational macrounit pool using a multiplexer array and a cross-connection switch matrix, and by controlling the multiplexer array and cross-connection switch matrix through level control of the register configuration instructions issued by the configuration register, the computational macrounits in the computational macrounit pool can be combined into dedicated computational pipelines required for the SM2, SM3, SM4, or SHA-256 algorithms commonly used in distributed energy trading, which helps to improve the efficiency of processing distributed energy trading tasks.
[0116] In one embodiment of this disclosure, sending the corresponding register configuration instruction to the corresponding reconfigurable computing core via the array shared configuration bus includes: The array shares a configuration bus to broadcast register configuration instructions to multiple reconfigurable computing cores.
[0117] Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
[0118] According to the technical solution provided in this disclosure, broadcasting register configuration instructions to multiple reconfigurable computing cores via an array shared configuration bus enables these cores to be reconfigured synchronously, avoiding the serial delay of individual configuration and improving the efficiency of reconfiguring reconfigurable computing cores in parallel scenarios. By sending register configuration instructions to corresponding reconfigurable computing cores via the array shared configuration bus, heterogeneous parallel computing can be achieved by configuring each reconfigurable computing core separately when different reconfigurable computing cores need to execute different algorithms.
[0119] In one embodiment of this disclosure, the array scheduling controller is further configured to: Obtain the current working status of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores.
[0120] If the overall computing load rate is greater than or equal to the first computing load rate threshold, a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores.
[0121] If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold.
[0122] The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: At least one reconfigurable computing core is identified as being in the startup state among the plurality of reconfigurable computing cores.
[0123] The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
[0124] In one implementation of this disclosure, the current operating state of the reconfigurable computing core refers to whether the reconfigurable computing core is currently in a startup state or a shutdown state. A startup state can be understood as the reconfigurable computing core being executed or configured to perform a computational task. A shutdown state can be understood as the clock of the reconfigurable computing core being turned off or its power supply being cut off.
[0125] In one implementation of this disclosure, the overall computing load rate can be understood as the proportion of the number of reconfigurable computing cores currently in a busy state to the total number of reconfigurable computing cores. By setting a dynamic power scheduling unit inside the array scheduler, this dynamic power scheduling unit can be responsible for real-time statistics on the current working state of each reconfigurable computing core and the overall computing load rate of multiple reconfigurable computing cores.
[0126] According to the technical solution provided in the embodiments of this disclosure, by dynamically monitoring the current working state and overall computing load rate of each reconfigurable computing core, and waking up the reconfigurable computing core in the off state when the computing load rate exceeds a first threshold, the parallel processing throughput can be increased; when the computing load rate is lower than a second threshold, the reconfigurable computing core that has been started but not assigned tasks can be shut down, which can avoid the power consumption waste caused by keeping each reconfigurable computing core constantly on.
[0127] This disclosure also provides a control method for a reconfigurable processor for distributed energy trading. The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0128] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0129] Figure 3 A flowchart illustrating a control method for a reconfigurable processor for distributed energy trading according to an embodiment of the present disclosure is shown. Figure 3 As shown, the control method for the reconfigurable processor used in distributed energy trading includes the following steps: In step S101, the array scheduling controller is controlled to obtain the current task queue of distributed energy trading, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0130] In step S102, the configuration register group is controlled to receive the register configuration instruction corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instruction, and output a level to the topology reconfiguration module according to the register configuration instruction.
[0131] In step S103, the multiplexer array is controlled to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group; the cross-connection switch matrix is controlled to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group; and the pipeline stage enable switch is controlled to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0132] In one embodiment of this disclosure, the reconfigurable computing core further includes a handshake interaction interface.
[0133] The control method further includes: The array scheduling controller allocates independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache.
[0134] The reconfigurable computing core is controlled to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, control the multiple computing macrounits to perform computing processing on the task data based on the computing pipeline, and control the reconfigurable computing core to write the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
[0135] In one embodiment of this disclosure, the corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the computation steps in the target task corresponding to the reconfigurable computing core.
[0136] In one embodiment of this disclosure, the plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm.
[0137] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm.
[0138] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm.
[0139] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
[0140] In one embodiment of this disclosure, the reconfigurable computing core further includes a key buffer.
[0141] The method further includes: The key buffer is controlled to store the public key, session key, and hash initialization vector.
[0142] In one embodiment of this disclosure, the arithmetic macrounit includes a modular multiplication macrounit, a modular exponentiation macrounit, a substitution box macrounit, a linear shift XOR macrounit, a hash compression iteration macrounit, a register file macrounit, and a pipelined latch macrounit.
[0143] In one embodiment of this disclosure, sending the corresponding register configuration instruction to the corresponding reconfigurable computing core via the array shared configuration bus includes: The array shares a configuration bus to broadcast register configuration instructions to multiple reconfigurable computing cores.
[0144] Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
[0145] In one embodiment of this disclosure, the method further includes: The array scheduling controller is instructed to perform the following steps: Obtain the current working status of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores.
[0146] If the overall computing load rate is greater than or equal to the first computing load rate threshold, a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores.
[0147] If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold.
[0148] The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: At least one reconfigurable computing core is identified as being in the startup state among the plurality of reconfigurable computing cores.
[0149] The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
[0150] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array-shared configuration bus that is physically isolated from the on-chip interactive cache. Since the configuration bus and data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
[0151] This disclosure also provides a control device for a reconfigurable processor for distributed energy trading. The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0152] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0153] Figure 4 A structural block diagram of a control device for a reconfigurable processor for distributed energy trading according to an embodiment of the present disclosure is shown. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both.
[0154] like Figure 4 As shown, the control device for the reconfigurable processor used in distributed energy trading includes: The instruction acquisition module 301 is configured to control the array scheduling controller to acquire the current task queue of distributed energy transactions, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0155] The level output module 302 is configured to control the configuration register group to receive the register configuration instruction corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instruction, and output a level to the topology reconfiguration module according to the register configuration instruction.
[0156] The unit reconfiguration module 303 is configured to control the multiplexer array to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group, control the cross-connection switch matrix to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group, and control the pipeline stage enable switch to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0157] According to the technical solution provided in this disclosure, the array scheduling controller obtains the current task queue and determines the operation type of each task, generates a corresponding register configuration instruction, and sends the instruction to the corresponding reconfigurable computing core via an array-shared configuration bus that is physically isolated from the on-chip interactive cache. Since the configuration bus and data transmission bus are physically isolated, forming two independent physical channels, the transmission of the register configuration instruction is not affected by data flow interference, resulting in a relatively complete signal and good timing stability for the register configuration instruction sent to the corresponding reconfigurable computing core. Furthermore, because the array scheduling controller can independently issue register configuration instructions at any time without waiting for the data transmission bus to be idle, the transmission latency of the register configuration instructions is low. After receiving and latching the register configuration instructions from the configuration register group within each reconfigurable computing core, the module directly outputs a level to the topology reconfiguration module. The multiplexer array in this module configures the input port connections of each computational macrounit based on this level, the cross-connection switch matrix configures the output port connections of each computational macrounit, and the inter-pipeline enable switch configures the clock of the pipeline latches in the computational pipeline. These three mechanisms work together to form a computational pipeline in the computational macrounit pool that matches the current computational logic. Through this scheme, the same group of computational macrounits can be reconfigured into the hardware paths required by different algorithms, without needing to pre-define independent fixed cores for each algorithm or replace hardware modules when switching algorithms. Furthermore, multiple independent reconfigurable computing cores can receive their respective configuration instructions in parallel and execute operations independently, thereby significantly improving the ability to process multiple transactions per unit time. In summary, this technical solution can effectively reduce chip area redundancy and improve the utilization rate of computing macrocells, thereby reducing the cost of processing distributed energy trading data. At the same time, since this solution does not involve frequent hardware module switching, it does not introduce additional bus arbitration and context save and restore overhead, thus improving the efficiency of processing distributed energy trading data.
[0158] This disclosure also discloses an electronic device, Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0159] like Figure 5 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to embodiments of the present disclosure.
[0160] This disclosure provides a control method for a reconfigurable processor for distributed energy trading. The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache.
[0161] The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells.
[0162] The control method includes: The array scheduling controller obtains the current task queue of distributed energy transactions, determines the operation type of each task in the current task queue, generates corresponding register configuration instructions according to the operation type of each task, and sends the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus.
[0163] The configuration register group is controlled to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions.
[0164] The multiplexer array is controlled to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group; the cross-connection switch matrix is controlled to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group; and the pipeline stage enable switch is controlled to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
[0165] In one embodiment of this disclosure, the reconfigurable computing core further includes a handshake interaction interface.
[0166] The control method further includes: The array scheduling controller allocates independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache.
[0167] The reconfigurable computing core is controlled to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, control the multiple computing macrounits to perform computing processing on the task data based on the computing pipeline, and control the reconfigurable computing core to write the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
[0168] In one embodiment of this disclosure, the corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the computation steps in the target task corresponding to the reconfigurable computing core.
[0169] In one embodiment of this disclosure, the plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm.
[0170] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm.
[0171] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm.
[0172] Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
[0173] In one embodiment of this disclosure, the reconfigurable computing core further includes a key buffer.
[0174] The method further includes: The key buffer is controlled to store the public key, session key, and hash initialization vector.
[0175] In one embodiment of this disclosure, the arithmetic macrounit includes a modular multiplication macrounit, a modular exponentiation macrounit, a substitution box macrounit, a linear shift XOR macrounit, a hash compression iteration macrounit, a register file macrounit, and a pipelined latch macrounit.
[0176] In one embodiment of this disclosure, sending the corresponding register configuration instruction to the corresponding reconfigurable computing core via the array shared configuration bus includes: The array shares a configuration bus to broadcast register configuration instructions to multiple reconfigurable computing cores.
[0177] Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
[0178] In one embodiment of this disclosure, the method further includes: The array scheduling controller is instructed to perform the following steps: Obtain the current working status of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores.
[0179] If the overall computing load rate is greater than or equal to the first computing load rate threshold, a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores.
[0180] If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold.
[0181] The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: At least one reconfigurable computing core is identified as being in the startup state among the plurality of reconfigurable computing cores.
[0182] The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
[0183] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown.
[0184] like Figure 6 As shown, the computer system includes a processing unit that can execute various methods described above based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0185] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processes via a network such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.
[0186] In particular, according to embodiments of this disclosure, the methods described above can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for performing the methods described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.
[0187] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0188] The units or modules described in the embodiments of this disclosure can be implemented in software or programmable hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.
[0189] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system described above; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to perform the methods described in this disclosure.
[0190] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A reconfigurable processor for distributed energy trading, characterized in that, It includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache, wherein the array shared configuration bus and the data transmission bus of the on-chip interactive cache are physically isolated, the array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache; The array scheduling controller is used to obtain the current task queue of distributed energy trading, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus. The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells. The configuration register group is used to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions; The multiplexer array is used to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group. The cross-connection switch matrix is used to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group. The pipeline stage enable switch is used to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
2. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The reconfigurable computing core also includes a handshake interaction interface; The array scheduling controller is also used to allocate independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache; The reconfigurable computing core is also used to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, and the multiple computing macro units perform computing processing on the task data based on the computing pipeline. The reconfigurable computing core writes the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
3. The reconfigurable processor for distributed energy trading according to claim 2, characterized in that, The corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the operation steps in the target task corresponding to the reconfigurable computing core.
4. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
5. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The reconfigurable computing core also includes a key buffer; The key buffer is used to store the public key, session key, and hash initialization vector.
6. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The computational macrounits include modular multiplication macrounits, modular exponentiation macrounits, substitution box macrounits, linear shift XOR macrounits, hash compression iteration macrounits, register file macrounits, and pipelined latch macrounits.
7. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: The array shared configuration bus broadcasts register configuration instructions to multiple reconfigurable computing cores. Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
8. The reconfigurable processor for distributed energy trading according to claim 1, characterized in that, The array scheduling controller is also configured to: Obtain the current working state of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores; If the overall computing load rate is greater than or equal to the first computing load rate threshold, then a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores; If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold. The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: Among the plurality of reconfigurable computing cores, at least one reconfigurable computing core is identified as being in the startup state; The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
9. A control method for a reconfigurable processor for distributed energy trading, characterized in that, The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache. The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells. The control method includes: The array scheduling controller is controlled to obtain the current task queue of distributed energy trading, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus. The configuration register group is controlled to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions; The multiplexer array is controlled to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group; the cross-connection switch matrix is controlled to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group; and the pipeline stage enable switch is controlled to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
10. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The reconfigurable computing core also includes a handshake interaction interface; The control method further includes: The array scheduling controller allocates independent read and write addresses to the corresponding reconfigurable computing cores in the on-chip interactive cache; The reconfigurable computing core is controlled to read corresponding task data from the on-chip interactive cache according to the corresponding read address through the handshake interaction interface, send the task data into the computing pipeline, control the multiple computing macrounits to perform computing processing on the task data based on the computing pipeline, and control the reconfigurable computing core to write the computing processing result into the on-chip interactive cache according to the corresponding write address through the handshake interaction interface.
11. The control method for a reconfigurable processor for distributed energy trading according to claim 10, characterized in that, The corresponding task data includes the task data of the target task corresponding to the reconfigurable computing core in the current task queue, or the task data of the operation steps in the target task corresponding to the reconfigurable computing core.
12. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The plurality of computational macrounits form a computational pipeline that matches the corresponding computational logic, including: The multiple computational macrounits form a computational pipeline that matches the SM2 asymmetric encryption algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM3 hash iteration algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SM4 block cipher algorithm; Alternatively, the plurality of computational macrounits may form a computational pipeline that matches the SHA-256 hash algorithm.
13. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The reconfigurable computing core also includes a key buffer; The method further includes: The key buffer is controlled to store the public key, session key, and hash initialization vector.
14. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The computational macrounits include modular multiplication macrounits, modular exponentiation macrounits, substitution box macrounits, linear shift XOR macrounits, hash compression iteration macrounits, register file macrounits, and pipelined latch macrounits.
15. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: The array shared configuration bus broadcasts register configuration instructions to multiple reconfigurable computing cores. Alternatively, register configuration instructions can be sent to the corresponding reconfigurable computing core via the array shared configuration bus.
16. The control method for a reconfigurable processor for distributed energy trading according to claim 9, characterized in that, The method further includes: The array scheduling controller is instructed to perform the following steps: Obtain the current working state of each of the plurality of reconfigurable computing cores and the overall computing load rate of the plurality of reconfigurable computing cores; If the overall computing load rate is greater than or equal to the first computing load rate threshold, then a start command is sent to the reconfigurable computing cores that are in the off state among the plurality of reconfigurable computing cores; If the overall computing load rate is less than or equal to the second computing load rate threshold, a shutdown command is sent to the reconfigurable computing cores that are in the startup state and currently have no assigned tasks, wherein the second computing load rate threshold is less than the first computing load rate threshold. The step of sending the corresponding register configuration instruction to the corresponding reconfigurable computing core through the array shared configuration bus includes: Among the plurality of reconfigurable computing cores, at least one reconfigurable computing core is identified as being in the startup state; The array shared configuration bus sends the corresponding register configuration instruction to the corresponding reconfigurable computing core in at least one of the reconfigurable computing cores that is in the startup state.
17. A control device for a reconfigurable processor used in distributed energy trading, characterized in that, The reconfigurable processor for distributed energy trading includes an array scheduling controller, multiple independent reconfigurable computing cores, an array shared configuration bus, and an on-chip interactive cache. The array shared configuration bus is physically isolated from the data transmission bus of the on-chip interactive cache. The array scheduling controller is connected to each of the multiple reconfigurable computing cores through the array shared configuration bus, and each reconfigurable computing core is connected to the on-chip interactive cache. The reconfigurable computing core includes a configuration register set, a topology reconstruction module, and a pool of computational macrocells. The pool of computational macrocells includes multiple computational macrocells. The configuration register set is connected to the topology reconstruction module. The topology reconstruction module includes a multiplexer array, a cross-connection switch matrix, and pipeline stage enable switches. The multiplexer array is connected to the input port of each of the multiple computational macrocells, and the cross-connection switch matrix is connected to the output port of each of the multiple computational macrocells. The control device includes: The instruction acquisition module is configured to control the array scheduling controller to acquire the current task queue of distributed energy transactions, determine the operation type of each task in the current task queue, generate corresponding register configuration instructions according to the operation type of each task, and send the corresponding register configuration instructions to the corresponding reconfigurable computing core through the array shared configuration bus. The level output module is configured to control the configuration register group to receive register configuration instructions corresponding to the reconfigurable computing core through the array shared configuration bus, latch the register configuration instructions, and output a level to the topology reconfiguration module according to the register configuration instructions; The unit reconfiguration module is configured to control the multiplexer array to configure the connection relationship of the input ports of each arithmetic macro unit according to the level output of the configuration register group, control the cross-connection switch matrix to configure the connection relationship of the output ports of each arithmetic macro unit according to the level output of the configuration register group, and control the pipeline stage enable switch to configure the clock of the pipeline latch in the arithmetic pipeline according to the level output of the configuration register group, so that the multiple arithmetic macro units form an arithmetic pipeline that matches the corresponding arithmetic logic.
18. An electronic device, characterized in that, It includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method of any one of claims 9-16.
19. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by a processor, the computer instructions implement the method of any one of claims 9-16.