Low-Storage-Overhead and High-Resource-Utilization Momentum Parallel Tempering Ising Processing Circuit

By designing a momentum parallel tempering Isin processing circuit of shared coefficient memory and time division multiplexing calculation circuit, the problems of high storage overhead and low resource utilization in the prior art are solved, and the effects of low storage overhead, high resource utilization and low hardware cost are achieved.

CN120069019BActive Publication Date: 2025-06-24SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510549440.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-24
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing Isin processing circuits have problems with high storage overhead and low resource utilization when dealing with combination optimization problems, especially under dynamically changing replica count requirements.

Method used

A momentum parallel tempered Ising processing circuit with low storage overhead and high resource utilization is designed to improve the utilization of hardware resources by sharing the same coefficient memory, time division multiplexing calculation circuit and storage circuit, flow processing spin update, local field compensation and replica state switching tasks.

Benefits of technology

It has achieved the reduction of storage resource overhead, the improvement of resource utilization and the reduction of hardware costs, and adapted to the combined optimization problem needs of different complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069019B_ABST
    Figure CN120069019B_ABST
Patent Text Reader

Abstract

The present invention discloses a momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization rate, and this solution is proposed in view of the problem of low hardware utilization rate in the prior art. The top-level control module decodes the input instructions through an instruction decoder; the parameter-copy mapping table records the relationship between annealing parameters and activated copies; the spin update and copy exchange module processes spin clusters one by one in a pipeline manner; the local field compensation calculation module processes one spin cluster each time; the linear feedback shift register array provides pseudo-random numbers for decision-making; the copy memory stores the spin states of each copy; the local field memory stores the unilateral local fields of each copy; the parameter memory stores the operating parameters of each copy; the external field and momentum coupling coefficient memory stores the external fields and momentum coupling coefficients of each spin; the spin coupling coefficient memory stores the interaction coefficients. The advantages are small storage resource overhead, high resource utilization rate, and low hardware cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Ising models, and in particular to a momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization rate. Background Art

[0002] Combinatorial optimization problems are a class of problems that seek the optimal solutions to combinatorial problems, and are involved in various fields such as economic management, industrial engineering, and communication networks, such as path planning problems, resource allocation problems, etc. Since most of these problems are non-deterministic polynomial problems, the corresponding solution space will grow exponentially with the increase of the problem scale. Therefore, modern von Neumann computers have huge resource overheads when dealing with combinatorial optimization problems and it is difficult to obtain the optimal solutions to combinatorial optimization problems quickly and efficiently.

[0003] The Ising model is one of the most classical models in statistical physics, which uses a lattice to describe the phase transition phenomenon of ferromagnetic substances. Among them, each lattice point is occupied by a spin in an up or down state, and the total energy of the system is jointly composed of the interactions between spins and the action of an external field. Due to the high abstraction of the Ising model, it can simulate a wide range of complex phenomena. Therefore, combinatorial optimization problems can be mapped to the Ising model for solution. Based on quantum computing, the Ising quantum annealing processing architecture can find the optimal solution to a problem quickly while maintaining high precision when solving combinatorial optimization problems. However, due to the strict requirements on the working environment temperature and being limited by the scale of the problem to be solved, its practical application faces insurmountable challenges. In contrast, with the maturity and development of semiconductor technology, the CMOS-based Ising annealing processor provides great potential for combinatorial optimization problems and has strong adaptability, low cost, and high stability. However, most Ising processors focus more on the locally connected Ising model. Due to the sparsity of its spin connections, it is greatly limited in application scenarios. Although the fully connected Ising model can be mapped to the locally connected Ising model through a certain algorithm, the cost is to use more spins, increasing the consumption of hardware resources and reducing the hardware implementation efficiency. In addition, the annealing processing architecture based on the fully connected Ising model can effectively solve many combinatorial optimization problems. However, since this architecture uses the traditional simulated annealing algorithm, it can only update a single spin in each iteration, greatly increasing the time cost. In addition, when the system is at a local minimum and the energy barrier is relatively large, the simulated annealing algorithm will greatly reduce the possibility of the system escaping from the local minimum at this time, which will make it difficult for the system to approach the optimal solution to the problem.

[0004] The parallel tempering algorithm, a Monte Carlo method with multi-Markov chain coupling, is used in many disciplinary fields such as statistics, biology, and materials science. The parallel tempering algorithm can simultaneously run M sets of replicas, with each replica corresponding to a temperature T, and the temperatures need to increase sequentially. Through replica exchange in the parallel tempering algorithm, replicas at high temperatures will have a higher probability of jumping out of local minima to explore more energy states, while replicas at low temperatures will have a higher possibility of approaching the global minimum.

[0005] The parallel tempering Ising processing circuit aims to quickly solve combinatorial optimization problems using the parallel tempering algorithm. The current distributed Ising processing circuit independently configures computing units and dedicated coupling coefficient memories for each replica, and under the momentum annealing algorithm, 2 independent spin state memories need to be deployed for the left and right spin groups respectively, which results in extremely high memory overhead. In addition, the current processing circuit only supports a pre-fixed number of replicas. However, the optimal number of replicas required for combinatorial optimization problems of different complexities is dynamically changing. When the actually required optimal number of replicas does not match the fixed number of replicas it supports, it will lead to low hardware resource utilization. Generally speaking, each replica processing circuit mainly consists of a spin update and a local field compensation sub-circuit. There is a mutual restriction between the spin update and local field compensation operations in the traditional replica circuit. When performing spin updates, the local field compensation unit is idle and must wait for the spin update to complete before starting the local field compensation operation; when performing local field compensation calculations, the spin update unit is also idle and must wait for the local field compensation to complete before performing the next spin update, which further reduces the hardware resource utilization. Summary of the Invention

[0006] The purpose of the present invention is to provide a momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization to solve the problems existing in the above-mentioned prior art.

[0007] The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization in the present invention includes the following modules:

[0008] The top-level control module decodes the input instructions through an instruction decoder to complete operations such as initial data loading, running debugging, and result return;

[0009] The parameter-replica mapping table is used to record the mapping relationship between different annealing parameters and the activated replicas;

[0010] The spin update and replica exchange module processes spin clusters one by one in a pipelined manner, with two modes: spin update and replica exchange, which are dynamically switched by the mode control signal provided by the top-level control module;

[0011] A local field compensation calculation module, which is used for the compensation calculation of the local field, processes one spin cluster each time, and loads the next pair of spin clusters after the current pair of spin clusters is processed;

[0012] A linear feedback shift register array, which is used to provide pseudo-random numbers for the random discard mechanism and the Metropolis criterion determination;

[0013] A replica memory, which is used to store the spin states of each replica, reads and outputs a pair of spin clusters on the left and right sides respectively each time ( ), and updates one side of the spin cluster each time it writes ; S is the side variable identifier, and its value is L for the left side or R for the right side;

[0014] A local field memory, which is used to store the unilateral local fields of each replica , and only performs read and write operations on the local field of one spin cluster each time;

[0015] A parameter memory, which is used to store the operating parameters of each replica;

[0016] An external field and momentum coupling coefficient memory, which is used to store the external fields of each spin and the momentum coupling coefficient ;

[0017] A spin coupling coefficient memory, which is used to store the interaction coefficients between each spin .

[0018] In the momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization described in the present invention, its advantages are as follows: small storage resource overhead: all spin replicas share the same coefficient memory, reducing the storage complexity of the fully connected spin coupling coefficient from to , greatly reducing the memory resource consumption. High resource utilization: all spin replicas time-division multiplex the same computing circuit and storage circuit. When the optimal number of replicas required by the actual problem changes, the computing circuit processes different replicas in a time-division manner, improving the hardware resource utilization. By pipelining the spin update, local field compensation, and replica state exchange tasks, the hardware resource utilization is further improved. Low hardware cost: all spin replicas are processed in the same computing circuit, eliminating the communication circuit between different spin replicas compared with the traditional distributed multi-replica computing circuit. Brief Description of the Drawings

[0019] Figure 1 is a schematic structural diagram of the momentum parallel tempering Ising processing circuit described in the present invention.

[0020] Figure 2It is a schematic diagram of the structure and processing flow of the local field compensation calculation module described in the present invention.

[0021] Figure 3 It is a schematic diagram of the structure of the system energy change - negative spin energy calculation unit described in the present invention.

[0022] Figure 4 It is a schematic diagram of the structure of the acceptance probability calculation - decision unit described in the present invention.

[0023] Figure 5 It is a schematic diagram of the structure of the exchange calculation unit described in the present invention. Detailed implementation manners

[0024] As Figures 1 to 5 shown, the low - storage - overhead and high - resource - utilization momentum parallel tempering Ising processing circuit described in the present invention includes multiple sub - circuits: a top - layer control module, a parameter - replica mapping table, a spin update and replica exchange module, a local field compensation calculation module, a linear feedback shift register array, a replica memory, a local field memory, a parameter memory, an external field and momentum coupling coefficient memory, and a spin coupling coefficient memory.

[0025] The top - layer control module consists of an instruction decoder, an interface, several communication - type state machines, a load - aware unit, and a counter bank. Among them, the communication - type state machines include a global state machine, an update and exchange state machine, and a compensation state machine. The top - layer control module decodes the input instructions through the instruction decoder to complete operations such as initial data loading, running debugging, and result return. Through the number of replicas required by the actual demand and the number of actually mapped and activated spin clusters within each replica to regulate multiple communication - type state machines, generate corresponding various control signals, and implement the spin ground - state configuration search process under the target replica exchange times and the number of iterations within a single replica. In this embodiment , .

[0026] The parameter - replica mapping table is mainly composed of a register bank, which records the mapping relationship between different annealing parameters and the activated replicas. The position of the register in the table represents the address index from the high - temperature parameter to the low - temperature parameter, and the data stored in the register represents the address index of the replica at this temperature parameter.

[0027] The spin update and replica exchange module processes spin clusters one by one in a pipelined manner, with two modes: spin update and replica exchange, which are dynamically switched by the mode control signal provided by the top-level control module. In the spin update mode, the spin clusters with updated states are written back to the replica memory in real time; in the replica exchange mode, the final exchange decision signal is sent to the parameter-replica mapping table, and the state exchange is achieved by exchanging the address indexes in the corresponding memories of adjacent temperature replicas while keeping the physical storage locations unchanged.

[0028] The local field compensation calculation module is used for the compensation calculation process of the local field. It processes one spin cluster each time, and loads the next pair of spin clusters after the current pair of spin clusters is processed. In particular, in the initialization stage, the local field register inside it serves as an intermediate register, and is coordinated by the top-level control module to perform initialization write operations for the random number seed and different memories.

[0029] The linear feedback shift register array provides pseudo-random numbers for the determination of the random dropout mechanism of the momentum coupling effect in the spin update process, the determination of the Metropolis criterion for whether to flip the spin, and the determination of the Metropolis criterion for whether to exchange the adjacent replica configurations in the replica exchange process.

[0030] The replica memory stores the spin states of 32 replicas. Each replica contains 16 pairs of spin clusters on both sides. Each time it reads out and outputs a pair of spin clusters on the left and right sides ( ), and each time it writes and updates one side of the spin cluster ; S is the side variable identifier, taking values of left L or right R. Among them, the local field compensation calculation module has access priority. In hardware implementation, the state values of the spin +1 or -1 are stored as 0 and 1 respectively. At initialization, according to the number of replicas actually mapped for processing , the number of spin clusters actually mapped and activated in each replica , the corresponding spin states are reset to +1, that is, the corresponding memories are reset to 0.

[0031] The local field memory stores the unilateral local fields of 32 replicas , and only reads and writes the local field of one spin cluster each time. This module also gives priority to serving the local field compensation calculation module. At initialization, through an external instruction, the left local fields of the actually processed spins are written into the local field storage locations corresponding to replicas in units of spin clusters.

[0032] The parameter memory stores the operating parameters of 32 replicas, including: the temperature required for the spin update process , the momentum scaling factor and random drop rate , and the corresponding parameters of the inverse temperature difference of the replicas required for the replica exchange process During initialization, external instructions are used to The parameters of the copies are written to memory.

[0033] The external field and momentum coupling coefficient memory stores the external fields of 2048 spins and momentum coupling coefficient , where the external field For the replica exchange process, the momentum coupling coefficient Used for spin update process. Supports reading the external field or momentum coupling coefficient corresponding to one spin cluster at a time. During initialization, external instructions are used to The fields and coefficients corresponding to the spin clusters are written to the specified locations.

[0034] The spin coupling coefficient memory stores the interaction coefficients between 2048 spins. , which is dedicated to the local field compensation calculation process. In order to achieve parallel computing, it is divided into The coefficient storage unit corresponding to each sub-spin cluster uses dual-port RAM to support synchronous reading. The coupling coefficients between the spins are written into the corresponding locations in the memory.

[0035] The main function of the local field compensation calculation module is to perform compensation calculation on the local field of the spin on the other side of the replica after the spin update on one side is completed. This calculation provides support for the update processing of the spin on the other side or the energy calculation of the replica. Its core calculation circuit is composed of a local field compensation accumulator array, an XOR gate calculation array, a result register array, and a mixed priority encoder. The local field compensation accumulator array adopts a pipeline processing method to periodically complete the local field compensation calculation of all replicas from the right spin to the left spin in the order from the high-temperature replica to the low-temperature replica. It does this by Based on the combination of the different spin states on both sides The compensation calculation can realize the local field value on the other side The specific circuit processing steps are as follows:

[0036] Step 1: Take the spin cluster as the basic processing unit and read the actual required The spin states corresponding to the spin clusters , contralateral spin state By sending a pair of two-sided spin clusters into the XOR gate calculation array, we can obtain a result register array representing spin pairs in different states. Obviously, the XOR operation result of the same-state spin pair is 0, and the XOR operation result of the different-state spin pair is 1.

[0037] Step 2: Use a hybrid priority encoder to search for the addresses of the "1"s in the result array respectively. Each bit hybrid priority encoder contains a most significant bit first encoder and a least significant bit first encoder, and can output the addresses of the most significant bit and least significant bit "1"s simultaneously.

[0038] Step 3: According to the obtained addresses of different state spin pairs, set the "1"s at the corresponding positions in the result register array to zero respectively to ensure the correctness of the address search in the next cycle. At the same time, according to these addresses, read the different state spin states of the corresponding side respectively , and access the corresponding coefficient storage unit to obtain the interaction coefficients of these spins with other spins .

[0039] Step 4: Send these corresponding spin states and spin coupling interaction coefficients to the activated local field compensation accumulators. First, multiply them through a spin multiplier, and then accumulate them based on the current unilateral local field, so as to realize the local field compensation calculation.

[0040] The spin update and replica exchange module consists of a system energy change - negative spin energy calculation unit, an acceptance probability calculation - decision unit, and an exchange calculation unit. The system energy change - negative spin energy calculation unit is the pre - core calculation unit in the spin state update and replica exchange process. This unit controls a two - way selector through the calculation mode signal calculation_mode, so as to support two calculation modes: spin update and replica exchange. In the spin update mode, first calculate the spin on the current update side . The relationship between the system energy change and the spin local field can be calculated by the following formula:

[0041]

[0042] Among them, the logical operation and the multiplication operation represent momentum regulation. The corresponding circuit calculation process is as follows: The random dropout rate is compared with the random number through a comparator. When , the selector is controlled to output , otherwise output 0. Contains two multiplication operations. The multiplication of the momentum scaling factor and ( )( ) is a conventional multiplication operation. For another multiplication operation, since the spin state can only take values of +1 or -1 and is stored as 0 or 1 in hardware respectively, it can be implemented with relatively low resources through a spin multiplier. Subsequently, when the calculation mode signal calculation_mode = 0, the control selector outputs term. Through the adder and added together, the result is multiplied by output by another selector, and finally the scaled system energy change is obtained. In the replica exchange mode, it is necessary to calculate the Hamiltonian corresponding to the right spin configuration in the original Ising model, and this Hamiltonian is used as the basis for measurement during the exchange process. The Hamiltonian of the system is jointly determined by the interaction energy between all spins and the energy of the external field acting on the spins. Among them: the external field acts directly on the spins, and this part of the energy can be directly regarded as what the spins possess; the interaction energy between any two spins can be evenly distributed to the two spins. Therefore, the final spin energy is calculated through the following equation:

[0043]

[0044] By summing up all the spin energies, the calculation of the Hamiltonian of the system can be realized. The corresponding circuit calculation process is as follows: when the calculation mode signal is 1, the control selector outputs , and through the adder and added together, then the result is multiplied by output by another selector, and finally the negative value of the energy possessed by the spin can be obtained.

[0045] The acceptance probability calculation - decision unit can output the final acceptance or rejection decision on whether to accept the spin state flip or replica configuration exchange. This circuit avoids complex logarithmic function operations and multiplication operations, greatly reducing resource consumption. This unit consists of a shifter and a comparator controlled by a random number. The approximate calculation of logarithms and multiplications is realized through the shifter controlled by a random number. The shifter controlled by a random number consists of a selector, multiple left shifters and right shifters.

[0046] The exchange calculation unit uses the Hamiltonian corresponding to all spin configurations on the right side of the copy in the original Ising model as the basis for measuring the Hamiltonian during the exchange process. In addition, the number of iterations for each copy is even: the first iteration is used to update the spins on the left side of all copies, and the last iteration is used to update the spins on the right side of all copies. When the last right-side spin update of all copies is completed, the locally field values after compensation from the high-temperature copy to the low-temperature copy will be calculated in chronological order. At this time, the system Hamiltonians from the high-temperature copy to the low-temperature copy can be calculated serially in sequence, and an attempt can be made to complete the spin configuration exchange of adjacent copies in odd or even copy pairs.

[0047] For those skilled in the art, various corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all such changes and deformations should fall within the protection scope of the claims of the present invention.

Claims

1. A momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization, characterized in that: Includes the following modules: The top-level control module decodes the input instructions through the instruction decoder to complete the initialization data loading, running debugging and result return operations; A parameter-replica mapping table, used to record the mapping relationship between different annealing parameters and activated replicas; A spin update and replica exchange module processes spin clusters one by one in a pipeline manner, has two modes: spin update and replica exchange, and is dynamically switched by a mode control signal provided by the top-level control module; The local field compensation calculation module is used for the local field compensation calculation. It processes one spin cluster at a time. After the current pair of spin clusters is processed, the next pair of spin clusters is loaded. A linear feedback shift register array for providing pseudo-random numbers for the random drop mechanism and the Metropolis criterion; The replica memory is used to store the spin state of each replica, and each time it reads and outputs a pair of spin clusters on the left and right sides ( ), each write updates one side of the spin cluster ; S is the side variable identifier, and its value is L on the left or R on the right; Local field memory, used to store the single-sided local field of each replica , only the local field of one spin cluster is read and written at a time; Parameter storage, used to store the operating parameters of each replica; External field and momentum coupling coefficient memory, used to store the external field of each spin and momentum coupling coefficient ; Spin coupling coefficient memory, used to store the interaction coefficients between individual spins .

2. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1 is characterized in that: The parameter-copy mapping table includes a register group; the position of the register in the table represents the address index from the high temperature parameter to the low temperature parameter, and the data stored in the register represents the address index of the copy under the corresponding temperature parameter.

3. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: The top-level control module activates the number of replicas according to actual demand , the number of spin clusters actually mapped and activated in each replica To regulate several communicating state machines, generate corresponding control signals, and realize the spin ground state configuration search processing under the target number of replica exchanges and the number of iterations within a single replica.

4. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: In the spin update mode, the spin update and replica exchange module writes the spin cluster that has completed the state update back to the replica memory in real time; in the replica exchange mode, the final exchange decision signal is sent to the parameter-replica mapping table, and the state exchange is realized by exchanging the address indexes in the corresponding memories of adjacent temperature replicas, while keeping the physical storage location unchanged.

5. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: During the initialization phase of the local field compensation calculation module, the internal local field register is used as an intermediate register and coordinated by the top-level control module to perform initialization write operations for random number seeds and different memories.

6. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: The determination of the linear feedback shift register array includes: random discard mechanism determination of momentum coupling in the spin update process, Metropolis criterion determination of spin flipping or not, and Metropolis criterion determination of whether adjacent replica configurations are exchanged or not in the replica exchange process.

7. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: In the replica memory, spin The state value +1 is stored as 0, and the state value -1 is stored as 1; at initialization, the number of copies processed by the actual mapping , the number of spin clusters actually mapped and activated in each replica The corresponding spin state is reset to +1, that is, the corresponding memory is reset to 0.

8. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to claim 1, characterized in that: When the local field memory is initialized, the actual processing The local field of the left spin Write in spin cluster units The local field storage location corresponding to each copy.

9. The momentum parallel tempering Ising processing circuit with low storage overhead and high resource utilization according to any one of claims 1 to 8, characterized in that: The local field compensation calculation process specifically includes the following steps: Step 1: Take the spin cluster as the basic processing unit and read the actual required The spin states corresponding to the spin clusters , contralateral spin state ; Send a pair of two-sided spin clusters into the XOR gate calculation array to obtain a result register array representing spin pairs in different states; the XOR operation result of the same-state spin pair is 0, and the XOR operation result of the heterogeneous spin pair is 1; Step 2: Use the mixed priority encoder to search the addresses of the results that are 1 in the result register array respectively; Step 3: According to the addresses of the obtained spin pairs in different states, the 1s at the corresponding positions in the result register array are set to zero respectively; at the same time, according to the addresses, the corresponding Different spin states on the side , and access the corresponding coefficient storage unit to obtain the corresponding spin action coefficient ; Step 4: The corresponding spin state and the spin coupling coefficient Enter the activated In the local field compensation accumulators, they are first multiplied by the spin multiplier and then accumulated on the basis of the current unilateral local field, thereby realizing the local field compensation calculation.

Citation Information

Patent Citations

  • Full-connection Isin model annealing processing circuit based on parallel tempering

    CN116151171A

  • Full-connection Isin model reconfigurable processing circuit supporting multiple algorithms

    CN117951427A