High-energy-efficiency digital-analog hybrid CMOS (complementary metal oxide semiconductor) simulated annealing Isin chip and application thereof
Through the combined computing architecture of delay chains and addition trees of a high-efficiency hybrid digital-analog CMOS simulated annealing Ising chip, the problems of insufficient flexible configuration and area utilization of CMOS simulated annealing Ising chips in the existing technology are solved, and a high-efficiency and efficient combined optimization solution is achieved.
Patent Information
- Application Number
- CN202510764643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-12
AI Technical Summary
Existing CMOS simulated annealing Ising chips have deficiencies in flexible configuration of computing resources and area utilization, making it difficult to efficiently solve large-scale combinatorial optimization problems.
A high-efficiency hybrid digital-analog CMOS simulated annealing Ising chip is used. Through the combined computing architecture of delay chains and addition trees, combined with the simulated annealing algorithm, flexible configuration and high-efficiency computing are achieved. Delay chains are used for coarse-grained computing, and then addition trees are used for fine-grained computing to find the optimal solution to the combinatorial optimization problem.
It achieves flexible adaptation to the solution needs of combinatorial optimization problems of different scales, improves the energy efficiency and area utilization of the chip, and reduces energy consumption.
Smart Images

Figure CN120633786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of analog processing, and in particular to a high-energy-efficiency digital-analog hybrid CMOS simulated annealing Ising chip and a method for using the chip to solve large-scale combinatorial optimization problems. Background Art
[0002] Combinatorial optimization problems are widespread in society and industrial production, such as route planning, logistics management, communication networking, drug discovery, chip design, and financial optimization. Solving these problems has always been key to improving efficiency and promoting innovation. However, these optimization problems often exhibit the characteristic of being non-deterministic and polynomially hard, meaning that exact solutions are difficult to find in polynomial time. This poses significant challenges to traditional computing.
[0003] Annealing algorithms based on the Ising model are considered an effective method for solving large-scale combinatorial optimization problems. The Ising model is a physical model that describes phase transitions in ferromagnetic systems. By mapping real-world combinatorial optimization problems into the Ising model and leveraging the natural annealing process of spin states to find the ground state (i.e., the lowest energy state) of the Hamiltonian, optimal or near-optimal solutions can be obtained. In recent years, quantum annealing computers, such as D-Wave, have achieved breakthroughs, demonstrating quantum advantages in solving complex combinatorial optimization problems. Google and NASA have released test results showing that the D-Wave quantum computer can perform computations 100 million times faster than classical computers for certain specific problems. However, quantum annealing computers, based on superconducting qubit technology, have extremely low operating temperatures (15mK) and high costs, which severely limit their large-scale application, particularly in resource-constrained edge applications. Therefore, developing highly parallel and energy-efficient CMOS Ising chips based on existing CMOS processes is crucial for efficiently solving combinatorial optimization problems.
[0004] Prior art document CN117572931A discloses a high-energy-efficiency CMOS simulated annealing Ising chip, comprising a processor core, SRAM peripheral read / write driver circuits connected to the processor core for reading and writing data, and a spin state readout circuit for reading spin state information. The processor core includes several processing units, which are arranged using a Wangtu layout, with adjacent processing units connected using a top-level splicing approach. However, due to its all-digital design, this document cannot achieve flexible configuration based on the computing resource requirements of the problem being solved, limiting the chip's energy efficiency. Furthermore, the chip's area utilization still needs to be improved. Summary of the Invention
[0005] In view of the above situation, the main purpose of the present invention is to propose a high-energy-efficiency digital-analog hybrid CMOS simulated annealing Ising chip to solve the above technical problems.
[0006] The present invention provides a high-energy-efficiency digital-analog hybrid CMOS simulated annealing Ising chip, comprising a processor core, an SRAM peripheral read / write drive circuit respectively connected to the processor core for reading and writing data, and a spin state readout circuit for reading spin state bit information. The processor core comprises a plurality of processing units, which are laid out in a King's graph manner, and adjacent processing units are connected in a top-level splicing manner. The processing units comprise a delay chain circuit for analog calculation, a sampling circuit for sampling delay chain calculation results, an addition tree circuit for digital calculation, a coefficient storage array circuit for coupling coefficient storage, a multiplication circuit for coupling coefficient multiplication operation, a spin storage circuit for spin state storage, and an interface and instruction peripheral control circuit for spin state connection. The coupling coefficient storage format is binary complement.
[0007] The present invention also provides a method for using this high-efficiency hybrid digital-analog CMOS simulated annealing Ising chip to solve large-scale combinatorial optimization problems. Specifically, in the initial stage of solving the large-scale combinatorial optimization problem, the method first performs a coarse-grained calculation of the Hamiltonian of the large-scale combinatorial optimization problem based on simulated calculations using a delay chain, thereby obtaining a local minimum for the search. After the Hamiltonian of the simulated calculation iteratively reaches a local minimum, a fine-grained calculation of the Hamiltonian is initiated using digital calculations using an additive tree, and the simulated annealing algorithm is used to search for an optimal solution or a near-optimal solution to the large-scale combinatorial optimization problem.
[0008] The details are as follows:
[0009] A high-efficiency digital-analog hybrid CMOS simulated annealing Ising chip comprises a processor core, an SRAM peripheral read / write drive circuit respectively connected to the processor core for reading and writing data, and a spin state readout circuit for reading spin state bit information. The processor core comprises a plurality of processing units, which are laid out in a king diagram manner, and adjacent processing units are connected in a top-level splicing manner. The processing units comprise a delay chain circuit for analog calculation, a sampling circuit for sampling delay chain calculation results, an addition tree circuit for digital calculation, a coefficient storage array circuit for coupling coefficient storage, a multiplication circuit for coupling coefficient multiplication operations, a spin storage circuit for spin state storage, and an interface and instruction peripheral control circuit for spin state connection. The coupling coefficient is stored in a binary complement format.
[0010] As a preference, the layout of the Wang diagram includes five connection relationships: left and right, up and down, left oblique angle, right oblique angle and local magnetic coupling strength.
[0011] Preferably, adjacent processing units are connected via a spin state interface.
[0012] Preferably, the analog calculation uses delay units connected step by step to form a delay chain, and the delay chain is divided into a calculation delay chain and a reference delay chain. The calculation result of the multiplication circuit is sent to the calculation delay chain to complete the Hamiltonian calculation.
[0013] Preferably, a plurality of processing units in a column multiplex a reference delay chain, and an output end of the reference delay chain provides reference delay information to the plurality of processing units simultaneously.
[0014] Preferably, the digital calculation is completed using an addition tree circuit, and the calculation result of the multiplication circuit is sent to the addition tree to complete the Hamiltonian calculation.
[0015] Preferably, the delay unit includes a first PMOS transistor, a second PMOS transistor, a first NMOS transistor, a second NMOS transistor, and a third NMOS transistor.
[0016] Preferably, the source of the first PMOS tube is connected to the drain of the first NMOS tube;
[0017] The source of the first NMOS tube is connected to the drain of the second NMOS tube;
[0018] The gate of the second NMOS tube is connected to the input analog voltage;
[0019] The second PMOS transistor gate and the third NMOS transistor gate are connected to the first PMOS source and the first NMOS drain;
[0020] The source of the second PMOS tube is connected to the drain of the third NMOS tube.
[0021] Preferably, 9 delay units are connected end to end in sequence to form a delay chain, and the 9 delay units correspond to 8 connection directions of the king diagram and a local magnetic coupling strength calculation.
[0022] Preferably, each delay unit can generate four different delay lengths according to the input analog voltage.
[0023] Preferably, the storage array for storing the coupling coefficient sends the coupling coefficient to the multiplication calculation circuit, the multiplication circuit for multiplication calculation multiplies the coupling coefficient with the external input spin state, the instruction peripheral control circuit for controlling the calculation mode sends the multiplication result to the delay chain or the addition tree according to the calculation mode, and the delay chain or the addition tree for Hamiltonian accumulation calculation completes the calculation and sends the final spin state to the spin state storage circuit.
[0024] A method for using the above-mentioned high-efficiency hybrid digital-analog CMOS simulated annealing Ising chip to solve large-scale combinatorial optimization problems is characterized in that, in the initial stage of solving the large-scale combinatorial optimization problem, a coarse-grained calculation of the Hamiltonian of the large-scale combinatorial optimization problem is first performed based on analog calculation of a delay chain, thereby obtaining a local minimum for the search; after the Hamiltonian of the analog calculation iteratively falls into the local minimum, a fine-grained calculation of the Hamiltonian is initiated based on digital calculation of an addition tree, and the optimal solution or approximate optimal solution of the large-scale combinatorial optimization problem is searched in conjunction with a simulated annealing algorithm.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. Each adjacent processing unit of the present invention is connected to each other through a spin state interface, ensuring the integrity of the overall system connection relationship. Due to the design of using multiple processing units at the top level, the processor can flexibly adapt to the solution requirements of combinatorial optimization problems of different scales.
[0027] 2. The present invention adopts a mixed digital-analog computing architecture, which can be flexibly configured according to the computing resource requirements for solving the problem, thereby improving the energy efficiency performance of the CMOS Ising chip.
[0028] 3. The reference delay chain multiplexing strategy adopted by the present invention multiplexes a reference delay chain in the column direction, thereby reducing chip area overhead and improving chip area utilization.
[0029] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a schematic diagram of the structure of a high-efficiency mixed digital-analog CMOS simulated annealing Ising chip proposed by the present invention;
[0031] Figure 2 Schematic diagram of the spin and coupling coefficient of the connection relationship of the processing unit diagram of the present invention;
[0032] Figure 3 It is a structural diagram of the delay unit of the present invention;
[0033] Figure 4 1 is a structural diagram of the delay chain and the delay chain sampling circuit of the present invention;
[0034] Figure 5 A schematic diagram of the vertical reference delay chain multiplexing of the present invention;
[0035] Figure 6A structural diagram of a digital domain calculation addition tree according to the present invention;
[0036] Figure 7 This is a diagram of the generation process of the example verification picture of the present invention.
[0037] In the figure, PE. processing unit, P1. first PMOS transistor, P2. second PMOS transistor, N1. first NMOS transistor, N2. second NMOS transistor, N3. third NMOS transistor. DETAILED DESCRIPTION
[0038] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0039] These and other aspects of the embodiments of the present invention will become clear with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0040] See also Figure 1 An embodiment of the present invention provides a high-efficiency digital-analog hybrid CMOS simulated annealing Ising chip, including a processor core, an SRAM peripheral read / write driver circuit connected to the processor core for reading and writing data, and a spin state readout circuit for reading spin state information. The processor core includes a plurality of processing units PE, and the plurality of processing units are laid out in a king diagram manner, and adjacent processing units are connected in a top-level splicing manner.
[0041] See also Figure 2 In the above scheme, this embodiment uses 784 processing units, with the layout of the king diagram including five connection relationships: left-right, top-bottom, left-angled, right-angled, and local magnetic coupling strength. Each adjacent processing unit is interconnected via a spin state interface, ensuring the integrity of the overall king diagram connection relationship of the system. The design of using multiple processing units spliced together at the top level allows the processor to flexibly adapt to the needs of solving combinatorial optimization problems of varying scales.
[0042] The processing unit includes a delay chain circuit for analog calculation, a sampling circuit for sampling the delay chain calculation results, an addition tree circuit for digital calculation, a coefficient storage circuit for coupling coefficient storage, a multiplication circuit for coupling coefficient multiplication operation, a spin storage circuit for spin state storage, and an interface and instruction peripheral control circuit for spin state connection. The coupling coefficient storage format is binary complement, and the bit width of each coupling coefficient storage calculation unit is 4 bits.
[0043] See also Figure 3 In the above scheme, the delay chain unit of the present invention is a two-stage inverter cascade with an additional NMOS tube, and its specific structure includes a first PMOS tube P1, a second PMOS tube P2, a first NMOS tube N1, a second NMOS tube N2, and a third NMOS tube N3.
[0044] The source of the first PMOS tube is connected to the drain of the first NMOS tube;
[0045] The source of the first NMOS tube is connected to the drain of the second NMOS tube;
[0046] The gate of the second NMOS tube is connected to the input analog voltage;
[0047] The second PMOS transistor gate and the third NMOS transistor gate are connected to the first PMOS source and the first NMOS drain;
[0048] The source of the second PMOS tube is connected to the drain of the third NMOS tube;
[0049] Among them, 9 delay units are connected end to end to form a delay chain, and the 9 delay units correspond to the 8 connection directions of the king diagram and a local magnetic coupling strength.
[0050] See also Figure 4 The delay chain sampling circuit is a trigger, wherein the clock end of the trigger is connected to the last stage of the reference delay chain, and the data input end of the trigger is connected to the last stage of the calculation delay chain. At the beginning of the calculation, the reference delay chain and the first stage of the calculation delay chain simultaneously input a level rising signal. When the delay time of the reference delay chain is greater than that of the calculation delay chain, the data output end of the trigger is sampled to obtain a high-level signal, otherwise the data output end of the trigger is sampled to obtain a low-level signal. The data output end of the trigger sampling circuit is used as the analog domain calculation value result of the spin state.
[0051] See also Figure 5 In order to reduce area overhead and improve chip area utilization, in the array composed of processing units, 28 processing units in a vertical column reuse a reference delay chain, and the output end of the reference delay chain provides reference delay information for 28 processing units at the same time.
[0052] See also Figure 6,The digital domain calculation uses an addition tree circuit to accumulate ,the calculation results of the multiplication circuit in the digital ,domain, and the sign bit of the calculation result is used as the ,digital domain calculation value result of the spin state.
[0053] See also Figure 7 In order to verify the effectiveness of the present invention, a maximum cut problem was tested. In the initial stage, all spin states are randomly generated. The "12AB" picture with a pixel size of 28*28 is used as the lowest energy state of the system. The coupling coefficient mapping method is that the coupling coefficient at the junction of the letters and the background is the minimum value of the 4-bit coupling coefficient -7, and the coupling coefficient of the remaining non-junction parts is +7, and the local magnetic coupling strength is mapped to 0. The spin state distribution in the initial stage and the spin state distribution after annealing are shown in the figure. The initially randomly distributed spin state gradually collapses to the specified lowest energy state of the system. Under the 65nm CMOS process, the power supply voltage is 0.6V, and the energy consumed for each spin state update is only 1.2fJ. It can be concluded that the present invention can greatly reduce energy consumption.
[0054] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0055] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0056] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A high-efficiency digital-analog hybrid CMOS simulated annealing Ising chip, characterized by: The invention comprises a processor core, an SRAM peripheral read / write drive circuit respectively connected to the processor core for reading and writing data, and a spin state readout circuit for reading spin state bit information. The processor core comprises a plurality of processing units, which are laid out in a king diagram manner, and adjacent processing units are connected in a top-level splicing manner. The processing units comprise a delay chain circuit for analog calculation, a sampling circuit for sampling delay chain calculation results, an addition tree circuit for digital calculation, a coefficient storage array circuit for coupling coefficient storage, a multiplication circuit for coupling coefficient multiplication operation, a spin storage circuit for spin state storage, and an interface and instruction peripheral control circuit for spin state connection. The coupling coefficient storage format is binary complement.
2. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 1, characterized in that: The layout of the Wangtu includes five connection relationships: left and right, up and down, left oblique angle, right oblique angle and local magnetic coupling strength.
3. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 1, characterized in that: The analog calculation uses delay units connected step by step to form a delay chain, and the delay chain is divided into two parts: a calculation delay chain and a reference delay chain. The calculation result of the multiplication circuit is sent to the calculation delay chain to complete the Hamiltonian calculation.
4. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 3, characterized in that: A plurality of processing units in a vertical column multiplex a reference delay chain, and an output end of the reference delay chain provides reference delay information to the plurality of processing units simultaneously.
5. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 1, characterized in that: Digital calculations are completed using an addition tree circuit, and the calculation results of the multiplication circuit are sent to the addition tree to complete the Hamiltonian calculation.
6. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 3, characterized in that: The delay unit includes a first PMOS transistor, a second PMOS transistor, a first NMOS transistor, a second NMOS transistor, and a third NMOS transistor.
7. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 6, characterized in that: The source of the first PMOS tube is connected to the drain of the first NMOS tube; The source of the first NMOS tube is connected to the drain of the second NMOS tube; The gate of the second NMOS tube is connected to the input analog voltage; The second PMOS transistor gate and the third NMOS transistor gate are connected to the first PMOS source and the first NMOS drain; The source of the second PMOS tube is connected to the drain of the third NMOS tube.
8. The high-efficiency mixed-analog CMOS simulated annealing Ising chip according to claim 7, characterized in that: The 9 delay units are connected end to end to form a delay chain. The 9 delay units correspond to the 8 connection directions of the Wang diagram and a local magnetic coupling strength calculation. Each delay unit can generate 4 different delay lengths according to the input analog voltage.
9. A high-efficiency mixed-analog CMOS simulated annealing Ising chip according to any one of claims 4 to 8, characterized in that: The storage array for storing the coupling coefficient sends the coupling coefficient to the multiplication calculation circuit; the multiplication circuit for multiplication calculation multiplies the coupling coefficient with the external input spin state; the instruction peripheral control circuit for controlling the calculation mode sends the multiplication result to the delay chain or the addition tree according to the calculation mode; the delay chain or the addition tree for Hamiltonian accumulation calculation completes the calculation and sends the final spin state to the spin state storage circuit.
10. A method for solving large-scale combinatorial optimization problems using a high-efficiency mixed-analog CMOS simulated annealing Ising chip according to any one of claims 1 to 9, characterized in that: In the initial stage of solving large-scale combinatorial optimization problems, the Hamiltonian of the large-scale combinatorial optimization problem is first coarse-grainedly calculated based on the simulation calculation of the delay chain, thereby obtaining a local minimum for the search; after the Hamiltonian of the iterative simulation calculation falls into the local minimum, the digital calculation based on the addition tree is started to perform fine-grained calculation of the Hamiltonian, and the simulated annealing algorithm is used to search for the optimal solution or approximate optimal solution of the large-scale combinatorial optimization problem.
Citation Information
Patent Citations
High-energy-efficiency CMOS simulated annealing Isin chip
CN117572931A