GHMC simulated annealing algorithm-based FPGA parallel Isin model optimization method and system

Parallel Ising model optimization on FPGA using the GHMC simulated annealing algorithm solves the problems of low convergence accuracy and low solution quality of traditional Monte Carlo methods, and achieves efficient global optimal solution finding, which is suitable for complex optimization scenarios such as power system dispatching.

CN121031808APending Publication Date: 2025-11-28YISI GYROMAGNETIC (JIAXING) ELECTRONICS CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511131357.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing FPGA-based Ising model solvers employ the traditional Monte Carlo method, which suffers from low convergence accuracy and low solution quality, making it difficult to efficiently find the global optimal solution in complex energy barrier scenarios.

Method used

An FPGA-based parallel Ising model optimization method based on the GHMC simulated annealing algorithm is adopted. The problem is mapped to the Hamiltonian parameters of the Ising model by the host computer. The solution is parallelized on the FPGA using a hybrid GHMC sampler and a temperature scheduling module. By combining Leapfrog integration and Metropolis acceptance mechanism, the system temperature is gradually reduced to drive the Ising model to converge to the global optimum.

Benefits of technology

It significantly improves the convergence speed and solution quality of the Ising model, achieves efficient global optimal solution finding, and combines advanced algorithm with high hardware efficiency, making it suitable for high real-time demand scenarios such as power systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121031808A_ABST
    Figure CN121031808A_ABST
Patent Text Reader

Abstract

The invention provides an FPGA parallel Isin model optimization method and system based on a GHMC simulated annealing algorithm, relates to the field of power system combination optimization, and solves the technical problem that an existing FPGA-based Isin model solver adopts a traditional Monte Carlo method and is low in convergence precision and solution quality. The method comprises the following steps: receiving a question input by a user through an upper computer, and mapping the received question into a Hamiltonian parameter of an Isin model; the Hamiltonian parameter is transmitted to an FPGA through a communication interface, and a mixed GHMC sampler and a temperature scheduling module are arranged in the FPGA; the hybrid GHMC sampler carries out parallel solution on Hamiltonian parameters; and the temperature scheduling module gradually reduces the system temperature according to a preset annealing curve, and drives the Isin model to converge to a global optimal solution. The method is used in the power system combination optimization process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power system combinatorial optimization, and in particular to an FPGA parallel Ising model optimization method and system based on the GHMC simulated annealing algorithm. Background Technology

[0002] Combinatorial optimization problems have important applications in power system dispatching, and their core lies in finding the globally optimal or near-optimal solution from a massive number of feasible solutions. As the problem scale increases, traditional solution methods such as simulated annealing and genetic algorithms face problems such as slow convergence speed and high computational resource consumption. While quantum annealing machines can achieve efficient solutions using the Ising model, their widespread deployment is limited by high cost, low-temperature operating conditions, and the maturity of chip technology. In recent years, FPGA-based hardware acceleration solutions have become the preferred path for implementing classical Ising machines due to their high parallelism, low power consumption, and reconfigurability. However, existing FPGA implementations mostly rely on the standard Metropolis Monte Carlo sampling algorithm, lacking efficient long-range state exploration capabilities, which leads to easy getting trapped in local optima in complex energy barrier scenarios, thus limiting convergence efficiency. Summary of the Invention

[0003] This application provides an FPGA parallel Ising model optimization method and system based on the GHMC simulated annealing algorithm, which solves the technical problems of low convergence accuracy and low solution quality in existing FPGA-based Ising model solvers that use the traditional Monte Carlo method.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] Firstly, a method for optimizing the parallel Ising model of FPGA based on the GHMC simulated annealing algorithm is provided, including:

[0006] The system receives user-input questions from a host computer and maps the received questions to Hamiltonian parameters of the Ising model. The questions include QUBO problems and combinatorial optimization examples.

[0007] The Hamiltonian parameters are transmitted to the FPGA via a communication interface. The FPGA is equipped with a hybrid GHMC sampler and a temperature scheduling module.

[0008] The hybrid GHMC sampler solves the Hamiltonian parameters in parallel.

[0009] The temperature scheduling module gradually reduces the system temperature according to the preset annealing curve, driving the Ising model to converge to the global optimum.

[0010] The global optimal solution is returned to the host computer via serial communication protocol for visualization.

[0011] Based on the above technical solutions, the FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm provided in this application achieves user-friendly problem input and result visualization through a host computer, efficiently mapping QUBO and combinatorial optimization problems to Hamiltonian parameters of the Ising model, and realizing high-performance solution with the help of an FPGA hardware acceleration platform. The hybrid GHMC sampler utilizes the long-range state transition capability guided by molecular dynamics to significantly improve sampling efficiency and convergence speed. Combined with parallel design, it can support the synchronous evolution of large-scale spin systems. The temperature scheduling module dynamically adjusts the system temperature according to a preset annealing curve, collaboratively driving the system to converge to the global optimum. The entire process achieves reliable interaction between the host and host computers through serial communication, significantly shortening the computation time while ensuring solution accuracy. It combines algorithmic advancement, hardware efficiency, and system scalability, providing a practical hardware solution for complex combinatorial optimization problems with low latency and high throughput.

[0012] In conjunction with the first aspect above, in one possible implementation, the hybrid GHMC sampler includes:

[0013] A state register array is used to store the state s of each spin node. i ∈{-1,1}; where i is the index of the spin node;

[0014] Momentum register, used to assign a random momentum p to each spin node. i ;

[0015] Leapfrog integrator unit is used for parallel computation of spin and momentum updates;

[0016] A random number generator is used to dynamically provide pseudo-random numbers required for momentum initialization, Metropolis determination, and annealing perturbation.

[0017] The Metropolis acceptance mechanism is used to determine whether to accept or reject the state at the end of a dynamic trajectory.

[0018] In conjunction with the first aspect above, in one possible implementation, the sampling process of the hybrid GHMC sampler includes:

[0019] Step S1: Assign a random momentum p to each spin node i , to initialize the system state; where p i ~N(0,1) or p i ~U(-1,1);

[0020] Step S2: Use the Leapfrog algorithm to update the spin state and random momentum through integration, and generate a new state through deterministic dynamic evolution;

[0021] Step S3: According to the preset annealing curve, gradually reduce the system temperature to guide the system towards the global optimal solution in a probabilistic manner;

[0022] Step S4: Calculate the Hamiltonian difference between the new state and the current state, and based on the current temperature, calculate the probability that the new state is accepted, and mark it as the acceptance probability;

[0023] Determine if the acceptance probability is greater than the probability threshold; if yes, write the new state into the spin state register array to complete the state update for the current iteration; otherwise, return to step S2.

[0024] Step S5: If the system meets the convergence condition, stop the iteration and output the result; otherwise, return to step S2 and repeat the process.

[0025] In conjunction with the first aspect above, in one possible implementation, the preset annealing curve includes:

[0026] T k+1 =T k ×γ;

[0027] Where k is the iteration number of the current annealing algorithm, and T k Let T be the temperature of the k-th iteration. k+1 The temperature for the (k+1)th iteration is γ, and γ is the annealing coefficient.

[0028] If the temperature register value T k If the temperature is below the preset temperature threshold, a termination signal is triggered; otherwise, the next sampling iteration continues.

[0029] In conjunction with the first aspect above, in one possible implementation, the hybrid GHMC sampler solves for the Hamiltonian parameters in parallel, including:

[0030] The spin network partitioning module is used to divide spin nodes into multiple non-conflicting sub-blocks and alternately update the spin state within the sub-blocks using a checkerboard strategy.

[0031] The parallel coupled computation unit performs parallel computation of the coupling term J between spin nodes via the FPGA's DSP module. i,j J i,j The coupling strength between spin nodes i and j;

[0032] The Leapfrog integrator is used to perform momentum p in parallel on the partitioned sub-blocks. i and spin state s i Hamiltonian dynamics update.

[0033] In conjunction with the first aspect above, in one possible implementation, calculating the probability that the new state is accepted includes:

[0034] The acceptance probability \(P_{accept}\) is calculated by the formula \(P_{accept}=\min(1, e\) -ΔH / T0 ) where \(\Delta H\) is the difference in Hamiltonian between the new state and the current state, and \(T_0\) is the current temperature;

[0035] Generate a random number \(u\sim U(0, 1)\). If \(u < P_{accept}\), accept the new state; otherwise, reject it and retain the original state.

[0036] Combined with the first aspect above, in a possible implementation, calculating the difference in Hamiltonian between the new state and the current state includes:

[0037] The difference in Hamiltonian \(\Delta H\) between the new state and the current state is calculated by the formula \(\Delta H=-0.5\times\sum\) i,j J i,j (s new,i \times s new,j -s old,i \times s old,j )-\sum i h i (s new,i -s old,i );

[0038] where \(i\) and \(j\) are the sequence numbers of the spin nodes, \(j\neq i\), \(s\) new,i is the spin value of the \(i\)-th spin node in the new state, \(s\) old,i is the spin value of the \(i\)-th spin node in the current state, \(s\) new,j is the spin value of the \(j\)-th spin node in the new state, \(s\) old,j is the spin value of the \(j\)-th spin node in the current state, \(h\) i is the external magnetic field strength received by the \(i\)-th spin node.

[0039] Combined with the first aspect above, in a possible implementation, the host computer includes problem modeling, parameter configuration, real-time monitoring, and interacts with the FPGA hardware through a serial communication interface.

[0040] Combined with the first aspect above, in a possible implementation, the core algorithm processing module includes the following sub-units:

[0041] Momentum initialization unit: Initialize the momentum of each spin node in the system;

[0042] Leapfrog integration unit: Implement efficient dynamic integration calculation and update the spin and momentum states;

[0043] Temperature annealing scheduling unit: Gradually reduce the system temperature according to a preset temperature curve to improve the global search ability;

[0044] Metropolis Acceptance Decision Unit: Based on energy differences and probability, it determines whether to accept the state update, thereby deciding whether to update the current state.

[0045] Secondly, an FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm is provided, including a communication unit and a processing unit;

[0046] The communication unit receives user-input questions via a host computer and maps the received questions to Hamiltonian parameters of the Ising model. The questions include QUBO problems and combinatorial optimization examples.

[0047] The Hamiltonian parameters are transmitted to the FPGA via a communication interface. The FPGA is equipped with a hybrid GHMC sampler and a temperature scheduling module.

[0048] The processing unit uses a hybrid GHMC sampler to solve the Hamiltonian parameters in parallel; the temperature scheduling module gradually reduces the system temperature according to a preset annealing curve, driving the Ising model to converge to the global optimum.

[0049] The global optimal solution is returned to the host computer via serial communication protocol for visualization.

[0050] Thirdly, this application provides an FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm, comprising: a processor and a storage medium; the storage medium includes instructions, and the processor is used to execute the instructions to implement the method described in the first aspect and any possible implementation thereof. This FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm can be an electronic device or a chip within an electronic device.

[0051] Fourthly, this application provides an FPGA parallel Ising model optimization system based on the GHMC simulated annealing algorithm, including: a host computer, a serial communication interface module, a core algorithm processing module, an Ising spin array module, and a result output module;

[0052] The host computer is used for user problem modeling, optimization algorithm parameter setting, temperature annealing curve definition, and visualization of the final optimization results;

[0053] The serial communication interface is used for bidirectional communication between the host computer and the FPGA hardware, for sending problem mapping parameters, annealing curve configuration data, and receiving the real-time status and final optimization results of the optimization process.

[0054] The core algorithm processing module is used to efficiently solve the global optimal solution or thermodynamic equilibrium state of the spin system by combining dynamic evolution and probabilistic sampling.

[0055] The Ising spin array is used to store and update the state of each spin node in the Ising model in real time, and supports high parallelism operation.

[0056] The result output module is used to return the final result of the Ising machine optimization to the host computer for display.

[0057] Fifthly, this application provides a computer-readable storage medium storing instructions that, when executed on an FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm, cause the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm to perform the method described in the first aspect and any possible implementation thereof.

[0058] In a sixth aspect, this application provides a computer program product containing instructions that, when run on an FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm, cause the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm to perform the method described in the first aspect and any possible implementation thereof.

[0059] This application provides an FPGA-based parallel Ising model optimization method and system based on the GHMC simulated annealing algorithm. It deeply integrates advanced GHMC simulated annealing algorithm innovation with FPGA parallel hardware architecture, achieving a balance between high accuracy, fast convergence, and engineering feasibility. Specifically, the host computer not only automatically maps the QUBO / combinatorial optimization problem to the Hamiltonian parameters of the Ising model but also enables flexible configuration of the FPGA through a standardized communication interface. The parallelized GHMC sampler integrated within the FPGA supports the simultaneous evolution of hundreds to tens of thousands of spin nodes, fully leveraging the advantages of hardware parallelism and overcoming the bottleneck of sampling efficiency in the traditional Metropolis method. Simultaneously, the temperature scheduling module works in conjunction with the dynamic integration process, using a simulated annealing strategy to gradually reduce the system temperature, guiding GHMC to perform global exploration at high temperatures and fine convergence at low temperatures, significantly improving the quality and stability of the solution. Finally, the optimal solution is transmitted back to the host computer via serial port for visualization analysis, forming a complete "modeling-solving-feedback" system link. This solution not only achieves a revolution in algorithm level (GHMC + annealing), but also builds a scalable, low-latency, high-throughput dedicated computing platform, providing comprehensive performance superior to traditional software solvers and general hardware solutions for complex optimization scenarios with high real-time requirements, such as power system scheduling.

[0060] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0061] Figure 1 A system architecture diagram of an FPGA parallel Ising model optimization system based on the GHMC simulated annealing algorithm is provided for embodiments of this application;

[0062] Figure 2 A flowchart illustrating an FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm, provided for an embodiment of this application;

[0063] Figure 3 A schematic flowchart of the sampling process method of the hybrid GHMC sampler provided in the embodiments of this application;

[0064] Figure 4 Comparison diagrams of the optimization process provided for embodiments of this application;

[0065] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0066] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0067] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0068] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0069] The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm provided in this application embodiment can be applied to, for example... Figure 1 In the optimization system 100 shown, such as Figure 1 As shown, the optimization system includes: a host computer, a serial communication interface module, a core algorithm processing module, an Ising spin array module, and a result output module.

[0070] The host computer is used for user problem modeling, optimization algorithm parameter setting, temperature annealing curve definition, and visualization of the final optimization results.

[0071] The serial communication interface is used for bidirectional communication between the host computer and the FPGA hardware. It is used to send problem mapping parameters, annealing curve configuration data, and receive the real-time status of the optimization process and the final optimization results.

[0072] The core algorithm processing module is used to efficiently solve for the global optimal solution or thermodynamic equilibrium state of a spin system by combining dynamic evolution and probabilistic sampling.

[0073] The core algorithm processing module includes a momentum initialization unit, a Leapfrog integration unit, a temperature annealing scheduling unit, and a Metropolis acceptance and decision unit;

[0074] Momentum initialization unit: Initializes the momentum of each spin node in the system.

[0075] Leapfrog Integrating Unit: Enables efficient dynamic integration calculations and updates spin and momentum states.

[0076] Temperature annealing scheduling unit: Gradually reduce the system temperature according to the preset temperature curve to improve global search capability.

[0077] Metropolis Acceptance Decision Unit: Based on energy differences and probability, it determines whether to accept the state update, thereby deciding whether to update the current state.

[0078] The Ising spin array is used to store and update the state of each spin node in the Ising model in real time, and supports high parallelism operations.

[0079] The results output module is used to return the final results of the Ising machine optimization to the host computer for display.

[0080] To address the technical issues of low convergence accuracy and solution quality in existing FPGA-based Ising model solvers employing the traditional Monte Carlo method, this application provides an FPGA-based parallel Ising model optimization method based on the GHMC simulated annealing algorithm. This method includes: receiving user-inputted problems via a host computer and mapping the received problems to Hamiltonian parameters of the Ising model; transmitting the Hamiltonian parameters to the FPGA via a communication interface; the FPGA being equipped with a hybrid GHMC sampler and a temperature scheduling module; the hybrid GHMC sampler performing parallel solution of the Hamiltonian parameters; the temperature scheduling module gradually reducing the system temperature according to a preset annealing curve, driving the Ising model to converge to the global optimum; and returning the global optimum solution to the host computer via a serial communication protocol for visualization. Based on this, large-scale spin-parallel evolution is achieved through the FPGA, significantly improving the convergence speed and solution quality of the Ising machine. The host computer completes problem modeling and parameter distribution, the FPGA efficiently executes optimization calculations, and the temperature scheduling mechanism guides the system towards the global optimum, with visualization achieved through serial port feedback. The overall architecture combines advanced algorithms with high hardware efficiency, solving the problems of slow convergence and poor scalability in complex combinatorial optimization of traditional methods, and is suitable for high real-time demand scenarios such as power systems.

[0081] like Figure 2 As shown in the embodiments of this application, the FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm includes:

[0082] S201. Receive user input questions via the host computer and map the received questions to Hamiltonian parameters of the Ising model;

[0083] The problems include the QUBO problem and combinatorial optimization examples.

[0084] The host computer is responsible for user problem modeling, optimizing algorithm parameter settings, defining temperature annealing curves, and visualizing the final optimization results.

[0085] S202. Transmit the Hamiltonian parameters to the FPGA through the communication interface.

[0086] The FPGA includes a hybrid GHMC sampler and a temperature scheduling module.

[0087] Hybrid GHMC sampler, including:

[0088] A state register array is used to store the state s of each spin node. i ∈{-1,1}; where i is the index of the spin node;

[0089] Momentum register, used to assign a random momentum p to each spin node. i ;

[0090] Leapfrog integrator unit is used for parallel computation of spin and momentum updates;

[0091] Random number generator, used to dynamically provide pseudo-random numbers required for momentum initialization, Metropolis determination and annealing perturbation;

[0092] The Metropolis acceptance mechanism is used to determine whether to accept or reject the state at the end of a dynamic trajectory.

[0093] S203. The hybrid GHMC sampler solves the Hamiltonian parameters in parallel.

[0094] S204. The temperature scheduling module gradually reduces the system temperature according to the preset annealing curve, driving the Ising model to converge to the global optimum.

[0095] S205. The global optimal solution is returned to the host computer for visualization via serial communication protocol.

[0096] Based on the above technical solutions, the FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm provided in this application firstly achieves user-friendly problem modeling and parameter configuration through a host computer, supporting flexible mapping of QUBO and various combinatorial optimization problems, greatly improving the system's ease of use and applicability. Secondly, the FPGA-based hybrid GHMC sampler adopts a fully hardware parallel architecture, achieving a computation speed tens of times faster than traditional CPU solutions through the collaborative design of a state register array, a parallel Leapfrog integration unit, and a distributed random number generator. Precise control of the temperature scheduling module, combined with the Metropolis receiving mechanism, effectively balances exploration and development capabilities, significantly improving the probability of obtaining the global optimum. Furthermore, the efficient data transmission and real-time visualization functions of the communication interface make the entire optimization process monitorable and verifiable. This system demonstrates excellent performance in computational efficiency, solution accuracy, and energy efficiency ratio, making it particularly suitable for industrial optimization scenarios requiring rapid response.

[0097] In one possible implementation of the embodiments of this application, such as Figure 3 As shown, the sampling process of the above-mentioned hybrid GHMC sampler includes:

[0098] Step S1: Assign a random momentum p to each spin node i , to initialize the system state; where p i Follows a Gaussian distribution, p i ~N(0,1) or p i Follows a uniform distribution, pi ~U(-1,1);

[0099] Step S2: Use the Leapfrog algorithm to update the spin state and random momentum through integration, and generate a new state through deterministic dynamic evolution;

[0100] Step S3: According to the preset annealing curve, gradually reduce the system temperature to guide the system towards the global optimal solution in a probabilistic manner;

[0101] Among them, the preset annealing curve is:

[0102] T k+1 =T k ×γ;

[0103] Where k is the iteration number of the current annealing algorithm, and T k Let T be the temperature of the k-th iteration. k+1 The temperature for the (k+1)th iteration is γ, and γ is the annealing coefficient.

[0104] Specifically, the current value T of the temperature register is obtained through a hardware multiplier or fixed-point arithmetic unit. k Multiply by the annealing factor γ, and update the temperature register to T. k+1 =T k ×γ;

[0105] If the temperature register value T k If the temperature is below the preset temperature threshold, a termination signal is triggered; otherwise, the next sampling iteration continues.

[0106] Based on the above technical solution, this annealing strategy achieves a gradual decrease in system temperature through a preset exponential cooling curve, effectively balancing the exploration and development capabilities during the search process: in the high-temperature stage, the system maintains a strong stochastic transition capability, which helps to escape local optima; as the temperature gradually decreases, the probability of accepting inferior solutions gradually decreases, guiding the system to converge towards a lower-energy global optimum. Efficient temperature updates are achieved through hardware multipliers or fixed-point arithmetic units, ensuring not only the real-time and deterministic nature of the annealing process but also applicability to resource-constrained scenarios such as FPGAs or dedicated accelerators; a termination mechanism is triggered when the temperature drops to a preset temperature threshold, ensuring the algorithm converges within a finite time, thus combining physical feasibility with optimization effectiveness.

[0107] Step S4: Calculate the Hamiltonian difference between the new state and the current state, and based on the current temperature, calculate the probability that the new state is accepted, and mark it as the acceptance probability;

[0108] The calculation of the probability that the new state is accepted includes:

[0109] Calculate the Hamiltonian difference between the new state and the current state, including:

[0110] The Hamiltonian difference ΔH between the new state and the current state is calculated by the formula ΔH = -0.5×∑ i,j J i,j (s new,i ×s new,j -s old,i ×s old,j ) - ∑ i h i (s new,i -s old,i ) The Hamiltonian difference ΔH between the new state and the current state is calculated by the formula ΔH = -0.5×∑

[0111] where i and j are the sequence numbers of the spin nodes, j≠i, s new,i is the spin value of the i-th spin node in the new state, s old,i is the spin value of the i-th spin node in the current state, s new,j is the spin value of the j-th spin node in the new state, s old,j is the spin value of the j-th spin node in the current state, h i is the external magnetic field intensity received by the i-th spin node.

[0112] The acceptance probability Paccept is calculated by the formula Paccept = min(1, e -ΔH / T0 ) where ΔH is the Hamiltonian difference between the new state and the current state, and T0 is the current temperature;

[0113] Generate a random number u ∼ U(0, 1). If u < Paccept, accept the new state; otherwise, reject it and retain the original state.

[0114] Judge whether the acceptance probability is greater than the probability threshold; if yes, write the new state into the spin state register array to complete the state update of the current iteration; if no, return to step S2;

[0115] Based on the above technical solution, this state update mechanism realizes the probabilistic control of the evolution process of the spin system by accurately calculating the Hamiltonian difference and combining the Metropolis acceptance criterion, ensuring that the sampling process strictly follows the Boltzmann distribution. And it evaluates the energy change brought by the state transformation, avoiding the repeated calculation of the energy of the entire system, significantly improving the calculation efficiency; calculating the acceptance probability in combination with the current temperature and deciding the state update after comparing the random numbers not only retains the characteristics of thermal fluctuations but also ensures the directional evolution to the low-energy state. This strategy has a simple hardware implementation and only requires a multiplier, an exponential function unit, and a comparator to complete the judgment, which is suitable for the efficient simulation of large-scale parallel spin arrays and has good convergence.

[0116] Step S5: When the system meets the convergence condition, stop the iteration and output the result; if not, return to step S2 and repeat the execution.

[0117] The hybrid GHMC sampler parallelizes the solution of Hamiltonian parameters, including:

[0118] The spin network partitioning module is used to divide spin nodes into multiple non-conflicting sub-blocks and alternately update the spin state within the sub-blocks using a checkerboard strategy.

[0119] The parallel coupled computation unit performs parallel computation of the coupling term J between spin nodes via the FPGA's DSP module. i,j J i,j The coupling strength between spin nodes i and j;

[0120] The Leapfrog integrator is used to perform momentum p in parallel on the partitioned sub-blocks. i and spin state s i Hamiltonian dynamics update.

[0121] Based on the above technical solution, this parallel architecture achieves synchronous updates of multiple spin states by dividing the spin network into multiple independent sub-blocks and adopting a checkerboard update strategy, significantly improving computational efficiency. Utilizing the FPGA's DSP units for parallel computation of coupling terms and a pipelined Leapfrog integrator significantly accelerates the Hamiltonian dynamics evolution process. A distributed random number generator and a global synchronization control mechanism ensure efficient collaboration among the computational units, guaranteeing algorithm accuracy while enabling the system to quickly escape local optima and converge to the global optimum. This design fully leverages the parallel computing advantages of FPGAs, giving the Ising model solver faster computation speed and lower energy consumption when handling large-scale combinatorial optimization problems.

[0122] In addition, this application provides an embodiment that analyzes the impact of setting different problem structures on the convergence efficiency of the optimization algorithm of this application.

[0123] For example, parameter settings such as the Ising mapping: assign J to each pair of nodes (i,j). i,j =-1, GHMC parameters: Leapfrog step count: 5, step size: 0.01, initial temperature: 5.0, decay factor: 0.995, platform: FPGACycloneVSoC, 100MHz, fixed-point Q8.8 operation, sampling iteration: 500 rounds.

[0124] like Figure 4The figure shows a comparison curve of the optimization process for four types of combinatorial optimization problems: MaxCut Dense, SK Model, NAE3SAT, and Spin Glass. The vertical axis represents normalized energy, and the horizontal axis represents the number of iterations. The initial normalized energy values ​​are all close to 1, indicating that the initial solution is in a high energy state. During the iteration process, the energy gradually decreases until it approaches 0, corresponding to the global or near-optimal solution.

[0125] In terms of convergence speed, there are significant differences among different problem types. The Spin Glass problem (red dashed line) experiences a sharp drop in energy in the early stages of iteration and converges by the 140th iter, with the normalized energy dropping to approximately 0, demonstrating extremely fast optimization efficiency. The MaxCut Dense problem (blue solid line) experiences a rapid energy drop in the first 200 iterations, then enters a gradual convergence phase, finally converging to the lowest energy state by the 498th iter. In contrast, the NAE3SAT problem (green dotted line) and SKModel (orange dashed line) exhibit significantly slower energy decreases and show obvious fluctuations during iteration, reflecting that the optimization process is more significantly affected by the complexity of the solution space.

[0126] From the final performance perspective, the normalized energy values ​​after convergence for all four problem types are close to 0, indicating that the optimization method can obtain relatively good solutions for different problem types. However, comparing the number of convergence iterations, Spin Glass converges approximately 358 iterations faster than MaxCut Dense, and even faster than NAE3SAT and SK Model, demonstrating the significant impact of different problem structures on the convergence efficiency of the optimization algorithm.

[0127] In summary, the results of this experiment show that although the optimization algorithm can converge on different problem types, its convergence speed and descent trajectory are significantly affected by the characteristics of the problem: problems with more regular solution spaces, such as Spin Glass and MaxCutDense, converge faster, while problems with complex structures or more constraints, such as NAE3SAT and SK Model, converge slower and fluctuate significantly.

[0128] The above primarily describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, such as the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm, includes at least one of the hardware structures and software modules corresponding to each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0129] This application embodiment can divide the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0130] When using integrated units, Figure 5 A possible structural schematic diagram of the FPGA parallel Ising model optimization device (referred to as electronic device 50) based on the GHMC simulated annealing algorithm involved in the above embodiments is shown. The electronic device 50 includes a processing unit 501 and a communication unit 502, and may also include a storage unit 503. Figure 5 The schematic diagram shown can be used to illustrate the structure of the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm involved in the above embodiments.

[0131] when Figure 5 The schematic diagram shown illustrates the structure of the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm involved in the above embodiments. The processing unit 501 is used to control and manage the operation of the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm. The communication unit 502 is used for the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm to communicate with other devices. The storage unit 503 is used to store the program code and data of the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm.

[0132] The communication unit 502 receives user input questions through the host computer and maps the received questions to Hamiltonian parameters of the Ising model. The questions include QUBO problems and combinatorial optimization examples.

[0133] The Hamiltonian parameters are transmitted to the FPGA via a communication interface. The FPGA is equipped with a hybrid GHMC sampler and a temperature scheduling module.

[0134] Processing unit 501 uses a hybrid GHMC sampler to solve the Hamiltonian parameters in parallel; the temperature scheduling module gradually reduces the system temperature according to the preset annealing curve, driving the Ising model to converge to the global optimum; the global optimum is returned to the host computer for visualization via serial communication protocol.

[0135] The processing unit 501 can be a processor or a controller, and the communication unit 502 can be a communication interface, transceiver, transceiver circuit, transceiver device, etc. The term "communication interface" is a general term and may include one or more interfaces. The storage unit 503 can be a memory. When the electronic device 50 is a chip, the processing unit 501 can be a processor or a controller, and the communication unit 502 can be an input interface and / or an output interface, pins, or circuits, etc. The storage unit 503 can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip (e.g., read-only memory (ROM), random access memory (RAM, etc.)).

[0136] The communication unit can also be called a transceiver unit. The antenna and control circuit with transceiver functions in the electronic device 50 can be considered as the communication unit 502 of the electronic device 50, and the processor with processing functions can be considered as the processing unit 501 of the electronic device 50. Optionally, the device in the communication unit 502 that implements the receiving function can be considered as a communication unit. The communication unit is used to execute the receiving steps in the embodiments of this application, and the communication unit can be a receiver, a receiver circuit, etc. The device in the communication unit 502 that implements the transmitting function can be considered as a transmitting unit. The transmitting unit is used to execute the transmitting steps in the embodiments of this application, and the transmitting unit can be a transmitter, a transmitter, a transmitting circuit, etc.

[0137] Figure 5If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0138] Figure 5 The units in the process can also be called modules; for example, a processing unit can be called a processing module.

[0139] This application also provides a hardware structure diagram of an FPGA parallel Ising model optimization device (denoted as electronic device 60) based on the GHMC simulated annealing algorithm, see [link to relevant documentation]. Figure 6 The electronic device 60 includes a processor 601, and optionally, a memory 602 connected to the processor 601.

[0140] In the first possible implementation, see Figure 6 The electronic device 60 also includes a transceiver 603. The processor 601, memory 602, and transceiver 603 are connected via a bus. The transceiver 603 is used to communicate with other devices or communication networks. Optionally, the transceiver 603 may include a transmitter and a receiver. The device in the transceiver 603 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 603 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0141] Based on the first possible implementation method Figure 6 The schematic diagram shown can be used to illustrate the structure of the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm involved in the above embodiments.

[0142] in, Figure 6 The diagram can also illustrate the system chip in the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm. In this case, the actions performed by the FPGA parallel Ising model optimization device based on the GHMC simulated annealing algorithm can be implemented by this system chip. The specific actions performed can be found above and will not be repeated here.

[0143] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0144] The processor in this application may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., and other computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor may be a standalone semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it may form a System-on-a-Chip (SoC) with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits), or it may be integrated as a built-in processor within an ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), or logic circuits that implement dedicated logic operations.

[0145] The memory in the embodiments of this application may include at least one of the following types: read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; or electrically erasable programmable-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0146] This application also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0147] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0148] This application also provides a chip including a processor and an interface circuit. The interface circuit is coupled to the processor, which is used to run computer programs or instructions to implement the above-described method. The interface circuit is used to communicate with other modules outside the chip.

[0149] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0150] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0151] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

[0152] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

Claims

1. An FPGA parallel Ising model optimization method based on GHMC simulated annealing algorithm, characterized in that, It includes: Receiving the problem input by the host computer and mapping the received problem to the Hamiltonian parameters of the Ising model, where the problem includes QUBO problems and combinatorial optimization instances; Transmitting the Hamiltonian parameters to the FPGA through the communication interface, and a hybrid GHMC sampler and a temperature scheduling module are set in the FPGA; The hybrid GHMC sampler performs parallel solution on the Hamiltonian parameters; The temperature scheduling module gradually reduces the system temperature according to the preset annealing curve, driving the Ising model to converge to the global optimal solution; Returning the global optimal solution to the host computer through the serial communication protocol for visual display.

2. The FPGA parallel Ising model optimization method based on GHMC simulated annealing algorithm according to claim 1, characterized in that, The hybrid GHMC sampler includes: A state register array is used to store the state s of each spin node. i ∈{-1,1}; where i is the index of the spin node; Momentum register, used to assign a random momentum p to each spin node. i ; A Leapfrog integration unit for parallelly calculating the updates of spin and momentum; A random number generator for dynamically providing the pseudo-random numbers required for momentum initialization, Metropolis determination, and annealing perturbation; A Metropolis acceptance mechanism for determining whether to accept or reject the state at the end of the dynamic trajectory.

3. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 1, characterized in that, The sampling process of the hybrid GHMC sampler includes: Step S1: Assign a random momentum p to each spin node i , to initialize the system state; where p i ~N(0,1) or p i ~U(-1,1); where i is the index of the spin node; Step S2: Using the Leapfrog algorithm to perform integral updates on the spin state and the random momentum, and generating a new state through deterministic dynamic evolution; Step S3: According to the preset annealing curve, gradually reduce the system temperature to probabilistically guide the system to converge to the global optimal solution; Step S4: Calculating the Hamiltonian difference between the new state and the current state, and calculating the probability of accepting the new state based on the current temperature, marked as the acceptance probability; Judging whether the acceptance probability is greater than the probability threshold; if yes, writing the new state into the spin state register array to complete the state update of the current iteration; if no, returning to Step S2; Step S5: When the system meets the convergence condition, stop the iteration and output the result; if not, return to Step S2 and repeat the execution.

4. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 3, characterized in that, The preset annealing curve includes: T k+1 =T k ×γ; Where k is the iteration number of the current annealing algorithm, and T k Let T be the temperature of the k-th iteration. k+1 The temperature for the (k+1)th iteration is γ, and γ is the annealing coefficient. If the temperature register value T k If the temperature is below the preset temperature threshold, a termination signal is triggered; otherwise, the next sampling iteration continues.

5. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 3, characterized in that, The parallel solution of the Hamiltonian parameters by the hybrid GHMC sampler includes: A spin network partitioning module for dividing spin nodes into multiple non-conflicting sub-blocks and alternately updating the spin states within the sub-blocks using the checkerboard strategy; The parallel coupled computation unit performs parallel computation of the coupling term J between spin nodes via the FPGA's DSP module. i,j J i,j The coupling strength between spin nodes i and j; The Leapfrog integrator is used to perform momentum p in parallel on the partitioned sub-blocks. i and spin state s i Hamiltonian dynamics update.

6. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 5, characterized in that, The calculation of the probability of accepting the new state includes: Using the formula Paccept=min(1,e) -ΔH / T0 The acceptance probability Paccept is calculated; where ΔH is the Hamiltonian difference between the new state and the current state, and T0 is the current temperature. Generating a random number u ∼ U(0,1), if u < Paccept, then accept the new state, otherwise reject and retain the original state.

7. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 5, characterized in that, The calculation of the Hamiltonian difference between the new state and the current state includes: The formula ΔH=-0.5×∑ i,j J i,j (s new,i ×s new,j -s old,i ×s old,j )-∑ i h i (s new,i -s old,i The Hamiltonian difference ΔH between the new state and the current state is calculated. Where i and j are the indices of the spin nodes, j ≠ i, s new,i Let s be the spin value of the i-th spin node in the new state. old,i Let s be the spin value of the i-th spin node in the current state. new,j Let s be the spin value of the j-th spin node in the new state. old,j Let h be the spin value of the j-th spin node in the current state. i Let be the intensity of the external magnetic field experienced by the i-th spin node.

8. The FPGA parallel Ising model optimization method based on the GHMC simulated annealing algorithm according to claim 1, characterized in that, The host computer includes problem modeling, parameter configuration, real-time monitoring, and hardware interaction with the FPGA through the serial communication interface.

9. An FPGA parallel Ising model optimization system based on the GHMC simulated annealing algorithm, characterized in that, The system includes: a host computer, a serial communication interface module, a core algorithm processing module, an Ising spin array module, and a result output module; The host computer is used for user problem modeling, optimization algorithm parameter setting, temperature annealing curve definition, and visual display of the final optimization result; The serial communication interface is used for two-way communication between the host computer and the FPGA hardware, for sending problem mapping parameters, annealing curve configuration data, and receiving the real-time state and the final optimization result of the optimization process; The core algorithm processing module is used to efficiently solve the global optimal solution or thermodynamic equilibrium state of the spin system by combining dynamic evolution and probabilistic sampling. The Ising spin array is used to store and update the state of each spin node in the Ising model in real time, and supports high parallelism operation. The result output module is used to return the final result of the Ising machine optimization to the host computer for display.

10. The FPGA parallel Ising model optimization system based on the GHMC simulated annealing algorithm according to claim 9, characterized in that, The core algorithm processing module includes the following sub-units: Momentum initialization unit: Initializes the momentum of each spin node in the system; Leapfrog Integrating Unit: Enables efficient dynamic integration calculations and updates spin and momentum states; Temperature annealing scheduling unit: Gradually reduces the system temperature according to the preset temperature curve, improving global search capability; Metropolis Acceptance Decision Unit: Based on energy difference and probability, it determines whether to accept the state update, and thus decides whether to update the current state; Ising Spin Array Module: Stores and updates the state of each spin node in the Ising model in real time, supporting high parallelism operations.

Citation Information

Cited By

  • Continuous time Isin model hardware solving system based on multi-chip interconnection

    CN121765178A

  • A continuous-time Ising model hardware solving system based on multi-chip interconnection

    CN121765178B