Data processing device, data processing method and program
By employing multiple replicas for parallel processing stages within the data processing device, the inefficiencies of sequential processing in existing MCMC methods are addressed, leading to enhanced computational resource utilization and improved solution search efficiency for combinatorial optimization problems.
Patent Information
- Application Number
- JP2021148257
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-09-13
AI Technical Summary
Existing data processing devices for solving combinatorial optimization problems using the MCMC method face inefficiencies due to sequential processing, which limits the effectiveness of increased parallelism in solution search processes.
A data processing device that utilizes multiple replicas to execute parallel processing stages for updating state variables, ensuring that different replicas are processed at the same timing to maintain the sequential processing principle of the MCMC method.
This approach allows for efficient use of computing resources, effectively parallelizing the search process while ensuring that appropriate solutions are obtained, thus improving the overall performance in solving combinatorial optimization problems.
Smart Images

Figure 0007678314000007 
Figure 0007678314000008 
Figure 0007678314000009
Abstract
Description
[Technical field]
[0001] The present invention relates to a data processing device, a data processing method, and a program. [Background technology]
[0002] Information processing devices are sometimes used to solve combinatorial optimization problems. The information processing device converts the combinatorial optimization problem into an energy function of an Ising model, which is a model that represents the behavior of spin in a magnetic material, and searches for a combination that minimizes the value of the energy function among combinations of state variable values included in the energy function. The combination of state variable values that minimizes the value of the energy function corresponds to the ground state or optimal solution represented by the set of state variables. Methods for obtaining an approximate solution to a combinatorial optimization problem in a practical amount of time include the Simulated Annealing (SA) method and the replica exchange method, which are based on the Markov Chain Monte Carlo (MCMC) method.
[0003] For example, there has been a proposal for an information processing device having multiple Ising devices. In this proposal, the Ising devices have multiple neuron circuits, each of which performs processing related to one bit. Each of the multiple Ising devices reflects the neuron state of the other Ising devices obtained via a router in its own neuron circuit.
[0004] There has also been proposed an optimization device that divides a combinatorial optimization problem into a number of subproblems, solves them, and obtains an overall solution based on the solutions of the subproblems. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2017-219948 A [Patent Document 2] Patent Publication No. 2021-5282 Summary of the Invention [Problem to be solved by the invention]
[0006] It is conceivable that the efficiency of finding a solution can be improved by increasing the parallelism of the search process for a solution to the problem using the resources of the computing unit of the device. Here, the principle of minimizing an Ising-type energy function using the MCMC method is that sequential processing is the rule, in which state variables are updated one by one for the problem. Therefore, even if the parallelism of the search process is increased, it is not possible to obtain an appropriate solution to the problem unless the principle of sequential processing in the MCMC method is observed.
[0007] In one aspect, the present invention has an object to provide a data processing device, a data processing method, and a program that effectively utilize computing resources. [Means for solving the problem]
[0008] In one embodiment, a data processing device is provided that solves a problem represented by an energy function including a plurality of state variables. The data processing device includes a storage unit and a processing unit. The storage unit holds a plurality of replicas, each of which indicates a plurality of state variables. The processing unit executes in parallel a first process for executing a plurality of stages for the plurality of replicas, including determining a first state variable to be updated and updating the value of the first state variable to be updated, according to a change amount of the value of the energy function when each of a plurality of first state variables belonging to a first index range, which is a range of indexes corresponding to the plurality of state variables, is set as an update candidate, and a second process for executing a plurality of stages for the plurality of replicas, including determining a second state variable to be updated and updating the value of the second state variable to be updated, according to a change amount of the value of the energy function when each of a plurality of second state variables belonging to a second index range that does not overlap with the first index range is set as an update candidate, and each stage included in the first process and the second process processes different replicas at the same timing.
[0009] Also, in one aspect, a data processing method is provided. Also, in one embodiment, a program is provided. Effect of the Invention
[0010] On the one hand, it enables more efficient use of computing resources. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating a data processing device according to a first embodiment. [Diagram 2] FIG. 11 illustrates an example of hardware of a data processing device according to a second embodiment. [Diagram 3] FIG. 2 illustrates an example of functions of a data processing device. [Figure 4] FIG. 13 illustrates an example of the functionality of local field updates in groups. [Diagram 5] FIG. 1 illustrates an example of pipeline processing. [Figure 6] FIG. 13 is a diagram illustrating an example of reading out weighting coefficients. [Figure 7] 11 is a flowchart illustrating an example of processing performed by the data processing device. [Figure 8] FIG. 13 illustrates an example of pipeline processing according to a third embodiment; [Figure 9] FIG. 13 is a diagram illustrating an example of a memory configuration for storing weighting coefficients. [Figure 10] FIG. 4 is a diagram illustrating an example of a weighting coefficient storage memory unit. [Figure 11] FIG. 2 illustrates an example of a function of stall control in a data processing device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The present embodiment will be described below with reference to the drawings. [First embodiment] A first embodiment will be described.
[0013] FIG. 1 is a diagram illustrating a data processing device according to a first embodiment. The data processing device 10 searches for a solution to a combinatorial optimization problem using the MCMC method, and outputs the searched solution. For example, the data processing device 10 uses an SA method based on the MCMC method, a parallel tempering (PT) method, or the like to search for a solution. The PT method is also called a replica exchange method. The data processing device 10 has a memory unit 11 and a processing unit 12.
[0014] The storage unit 11 may be a volatile storage device such as a random access memory (RAM) or a non-volatile storage device such as a flash memory. The storage unit 11 may include an electronic circuit such as a register. The processing unit 12 may be an electronic circuit such as a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a graphics processing unit (GPU). The processing unit 12 may be a processor that executes a program. The "processor" may include a collection of multiple processors (multiprocessor).
[0015] A combinatorial optimization problem is formulated by an Ising-type energy function, and is replaced by, for example, a problem of minimizing the value of the energy function. The energy function is sometimes called an objective function or an evaluation function. The energy function includes multiple state variables. The state variables are binary variables that take on values of 0 or 1. The state variables may be expressed as bits. A solution to a combinatorial optimization problem is represented by the values of the multiple state variables. A solution that minimizes the value of the energy function represents the ground state of the Ising model, and corresponds to the optimal solution to the combinatorial optimization problem. The value of the energy function is expressed as energy.
[0016] The Ising-type energy function is expressed by equation (1).
[0017]
number
[0018] The state vector x has multiple state variables as elements and represents the state of the Ising model. Equation (1) is an energy function formulated in the QUBO (Quadratic Unconstrained Binary Optimization) format. Note that in the case of a problem of maximizing energy, the sign of the energy function can be reversed.
[0019] The first term on the right side of equation (1) is the sum of the values of the two state variables and the weighting coefficients for all combinations of two state variables that can be selected from all state variables, without omissions or duplications. The subscripts i and j are the indices of the state variables. i is the i-th state variable. x j is the jth state variable. W ij is the weighting coefficient indicating the strength of the connection between the i-th state variable and the j-th state variable. W ij =W ji And W ii =0.
[0020] The second term on the right-hand side of equation (1) is the sum of the products of the biases of all state variables and the values of the state variables. i indicates the bias for the i-th state variable. Problem information including the weighting coefficient and bias included in the energy function is stored in the storage unit 11.
[0021] State variable x i The value of 1-x changes i Then, the state variable x i The increase in is δx i =(1-x i )-x i =1-2x i Therefore, for the energy function E(x), the state variable x i The change in energy due to the change in i is expressed by equation (2).
[0022]
number
[0023] h i is called a local field and is expressed by Equation (3). The local field may also be called a local field (LF).
[0024]
number
[0025] State variable x j When changes in the local field h i Change in δh i (j) is expressed by equation (4).
[0026]
number
[0027] The memory unit 11 stores local fields h i The processing unit 12 holds the state variable x j When the value of i (j) h i By adding to the bit-reversed state, h i get.
[0028] The processing unit 12 detects the energy change ΔE i That is, the state transition of the state variable x i The Metropolis method and the Gibbs method are used to determine whether to allow a change in the value of ΔE. Specifically, in a neighborhood search that searches for a transition from a certain state to another state with lower energy than the state, the processing unit 12 probabilistically allows not only a state with lower energy but also a transition to a state with higher energy. For example, the probability A of accepting a change in the value of the state variable of the energy change ΔE is expressed by Equation (5).
[0029] [Mathematics]
[0030] β is the reciprocal of the temperature value T (T > 0) (β = 1 / T), and is called the inverse temperature. The min operator indicates taking the minimum value among the arguments. The upper right side of the right side of Equation (5) corresponds to the Metropolis method. The lower right side of the right side of Equation (5) corresponds to the Gibbs method. The processing unit 12 compares a uniform random number u where 0 < u < 1 with A for a certain index i, and if u < A, it accepts the change in the value of the state variable x i and changes the value of the state variable x i If u ≥ A, the processing unit 12 does not accept the change in the value of the state variable x i and does not change the value of the state variable x i . According to Equation (5), the larger the value of ΔE, the smaller A becomes. Also, the smaller β is, that is, the larger T is, the more easily a state transition with a large ΔE is allowed. For example, when the Metropolis method is used, the processing unit 12 may perform transition determination using Equation (6) obtained by transforming Equation (5).
[0031] [Mathematics]
[0032] That is, for a uniform random number u (0 < u ≤ 1), when the energy change ΔE satisfies Equation (6), the processing unit 12 allows the change in the value of the corresponding state variable. When the energy change ΔE does not satisfy Equation (6) for the uniform random number u, the processing unit 12 does not allow the change in the value of the corresponding state variable.
[0033] The processing unit 12 can speed up the solution search by determining state variables whose values are changed by parallel trials for a plurality of state variables. For example, the processing unit 12 calculates ΔE in parallel for each index belonging to a predetermined index range. Then, the processing unit 12 selects an index of a state variable whose value is changed from among the indexes that satisfy the formula (6) for ΔE, using a random number or the like. The predetermined index range is a partial index range of the entire index range. In this example, the processing unit 12 performs parallel trials for the index range, i.e., partial parallel trials, for each index range.
[0034] Furthermore, the processing unit 12 parallelizes the search process for a solution to the problem by using a plurality of replicas, each of which indicates a plurality of state variables. For example, the processing unit 12 may execute the SA method in parallel for each of the plurality of replicas. Alternatively, the processing unit 12 may execute the replica exchange method by using the plurality of replicas.
[0035] For example, the storage unit 11 stores replicas R0, R1, R2, and R3. R0={x01, x02, ..., x0 N}. x0 i indicates the state variables that belong to replica R0. N is the number of state variables. The total range of the index is 1 to N. R1 = {x11, x12, …, x1 N}. x1 i denotes the state variables belonging to replica R1. R2 = {x21, x22, …, x2 N}x2 i denotes the state variables belonging to replica R2. R3 = {x31, x32, …, x3 N}x3 i denotes the state variables belonging to replica R3.
[0036] The processing unit 12 executes the partial parallel trials described above in parallel for each of the replicas R0 to R3 using multiple pipelines. A pipeline is a process of sequentially executing a series of stages belonging to the pipeline for each replica. A replica to be processed can be input to the pipeline for each cycle corresponding to the execution time of one stage. The number of pipelines matches the number of index ranges that are the targets of the partial parallel trials. For one replica, the processing unit 12 executes partial parallel trials for each of the multiple index ranges in sequence, such as the first index range, the second index range, .... After executing the partial parallel trial for the last index range for a certain replica, the processing unit 12 returns to the partial parallel trial for the first index range for that replica.
[0037] In one example, the processing unit 12 executes in parallel a process P1 corresponding to a first pipeline and a process P2 corresponding to a second pipeline. The process P1 corresponds to the first pipeline, i.e., a series of stages in the first pipeline. The process P2 corresponds to the second pipeline, i.e., a series of stages in the second pipeline. The process P1 is assigned an index range {1 to m}. That is, the process P1 is assigned an index range {x1 to x} corresponding to a group of state variables. m More specifically, the process P1 is assigned {x01 to x0 m}, {x11~x1 m}, {x21~x2 m}, {x31~x3 m} is assigned to the process P2. In addition, the index range {m+1 to N} is assigned to the process P2. That is, the group of state variables {x m+1 ~x N More specifically, the process P2 is assigned {x0 m+1 ~x0 N}, {x1 m+1 ~x1 N}, {x2 m+1 ~x2 N}, {x3 m+1 ~x3 N} is assigned. The index range assigned to process P1 and the index range assigned to process P2 do not overlap.
[0038] Each of the processes P1 and P2 has the same number of stages. The stages include each procedure in the partial parallel trial. In one example, the number of stages is two. In this case, for example, the first stage is determining which one state variable is to be updated according to the amount of change in energy for each update candidate when each state variable belonging to the corresponding index range is set as an update candidate. The second stage is updating the one state variable to be updated obtained by the determination. The update of the state variable involves the update of the local field described above. Furthermore, these stages can be further subdivided.
[0039] Process P1 includes stages P1-1 and P1-2. Stage P1-1 is the first stage of process P1. Stage P1-2 is the second stage of process P2. The second stage is executed after the first stage. In stages P1-1 and P1-2, state variables belonging to the index range {1 to m} are processed. Process P2 includes stages P2-1 and P2-2. Stage P2-1 is the first stage of process P2. Stage P2-2 is the second stage of process P2. In stages P2-1 and P2-2, state variables belonging to the index range {m+1 to N} are processed.
[0040] The processing unit 12 processes different replicas at the same timing in each stage included in the processes P1 and P2. For example, the processing unit 12 has an arithmetic circuit that executes each stage included in the processes P1 and P2. For example, the processing unit 12 processes the replicas R0 to R3 at each of the times t1, t2, t3, t4, and t5 as follows. Note that, among the times t1 to t5, the direction from time t1 to time t5 is the positive direction of time. Also, time t1 is assumed to be a timing when a certain amount of time has elapsed since the processing unit 12 started searching for a solution using the replicas R0 to R3.
[0041] At time t1, the processing unit 12 processes replica R0 at stage P1-1, and processes replica R1 at stage P1-2. The processing unit 12 also processes replica R2 at stage P2-1, and processes replica R3 at stage P2-2.
[0042] At time t2, the processing unit 12 processes replica R3 at stage P1-1 and replica R0 at stage P1-2, and also processes replica R1 at stage P2-1 and replica R2 at stage P2-2.
[0043] At time t3, the processing unit 12 processes replica R2 at stage P1-1, and processes replica R3 at stage P1-2. The processing unit 12 also processes replica R0 at stage P2-1, and processes replica R1 at stage P2-2.
[0044] At time t4, the processing unit 12 processes replica R1 at stage P1-1, and processes replica R2 at stage P1-2. The processing unit 12 also processes replica R3 at stage P2-1, and processes replica R0 at stage P2-2.
[0045] At time t5, the processing unit 12 processes replica R0 at stage P1-1, and processes replica R1 at stage P1-2. The processing unit 12 also processes replica R2 at stage P2-1, and processes replica R3 at stage P2-2.
[0046] The processing unit 12 performs a round of partial parallel trials for the entire index range of each of the replicas R0 to R3 at times t1 to t4, and repeats the procedure of the round from time t5 onward. When the processing unit 12 completes the search by the SA method or the replica exchange method by repeatedly executing the procedure, it outputs a set of values of multiple state variables indicated by each of the replicas R0 to R3 as a solution. For example, the processing unit 12 may output the solution with the smallest energy as the best solution among the four solutions obtained for the replicas R0 to R3.
[0047] In this way, according to the data processing device 10, the first process and the second process are executed in parallel. In the first process, a plurality of stages are executed for a plurality of replicas. The plurality of stages include determining a first state variable to be updated according to an amount of change in energy when each of a plurality of first state variables belonging to a first index range is set as an update candidate, and updating the value of the first state variable to be updated. In the second process, a plurality of stages are executed for a plurality of replicas. The plurality of stages include determining a second state variable to be updated according to an amount of change in energy when each of a plurality of second state variables belonging to a second index range is set as an update candidate, and updating the value of the second state variable to be updated. The second index range does not overlap with the first index range. In the stages included in the first process and the second process, different replicas are processed at the same timing.
[0048] This allows the data processing device 10 to effectively utilize computing resources. In particular, the data processing device 10 shifts the processing timing of each replica in each stage included in the first process and the second process so that different replicas are processed at the same timing. This ensures that the principle of sequential processing in the MCMC method is observed for each replica. Therefore, the data processing device 10 can appropriately parallelize the search by partial parallel trials for the replicas for multiple replicas. Furthermore, the data processing device 10 can appropriately obtain a solution using each replica. In this way, the data processing device 10 can effectively utilize the computational resources of the processing unit 12 to efficiently find a solution.
[0049] Furthermore, the data processing device 10 only needs to hold one set of all weighting coefficients for multiple replicas, which means that the data processing device 10 does not need to increase the memory capacity for holding the weighting coefficients even if the number of replicas increases.
[0050] Although two pipelines are exemplified in the example of the first embodiment, the data processing device 10 may execute three or more pipelines in parallel. Also, the number of stages in one pipeline may be three or more. The number of replicas may be any number other than four. For example, when the processing unit 12 executes three or more pipelines in parallel, any two pipelines correspond to the first process and the second process.
[0051] [Second embodiment] Next, a second embodiment will be described. FIG. 2 illustrates an example of hardware of a data processing device according to the second embodiment.
[0052] The data processing device 20 searches for a solution to a combinatorial optimization problem using the MCMC method and outputs the searched solution. The data processing device 20 has a CPU 21, a RAM 22, a HDD (Hard Disk Drive) 23, a GPU 24, an input interface 25, a media reader 26, a NIC (Network Interface Card) 27, and an accelerator card 28.
[0053] The CPU 21 is a processor that executes program instructions. The CPU 21 loads at least a part of the program and data stored in the HDD 23 into the RAM 22 and executes the program. The CPU 21 may include multiple processor cores. The data processing device 20 may have multiple processors. The processes described below may be executed in parallel using multiple processors or processor cores. A set of multiple processors may also be called a "multiprocessor" or simply a "processor."
[0054] The RAM 22 is a volatile semiconductor memory that temporarily stores programs executed by the CPU 21 and data used in calculations by the CPU 21. Note that the data processing device 20 may include a type of memory other than a RAM, or may include multiple memories.
[0055] The HDD 23 is a non-volatile storage device that stores software programs such as an OS (Operating System), middleware, and application software, as well as data. The data processing device 20 may include other types of storage devices, such as a flash memory or an SSD (Solid State Drive), or may include multiple non-volatile storage devices.
[0056] The GPU 24 outputs an image to a display 101 connected to the data processing device 20 in accordance with an instruction from the CPU 21. As the display 101, any type of display can be used, such as a CRT (Cathode Ray Tube) display, a liquid crystal display (LCD: Liquid Crystal Display), a plasma display, or an organic EL (OEL: Organic Electro-Luminescence) display.
[0057] The input interface 25 acquires an input signal from an input device 102 connected to the data processing device 20, and outputs the signal to the CPU 101. As the input device 102, a pointing device such as a mouse, a touch panel, a touch pad, or a trackball, a keyboard, a remote controller, a button switch, or the like can be used. In addition, a plurality of types of input devices may be connected to the data processing device 20.
[0058] The medium reader 26 is a reading device that reads programs and data recorded on the recording medium 103. For example, a magnetic disk, an optical disk, a magneto-optical disk (MO: Magneto-Optical disk), a semiconductor memory, etc. can be used as the recording medium 103. Magnetic disks include flexible disks (FD: Flexible Disks) and HDDs. Optical disks include compact discs (CDs) and digital versatile discs (DVDs).
[0059] The medium reader 26 copies the program or data read from the recording medium 103 to another recording medium such as the RAM 22 or the HDD 23. The read program is executed by the CPU 21, for example. The recording medium 103 may be a portable recording medium, and may be used for distributing the program or data. The recording medium 103 and the HDD 23 may also be referred to as computer-readable recording media.
[0060] The NIC 27 is an interface that is connected to the network 104 and communicates with other computers via the network 104. The NIC 27 is connected to a communication device such as a switch or a router via a cable. The NIC 27 may be a wireless communication interface.
[0061] The accelerator card 28 is a hardware accelerator that searches for a solution to a problem expressed by the Ising-type energy function of formula (1) using the MCMC method. The accelerator card 28 can be used as a sampler that samples a state following the Boltzmann distribution at a given temperature by performing the MCMC method at a constant temperature or the replica exchange method that exchanges the state of the Ising model between multiple temperatures. To solve a combinatorial optimization problem, the accelerator card 28 executes annealing processes such as the replica exchange method or the SA method that gradually lowers the temperature value.
[0062] The SA method is a method for efficiently finding an optimal solution by sampling the state according to the Boltzmann distribution at each temperature value and lowering the temperature value used for sampling from a high temperature to a low temperature, i.e., increasing the inverse temperature β. Even on the low temperature side, i.e., when β is large, the state changes to some extent, so that it is more likely that a good solution can be found even if the temperature value is lowered quickly. For example, when the SA method is used, the accelerator card 28 repeats the operation of trying state transitions at a constant temperature value a certain number of times and then lowering the temperature value.
[0063] The replica exchange method is a technique in which the MCMC method is executed independently using multiple temperature values, and the temperature values are appropriately exchanged for the states obtained at each temperature value. A good solution can be found efficiently by searching a narrow range of the state space by MCMC at low temperatures and searching a wide range of the state space by MCMC at high temperatures. For example, when using the replica exchange method, the accelerator card 28 performs parallel trials of state transitions at each of multiple temperature values, and repeats the operation of exchanging the temperature values for the states obtained at each temperature value with a predetermined exchange probability every time a certain number of trials are performed.
[0064] The accelerator card 28 has an FPGA 28a. The FPGA 28a realizes a search function in the accelerator card 28. The search function may be realized by other types of electronic circuits such as a GPU or an ASIC. The FPGA 28a has a memory 28b. The memory 28b holds data such as problem information used in the search in the FPGA 28a and a solution searched by the FPGA 28a. The FPGA 28a may have a plurality of memories including the memory 28b. The FPGA 28a is an example of the processing unit 12 of the first embodiment. The memory 28b is an example of the storage unit 11 of the first embodiment. The accelerator card 28 may have a RAM outside the FPGA 28a, and data stored in the memory 28b may be temporarily saved to the RAM in response to the processing of the FPGA 28a.
[0065] A hardware accelerator that searches for a solution to an Ising-type problem, such as the accelerator card 28, is sometimes called an Ising machine or a Boltzmann machine. The accelerator card 28 performs a solution search in parallel using multiple replicas. The replicas indicate multiple state variables included in the energy function. In the following description, the state variables are represented as bits. Each bit included in the energy function is associated with an integer index and is identified by the index.
[0066] FIG. 3 is a diagram illustrating an example of functions of the data processing device. The data processing device 20 includes memory units 30, 30a, 30b, and 30c, a read unit 31, h calculation units 32a1 to 32aN, ΔE calculation units 33a1 to 33aN, and selectors 34, 34a, 34b, and 34c. N is the number of bits in one replica. In one example, N=1024. In this case, for example, each bit is identified by an index of 1 to 1024.
[0067] The h calculation units 32a1 to 32aN refer to the h calculation units 32a1, 32a2, ..., 32a(N-1), and 32aN. The illustration of the h calculation units 32a2 to 32a(N-1) is omitted. The ΔE calculation units 33a1 to 33aN refer to the ΔE calculation units 33a1, 33a2, ..., 33a(N-1), and 33aN. The illustration of the ΔE calculation units 33a2, ..., 33a(N-1) is omitted.
[0068] For example, the memory units 30 to 30c are realized by a plurality of memories including the memory 28b in the FPGA 28a. The read unit 31, the h calculation units 32a1 to 32aN, the ΔE calculation units 33a1 to 33aN, and the selectors 34, 34a, 34b, and 34c are realized by electronic circuits in the FPGA 28a.
[0069] In Fig. 3, the h calculation units 32a1 to 32aN are named with the subscript n, such as "hn" calculation unit, so that it is easy to understand that they correspond to the nth bit. Also, in Fig. 3, the ΔE calculation units 33a1 to 33aN are named with the subscript n, such as "ΔEn" calculation unit, so that it is easy to understand that they correspond to the nth bit.
[0070] For example, h calculation unit 32a1 and ΔE calculation unit 33a1 perform calculations on the first bit of N bits. Also, h calculation unit 32ai and ΔE calculation unit 33ai perform calculations on the i-th bit. Similarly, the number n at the end of the code such as "32an" or "33an" indicates that a calculation corresponding to the n-th bit is to be performed.
[0071] Here, the data processing device 20 divides the entire index into a plurality of index ranges, and for each index range, performs parallel trials of inversion of each bit corresponding to the index belonging to the index range, that is, partial parallel trials. As an example, the data processing device 20 divides the entire index into four index ranges. The first index range is 1 to i. The second index range is i+1 to j. The third index range is j+1 to k. The fourth index range is k+1 to N. When N=1024, the number of indexes belonging to each index range may be 256. The above circuits in the FPGA 28a are divided into four groups G0, G1, G2, and G3 as follows.
[0072] The memory unit 30, the h calculation units 32a1-32ai, the ΔE calculation units 33a1-33ai, and the selector 34 belong to group G0. The memory unit 30a, the h calculation units 32a(i+1)-32aj, the ΔE calculation units 33a(i+1)-33aj, and the selector 34a belong to group G1. The memory unit 30b, the h calculation units 32a(j+1)-32ak, the ΔE calculation units 33a(j+1)-33ak, and the selector 34b belong to group G2. The memory unit 30c, the h calculation units 32a(k+1)-32aN, the ΔE calculation units 33a(k+1)-33aN, and the selector 34c belong to group G3.
[0073] A decision as to whether to invert any bit included in a state vector in one group and an inversion of the bit according to the decision result correspond to one trial of searching for a solution in the group. However, one trial may not result in inversion of a bit. The one trial is repeatedly executed. In each group, the h calculation unit and the ΔE calculation unit perform partial parallel trials in parallel for each bit belonging to the group, thereby speeding up the calculation. As described later, the data processing device 20 performs partial parallel trials for replicas in parallel for multiple replicas using multiple pipelines, thereby efficiently utilizing the calculation resources of the FPGA 28a. Specifically, the data processing device 20 processes multiple replicas in parallel using four pipelines corresponding to groups G0 to G3. In this example, the number of replicas is 16. The 16 replicas are denoted as replicas R0, R1, ..., R15.
[0074] Here, the information stored in the memory units 30 to 30c will be described. Each of the memory units 30 to 30c stores a weighting factor W={W γ,δ When the number of bits of the state vector is N, the total number of weight coefficients is N 2 It becomes. W γ,δ =W δ,γ It is. γ,γ =0.
[0075] The memory unit 30 stores the weighting coefficient W 1,1 ~W 1,N ,W 2,1 ~W 2,N ,…,W i,1 ~W i,N For example, the weighting factor W 1,1 ~W 1,N is used in the calculation corresponding to the 1st bit. The total number of weighting coefficients stored in the memory unit 30 is i×N.
[0076] The memory unit 30a stores the weighting coefficient W i+1,1 ~W i+1,N ,W i+2,1 ~W i+2,N ,…,Wj,1 ~W j,N The total number of weighting coefficients stored in the memory unit 30a is (ji)×N.
[0077] The memory unit 30b stores the weighting coefficient W j+1,1 ~W j+1,N ,W j+2,1 ~W j+2,N ,…,W k,1 ~W k,N The total number of weighting coefficients stored in the memory unit 30b is (kj)×N.
[0078] The memory unit 30c stores the weighting coefficient W k+1,1 ~W k+1,N ,W k+2,1 ~W k+2,N ,…,W N,1 ~W N,N The total number of weighting coefficients stored in the memory unit 30c is (Nk)×N.
[0079] In the following, the h calculation unit 32a1 and ΔE calculation unit 33a1 corresponding to the first bit will be mainly described as an example. The h calculation units 32a2 to 32aN and ΔE calculation units 33a2 to 33aN with the same names have the same functions.
[0080] The readout unit 31 reads out the weighting coefficients W corresponding to the indexes supplied by the selectors 34 to 34c from the memory units 30 to 30c, and outputs them to the h calculation units 32a1 to 32aN. Since the number of the selectors 34 to 34c is four, a maximum of four indexes are simultaneously supplied to the readout unit 31. That is, the readout unit 31 outputs a maximum of four weighting coefficients simultaneously to each of the h calculation units 32a1 to 32aN. The four weighting coefficients correspond to the four replicas processed in parallel by the four selectors 34 to 34c. The readout unit 31 reads out the maximum of four weighting coefficients W corresponding to the indexes supplied by the selectors 34 to 34c from the memory units 30 to 30c, and outputs them to the h calculation unit 32a1 to 32aN. 1,1 ~W 1,NThe reading unit 31 is an address decoder that converts the supplied index into an address in the memory units 30 to 30c and reads out the weighting coefficient at that address. The reading unit 31 may be provided separately for each of the groups G0 to G3.
[0081] The h calculation unit 32a1 calculates the local field h1 for each of the four replicas processed in parallel based on the formulas (3) and (4) using the weighting coefficients supplied from the readout unit 31. For example, the h calculation unit 32a1 has a register that holds the local field h1 previously calculated for the replica, and updates the h1 of the replica stored in the register by multiplying the Δh1 of the replica by the h1. Note that a signal indicating the inversion direction of the bit indicated by the index to be inverted for each replica may be supplied from the selectors 34 to 34c to the h calculation unit 32a1. Alternatively, the readout unit 31 may receive a signal indicating the inversion direction and determine the sign of the weighting coefficient to be supplied to the h calculation unit 32a1 according to the inversion direction. The initial value of h1 is calculated in advance according to the formula (3) according to b1 according to the problem, and is set in advance in the register of the h calculation unit 32a1.
[0082] The ΔE calculation unit 33a1 calculates the energy change ΔE1 according to the inversion of the own bit in one replica, which is the next processing target and is held in the h calculation unit 32a1, based on the formula (2). The ΔE calculation unit 33a1 can determine the inversion direction of the own bit from the current value of the own bit of the corresponding replica, for example. For example, if the current value of the own bit is 0, the inversion direction is from 0 to 1, and if the current value of the own bit is 1, the inversion direction is from 1 to 0. The ΔE calculation unit 33a1 supplies the calculated ΔE1 to the selector 34.
[0083] The selector 34 performs a judgment of the formula (6) for each ΔE simultaneously supplied from the ΔE calculation units 33a1 to 33ai, and decides whether or not to invert the corresponding bit. For example, the selector 34 judges whether or not to permit inversion of the bit with index=1 for the energy change ΔE1 calculated by the ΔE calculation unit 33a1, based on the formula (6). Specifically, the selector 34 judges whether or not to invert the corresponding bit for the corresponding replica, based on a comparison between -ΔE1 and thermal noise corresponding to the temperature value T. The thermal noise corresponds to the product of the natural logarithm value of the uniform random number u in the formula (6) and the temperature value T.
[0084] Furthermore, the selector 34 randomly selects one of the bits determined to be invertible based on the formula (6) based on a random number, and supplies the readout unit 31 with an index corresponding to the selected bit.
[0085] The selector 34 may update the bit corresponding to the index by supplying the index to a storage unit that holds the bit corresponding to the replica. The selector 34 may update the energy of the replica by adding ΔE corresponding to the index to an energy storage unit that holds the energy corresponding to the current bit corresponding to the replica. In FIG. 3, the storage unit that holds the current bit corresponding to each replica and the energy storage unit that holds the energy corresponding to the current bit of each replica are omitted. The storage unit and the energy storage unit may be realized, for example, by a storage area of the memory 28b in the FPGA 28a or by a register.
[0086] The selectors 34a to 34c also function in the same manner as the selector 34 with respect to the bits of their own group. FIG. 4 is a diagram illustrating an example of the function of local field updates in a group.
[0087] The memory unit 30 includes memories 30p1, 30p2, 30p3, and 30p4. The readout unit 31 is omitted in FIG. 4. The memory 30p1 stores a weighting factor W 1,1 ~W 1,i ,W2,1 ~W 2,i ,…,W i,1 ~W i,i The memory 30p2 stores W 1,i+1 ~W 1,j ,W 2,i+1 ~W 2,j ,…,W i,i+1 ~W i,j The memory 30p3 stores W 1,j+1 ~W 1,k ,W 2,j+1 ~W 2,k ,…,W i,j+1 ~W i,k The memory 30p4 stores W 1,k+1 ~W 1,N ,W 2,k+1 ~W 2,N ,…,W i,k+1 ~W i,N Remember.
[0088] The read unit 31 simultaneously reads out four weighting coefficients corresponding to bit updates in up to four groups for one h calculation unit from the memories 30p1, 30p2, 30p3, and 30p4, and supplies the weighting coefficients to the h calculation unit. For example, when the read unit 31 receives four indexes of bits to be updated from the selectors 34 to 34c, the read unit 31 reads out four weighting coefficients each from the memories 30p1 to 30p4 for the h calculation units 31a1 to 32ai, and supplies the weighting coefficients to the h calculation units 31a1 to 32ai.
[0089] For example, the W held in the memory 30p1 for the index of the bit to be updated output by the selector 34 is 1,1 ~W 1,i One weighting factor is read out from the memory 30p2 and supplied to the h calculation unit 32a1. 1,i+1 ~W 1,j One weighting factor is read out from the memory 30p3 and supplied to the h calculation unit 32a1. 1,j+1 ~W 1,kOne weighting factor is read out from the memory 30p4 and supplied to the h calculation unit 32a1. 1,k+1 ~W 1,N One weighting factor is read out from the weighting factor register 32a1 and supplied to the h calculation unit 32a1. Similarly, up to four weighting factors are supplied to the other h calculation units at the same time.
[0090] Each of the h calculation units 32a1 to 32ai uses up to four weighting coefficients provided to update the local fields corresponding to up to four replicas of its own bits in parallel based on the formulas (3) and (4). For example, the h calculation unit 32a1 has an h holding unit r1, selectors s11, s12, and s13, and adders c1, c2, c3, and c4.
[0091] The h-holding unit r1 holds the local field of the own bit corresponding to each of the 16 replicas. The h-holding unit r1 may be composed of a flip-flop or may be composed of four RAMs that read one word per read. The own bit in the h-calculation unit 32a1 is the bit with index=1.
[0092] The selector s11 reads out the local fields of the replicas to be updated in each group from the h-holding unit r1 and supplies them to the adders c1, c2, c3, and c4. The maximum number of local fields that the selector s11 reads out simultaneously from the h-holding unit r1 is four.
[0093] The adders c1, c2, c3, and c4 update the local field supplied from the selector s11 by adding the weighting coefficients read from the memories 30p1 to 30p4, respectively, and supply the updated local field to the selector s12. As described above, the sign of the weighting coefficient may be determined by the read unit 31 or by the h calculation unit 32a1 according to the bit inversion direction. The adder c1 updates the local field related to the replica being processed in the group G0. The adder c2 updates the local field related to the replica being processed in the group G1. The adder c3 updates the local field related to the replica being processed in the group G2. The adder c4 updates the local field related to the replica being processed in the group G3.
[0094] The selector s12 stores the local field of the relevant replica updated by the adders c1 to c4 in the h-holding unit r1. The selector s13 reads out the local field of its own bit in the replica to be processed next in group G0 from the h holding unit r1, and supplies it to the ΔE calculation unit 33a1.
[0095] In this way, the h calculation unit 32a1 can simultaneously update the local field corresponding to index=1 for a maximum of four replicas by using the selectors s11 and s12 and the adders c1, c2, c3, and c4.
[0096] Other h calculation units have the same function as the h calculation unit 32a1. For example, the h calculation unit 32ai has an h holding unit ri, selectors si1, si2, si3, and adders c5, c6, c7, and c8. The h holding unit ri holds the local field of the own bit corresponding to each of the 16 replicas. The own bit in the h calculation unit 32ai is the bit of index=i.
[0097] The selector si1 reads the local fields of the replicas to be updated in each group from the h-holding unit ri and supplies them to the adders c5, c6, c7, and c8. The maximum number of local fields that the selector si1 simultaneously reads from the h-holding unit ri is 4, the same as the h-holding unit r1.
[0098] The adders c5, c6, c7, and c8 update the local field supplied from the selector si1 by adding the weighting coefficients read from the memories 30p1 to 30p4, respectively, and supply the updated local field to the selector si2. As described above, the sign of the weighting coefficient may be determined by the read unit 31 or by the h calculation unit 32ai according to the bit inversion direction. The adder c5 updates the local field related to the replica being processed in the group G0. The adder c6 updates the local field related to the replica being processed in the group G1. The adder c7 updates the local field related to the replica being processed in the group G2. The adder c9 updates the local field related to the replica being processed in the group G3.
[0099] The selector si2 stores the local field of the relevant replica updated by the adders c5 to c8 in the h-holding unit ri. The selector si3 reads out the local field of its own bit in the replica to be processed next in the group G0 from the h-holding unit ri, and supplies it to the ΔE calculation unit 33ai.
[0100] Groups G1 to G3 also have the same local field update function as group G0. With the above configuration, the data processing device 20 executes four pipelines in parallel for 16 replicas.
[0101] FIG. 5 is a diagram illustrating an example of pipeline processing. As an example, the number of stages in one pipeline, that is, the number of stages, is assumed to be four. The first stage is ΔE calculation. The ΔE calculation is a process of calculating in parallel, in each group, ΔE for each bit belonging to the group. The second stage is flip determination. The flip determination is a process of selecting one bit to be inverted for ΔE of each bit calculated in parallel. The third stage is W Read. The W Read is a process of reading weight coefficients from the memory units 30 to 30c. The fourth stage is h update. The h update is a process of updating the local field related to the corresponding replica based on the read weight coefficient. In parallel with the h update stage, the inversion of the bit to be inverted in the corresponding replica is performed. Therefore, the h update stage can also be said to be a bit update stage.
[0102] Time charts 201, 202, 203, and 204 show replicas processed by four pipelines at each timing by stage. Time chart 201 shows the ΔE calculation stage. Time chart 202 shows the flip determination stage. Time chart 203 shows the W Read stage. Time chart 204 shows the h update stage. The positive direction of time is from left to right in the diagram.
[0103] G0, G1, G2, and G3 attached to each row of time charts 201, 202, 203, and 204 identify the pipeline to which the corresponding row belongs. In the pipeline corresponding to G0, an operation is performed on bits corresponding to index range 1 to i. In the pipeline corresponding to G1, an operation is performed on bits corresponding to index range i+1 to j. In the pipeline corresponding to G2, an operation is performed on bits corresponding to index range j+1 to k. In the pipeline corresponding to G3, an operation is performed on bits corresponding to index range k+1 to N.
[0104] The data processing device 20 starts processing the replica at a timing shifted by four pipeline stages or four or more stages so that the replica being processed in group G1 is processed after the h update of the replica being processed in group G0 is completed. This allows ΔE calculation to be performed in each replica using a local field that reflects the previous bit update, thereby observing the principle of sequential processing of MCMC.
[0105] Here, the update of the local field needs to be reflected in all bits of the corresponding replica. For this reason, the weighting coefficients are read out simultaneously for all bits of the four replicas. As illustrated in FIG. 5, the data processing device 20 divides the memory that holds the weighting coefficients corresponding to each group into, for example, memories 30p1 to 30p4. Therefore, accesses corresponding to multiple replicas do not overlap with each other to the same memory. For example, in step 203a in the time chart 203, the weighting coefficients are read out as follows.
[0106] FIG. 6 is a diagram showing an example of reading out the weighting coefficients. For example, the memory unit 30 of group G0 holds the weighting coefficients W0(G0), W0(G1), W0(G2), and W0(G3) in memories 30p1, 30p2, 30p3, and 30p4, respectively. The memory unit 30a of group G1 holds the weighting coefficients W1(G0), W1(G1), W1(G2), and W1(G3) in four memories. The memory unit 30b of group G2 holds the weighting coefficients W2(G0), W2(G1), W2(G2), and W2(G3) in four memories. The memory unit 30c of group G3 holds the weighting coefficients W3(G0), W3(G1), W3(G2), and W3(G3) in four memories.
[0107] The weighting factor W0(G0) is the weighting factor W0 corresponding to the G0 allocation, i.e., the update of bits in the index range 1 to i. 1,1 ~W 1,i ,W 2,1 ~W 2,i ,…,W i,1 ~W i,i It is.
[0108] Weighting factor W0(G1) is the weighting factor W corresponding to the update of the bits in the G1 allocation, that is, the bits in the index range i+1 to j. 1,i+1 ~W 1,j ,W 2,i+1 ~W 2,j ,…,W i,i+1 ~W i,j It is.
[0109] The weighting factor W0(G2) is the weighting factor W corresponding to the update of the bits of the G2 allocation, that is, the bits in the index range j+1 to k. 1,j+1 ~W 1,k ,W 2,j+1 ~W 2,k ,…,W i,j+1 ~W i,k It is.
[0110] The weighting factor W0(G3) is the weighting factor W corresponding to the update of the bits of the G3 allocation, that is, the bits in the index range k+1 to N. 1,k+1 ~W 1,N ,W 2,k+1 ~W 2,N ,…,W i,k+1 ~W i,N It is.
[0111] The weighting factor W1(G0) is the weighting factor W corresponding to the update of the bits in the G0 allocation. i+1,1 ~W i+1,i ,W i+2,1 ~W i+2,i ,…,W j,1 ~W j,i It is. The weighting factor W1(G1) is the weighting factor W corresponding to the update of the bits in the G1 allocation. i+1,i+1 ~W i+1,j ,W i+2,i+1 ~W i+2,j ,…,W j,i+1 ~W j,j It is.
[0112] The weighting factor W1(G2) is the weighting factor W corresponding to the update of the bits in the G2 allocation. i+1,j+1 ~W i+1,k ,W i+2,j+1 ~W i+2,k ,…,Wj,j+1 ~W j,k It is.
[0113] Weighting factor W1(G3) is the weighting factor W corresponding to the update of the bits in the G3 allocation. i+1,k+1 ~W i+1,N ,W i+2,k+1 ~W i+2,N ,…,W j,k+1 ~W j,N It is.
[0114] The weighting factor W2(G0) is the weighting factor W corresponding to the update of the bits in the G0 allocation. j+1,1 ~W j+1,i ,W j+2,1 ~W j+2,i ,…,W k,1 ~W k,i It is. The weighting factor W2(G1) is the weighting factor W corresponding to the update of the bits in the G1 allocation. j+1,i+1 ~W j+1,j ,W j+2,i+1 ~W j+2,j ,…,W k,i+1 ~W k,j It is.
[0115] Weighting factor W2(G2) is the weighting factor W corresponding to the update of the bits in the G2 allocation. j+1,j+1 ~W j+1,k ,W j+2,j+1 ~W j+2,k ,…,W k,j+1 ~W k,k It is.
[0116] Weighting factor W2(G3) is the weighting factor W corresponding to the update of the bits in the G3 allocation. j+1,k+1 ~W j+1,N ,W j+2,k+1 ~W j+2,N ,…,W k,k+1 ~W k,N It is.
[0117] Weighting factor W3(G0) is the weighting factor W3(G0) corresponding to the update of the bits in the G0 allocation. k+1,1 ~W k+1,i ,W k+2,1 ~W k+2,i ,…,W N,1 ~W N,i It is. The weighting factor W3(G1) is the weighting factor W3(G1) corresponding to the update of the bits in the G1 allocation. k+1,i+1 ~W k+1,j ,W k+2,i+1 ~W k+2,j ,…,W N,i+1 ~W N,j It is.
[0118] Weighting factor W3(G2) is the weighting factor W corresponding to the update of the bits in the G2 allocation. k+1,j+1 ~W k+1,k ,W k+2,j+1 ~W k+2,k ,…,W N,j+1 ~W N,k It is.
[0119] Weighting factor W3(G3) is the weighting factor W corresponding to the update of the bits of the G3 allocation. k+1,k+1 ~W k+1,N ,W k+2,k+1 ~W k+2,N ,…,W N,k+1 ~W N,N It is.
[0120] In this way, the memory units 30 to 30c store weighting coefficients in separate memories for each index range corresponding to the groups G0 to G3. In step 203a, for the bit update of the replica R0 in the group G0, weighting coefficients are read from the memories storing W0(G0) to W3(G0) of the groups G0 to G3. For the bit update of the replica R12 in the group G1, weighting coefficients are read from the memories storing W0(G1) to W3(G1) of the groups G0 to G3. For the bit update of the replica R8 in the group G2, weighting coefficients are read from the memories storing W0(G2) to W3(G2) of the groups G0 to G3. For the bit update of the replica R4 in the group G3, weighting coefficients are read from the memories storing W0(G3) to W3(G3) of the groups G0 to G3.
[0121] Therefore, the data processing device 20 can simultaneously read out the weighting coefficients corresponding to the update of the G0 allocation bits of replica R0, the update of the G1 allocation bits of replica R12, the update of the G2 allocation bits of replica R8, and the update of the G3 allocation bits of replica R4 in step 203a, and can update the local fields of all the bits in each of replicas R0, R12, R8, and R4 in parallel. At this time, the accesses to read out the weighting coefficients do not overlap with each other in the same memory.
[0122] Next, the processing procedure of the data processing device 20 will be described. FIG. 7 is a flowchart showing an example of processing by the data processing device. (S10) The CPU 21 sets operation parameters in the FPGA 28a. For example, the operation parameters include the number of divided groups of the partial region, that is, the number of groups to be created by dividing the entire index range, and the number of replicas M. For example, the number of replicas M=16. The operation parameters also include the replica interval between groups. For example, the replica interval is 4. The replica interval is set to the number of stages in the pipeline or a value equal to or greater than the number of stages. In this case, for the execution count i of the following loop process, groups G0 to G3 process replicas identified by the following numbers:
[0123] The number G0(i) of the replica processed by group G0 is G0(i) = i mod M. The number G1(i) of the replica processed by group G1 is G1(i) = i + 12 mod M. The number G2(i) of the replica processed by group G2 is G2(i) = i + 8 mod M. The number G3(i) of the replica processed by group G3 is G3(i) = i + 4 mod M. For integers a and b, "a mod b" indicates the remainder when a is divided by b. The +4, +8, etc. included in a are determined according to the replica interval.
[0124] (S11) FPGA 28a performs loop processing for the number of replicas. In the loop processing, FPGA 28a operates four groups G0 to G3 in parallel with a shift by the replica interval between groups. FPGA 28a sets the initial value of the execution count i of the loop processing to 0 and increments the execution count i until i < M - 1 is satisfied.
[0125] FPGA 28a executes the following steps S12 to S12c in parallel. (S12) ΔE calculation units 33a1 to 33ai calculate ΔE1 to ΔE i for group G0 of replica R(G0(i)) and output ΔE1 to ΔE i to selector 34. Note that the i at the end of the symbol of the ΔE calculation unit and the subscript i of ΔE indicate the end of the index of group G0.
[0126] (S12a) ΔE calculation units 33a(i + 1) to 33aj calculate ΔE i+1 to ΔE j for group G1 of replica R(G1(i)) and output ΔE i+1 to ΔE j to selector 34a.
[0127] (S12b) ΔE calculation units 33a(j + 1) to 33ak calculate ΔE j+1 to ΔE k for group G2 of replica R(G2(i)) and output ΔE j+1 to ΔE k to selector 34b.
[0128] (S12c) ΔE calculation units 33a(k + 1) to 33aN calculate ΔE k+1 to ΔE N for group G3 of replica R(G3(i)) and output ΔE k+1 to ΔE N to selector 34c.
[0129] FPGA 28a executes the following steps S13 to S13c in parallel. (S13) The selector 34 performs a flip decision for the group G0 of the replica R (G0(i)). i Among the bits that can be inverted based on the formula (6), a process is performed to select one bit as a flip target, and it is determined whether or not the flip target bit has been selected. For example, ΔE1 to ΔE i And if there are no bits that can be flipped based on equation (6), the selector 34 does not select any bits to be flipped.
[0130] (S13a) The selector 34a performs a flip decision for the group G1 of the replica R (G1(i)). Specifically, the selector 34a determines whether or not ΔE i+1 ~ΔE j A process is performed to select one bit as a flip target from among the bits that can be inverted based on equation (6), and it is determined whether or not the bit to be flipped has been selected.
[0131] (S13b) The selector 34b performs a flip decision for the group G2 of the replica R (G2(i)). Specifically, the selector 34b determines whether or not ΔE j+1 ~ΔE k A process is performed to select one bit as a flip target from among the bits that can be inverted based on equation (6), and it is determined whether or not the bit to be flipped has been selected.
[0132] (S13c) The selector 34c performs a flip decision for the group G3 of the replica R (G3(i)). k+1 ~ΔE N A process is performed to select one bit as a flip target from among the bits that can be inverted based on equation (6), and it is determined whether or not the bit to be flipped has been selected.
[0133] The FPGA 28a executes the following steps S14 to S14c in parallel. (S14) If the selector 34 selects a bit to be flipped in the determination in step S13, it outputs the index of the selected bit to the readout unit 31 and proceeds to step S15. If the selector 34 does not select a bit to be flipped in the determination in step S13, it skips steps S15 and S16 below and proceeds to step S17. When steps S15 and S16 are skipped, group G0 waits without performing W Read and h update for replica R(G0(i)) while the other groups perform steps corresponding to steps S15 and S16.
[0134] (S14a) If the selector 34a selects a bit to be flipped in the determination of step S13a, it outputs the index of the selected bit to the readout unit 31 and proceeds to step S15a. If the selector 34a does not select a bit to be flipped in the determination of step S13a, it skips steps S15a and S16a below and proceeds to step S17. When steps S15a and S16a are skipped, group G1 waits without performing W Read and h update for replica R (G1(i)) while the other groups perform steps corresponding to steps S15a and S16a.
[0135] (S14b) If the selector 34b selects a bit to be flipped in the determination of step S13b, it outputs the index of the selected bit to the readout unit 31 and proceeds to step S15b. If the selector 34b does not select a bit to be flipped in the determination of step S13b, it skips steps S15b and S16b below and proceeds to step S17. When steps S15b and S16b are skipped, group G2 waits without executing W Read and h update for replica R(G2(i)) while the other groups perform steps corresponding to steps S15b and S16b.
[0136] (S14c) If the selector 34c selects a bit to be flipped in the determination of step S13c, it outputs the index of the selected bit to the readout unit 31 and proceeds to step S15c. If the selector 34c does not select a bit to be flipped in the determination of step S13c, it skips steps S15c and S16c below and proceeds to step S17. When steps S15c and S16c are skipped, the group G3 waits without performing W Read and h update for replica R(G3(i)) while the other groups perform steps corresponding to steps S15c and S16c.
[0137] The FPGA 28a executes the following steps S15 to S15c in parallel. (S15) The reading unit 31 reads out weighting coefficients for all groups of the replica R (G0(i)) based on the indexes supplied from the selectors 34 to 34c.
[0138] (S15a) The reading unit 31 reads out weighting coefficients for all groups of the replica R (G1(i)) based on the indexes supplied from the selectors 34 to 34c.
[0139] (S15b) The reading unit 31 reads out weighting coefficients for all groups of the replica R(G2(i)) based on the indexes supplied from the selectors 34 to 34c.
[0140] (S15c) The reading unit 31 reads out weighting coefficients for all groups of the replica R (G3(i)) based on the indexes supplied from the selectors 34 to 34c.
[0141] The FPGA 28a executes the following steps S16 to S16c in parallel. (S16) The calculation units 32a1 to 32aN perform LF update for all groups of replica R(G0(i)), that is, update of the local field.
[0142] (S16a) The h calculation units 32a1 to 32aN perform LF updates for all groups of the replica R(G1(i)), that is, updates of the local field. (S16b) The h calculation units 32a1 to 32aN perform LF updates for all groups of the replica R(G2(i)), that is, updates of the local field.
[0143] (S16c) The h calculation units 32a1 to 32aN perform LF updates for all groups of the replica R(G3(i)), that is, updates of the local field. (S17) The FPGA 28a repeatedly executes steps S12 to S16, S12a to S16a, S12b to S16b, and S12c to S16c until the number of execution times i of the loop process satisfies i < M - 1. When i < M - 1 is satisfied, the FPGA 28a exits the loop process and proceeds to step S18.
[0144] (S18) The FPGA 28a determines whether the search has ended. If the search has ended, the FPGA 28a ends the process. If the search has not ended, the FPGA 28a advances the process to step S11.
[0145] In the search for a solution by the FPGA 28a, the SA method or the replica exchange method is used. When the SA method is used, the FPGA 28a performs a process of decreasing the temperature value used for the flip determination of each replica at a predetermined timing. When the replica exchange method is used, the FPGA 28a performs a process of exchanging the temperature values used for each replica among the replicas at a predetermined timing. Also, in step S16, the FPGA 28a also performs in parallel the update of the bits to be flipped for the corresponding replica based on the indexes of the bits to be inverted output by the selectors 34 to 34c in steps S14 to S14c.
[0146] When the process is completed, the FPGA 28a outputs bit strings corresponding to each finally obtained replica as a solution to the CPU 21. The FPGA 28a may output energy corresponding to each replica together with the bit string to the CPU 21. The FPGA 28a may output a solution with the lowest energy among the solutions obtained by the search to the CPU 21 as a final solution.
[0147] In this way, the data processing device 20 of the second embodiment executes four pipelines in parallel, each of which performs partial parallel trials of multiple replicas, using groups G0 to G3. This makes it possible to improve the performance of solving relatively large-scale problems by effectively utilizing the resources of the computing unit such as the FPGA 28a while adhering to the principle of sequential processing of MCMC and ensuring the convergence of the solution.
[0148] [Third embodiment] Next, a third embodiment will be described. Differences from the second embodiment will be mainly described, and descriptions of common points will be omitted.
[0149] In the second embodiment, all weighting coefficients including zero coefficients are stored in memory in groups, so that the data processing device 20 can process all cases including cases where all weighting coefficients are non-zero in a certain time. On the other hand, not all weighting coefficients are always non-zero. Depending on the problem, some weighting coefficients may be non-zero, while other weighting coefficients may be zero. In such a case, from the viewpoint of reducing the memory capacity for weighting coefficients, it may be better to configure the memory capacity to be reduced by not storing weighting coefficients with a value of zero in memory, but storing address information indicating the positions of the weighting coefficients and the values of the weighting coefficients in memory.
[0150] Therefore, in the third embodiment, the data processing device 20 provides a function of not holding weighting coefficients with a value of 0 in the memory, thereby reducing the memory capacity used compared to the second embodiment. In the data processor 20 of the third embodiment, the read time of the weight coefficients for each group varies depending on the number of non-zero weight coefficients to be read in association with bit updates. For this reason, in the third embodiment, the data processor 20 has a mechanism for stalling the pipeline in accordance with the read time of the weight coefficients.
[0151] In the third embodiment, as in the second embodiment, as an example, the data processing device 20 executes four pipelines. The number of stages in each pipeline is four. The number of replicas is sixteen.
[0152] FIG. 8 illustrates an example of pipeline processing according to the third embodiment. Time charts 211, 212, 213, and 214 show replicas processed by four pipelines at each timing by stage. Time chart 211 shows the ΔE calculation stage. Time chart 212 shows the Flip determination stage. Time chart 213 shows the W Read stage. Time chart 214 shows the h update stage. The direction from left to right in the figure is the positive direction of time. G0, G1, G2, and G3 attached to each row of time charts 211, 212, 213, and 214 identify the pipeline to which the corresponding row belongs.
[0153] For example, in step 213a of the time chart 213, a memory read conflict occurs in reading weight coefficients for updating the local fields of replicas R2, R6, R10, and R14, and the reading of weight coefficients for replicas R10 and R14 is delayed. In this case, groups G0 to G3 stall the pipeline once at the W Read stage, and propagate the stall to the h update stage immediately thereafter. Similarly, groups G0 to G3 also propagate pipeline stalls to the Δ calculation stage and the Flip decision stage immediately after the h update. This allows the data processing device 20 to maintain the principle of sequential processing of MCMC in all replicas.
[0154] Next, a memory configuration for storing weighting coefficients taking into account the sparsity of the weighting coefficients will be described. FIG. 9 is a diagram illustrating an example of a memory configuration for storing weighting coefficients.
[0155] When the number of bits in one replica is 1024, the entire weighting coefficient C1 is represented by a matrix with 1024 row elements and 1024 column elements. When the 1024 bits are divided into four groups of 256 bits each, the weighting coefficient assigned to one group, for example, group G0, is a portion C1a of the entire C1. The portion C1a has 256 row elements and 1024 column elements.
[0156] The portion C1a includes weighting factors W00, W10, W20, and W30. The weighting factor W00 is a weighting factor corresponding to an update of bits assigned to group G0. The weighting factor W10 is a weighting factor corresponding to an update of bits assigned to group G1. The weighting factor W20 is a weighting factor corresponding to an update of bits assigned to group G2. The weighting factor W30 is a weighting factor corresponding to an update of bits assigned to group G3.
[0157] The weighting coefficients included in the portion C1a are stored in address storage memories 41, 42, 43, 44 and a weighting coefficient storage memory section 50. The address storage memories 41, 42, 43, 44 and the weighting coefficient storage memory section 50 are realized by a plurality of memories in the FPGA 28a including the memory 28b.
[0158] The address storage memories 41-44 are addresses indicating storage positions of non-zero weight coefficients among the 256×1024 weight coefficients, and hold addresses in the weight coefficient storage memory unit 50. The number of address storage memories 41-44 is four, corresponding to the index range for the groups. This makes it possible to simultaneously access the address storage memories 41-44 for a maximum of four update bits in four groups.
[0159] As an example, the size of one weighting coefficient is 2 bytes. In this case, the four address storage memories 41 to 44 as a whole are 3 bytes x 1024 words. One address storage memory is 256 words. The position of one word in the address storage memory, that is, the row position of the address storage memories 41 to 44 in FIG. 9, corresponds to the index of the update bit. One address storage memory holds the logical address of the weighting coefficient storage memory unit 50 in which up to 256 weighting coefficients are stored in one group, corresponding to the update bit, and the number of words in the weighting coefficient storage memory unit 50. The logical address is 2 bytes, and the number of words is 1 byte.
[0160] The weighting coefficient storage memory unit 50 stores the substance of the weighting coefficient. The weighting coefficient storage memory unit 50 is 32 Bytes×8 Kwords in total. As an example, the weighting coefficient storage memory unit 50 is configured with 256 words×32 memories.
[0161] For example, to hold all weighting coefficients including a weighting coefficient of 0 in memory, a memory capacity of 2 Bytes x 1K x 1K = 2MBytes is required, but with the memory configuration of FIG. 9, the memory capacity can be reduced to approximately half.
[0162] FIG. 10 is a diagram illustrating an example of the weighting coefficient storage memory unit. In the weighting coefficient storage memory unit 50, one weighting coefficient is composed of a position index indicating a position within a row and the value of the weighting coefficient. The position index takes values from 0 to 255, so it is 1 byte. As mentioned above, the value of a weighting coefficient is 2 bytes. Also, each row has the number of non-zero weighting coefficients contained in that row. This number takes values from 0 to 256, so it is 2 bytes.
[0163] The weighting coefficient storage memory unit 50 is divided into 32 physical memories, that is, physical memories, in an interleaved manner by the lowest 5 bits of the logical address. When a read access occurs from each of the four groups to the weighting coefficient storage memory unit 50 for one row of weighting coefficients, the accesses can be read simultaneously if they do not overlap with the same memory. The physical memories are identified by physical memory numbers. The physical memory numbers take values from 0 to 31. The line numbers in FIG. 10 correspond to the rows of the address storage memories 41 to 44, that is, the indexes of the update bits. The line numbers take values from L0 to L1023.
[0164] When a weighting factor is read from the logical address and the number of words held in the address storage memory, first, one row in FIG. 10 is specified by the logical address of the weighting factor storage memory unit 50. Specifically, the FPGA 28a specifies the physical memory number of the physical memory by the lower 5 bits of the logical address. Also, the FPGA 28a specifies the physical memory address in the physical memory by, for example, the lower 6 bits or more of the logical address. Also, when the number of words held in the address storage memory is more than one, the FPGA 28a specifies another physical memory and the physical memory address in the other physical memory according to the number of words. Then, the FPGA 28a reads the corresponding word from the weighting factor storage memory unit 50, and converts the position index of the read word into the index of the bit belonging to the corresponding group. The FPGA 28a sets the weighting factor to 0 for an index in which no weighting factor exists in the read word.
[0165] Factors that may cause the stall illustrated in FIG. 8 include, for example, a line in the weighting coefficient storage memory unit 50 consisting of multiple words, or multiple groups or multiple replicas simultaneously accessing one physical memory.
[0166] The data processing device 20 may have the following stall control function for the above memory configuration. FIG. 11 is a diagram illustrating an example of a function of stall control in a data processing device.
[0167] 11, group G0 in data processing device 20 will be mainly illustrated, but groups G1 to G3 in data processing device 20 have the same functions as group G0. Memory unit 30 of group G0 includes address storage memories 41 to 44 and weight coefficient storage memories 50a1, 50a2, ..., 50a32. Group G0 in the third embodiment also has physical memory address generation units 61, 62, 63, 64, conflict detection and arbitration units 65a1, 65a2, ..., 65a32, weight coefficient restoration unit 66, selectors 67a1 to 67ai, stall signal generation unit 68, and selectors 69a1 to 69ai in addition to the functions illustrated in FIGS.
[0168] The physical memory address generation units 61, 62, 63, 64, the conflict detection and arbitration units 65a1, 65a2, . . . , 65a32, the weighting coefficient restoration unit 66, the selectors 67a1 to 67ai, the stall signal generation unit 68, and the selectors 69a1 to 69ai are realized by electronic circuits included in the FPGA 28a.
[0169] The selectors 69a1 to 69ai are provided in place of the selectors s13 to si3 of the h calculation units 32a1 to 32ai. In the third embodiment, the weighting coefficient read function in the portion surrounded by the dotted line in FIG. 11 corresponds to the read unit 31 in FIG.
[0170] The address storage memories 41 to 44 hold the logical addresses and the number of words of the weighting coefficient storage memory unit 50 in which non-zero weighting coefficients are stored for the update bits of each group. Information on the update bits from each group is passed to each address storage memory, and the logical addresses and the number of words are read out in parallel from each address storage memory.
[0171] The weighting coefficient storage memories 50a1 to 50a32 are 32 physical memories that hold the values of the weighting coefficients. The weighting coefficient storage memories 50a1, 50a2, ..., 50a32 are included in the weighting coefficient storage memory unit 50. The weighting coefficient storage memories 50a1 to 50a32 are identified by physical memory numbers.
[0172] The physical memory address generation unit 61 obtains a logical address and a word count corresponding to the update bit of the replica processed in group G0 from the address storage memory 41. The physical memory address generation unit 61 generates a physical memory number and a physical memory address of the access destination based on the logical address and the word count. The physical memory address generation unit 61 outputs the generated physical memory address to the conflict detection and arbitration unit corresponding to the generated physical memory number.
[0173] The physical memory address generation unit 62 obtains a logical address and a word count corresponding to the update bit of the replica processed in group G1 from the address storage memory 42. The physical memory address generation unit 62 generates a physical memory number and a physical memory address of the access destination based on the logical address and the word count. The physical memory address generation unit 62 outputs the generated physical memory address to the conflict detection and arbitration unit corresponding to the generated physical memory number.
[0174] The physical memory address generation unit 63 obtains a logical address and a word count corresponding to the update bit of the replica processed in group G2 from the address storage memory 43. The physical memory address generation unit 63 generates a physical memory number and a physical memory address of the access destination based on the logical address and the word count. The physical memory address generation unit 63 outputs the generated physical memory address to the conflict detection and arbitration unit corresponding to the generated physical memory number.
[0175] The physical memory address generation unit 64 obtains a logical address and a word count corresponding to the update bit of the replica processed in group G3 from the address storage memory 44. The physical memory address generation unit 64 generates a physical memory number and a physical memory address of the access destination based on the logical address and the word count. The physical memory address generation unit 64 outputs the generated physical memory address to the conflict detection and arbitration unit corresponding to the generated physical memory number.
[0176] When reading multiple words, the physical memory address generation units 61-64 generate physical memory addresses in multiple cycles. However, when reading multiple words, the physical memory address generation units 61-64 can also access multiple physical memories simultaneously in a single cycle and leave the arbitration to the conflict detection arbitration units 65a1-65a32.
[0177] The conflict detection and arbitration units 65a1-65a32 are provided in a one-to-one correspondence with the weighting coefficient storage memories 50a1-50a32. The conflict detection and arbitration units 65a1-65a32 detect the presence or absence of conflict in access to the weighting coefficient storage memories 50a1-50a32 based on the physical memory addresses supplied from the physical memory address generation units 61-64, and arbitrate the conflicting access. For example, the conflict detection and arbitration unit 65a1 detects conflict in access to the weighting coefficient storage memory 50a1 based on the physical memory addresses supplied from the physical memory address generation units 61-64.
[0178] When accesses from each group conflict, the conflict detection arbitration units 65a1-65a32 put one of the accesses on hold according to the priority. For example, the conflict detection arbitration units 65a1-65a32 determine the priority by a method such as giving priority to a physical memory address with a smaller number or to an access with a larger number of words to be accessed. The conflict detection arbitration units 65a1-65a32 supply the physical memory address of the access destination to the weight coefficient storage memories 50a1-50a32 and cause the weight coefficient storage memories 50a1-50a32 to output the weight coefficient of the access destination to the weight coefficient restoration unit 66.
[0179] Furthermore, when the conflict detection and arbitration units 65a1-65a32 detect an access conflict, they output a signal indicating the detection of the access conflict to the stall signal generation unit 68. The conflict detection and arbitration units 65a1-65a32 may determine the number of cycles for stalling the pipeline according to the read time of the weighting coefficient associated with the access conflict, and notify the stall signal generation unit 68 of the number of cycles.
[0180] The weighting coefficient restoring unit 66 restores the weighting coefficients whose values are 0 based on the position indexes included in the words read out from the weighting coefficient storage memories 50a1 to 50a32. The selectors 67a1 to 67ai are provided in a one-to-one correspondence with the h calculation units 32a1 to 32ai. The selectors 67a1 to 67ai supply the weighting coefficients restored by the weighting coefficient restoration unit 66 to the h calculation units 32a1 to 32ai.
[0181] When the stall signal generating unit 68 detects the occurrence of an access conflict based on the OR of the signals from the conflict detection and arbitration units 65a1 to 65a32, it generates a stall signal in the group G0. The stall signal generating unit 68 outputs the generated stall signal to the selectors 69a1 to 69ai and the stall signal generating units of the other groups.
[0182] Furthermore, when the stall signal generating unit 68 detects the occurrence of an access conflict in another group based on the OR of the stall signals from the groups G1 to G3, it generates a stall signal and outputs the generated stall signal to the selectors 69a1 to 69ai.
[0183] The selectors 69a1-69ai are provided in a one-to-one correspondence with the ΔE calculation units 33a1-33ai. The selectors 69a1-69ai acquire the local fields of the replicas to be processed next from the h calculation units 32a1-32ai, and output them to the ΔE calculation units 33a1-33ai. The selectors 69a1-69ai also stall the pipeline based on a stall signal supplied from the stall signal generation unit 68. For example, the selectors 69a1-69ai delay the supply of the local fields from the h calculation units 32a1-32ai to the ΔE calculation units 33a1-33ai by the amount of read time due to access contention based on the stall signal.
[0184] In addition, when an access conflict occurs, a delay occurs in reading the weight coefficient of a certain replica, but during the delay, the pipeline executes the ΔE calculation and flip decision stages of other replicas. Therefore, each of the groups G0 to G3 may have a buffer that holds the results of the flip decision performed in each group after the read delay caused by the access conflict occurs. Then, after the read of the weight coefficient involving the access conflict is completed, each of the groups G0 to G3 may sequentially read the results of the flip decision from the buffer and perform the h update.
[0185] In this way, the data processing device 20 of the third embodiment can save the memory capacity of the physical memory by storing only non-zero weight coefficients in the physical memory such as the memory 28b in the FPGA 28a and not storing weight coefficients with a value of 0 in the physical memory. In addition, the data processing device 20 uses groups G0 to G3 to execute four pipelines in parallel that perform partial parallel trials of multiple replicas, and allows pipeline stalls if contention occurs in access to the physical memory when reading out the weight coefficients. As a result, the data processing device 20 can improve the solution performance for relatively large-scale problems by effectively utilizing the resources of the arithmetic unit such as the FPGA 28a while adhering to the principle of sequential processing of MCMC and ensuring the convergence of the solution.
[0186] Furthermore, the data processing device 20 only needs to hold one set of all weighting coefficients for multiple replicas, which means that the data processing device 20 does not need to increase the memory capacity for holding the weighting coefficients even if the number of replicas increases.
[0187] The data processing device 20 of the second and third embodiments divides a problem into multiple domains, has multiple modules that perform partial parallel trials, and performs pipeline parallel processing of multiple replicas. Between the modules of each partial parallel trial, when one module is processing a certain replica, other modules do not perform parallel trial processing for that replica until the trial / update processing for that replica is completed, and the pipeline processing timing is shifted so that other replicas are processed during that time. This makes it possible to effectively utilize computational resources while adhering to the principle of sequential processing of the MCMC method.
[0188] In addition, in each parallel trial module, in order to alleviate the bottleneck of the local field update process, i.e., the bottleneck of the weighting coefficient reading process from memory, the data processing device 20 has either the first or second mechanism or both of the following.
[0189] First, the data processing device 20 divides the memory of the weighting coefficients so that the weights corresponding to the partial areas handled by each module are stored in separate memories (separate ports), and the weighting coefficients corresponding to the update bit information of the partial areas received from each module can be simultaneously read. This makes it possible to prevent contention for accessing the memory when reading the weighting coefficients. Note that by simply dividing the memory, the total capacity is the same as if it were not divided.
[0190] Secondly, the data processing device 20 has a mechanism for determining whether the coefficient value is 0, and does not read from memory weight coefficients whose coefficient value is 0, but reads only non-zero weight coefficients, thereby reducing the number of reads required for local field update processing. In this case, the number of cycles required to read the weight coefficients varies depending on the degree of sparseness of the weight coefficients, but the data processing device 20 stalls the pipeline when the number of cycles is longer than a specified number. This makes it possible to reduce the memory capacity used to store the weight coefficients.
[0191] In the second and third embodiments, the number of pipelines is four as an example, but the number of pipelines may be a number other than four. The number of pipeline stages may be a number other than four. Furthermore, the number of replicas may be a number other than 16. For example, for four pipelines with four stages, the number of replicas may be less than or greater than 16.
[0192] The data processing device 20 described above executes the following processes, for example. The data processing device 20 solves a problem represented by an energy function including a plurality of state variables. The data processing device 20 holds a plurality of replicas, each of which indicates a plurality of state variables, in a storage unit. The data processing device 20 executes a first pipeline and a second pipeline in parallel. The first pipeline is a process for executing a plurality of stages for a plurality of replicas, including determining a first state variable to be updated and updating the value of the first state variable to be updated, according to a change amount of a value of an energy function when each of a plurality of first state variables belonging to a first index range, which is a range of indexes corresponding to each of the plurality of state variables, is set as an update candidate. The second pipeline is a process for executing a plurality of stages for a plurality of replicas, including determining a second state variable to be updated and updating the value of the second state variable to be updated, according to a change amount of a value of an energy function when each of a plurality of second state variables belonging to a second index range that does not overlap with the first index range is set as an update candidate. The data processing device 20 processes different replicas at the same timing in each stage included in the first pipeline and the second pipeline.
[0193] This allows the data processing device 20 to effectively utilize the resources of the arithmetic unit such as the FPGA 28a while maintaining the principle of sequential processing of MCMC. The first pipeline may be referred to as the first processing. The second pipeline may be referred to as the second processing.
[0194] The processing for each replica in the data processing device 20 may be executed by the FPGA 28a, or may be executed by other computing units such as the CPU 101 or a GPU. The computing units such as the FPGA 28a and the CPU 101 are examples of processing units in the data processing device 20. Also, a storage unit that holds a plurality of replicas may be realized by the memory 28b or a register as described above, or may be realized by the RAM 102. Furthermore, the accelerator card 28 can also be said to be an example of a "data processing device".
[0195] Moreover, the data processing device 20 holds in the storage unit information on the local field used to calculate the amount of change in the value of the energy function for each state variable included in each of the multiple replicas. The local field is calculated based on a weighting coefficient indicating the weight for a pair of state variables included in the multiple state variables. For example, the data processing device 20 or the processing unit of the data processing device 20 has a first arithmetic circuit and a second arithmetic circuit. The first arithmetic circuit executes a first pipeline, i.e., a first process. The second arithmetic circuit executes a second pipeline, i.e., a second process. The first arithmetic circuit updates the first local field for each state variable belonging to the first index range of the first replica in response to an update of the value of the first state variable in the first replica, and updates the second local field for each state variable belonging to the first index range of the second replica in response to an update of the value of the second state variable in the second replica. The second arithmetic circuit updates the third local field for each state variable belonging to the second index range of the first replica in response to an update of the value of the first state variable in the first replica, and updates the fourth local field for each state variable belonging to the second index range of the second replica in response to an update of the value of the second state variable in the second replica.
[0196] In this way, the data processing device 20 can speed up the calculation by updating the first local field and the third local field in accordance with the update of the value of the first state variable, and updating the second local field and the fourth local field in accordance with the update of the value of the second state variable, in parallel. In the second and third embodiments, the group G0 is an example of a first arithmetic circuit. The group G1 is an example of a second arithmetic circuit. Alternatively, any two of the groups G0 to G3 can be said to be an example of the first arithmetic circuit and the second arithmetic circuit.
[0197] The data processing device 20 may also have a first memory, a second memory, a third memory, and a fourth memory. The first memory holds a first weighting coefficient indicating a weight for a pair of state variables belonging to a first index range and used to update the first local field. The second memory holds a second weighting coefficient indicating a weight for a pair of a state variable belonging to the first index range and a state variable belonging to a second index range and used to update the second local field. The third memory holds a third weighting coefficient indicating a weight for a pair of a state variable belonging to the second index range and a state variable belonging to the first index range and used to update the third local field. The fourth memory holds a fourth weighting coefficient indicating a weight for a pair of state variables belonging to the second index range and used to update the fourth local field.
[0198] This allows the data processing device 20 to prevent memory access contention that occurs when the weight coefficients are read out when updating the first to fourth local fields. In the second embodiment, the memory 30p1 is an example of a first memory. The memory 30p2 is an example of a second memory. For example, in the second embodiment, the memory unit 30a of the group G1 includes a total of four memories including two memories corresponding to a third memory and a fourth memory.
[0199] Alternatively, the data processing device 20 may have a first weighting coefficient storage memory unit, a first address storage memory, a second address storage memory, a second weighting coefficient storage memory unit, a third address storage memory, and a fourth address storage memory. The first weighting coefficient storage memory unit holds a non-zero weighting coefficient among weighting coefficients indicating weights for pairs of a state variable belonging to the first index range and a state variable belonging to the entire index range. The first address storage memory holds a storage destination address in the first weighting coefficient storage memory unit of a weighting coefficient to be read out in response to an update of a state variable belonging to the first index range. The second address storage memory holds a storage destination address in the first weighting coefficient storage memory unit of a weighting coefficient to be read out in response to an update of a state variable belonging to the second index range. The second weighting coefficient storage memory unit holds a non-zero weighting coefficient among weighting coefficients indicating weights for pairs of a state variable belonging to the second index range and a state variable belonging to the entire index range. The third address storage memory holds storage addresses in the second weighting coefficient storage memory unit that are storage addresses of weighting coefficients to be read in response to updates of state variables belonging to the first index range, and the fourth address storage memory holds storage addresses in the second weighting coefficient storage memory unit that are storage addresses of weighting coefficients to be read in response to updates of state variables belonging to the second index range.
[0200] In this way, by storing only non-zero weighting coefficients in memory, the data processing device 20 can reduce the memory capacity required to store weighting coefficients, compared to storing all weighting coefficients, including weighting coefficients that are 0, in memory.
[0201] In the third embodiment, the weighting coefficient storage memory unit 50 is an example of a first weighting coefficient storage memory unit. For example, in the third embodiment, a weighting coefficient storage memory unit corresponding to the second weighting coefficient storage memory unit is provided for the group G1. In the third embodiment, the address storage memory 41 is an example of a first address storage memory. Also, the address storage memory 42 is an example of a second address storage memory. For example, in the third embodiment, a total of four address storage memories including two memories corresponding to the third address storage memory and the fourth address storage memory are provided for the group G1.
[0202] In this case, the first calculation circuit acquires the first storage address of the first weighting coefficient in response to the update of the value of the first state variable in the first replica from the first address storage memory, and acquires the first weighting coefficient from the first weighting coefficient storage memory unit based on the first storage address. At the same time, the first calculation circuit acquires the second storage address of the second weighting coefficient in response to the update of the value of the second state variable in the second replica from the second address storage memory, and acquires the second weighting coefficient from the first weighting coefficient storage memory unit based on the second storage address. Then, the first calculation circuit updates the first local field by the first weighting coefficient and the second local field by the second weighting coefficient. In addition, the second calculation circuit acquires the third storage address of the third weighting coefficient in response to the update of the value of the first state variable in the first replica from the third address storage memory, and acquires the third weighting coefficient from the second weighting coefficient storage memory unit based on the third storage address. At the same time, the second arithmetic circuit acquires the fourth storage address of the fourth weighting coefficient in response to the update of the value of the second state variable in the second replica from the fourth address storage memory and acquires the fourth weighting coefficient from the second weighting coefficient storage memory unit based on the fourth storage address. Then, the second arithmetic circuit updates the third local field with the third weighting coefficient and updates the fourth local field with the fourth weighting coefficient.
[0203] In this way, the data processing device 20 can speed up calculations by performing, in parallel, the update of the first local field and the third local field accompanying the update of the value of the first state variable, and the update of the second local field and the fourth local field accompanying the update of the value of the second state variable.
[0204] Furthermore, the first weighting coefficient storage memory unit includes a plurality of first memories. Also, the second weighting coefficient storage memory unit includes a plurality of second memories. The first arithmetic circuit detects an access conflict to any one of the plurality of first memories based on the first storage destination address and the second storage destination address. Then, the first arithmetic circuit outputs a stall signal that stalls the first pipeline and the second pipeline according to the read time of the weighting coefficient. That is, the first arithmetic circuit outputs a stall signal that stalls the first process and the second process according to the read time of the weighting coefficient. Also, when the second arithmetic circuit detects an access conflict to any one of the plurality of second memories based on the third storage destination address and the fourth storage destination address, it outputs the stall signal.
[0205] This allows the data processing device 20 to properly maintain the principle of sequential processing of MCMC for each replica even when the reading of the weight coefficients is delayed due to access contention. The first and second arithmetic circuits may specify the read time of the weight coefficients according to, for example, the number of words to be read from the memory in which the access contention occurs, and may determine the stall time according to the read time. For example, when both the first and second arithmetic circuits output a stall signal, the first and second arithmetic circuits may determine the time to stall the first and second pipelines according to the longest read time due to access contention.
[0206] Furthermore, when the data processing device 20 starts processing the first replica through the first pipeline, the data processing device 20 starts processing the first replica through the second pipeline after the first pipeline has completed the processing on the first replica. That is, when the data processing device 20 starts the first processing on the first replica, the data processing device 20 starts the second processing on the first replica after the first processing on the first replica has completed.
[0207] In this way, the data processing device 20 shifts the timing of inputting each replica into the first pipeline and the second pipeline so that different replicas are processed at the same timing in the first and second pipelines, i.e., in each stage included in the first and second processes. This allows the data processing device 20 to effectively utilize the resources of the arithmetic unit such as the FPGA 28a while appropriately maintaining the principle of sequential processing of MCMC.
[0208] Furthermore, as illustrated in the second and third embodiments, the data processing device 20 may execute three or more pipelines that execute multiple stages for three or more non-overlapping index ranges, i.e., three or more processes that execute the multiple stages, in parallel. The three or more pipelines include the above-mentioned first and second pipelines. That is, the three or more processes include the first process and the second process. The data processing device 20 processes different replicas at the same timing in each stage included in the three or more pipelines, i.e., the three or more processes.
[0209] This allows the data processing device 20 to effectively utilize the resources of computing elements such as the FPGA 28a while maintaining the principle of sequential processing of MCMC. The information processing in the first embodiment may be realized by causing the processing unit 12 to execute a program. The information processing in the second embodiment may be realized by causing the CPU 21 to execute a program. The program can be recorded in a computer-readable recording medium 103.
[0210] For example, the program can be distributed by distributing the recording medium 103 on which it is recorded. The program may also be stored in another computer and distributed via a network. The computer may store (install) the program recorded in the recording medium 103 or a program received from another computer in a storage device such as the RAM 22 or HDD 23, and read and execute the program from the storage device. [Explanation of symbols]
[0211] 10 Data Processing Device 11 Storage section 12 Processing section
Claims
1. A data processing device for solving a problem represented by an energy function including a plurality of state variables, a storage unit that holds a plurality of replicas each representing the plurality of state variables; a first process for executing a plurality of stages for the plurality of replicas, including determining which first state variables are to be updated and updating values of the first state variables to be updated, in accordance with an amount of change in a value of the energy function in a case where each of a plurality of first state variables belonging to a first index range, which is a range of indexes corresponding to each of the plurality of state variables, is set as an update candidate; a second process of executing, for the plurality of replicas, the plurality of stages including determining which second state variables are to be updated and updating values of the second state variables to be updated in accordance with an amount of change in the value of the energy function in a case where each of a plurality of second state variables belonging to a second index range that does not overlap with the first index range is set as an update candidate; Run in parallel, In each stage included in the first process and the second process, different replicas are processed at the same timing. A processing section; A data processing device having:
2. the storage unit holds information on a local field used in calculating the amount of change for each state variable included in each of the plurality of replicas, the local field being calculated based on a weighting coefficient indicating a weight for a pair of state variables included in the plurality of state variables; the processing unit has a first arithmetic circuit that executes the first processing and a second arithmetic circuit that executes the second processing, the first arithmetic circuit updates a first local field for each of the state variables belonging to the first index range of the first replica in response to an update of a value of the first state variable in the first replica, and updates a second local field for each of the state variables belonging to the first index range of the second replica in response to an update of a value of the second state variable in the second replica; the second arithmetic circuit updates a third local field for each of the state variables belonging to the second index range of the first replica in response to an update of the value of the first state variable in the first replica, and updates a fourth local field for each of the state variables belonging to the second index range of the second replica in response to an update of the value of the second state variable in the second replica.
2. The data processing device according to claim 1.
3. a first memory for storing first weighting coefficients for the pair of state variables belonging to the first index range, the first weighting coefficients being used to update the first local field; a second memory configured to hold a second weighting factor for a pair of the state variable belonging to the first index range and the state variable belonging to the second index range, the second weighting factor being used to update the second local field; a third memory configured to hold a third weighting factor for a pair of the state variable belonging to the second index range and the state variable belonging to the first index range, the third weighting factor being used to update the third local field; a fourth memory for holding a fourth weighting factor for the pair of state variables belonging to the second index range, the fourth weighting factor being used to update the fourth local field; 3. The data processing apparatus according to claim 2, further comprising:
4. a first weighting coefficient storage memory unit that holds non-zero weighting coefficients among the weighting coefficients for pairs of the state variable belonging to the first index range and the state variable belonging to an entire index range; a first address storage memory that holds storage addresses of the weighting coefficients to be read out in response to updates of the state variables belonging to the first index range, the storage addresses being in the first weighting coefficient storage memory unit; a second address storage memory that holds the storage address of the weighting coefficient to be read out in response to an update of the state variable belonging to the second index range, the storage address being in the first weighting coefficient storage memory unit; a second weighting coefficient storage memory unit that holds non-zero weighting coefficients among the weighting coefficients for pairs of the state variable belonging to the second index range and the state variable belonging to the entire index range; a third address storage memory that holds the storage address of the weighting coefficient to be read in response to an update of the state variable belonging to the first index range, the storage address being in the second weighting coefficient storage memory unit; a fourth address storage memory that holds the storage address of the weighting coefficient to be read out in response to an update of the state variable belonging to the second index range, the storage address being in the second weighting coefficient storage memory unit; 3. The data processing apparatus according to claim 2, further comprising:
5. the first arithmetic circuit acquires from the first address storage memory a first storage address of a first weighting coefficient in response to an update of a value of the first state variable in the first replica, and acquires the first weighting coefficient from the first weighting coefficient storage memory unit based on the first storage address, acquires from the second address storage memory a second storage address of a second weighting coefficient in response to an update of a value of the second state variable in the second replica, and acquires the second weighting coefficient from the first weighting coefficient storage memory unit based on the second storage address, thereby updating the first local field with the first weighting coefficient and updating the second local field with the second weighting coefficient; the second arithmetic circuit acquires a third storage address of a third weighting coefficient in response to an update of a value of the first state variable in the first replica from the third address storage memory and acquires the third weighting coefficient from the second weighting coefficient storage memory unit based on the third storage address, acquires a fourth storage address of a fourth weighting coefficient in response to an update of a value of the second state variable in the second replica from the fourth address storage memory and acquires the fourth weighting coefficient from the second weighting coefficient storage memory unit based on the fourth storage address, and updates the third local field with the third weighting coefficient and the fourth local field with the fourth weighting coefficient.
5. The data processing device according to claim 4.
6. the first weighting coefficient storage memory unit includes a plurality of first memories; the second weighting coefficient storage memory unit includes a plurality of second memories; when detecting an access conflict to any one of the plurality of first memories based on the first storage destination address and the second storage destination address, the first arithmetic circuit outputs a stall signal for stalling the first process and the second process in accordance with a read time of the weight coefficient; the second arithmetic circuit outputs the stall signal when detecting an access conflict to any one of the plurality of second memories based on the third storage address and the fourth storage address.
6. The data processing device according to claim 5.
7. the processing unit starts the first process on a first replica, and then starts the second process on the first replica after the first process on the first replica is completed.
2. The data processing device according to claim 1.
8. the processing unit executes three or more processes for executing the plurality of stages on three or more index ranges that do not overlap with each other, the three or more processes including the first process and the second process in parallel, and processes different replicas at the same timing in each stage included in the three or more processes; 2. The data processing device according to claim 1.
9. A data processing method for solving a problem represented by an energy function including a plurality of state variables, comprising: a first process for executing a plurality of stages, including determining which first state variables are to be updated and updating values of the first state variables to be updated, in accordance with an amount of change in a value of the energy function when each of a plurality of first state variables belonging to a first index range, which is a range of indexes corresponding to each of the plurality of state variables, is set as an update candidate, for a plurality of replicas each indicating the plurality of state variables; a second process of executing, for the plurality of replicas, the plurality of stages including determining which second state variables are to be updated and updating values of the second state variables to be updated in accordance with an amount of change in the value of the energy function in a case where each of a plurality of second state variables belonging to a second index range that does not overlap with the first index range is set as an update candidate; Run in parallel, In each stage included in the first process and the second process, different replicas are processed at the same timing. Data processing methods.
10. A program for solving a problem represented by an energy function including a plurality of state variables, the program comprising: a first process for executing a plurality of stages, including determining which first state variables are to be updated and updating values of the first state variables to be updated, in accordance with an amount of change in a value of the energy function when each of a plurality of first state variables belonging to a first index range, which is a range of indexes corresponding to each of the plurality of state variables, is set as an update candidate, for a plurality of replicas each indicating the plurality of state variables; a second process of executing, for the plurality of replicas, the plurality of stages including determining which second state variables are to be updated and updating values of the second state variables to be updated in accordance with an amount of change in the value of the energy function in a case where each of a plurality of second state variables belonging to a second index range that does not overlap with the first index range is set as an update candidate; Run in parallel, In each stage included in the first process and the second process, different replicas are processed at the same timing. A program that executes a process.
Citation Information
Patent Citations
Information processing device, ising device and information processing device control method
JP2017219948A
Optimization device and control method thereof
JP2020086821A
Optimization device and method for controlling optimization device
JP2020087009A
Optimization device and method for controlling optimization device
JP2020173661A
Optimization device, control method of the optimization device, and control program of the optimization device
JP2021005282A