Probability calculation acceleration card, probability calculation acceleration method, device and medium

By storing the node weights of the probability graph on the probability calculation acceleration card and using full addition or full subtraction calculations, combined with the time division multiplexing mechanism, the hardware solution problem of high-connection large-scale probability graphs is solved, and high-efficiency and low-complexity hardware implementation is achieved.

CN118964284BActive Publication Date: 2025-07-18ICY TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410986573.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-07-18
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently implement hardware solutions for large-scale probability maps with high connectivity, traditional circuit wiring technology is difficult to provide sufficient interconnection resources, and optical chips are costly and immature, which limits system scale and performance.

Method used

A storage unit specially used to store the weight sum of the probability graph nodes is set up on the probability calculation acceleration card, and the initial weight sum of the nodes is received directly from the upper computer. The weight sum of the neighbor nodes is updated using full addition or full subtraction calculations, and the time division multiplexing mechanism is used to reduce the hardware connection complexity.

Benefits of technology

It realizes efficient solution to large-scale probability maps with high connection degrees, reduces the complexity of hardware connection implementation, expands the scale of Ising machines that can be accommodated by a single chip, improves the system's operating frequency and performance, and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964284B_ABST
    Figure CN118964284B_ABST
Patent Text Reader

Abstract

The present disclosure provides a probability calculation acceleration card, a probability calculation acceleration method, a device, and a medium, relating to the field of data processing. The probability calculation acceleration card includes: an interface configured to receive the initial states and initial weight sums of multiple nodes in a probability graph and the edge weights between the multiple nodes; a state update unit configured to determine whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node; a first storage control unit configured to, in response to determining to update the state of the target node, fetch the edge weights between the target node and adjacent nodes from a first storage unit; an intermediate processing unit configured to process the fetched edge weights to obtain an intermediate result; and a weight sum update unit configured to perform a full addition or full subtraction calculation on the weight sum of the adjacent nodes of the target node received from a second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a probability calculation acceleration card, a probability calculation acceleration method, a probability calculation acceleration device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Probability calculation is an emerging computing paradigm that uses probability models and stochastic processes to provide new ideas and methods for solving combinatorial optimization problems.

[0003] A calculation acceleration card is a dedicated hardware device used to accelerate certain specific types of computing tasks. Compared with a general-purpose CPU, an acceleration card usually has more arithmetic units and higher memory bandwidth, and can achieve higher parallel computing capabilities.

[0004] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered to be prior art merely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] It would be advantageous to provide a mechanism that alleviates, mitigates, or even eliminates one or more of the above problems.

[0006] According to one aspect of the present disclosure, there is provided a probability calculation acceleration card, including: an interface configured to receive, from a host computer, the initial states and initial weight sums of multiple nodes in a probability graph and the edge weights between the multiple nodes, and to send back to the host computer the final states of the multiple nodes, wherein the initial weight sum of each node among the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes; a first storage unit configured to store the edge weights between the multiple nodes; a second storage unit configured to store the weight sums of the multiple nodes; a state update unit configured to determine whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes; a first storage control unit configured to, in response to determining to update the state of the target node, retrieve the edge weights between the target node and its adjacent nodes from the first storage unit; an intermediate processing unit configured to process the edge weights retrieved by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and a weight sum update unit configured to perform a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.

[0007] According to another aspect of the present disclosure, a method for accelerating probability calculation is provided, including: receiving, via an interface, the initial state and initial weight sum of each of multiple nodes in a probability graph and the edge weights between the multiple nodes from a host computer, wherein the edge weights between the multiple nodes are stored in a first storage unit, the weight sums of the multiple nodes are stored in a second storage unit, and wherein the initial weight sum of each of the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; determining, by a state update unit, whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes; in response to determining to update the state of the target node, retrieving, by a first storage control unit, the edge weights between the target node and its adjacent nodes from the first storage unit; processing, by an intermediate processing unit, the edge weights retrieved by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and performing, by a weight sum update unit, a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then writing the new weight sum back to the second storage unit.

[0008] According to yet another aspect of the present disclosure, a device for accelerating probability calculation is provided, including: a receiving unit configured to receive, via an interface, the initial state and initial weight sum of each of multiple nodes in a probability graph and the edge weights between the multiple nodes from a host computer, wherein the edge weights between the multiple nodes are stored in a first storage unit, the weight sums of the multiple nodes are stored in a second storage unit, and wherein the initial weight sum of each of the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; a state update unit configured to determine, by the state update unit, whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes; a first storage control unit configured to retrieve, in response to determining to update the state of the target node, the edge weights between the target node and its adjacent nodes from the first storage unit; an intermediate processing unit configured to process the edge weights retrieved by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and a weight sum update unit configured to perform a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.

[0009] According to still another aspect of the present disclosure, a computer device is provided, including the above-mentioned probability calculation acceleration card.

[0010] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor is caused to execute the above method.

[0011] According to another aspect of the present disclosure, there is provided a computer program product including a computer program. When the computer program is executed by a processor, the processor is caused to execute the above method.

[0012] According to one or more embodiments of the present disclosure, by providing a storage unit specifically for storing the sum of weights of multiple nodes in a probability graph on a probability calculation acceleration card, and directly receiving the initial sum of weights of multiple nodes from a host computer, and then after updating the state of a target node, using simple addition or subtraction calculations to update the sum of weights of the neighbor nodes of the target node, it is made unnecessary to recalculate the sum of weights each time it is determined whether to update the node state, and thus it is not necessary to reserve connections for each edge in the probability graph, reducing the complexity of hardware connection implementation, and enabling efficient solution of large-scale probability graphs with high connectivity.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present disclosure are disclosed. In the drawings:

[0015] Figure 1 A block diagram of a probability calculation acceleration card according to an exemplary embodiment of the present disclosure is shown;

[0016] Figure 2 A block diagram of a probability calculation acceleration card according to an exemplary embodiment of the present disclosure is shown;

[0017] Figure 3 A flowchart of a probability calculation acceleration method according to an exemplary embodiment of the present disclosure is shown;

[0018] Figure 4 A flowchart of determining whether to update the state of a target node according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 5 A block diagram of a probability calculation acceleration device according to an exemplary embodiment of the present disclosure is shown; and

[0020] Figure 6 A block diagram of an exemplary computer device that can be applied to an exemplary embodiment is shown. Detailed implementation manners

[0021] In this disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or relative importance of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0022] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. As used herein, the term "plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". In addition, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.

[0023] Some concepts and terms involved in this disclosure will be introduced below.

[0024] Probability calculation is a mathematical and statistical method used to deal with uncertainty and random phenomena. In many applications, probability calculation is used for prediction and inference. In probability calculation, a probability graph is a structure composed of nodes and edges. Nodes represent random variables or states, and edges represent the strength or probability of interaction between two nodes. In addition, the weight sum of each node represents the sum of the interactions between the node and all its adjacent nodes (considering the states of the adjacent nodes). The weight sum can be used to measure the local energy of the node.

[0025] Simulated annealing is a probabilistic algorithm for global optimization problems, inspired by the annealing process in physics. The algorithm searches for the lowest energy state of the system, i.e., the optimal solution, by gradually reducing the "temperature". The "temperature" is the control parameter in simulated annealing, which determines the "chaos" degree of the system. At high temperatures, the system can accept worse solutions and thus jump out of local optima; at low temperatures, the system tends to stabilize near the global optimum. The temperature sequence represents the sequence or strategy for gradually reducing the control temperature, usually starting from a high temperature and gradually cooling down. The temperature sequence includes multiple temperature parameters, and each temperature parameter represents a specific temperature value or a parameter for temperature change. When performing the simulated annealing algorithm to solve the node states in the probability graph, it is necessary to sequentially take out the temperature parameters from high to low in the temperature sequence and perform one or more iterative solutions for the states of all nodes based on each taken-out temperature parameter. Finally, after all iterations of all temperature parameters are completed, the optimal solution of the probability graph is obtained.

[0026] In the related art, the hardware implementation of existing combinatorial optimization problems faces severe challenges, especially the interconnection problem in the hardware implementation of probability calculations (e.g., the Ising model). Since a fully connected topology requires a huge number of physical links, traditional circuit wiring techniques are difficult to provide sufficient interconnection resources, which limits the scale and performance of the system. Although emerging solutions such as optical chips may provide high-density interconnections, their high cost and immaturity make it difficult to meet the requirements of practical applications.

[0027] To solve the above problems, the present disclosure sets a storage unit on the probability calculation acceleration card specifically for storing the sum of weights of each of the multiple nodes in the probability graph, and directly receives the initial sum of weights of each of the multiple nodes from the host computer. Then, after updating the state of the target node, a simple full addition or subtraction calculation is used to update the sum of weights of the neighbor nodes of the target node, so that it is not necessary to recalculate the sum of weights every time it is determined whether to update the node state. Therefore, it is not necessary to reserve connections for each edge in the probability graph, reducing the complexity of the hardware connection implementation and enabling efficient solution of large-scale probability graphs with high connectivity.

[0028] The exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0029] According to one aspect of the present disclosure, a probability calculation acceleration card is proposed. Figure 1 The structural block diagram of the probability calculation acceleration card according to an exemplary embodiment of the present disclosure is shown. Referring to Figure 1 , the probability calculation acceleration card 100 includes: an interface 102 configured to receive the initial states and initial sum of weights of each of the multiple nodes in the probability graph and the edge weights between the multiple nodes from the host computer 104, and transmit back the final states of the multiple nodes to the host computer, where the initial sum of weights of each node in the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes; a first storage unit 106 configured to store the edge weights between the multiple nodes; a second storage unit 108 configured to store the sum of weights of each of the multiple nodes; a state update unit 110 configured to determine whether to update the state of the target node based on the target temperature parameter and the sum of weights of the target node among the multiple nodes; a first storage control unit 112 configured to, in response to determining to update the state of the target node, retrieve the edge weights between the target node and its adjacent nodes from the first storage unit; an intermediate processing unit 114 configured to process the edge weights retrieved by the first storage control unit to obtain an intermediate result, where the intermediate result represents the change amount of the sum of weights of the adjacent nodes of the target node before and after the state update of the target node; and a sum of weights update unit 116 configured to perform a full addition or subtraction calculation on the sum of weights of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new sum of weights back to the second storage unit.

[0030] Thus, by providing a storage unit specifically for storing the respective weights and of multiple nodes in a probability graph on a probability calculation acceleration card, and directly receiving the respective initial weights and of multiple nodes from a host computer, and then, after updating the state of a target node, using simple addition or subtraction calculations to update the weights and of the neighbor nodes of the target node, it is possible to avoid recalculating the weights and each time it is determined whether to update the node state, and thus there is no need to reserve connections for each edge in the probability graph, reducing the complexity of hardware connection implementation, and enabling efficient solution of large-scale probability graphs with high connectivity.

[0031] More specifically, when determining whether to update the state of a target node among multiple nodes, it is necessary to use a target temperature parameter and the weights and of the target node for judgment. Generally speaking, if the weights and of the target node are required, it is necessary to sum the edge weights between the target node and its adjacent nodes based on the states of the adjacent nodes of the target node. In some embodiments, the above process can be implemented by reserving dedicated lines for each edge in the probability graph (for example, hardware wiring or direct physical connections such as photons, optics, optoelectronics, superconducting qubits, other electrical devices, etc.), but this method will increase the complexity of hardware connection implementation, especially for large-scale probability graphs with high connectivity.

[0032] In the probability calculation acceleration card and probability calculation method provided by the present disclosure, a second storage unit specifically for storing the respective weights and of multiple nodes is provided in the probability calculation acceleration card, and a unit for updating the weights and of each node is provided. After the state of a target node (hereinafter referred to as node A) is updated, the weights and of the adjacent nodes (hereinafter referred to as node B) of the target node will change due to the change in the state of node A, and the change amount of the weights and is related to both the change situation of the state of node A and the edge weights between node A and node B. Therefore, after determining to update the state of node A, it is possible to instruct the first storage control unit to retrieve the edge weights between node A and node B from the first storage unit (for storing the weights of multiple edges), and then instruct the intermediate processing unit to process the edge weights to obtain an intermediate result, which represents the change amount of the above-mentioned weights and. Furthermore, it is possible to instruct the weights and update unit to perform simple addition or subtraction calculations and write the new weights and of node B back to the second storage unit. The new weights and of node B can be used when subsequently determining whether to update the state of node B, or can be further updated after updating the states of other adjacent nodes of node B.

[0033] In addition, when the target node has multiple adjacent nodes, the weight sum of each adjacent node can be updated in turn. In this way, time division multiplexing of the line can be achieved. This time division multiplexing mechanism greatly reduces the complexity of hardware connection implementation, allowing the highly connected Ising machine structure to be implemented with limited hardware connection resources.

[0034] Compared with traditional spatial multiplexing and optical interconnection solutions, the time division multiplexing solution disclosed in the present invention has significant advantages. First, by sharing the transmission path, time division multiplexing can support a logical connectivity that is much higher than the physical connectivity of the circuit without increasing the hardware area. This expands the scale of Ising machines that can be accommodated on a single chip and breaks through the bottleneck of the prior art. Secondly, since the data of different links are staggered in time transmission, the crosstalk and competition between links are greatly reduced. This increases the operating frequency of the system and reduces energy consumption. In addition, the orderly transmission of time division multiplexing is conducive to the caching and reuse of data, reduces unnecessary repeated access, and further improves performance and efficiency.

[0035] According to some embodiments, the probability computing acceleration card and probability computing acceleration method provided by the present disclosure can be used to solve the complete graph. That is, multiple nodes in the probability graph can have a fully connected relationship. According to the universal approximation theorem, a sufficiently large fully connected network can approximate any continuous function. Non-sparse fully connected neural networks can learn more complex functions and patterns, so that they can better fit high-dimensional data and have better performance in tasks such as image recognition and natural language processing. However, the inflated number of weights brought about by full connection also limits the further expansion of the network scale. In the probability graph solution, more fully connected nodes bring higher accuracy to the solution of more complex problems.

[0036] The probability calculation acceleration card provided by the present disclosure is used to perform probability calculation on the probability graph received from the host computer. More specifically, the probability calculation acceleration card receives the initial state and initial weight of each of the multiple nodes in the probability graph and the edge weights between the multiple nodes from the host computer (in some embodiments, it is also necessary to receive the temperature sequence and / or the number of loop solutions), and transmits the solution result of the probability graph (the final state of the multiple nodes) back to the host computer. It should be noted that Figure 1 The host computer 104 is not a component or submodule of the probability computing acceleration card 100.

[0037] The probability computing acceleration card proposed in the present disclosure can be implemented as an application-specific integrated circuit (ASIC) core or a field programmable gate array (FPGA). These methods provide flexible and high-performance hardware acceleration solutions to meet the needs of different application scenarios.

[0038] The following is a detailed introduction to each part of the probability calculation accelerator card.

[0039] In some embodiments, the interface 102 may adopt a PCIe (Peripheral Component Interconnect Express) interface. PCIe is a high-speed computer expansion bus standard widely used in various computer systems to connect the motherboard and a variety of peripheral devices, such as graphics cards, network cards, and storage devices. The PCIe interface supports high-bandwidth and low-latency data transmission. Its architecture is based on a point-to-point communication protocol, and each connected device has an independent signal channel. This design can effectively avoid bus conflicts and data transmission bottlenecks, making PCIe very suitable for application scenarios that require high-speed data transmission.

[0040] In addition to the PCIe interface, the probability calculation acceleration card provided by the present disclosure may also use other interfaces to communicate with the host computer, which is not limited herein.

[0041] In some embodiments, the first storage unit may adopt a main memory (System Memory). When processing large-scale graphs with high connectivity, especially probability graphs with a fully connected topology, the data volume of the edge weights between multiple nodes is large and the utilization rate is relatively low (only after determining to update the state of the target node, it is necessary to fetch the edge weights between the target node and its adjacent nodes). Therefore, it can be stored in the main memory with relatively high access latency but large capacity.

[0042] In an exemplary embodiment, the first storage unit may adopt DDR3 SDRAM. DDR3 SDRAM (Double-Data-Rate Three Synchronous Dynamic Random Access Memory) is the third-generation product of DDR SDRAM. Compared with DDR2, DDR3 has higher operating performance and lower voltage. DDR SDRAM is developed and improved on the basis of SDRAM technology. Compared with traditional SDRAM, the biggest feature of DDR SDRAM is double-edge triggering, that is, data acquisition and transmission can be performed on both the rising edge and the falling edge of the clock. With the same working clock, the read and write speed of DDR SDRAM can be twice as fast as that of traditional SDRAM.

[0043] Accordingly, the first storage control unit 112 may be a DDR3 controller. In some embodiments, the first storage control unit may also be configured to store the edge weights between multiple nodes received via the interface into the first storage unit. Thus, the DDR3 controller may store all the weight data into the first storage unit when the host computer provides the edge weights, and retrieve the corresponding edge weights when the weight sum needs to be updated. In terms of timing design, the first storage control unit may add a delay to prevent calculation from starting when the weight data has not been retrieved. In some embodiments, the DDR3 controller may be controlled by a DDR3 read / write state machine to write and read weight data to / from the DRAM.

[0044] In some embodiments, the second storage unit may be a cache or a register. The usage rate of the weight sum is higher than that of the edge weights. In addition to retrieving the old weight sum and writing the new weight sum when updating the weight sum of a node, the weight sum of the target node also needs to be retrieved when determining whether to update the state of the target node. Moreover, the space complexity of the weight sums of multiple nodes is O(N), where N is the number of nodes in the probability graph, while the space complexity of the edge weights between fully connected multiple nodes is O(N 2 ), that is to say, the data volume of the weight sums of multiple nodes is significantly less than that of the edge weights between multiple nodes. Therefore, the weight sums of multiple nodes can be stored in a cache or a register with a lower access latency but a smaller capacity.

[0045] The target node represents the node that needs to be solved currently, and the target temperature parameter represents the temperature parameter used in the solution process (e.g., simulated annealing). The target node may be determined from multiple nodes in any manner. The target temperature parameter may be pre-determined, or received from the host computer via the interface, or calculated by a corresponding module in the probability calculation acceleration card, which is not limited herein.

[0046] In some embodiments, for the target node that needs to be solved currently, the state update unit may calculate the energy difference before and after the state update of the target node based on the target temperature parameter and the weight sum of the target node, and determine whether to update the state of the target node based on the energy difference.

[0047] Considering that the calculation complexity of the energy difference is relatively high, the present disclosure provides a simplified implementation. According to some embodiments, the state update unit may be configured to: map the product of the weight sum of the target node and the target temperature parameter to a state update probability; and determine whether to update the state of the target node based on the comparison result between the state update probability and a random number.

[0048] Thus, by calculating the product of the weight sum of the target node and the target temperature parameter and mapping this product to a probability value (i.e., the state update probability), and then comparing this probability value with a random number and determining whether to update the state of the target node according to the comparison result, the calculation process is simplified, and the state update can be implemented with a simpler hardware structure. In addition, the above probability decision method can achieve efficient state update judgment.

[0049] In some embodiments, the product of the weight sum of the target node and the target temperature parameter can be processed by using tanh, sigmoid or other activation functions to map it to the state update probability. Regarding the hardware implementation of the mapping process (activation function), a dedicated hardware circuit can be used, or a piecewise linear approximation method can be used, or other methods can be used, which are not limited herein.

[0050] In some embodiments, the judgment rules and the update logic of the node state can be preset for the possible comparison results and different states of the target node. Exemplary embodiments will be given below.

[0051] The solution of the present disclosure can implement update algorithm strategies such as simulated annealing and stochastic gradient descent to optimize the overall energy / target function.

[0052] Figure 2 FIG. shows a structural block diagram of a probability calculation acceleration card according to an exemplary embodiment of the present disclosure. Refer to Figure 2 , the probability calculation acceleration card 200 includes: an interface 202 configured to receive the initial state and the initial weight sum of each of the multiple nodes in the probability graph and the edge weights between the multiple nodes from the host computer 204 and send back the final states of the multiple nodes to the host computer 204; a first storage unit 206; a second storage unit 208; a state update unit 210; a first storage control unit 212; an intermediate processing unit 214; a weight sum update unit 216; and a fourth storage unit 218 configured to store the initial state of each of the multiple nodes and receive the updated state of the target node in response to determining to update the state of the target node.

[0053] It can be understood that the operations and functions of the units 202 - 216 in the probability calculation acceleration card 200 can refer to the description of the units 102 - 116 in the probability calculation acceleration card 100 above, and will not be elaborated herein.

[0054] In some embodiments, in response to determining to update the state of the target node, the state update unit may send the updated state to the fourth storage unit.

[0055] In some embodiments, in response to determining to update the state of a target node, the first storage control unit may retrieve the edge weights between the target node and its adjacent nodes from the first storage unit. Subsequently, the intermediate processing unit may process the edge weights retrieved by the first storage control unit to obtain an intermediate result. This intermediate result represents the change in the sum of the weights of the adjacent nodes of the target node before and after the state update of the target node.

[0056] After obtaining the above intermediate result, the intermediate result may be sent to the weight sum update unit. In addition, the weight sum update unit may receive the sum of the weights of the adjacent nodes of the target node from the second storage unit. In some embodiments, the weight sum update unit may include a full adder / subtractor and may be configured to determine to perform a full addition or subtraction calculation on the sum of the weights of the adjacent nodes of the target node and the intermediate result based on the state of the target node before and after the update. After the calculation is completed, the weight sum update unit writes the new weight back to the second storage unit. Through the above method, efficient and accurate update of the sum of the weights of multiple nodes can be achieved.

[0057] In some embodiments, if the target node includes multiple adjacent nodes, the above update operation may be sequentially performed on the sum of the weights of each adjacent node. In an exemplary embodiment, target node A includes two adjacent nodes B and C. Then, the first storage control unit may first retrieve the edge weight b between node A and node B. The intermediate processing unit processes the edge weight b to obtain an intermediate result Δb, and then this intermediate result Δb is sent to the weight sum update unit. In addition, the second storage unit sends the sum of the weights of node B to the weight sum update unit. The weight sum update unit performs a full addition or subtraction calculation on the intermediate result Δb and the sum of the weights of node B to obtain the new weight of node B, and writes the new weight of node B back to the second storage unit. Then, the first storage control unit retrieves the edge weight c between node A and node C. The intermediate processing unit processes the edge weight c to obtain an intermediate result Δc, and then this intermediate result Δc is sent to the weight sum update unit. In addition, the second storage unit sends the sum of the weights of node C to the weight sum update unit. The weight sum update unit performs a full addition or subtraction calculation on the intermediate result Δc and the sum of the weights of node C to obtain the new weight of node C, and writes the new weight of node C back to the second storage unit.

[0058] In some embodiments, multiple full adder / subtractors may be provided in the weight sum update unit, so that it supports parallel execution of the above update operation on the sum of the weights of multiple adjacent nodes of the target node.

[0059] According to some embodiments, the probability calculation acceleration card 200 may further include: a third storage unit 220 configured to store a temperature sequence; and a top layer controller 222 configured to instruct the status update unit to sequentially use each temperature parameter included in the temperature sequence as a target temperature parameter to perform at least one round of solution for the statuses of multiple nodes, wherein, in each round of solution, the multiple nodes are sequentially determined as target nodes.

[0060] In some embodiments, the temperature sequence may be received from a host computer via an interface and the received temperature sequence may be stored in the third storage unit. In some other embodiments, a temperature sequence generation unit may be provided in the probability calculation acceleration card, and the temperature sequence generated by the temperature sequence generation unit may be stored in the third storage unit. The temperature sequence may include multiple entries, and each entry contains a temperature parameter (temperature value) and a corresponding number of iterations. In some embodiments, the top layer controller may automatically index the stored temperature sequence according to the current iteration step and send the corresponding temperature parameter to the status update unit.

[0061] After the solution of the probability graph is completely finished (for example, after the entire temperature sequence calculation is completed), the final statuses of the multiple nodes may be retrieved from the fourth storage unit and the final statuses of the multiple nodes may be sent back to the host computer via the interface.

[0062] The method of the present disclosure may be used to solve the Ising model. According to some embodiments, the probability graph solved by the present disclosure may be based on the Ising model. The status of a node may represent a spin variable, and its value may be +1 or -1. In some embodiments, the probability graph may also be based on a multi-layer neural network, a restricted Boltzmann machine or other models, which are not limited herein.

[0063] Compared with multi-valued or continuous-valued statuses, operations on binary statuses (such as multiplication and addition) are simpler and more efficient. When updating the weights and of a node, only simple addition and multiplication operations are required, greatly simplifying the calculation complexity and having a lower hardware implementation cost, enabling efficient solution and optimization.

[0064] According to some embodiments, the intermediate processing unit is configured to perform a shift process on the edge weight taken out by the first storage control unit so that the value of the intermediate result is twice the edge weight before the shift.

[0065] Shift processing (e.g., shifting edge weights to double their values) is a very efficient operation in hardware. Compared to multiplication, shift operations generally require fewer hardware resources and shorter processing times. By implementing efficient edge weight shift processing in hardware, the computational latency can be significantly reduced, the processing speed and efficiency can be improved, the hardware design can be simplified, resource utilization can be optimized, and power consumption can be reduced. In addition, shift processing can reduce numerical errors compared to floating-point multiplication and improve the accuracy of the calculation results. These effects make this solution have significant advantages in hardware acceleration applications that require high performance and low power consumption.

[0066] According to some embodiments, the weight and update unit is configured to: perform a full addition calculation in response to determining that the state of the target node is updated from -1 to +1; and perform a full subtraction calculation in response to determining that the state of the target node is updated from +1 to -1.

[0067] Since the initial weight sum of each node among multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node, when the state of the target node is updated from -1 to +1, the weight sum of the neighbor nodes of the target node increases by twice the edge weight between the target node and the neighbor nodes, that is, perform a full addition calculation on the intermediate result of the weight sum of the neighbor nodes (twice the edge weight). Similarly, when the state of the target node is updated from +1 to -1, the weight sum of the neighbor nodes of the target node decreases by twice the edge weight between the target node and the neighbor nodes, that is, perform a full subtraction calculation on the intermediate result of the weight sum of the neighbor nodes (twice the edge weight).

[0068] According to some embodiments, determining whether to update the state of the target node based on the comparison result of the state update probability and the random number may include: flipping the state of the target node in response to determining that the state update probability is greater than the random number. The random number can be a random number between 0 and 1. Flipping may include flipping the state of the target node from -1 to +1 and flipping the state of the target node from +1 to -1.

[0069] In some embodiments, the sources of the random number include but are not limited to: quantum random numbers, randomness of magnetic flipping of spintronic devices, thermal noise, linear feedback shift register algorithms, and so on.

[0070] In some embodiments, the state implementation methods include but are not limited to spintronic devices such as magnetic tunnel junctions, magnetic multilayers, spin oscillators, spin valves, magnons, nanomagnets, artificial spin ice, magnetic skyrmions, etc. In the magnetic device integration stage (e.g., for MTJ arrays), the above magnetic devices such as magnetic tunnel junctions and magnetic multilayers can be directly hetero-integrated (deposited) above the CMOS computing control structure, improving the integration density, reducing the required number of transistors, reducing energy consumption, increasing speed, and physically and inherently implementing random numbers or random sampling of the activation function.

[0071] In some embodiments, the write / flipping mechanism may include, but is not limited to, SOT spin-orbit torque, STT spin-transfer torque, VCMA voltage-controlled magnetization flipping, etc., and the read mechanism may include, but is not limited to, TMR tunneling magnetoresistance, GMR giant magnetoresistance, various Hall effects, etc. In addition, activation functions (such as linear / nonlinear processes like tanh, softmax, sigmoid, etc.) can be physically and intrinsically implemented using the above spintronic devices and physical mechanisms.

[0072] In some embodiments, determining whether to update the state of a target node may further include: retaining the current state of the target node in response to determining that the state update probability is not greater than a random number.

[0073] In some embodiments, the probability calculation acceleration card may further include a random sequence generator (not shown in the figure). The random sequence generator is configured to generate a random number sequence based on a set of loaded random seeds. The random number used for comparison with the state update probability is selected from the random number sequence.

[0074] According to another aspect of the present disclosure, a probability calculation acceleration method is provided. Figure 3 A flowchart of the probability calculation acceleration method according to an exemplary embodiment of the present disclosure is shown. Referring to Figure 3 , the probability calculation acceleration method 300 includes: Step S301, receiving, via an interface, the initial state and initial weight sum of each of multiple nodes in a probability graph and the edge weights between the multiple nodes from a host computer, wherein the edge weights between the multiple nodes are stored in a first storage unit, the weight sums of each of the multiple nodes are stored in a second storage unit, and wherein the initial weight sum of each node in the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; Step S302, determining, by a state update unit, whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes; Step S303, in response to determining to update the state of the target node, fetching, by a first storage control unit, the edge weights between the target node and its adjacent nodes from the first storage unit; Step S304, processing, by an intermediate processing unit, the edge weights fetched by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and Step S305, after performing an all-addition or all-subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit by a weight sum update unit, writing the new weight sum back to the second storage unit.

[0075] It can be understood that the operations of steps S301 - S305 in method 300 can respectively refer to the descriptions of interface 102, the first storage unit 106, the second storage unit 108, the state update unit 110, the first storage control unit 112, the intermediate processing unit 114, and the weight sum update unit 116 in the probability calculation acceleration card above, which will not be elaborated here.

[0076] Thus, by setting a storage unit specifically for storing the weight sums of multiple nodes in the probability graph on the probability calculation acceleration card, and directly receiving the initial weight sums of multiple nodes from the host computer, and then, after updating the state of the target node, using simple full addition or full subtraction calculations to update the weight sums of the neighbor nodes of the target node, it is possible to avoid recalculating the weight sums each time it is determined whether to update the node state, and thus there is no need to reserve a circuit for each edge in the probability graph, reducing the complexity of hardware connection implementation, and enabling efficient solution of large-scale probability graphs with high connectivity.

[0077] According to some embodiments, as Figure 4 shown, step S302, determining whether to update the state of the target node by the state update unit based on the target temperature parameter and the weight sum of the target node among multiple nodes may include: step S401, mapping the product of the weight sum of the target node and the target temperature parameter to a state update probability; and step S402, determining whether to update the state of the target node based on the comparison result between the state update probability and a random number.

[0078] According to some embodiments, the initial states of multiple nodes can be stored in the fourth storage unit. The probability calculation acceleration method 300 may further include (not shown in the figure): in response to determining to update the state of the target node, storing the updated state of the target node in the fourth storage unit.

[0079] According to some embodiments, a temperature sequence may be stored in the third storage unit. The probability calculation plus method 300 may further include (not shown in the figure): instructing the state update unit by the top controller to sequentially use each temperature parameter included in the temperature sequence as the target temperature parameter to perform at least one round of solution for the states of multiple nodes, where, in each round of solution, multiple nodes are sequentially determined as the target node.

[0080] According to some embodiments, the probability graph may be based on the Ising model, and the state of a node may represent a spin variable, taking values of +1 or -1.

[0081] According to some embodiments, step S305, after the weight sum updating unit performs an addition or subtraction operation on the weight sum of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, writing the new weight sum back to the second storage unit may include: performing an addition operation in response to determining that the state of the target node is updated from -1 to +1; and performing a subtraction operation in response to determining that the state of the target node is updated from +1 to -1.

[0082] According to some embodiments, step S402, determining whether to update the state of the target node based on the comparison result between the state update probability and the random number includes: flipping the state of the target node in response to determining that the state update probability is greater than the random number.

[0083] According to another aspect of the present disclosure, a probability calculation acceleration device is provided. As Figure 5 shown, the device 500 includes: a receiving unit 510 configured to receive, via an interface, the initial state and the initial weight sum of each of a plurality of nodes in a probability graph and the edge weights between the plurality of nodes from a host computer, wherein the edge weights between the plurality of nodes are stored in a first storage unit, the weight sums of the plurality of nodes are stored in a second storage unit, and wherein the initial weight sum of each of the plurality of nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; a state update unit 520 configured to determine whether to update the state of the target node based on the target temperature parameter and the weight sum of the target node among the plurality of nodes; a first storage control unit 530 configured to, in response to determining to update the state of the target node, retrieve the edge weights between the target node and its adjacent nodes from the first storage unit; an intermediate processing unit 540 configured to process the edge weights retrieved by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and a weight sum updating unit 550 configured to perform an addition or subtraction operation on the weight sum of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.

[0084] It can be understood that the operations of units 510 - 550 in the device 500 may refer to the descriptions of steps S301 - S305 in the probability calculation acceleration method 300 above, and will not be elaborated here.

[0085] According to another aspect of the present disclosure, a computer device is provided, including the above probability calculation acceleration card 100 or 200.

[0086] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor is caused to execute the above-described probability calculation acceleration method 300.

[0087] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, causes the processor to execute the above-described probability calculation acceleration method 300.

[0088] In the following, illustrative examples of such computer devices, computer-readable storage media, and computer program products will be described in conjunction with Figure 6 description of such computer devices, computer-readable storage media, and computer program products.

[0089] Figure 6 FIG. shows an example configuration of a computer device 600 that can be used to implement the methods described herein.

[0090] The computer device 600 can be of various different types. Examples of the computer device 600 include but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablets, cellular or other wireless phones (e.g., smartphones), notepad computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, gaming consoles), televisions or other display devices, automotive computers, and the like.

[0091] The computer device 600 can include at least one processor 602, a memory 604, (one or more) communication interfaces 606, a display device 608, other input / output (I / O) devices 610, and one or more mass storage devices 612 that can communicate with each other, such as via a system bus 614 or other suitable connections.

[0092] The processor 602 can be a single processing unit or multiple processing units, and all processing units can include a single or multiple computing units or multiple cores. The processor 602 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operation instructions. Among other capabilities, the processor 602 can be configured to obtain and execute computer-readable instructions stored in the memory 604, the mass storage device 612, or other computer-readable media, such as program code of an operating system 616, program code of an application 618, program code of other programs 620, and the like.

[0093] Memory 604 and mass storage device 612 are examples of computer-readable storage media for storing instructions that are executed by processor 602 to implement the various functions described above. For example, memory 604 generally can include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). In addition, mass storage device 612 generally can include a hard disk drive, a solid state drive, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical discs (e.g., CD, DVD), storage arrays, network attached storage, storage area network, and the like. Memory 604 and mass storage device 612 can both be collectively referred to herein as memory or computer-readable storage media, and can be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that can be executed by processor 602 as a particular machine configured to implement the operations and functions described in the examples herein.

[0094] Multiple programs can be stored on mass storage device 612. These programs include operating system 616, one or more application programs 618, other programs 620, and program data 622, and they can be loaded into memory 604 for execution.

[0095] Although illustrated as being stored in memory 604 of computer device 600 in Figure 6 , modules 616, 618, 620, and 622, or portions thereof, can be implemented using any form of computer-readable medium accessible by computer device 600. As used herein, "computer-readable medium" includes at least two types of computer-readable media, namely computer-readable storage media and communication media.

[0096] Computer-readable storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism. Computer-readable storage media as defined herein does not include communication media.

[0097] One or more communication interfaces 606 are used to exchange data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., network interface card (NIC)), wired or wireless (such as IEEE 802.6 wireless LAN (WLAN)) interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, BluetoothTM interface, Near Field Communication (NFC) interface, etc. The communication interface 606 can facilitate communication within a variety of network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. The communication interface 606 can also provide communication with external storage devices (not shown) such as in storage arrays, network-attached storage, storage area networks, etc.

[0098] In some examples, a display device 608 such as a monitor can be included for displaying information and images to a user. Other I / O devices 610 can be devices that receive various inputs from the user and provide various outputs to the user, and can include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, etc.

[0099] The techniques described herein can be supported by these various configurations of the computer device 600 and are not limited to the specific examples of the techniques described herein. For example, the functionality can also be implemented in whole or in part on a “cloud” using a distributed system. The cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud. Resources can include applications and / or data that can be used when performing computational processing on servers remote from the computer device 600. Resources can also include services provided via the Internet and / or via a subscriber network such as a cellular or Wi-Fi network. The platform can abstract the resources and functionality to connect the computer device 600 with other computer devices. Thus, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality can be implemented partially on the computer device 600 and partially via a platform that abstracts the functionality of the cloud.

[0100] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.

[0101] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A probability calculation acceleration card, comprising: An interface configured to receive, from a host computer, the initial state and initial weight sum of each of a plurality of nodes in a probability graph and the edge weights between the plurality of nodes, and to send back to the host computer the final states of the plurality of nodes, wherein the initial weight sum of each node in the plurality of nodes represents an accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; A first storage unit configured to store the edge weights between the plurality of nodes; A second storage unit configured to store the weight sums of the plurality of nodes respectively; A state update unit configured to determine whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the plurality of nodes; A first storage control unit configured to, in response to determining to update the state of the target node, fetch the edge weights between the target node and its adjacent nodes from the first storage unit; An intermediate processing unit configured to process the edge weights fetched by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and A weight sum update unit configured to perform an all-addition or all-subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.

2. The probability calculation acceleration card according to claim 1, wherein The plurality of nodes have a fully connected relationship, and the state of the node represents a spin variable, taking values of +1 or -1.

3. The probability calculation acceleration card according to claim 2, wherein, The intermediate processing unit is configured to perform a shift processing on the edge weights fetched by the first storage control unit, so that the value of the intermediate result is twice the edge weight before the shift.

4. The probability calculation acceleration card according to claim 2, wherein, The weight sum update unit is configured to: Perform an all-addition calculation in response to determining that the state of the target node is updated from -1 to +1; and Perform an all-subtraction calculation in response to determining that the state of the target node is updated from +1 to -1.

5. The probability calculation acceleration card according to claim 2, wherein, The state update unit is configured to: Map the product of the weight sum of the target node and the target temperature parameter to a state update probability; and Determine whether to update the state of the target node based on the comparison result between the state update probability and a random number.

6. The probability calculation acceleration card according to claim 5, wherein, Determining whether to update the state of the target node based on the comparison result between the state update probability and a random number includes: Flipping the state of the target node in response to determining that the state update probability is greater than the random number.

7. The probability calculation acceleration card according to any one of claims 1-6, further comprising: A third storage unit configured to store a temperature sequence; And A top-level controller configured to instruct the state update unit to sequentially use each temperature parameter included in the temperature sequence as the target temperature parameter to perform at least one round of solution on the states of the plurality of nodes, wherein in each round of solution, the plurality of nodes are sequentially determined as the target nodes.

8. The probability calculation acceleration card according to any one of claims 1-6, further comprising: A fourth storage unit, configured to store the initial states of the multiple nodes respectively, and upon determining to update the state of the target node, receive the updated state of the target node.

9. A method for accelerating probability calculation, comprising: Receiving, via an interface, the initial states, initial weight sums of the multiple nodes in a probability graph, and the edge weights between the multiple nodes from a host computer, wherein the edge weights between the multiple nodes are stored in a first storage unit, the weight sums of the multiple nodes are stored in a second storage unit, and wherein the initial weight sum of each node in the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; Determining, by a state update unit, whether to update the state of a target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes; Upon determining to update the state of the target node, fetching, by a first storage control unit, the edge weights between the target node and its adjacent nodes from the first storage unit; Processing, by an intermediate processing unit, the edge weights fetched by the first storage control unit to obtain an intermediate result, wherein the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and Performing a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit by a weight sum update unit, and then writing the new weight sum back to the second storage unit.

10. The method according to claim 9, wherein, The state of the node represents a spin variable, taking values of +1 or -1.

11. The method according to claim 10, wherein, The processing, by the intermediate processing unit, of the edge weights fetched by the first storage control unit includes: Performing a shift processing on the edge weights fetched by the first storage control unit so that the value of the intermediate result is twice the edge weight before the shift.

12. The method according to claim 10, wherein, The performing, by the weight sum update unit, of a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then writing the new weight sum back to the second storage unit includes: Performing a full addition calculation upon determining that the state of the target node is updated from -1 to +1; and Performing a full subtraction calculation upon determining that the state of the target node is updated from +1 to -1.

13. The method according to claim 10, wherein, The determining, by the state update unit, whether to update the state of the target node based on a target temperature parameter and the weight sum of the target node among the multiple nodes includes: Mapping the product of the weight sum of the target node and the target temperature parameter to a state update probability; and Determining whether to update the state of the target node based on the comparison result between the state update probability and a random number.

14. The method according to claim 13, wherein The determining whether to update the state of the target node based on the comparison result between the state update probability and a random number includes: Flipping the state of the target node upon determining that the state update probability is greater than the random number.

15. The method according to any one of claims 9-14, wherein A temperature sequence is stored in a third storage unit, and the method further includes: The top-level controller instructs the state update unit to sequentially use each temperature parameter included in the temperature sequence as a target temperature parameter to perform at least one round of solution for the states of the multiple nodes. Wherein, in each round of solution, the multiple nodes are sequentially determined as target nodes.

16. The method according to any one of claims 9-14, wherein The initial states of the multiple nodes are stored in the fourth storage unit, and the method further includes: In response to determining to update the state of the target node, storing the updated state of the target node in the fourth storage unit.

17. A probability calculation acceleration device, comprising: A receiving unit, configured to receive, via an interface, the initial states, initial weight sums, and edge weights between the multiple nodes in a probability graph from a host computer. Wherein, the edge weights between the multiple nodes are stored in a first storage unit, the weight sums of the multiple nodes are stored in a second storage unit, and wherein the initial weight sum of each node in the multiple nodes represents the accumulated result of the edge weights between the node and its adjacent nodes based on the initial states of the adjacent nodes of the node; A state update unit, configured to determine whether to update the state of the target node based on the target temperature parameter and the weight sum of the target node among the multiple nodes; A first storage control unit, configured to, in response to determining to update the state of the target node, retrieve the edge weights between the target node and its adjacent nodes from the first storage unit; An intermediate processing unit, configured to process the edge weights retrieved by the first storage control unit to obtain an intermediate result, where the intermediate result represents the change amount of the weight sums of the adjacent nodes of the target node before and after the state update of the target node; and A weight sum update unit, configured to perform a full addition or full subtraction calculation on the weight sums of the adjacent nodes of the target node received from the second storage unit and the intermediate result received from the first storage control unit, and then write the new weight sum back to the second storage unit.

18. A computer device, comprising: The probability calculation acceleration card according to any one of claims 1-8.

19. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor, causes the processor to execute the method according to any one of claims 9-16.

20. A computer program product, comprising a computer program, which when executed by a processor, causes the processor to execute the method according to any one of claims 9-16.

Citation Information

Patent Citations

  • Information Processing Apparatus and Information Processing Method

    US20200019885A1

  • Graph inference calculator

    WO2017039595A1