Inference acceleration system and method of probability model

Through the combination of subgraph decomposition, node reordering and a variety of hardware technologies, an efficient probability model inference acceleration system was designed, which solved the problem of slow computing speed in the inference process of probability circuits, and realized the ability to quickly and parallelly handle complex probability distributions.

CN120146183APending Publication Date: 2025-06-13PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200682.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When using probability circuits for probability reasoning, the prior art faces the problems of slow computing speed, high bandwidth usage and low computing utilization, especially in complex probability distribution processing.

Method used

The algorithm optimization process such as subgraph decomposition and node reordering is adopted, and combined with data compression, local storage, storage and computing integration and computing concurrency technology, an efficient probability model inference acceleration system is designed. Through modular and parameterized design, the system can easily deploy different types of probability circuits, improving computing speed and hardware utilization.

Benefits of technology

It realizes hardware acceleration of probability inference, and can quickly disassemble complex probability distributions in parallel and quickly in hardware, improves computing speed and hardware utilization, and is suitable for various types of probability inference environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146183A_ABST
    Figure CN120146183A_ABST
Patent Text Reader

Abstract

The invention provides a reasoning acceleration system and method for a probability model, and belongs to the field of probabilistic reasoning acceleration in an artificial intelligence algorithm. An algorithm part of the system comprises a sub-graph decomposition process and a node reordering process, a software part comprises a graph-matrix conversion module and a compiling module, and a hardware part comprises a probability storage module and a probability reasoning module. The method comprises preprocessing of an algorithm and a software part and operation of a hardware part, wherein the operation comprises model loading, input probability transmission, module probability distribution, operation module activation / closing, operation module probabilistic reasoning, module probability receiving and reasoning result sending; according to the method, a local storage technology, a storage and calculation integrated technology and an operation concurrency technology are used, hardware acceleration of probabilistic reasoning is achieved, and disassembly operation can be conducted on complex probability distribution in parallel and rapidly in hardware; the modular and parameterized design is adopted, the method can be applied to various types of probabilistic reasoning environments and various types of probabilistic circuit algorithm structures, and the universality is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the accelerated design of probabilistic models in artificial intelligence algorithms, and particularly to the inference acceleration of probabilistic circuits. Background Art

[0002] A probabilistic model is a special representation of a probability distribution. As Figure 1 shown, it is the core of modern machine learning (ML) and artificial intelligence (AI), providing a principled and almost universally adopted mechanism for decision-making under uncertainty. In machine learning, we assume that data information comes from an unknown complex probability distribution and simplify many machine learning tasks to simply performing probabilistic inference. Similarly, many forms of model-based artificial intelligence also attempt to directly represent the mechanisms governing the world around us as probability distributions in some form.

[0003] A probabilistic circuit (PC), as Figure 2 shown, is a general and unified computational framework for modeling operational probabilistic models, and its important elements include:

[0004] Computing Graph: A probabilistic circuit encodes a probability distribution in a recursive manner through a directed acyclic graph structure. A probabilistic circuit is a computational graph that encodes a descriptive distribution function (such as a probability mass function or a probability density function). Each initial definition or a probability distribution generated by compounding different distributions is represented as a node in the computing graph. By evaluating the input of the descriptive distribution function, a probabilistic circuit can encode the calculation of the descriptive distribution function of the target input for probabilistic inference.

[0005] Node: Each node in the computing graph of a probabilistic circuit represents a specific probability distribution. Three types of nodes are required to construct the graph: input node, product node, and sum node. An input node refers to a simple probability distribution defined initially, such as a Gaussian distribution, etc. The distributions of product nodes and sum nodes are defined by the distributions of the source nodes pointing to them. A product node indicates that the distribution of this node can be factorized into multiple distributions of the source nodes, and computationally it is reflected as a probability product; a sum node indicates that the distribution of this node is a mixture representation of the source node distributions, and computationally it is reflected as a weighted product and addition; finally, a product / sum node without any extra outgoing edges is called a root node, which is the end point of the composite distribution obtained by probabilistic inference.

[0006] Edge: Nodes in the computational graph are connected by edges. In terms of probability, the connection of edges represents the correlation of node distributions. Specifically, edges targeting addition nodes are assigned additional parameters for weighted operations in the probability mixture representation.

[0007] In probabilistic inference, a probabilistic circuit converts explicit evidence into input probabilities, then traverses and calculates each node in sequence according to the order defined by the directed acyclic graph to obtain the defined composite probability distribution, and performs probabilistic inference based on the finally obtained result. Therefore, when using a probabilistic circuit for probabilistic inference, there is a problem of slow operation speed due to the complexity of the structure. In existing systems, the operation mode of the GPU is to directly expand the directed acyclic graph into a matrix and perform operations directly in a dense form. Although this operation mode is simple, it introduces challenges of high additional bandwidth occupancy and low operation utilization. The CPU-based operation system directly expands each edge into a separate operation and performs sequential operations on the complete computational graph in series. Although this operation mode has high hardware utilization, the parallelism is too low to significantly improve the operation speed. Both systems are deployed to a general hardware platform for operation through software compilation, but there are common problems. In terms of algorithms and software, both methods lack optimization for the operation process, and the operation speed is severely lagged. At the same time, due to the separation of the storage and inference of the probability model in hardware, the additional data storage and transmission overhead is very high. Summary of the Invention

[0008] In view of the above problems existing in the prior art, the present invention proposes an efficient inference acceleration system and method for probability models. For the inference process of the probabilistic circuit, it uses various techniques such as data compression, local storage technology, in-memory computing technology, and operation concurrency technology to achieve an efficient inference process for the probability model. Since the probabilistic circuit has different directed acyclic graph definitions in different application scenarios, the present invention uses design techniques such as modularization and parameterization to achieve the effect of easily deploying different types of probabilistic circuits.

[0009] The technical solution of the present invention is as follows:

[0010] An inference acceleration system for a probability model, characterized in that it consists of an algorithm part, a software part and a hardware part; the algorithm part optimizes the probability circuit operation flow at the algorithm structure level, including the subgraph decomposition and node reordering processes; the subgraph decomposition process is used to disassemble the directed acyclic graph of the complete probability circuit operation to be deployed on multiple hardware modules for parallel operation to improve the operation speed; the software part transforms the directed acyclic graph structure into operation code readable by the hardware, including a graph-matrix conversion module and a compilation module; the hardware part is the hardware implementation at the system level for implementing the operation of the probability circuit, including a probability storage module and a probability inference module, and the probability inference module includes a multiplication node operation module, an addition node operation module, a node operation control module and a top-level control module.

[0011] Further, in the above-mentioned inference acceleration system for a probability model, in the subgraph decomposition process of the algorithm part, the rule of subgraph decomposition is as follows: each node on the operation graph is first layered according to its longest distance to the input node, and then according to the layering result, the nodes in the same layer are planned into the same subgraph, and then these subgraphs are connected in series before and after pairwise according to the operation dependence to form a large subgraph; thus, the operation node types in each layer of the large subgraph are the same, and the layers are directly alternately distributed. At the same time, each large subgraph is divided into multiple small subgraphs for hardware deployment to improve flexibility and configurability; the longest distance x from the root node of each small subgraph to the input node is saved as a position marker; in the node reordering process of the algorithm part, the nodes of the small subgraphs decomposed by the subgraph decomposition process are rearranged, and the rule of node reordering is: fix the input layer order of each small subgraph, and on this basis, reorder the nodes in the next layer according to the satisfaction order of the operation dependence; after the sorting of one layer is completed, sort the nodes in the next layer again until each node of the last small subgraph is traversed.

[0012] Further, in the above-mentioned inference acceleration system for a probability model, the graph-matrix conversion module of the software part is used to realize the conversion of the operation format of any directed acyclic graph structure to the matrix structure operation format, and the processed directed acyclic graph structure can or cannot pass through the sorting optimization process described in the algorithm part; the rule of graph-matrix conversion is: in an arbitrary directed acyclic graph structure, its internal connections are first layered according to the longest distance of each node to the input node, and the connections between nodes in any two layers are represented in the form of an adjacency matrix, and at the same time, the adjacency matrix is used in the form of row compression or column compression to reduce the operation and storage overhead; the compilation module of the software part converts the completed operation structure into assembly language or machine code supported by the hardware, and this process is the same as the normal compilation process after the structure operation format is converted by the graph-matrix conversion.

[0013] Further, in the inference acceleration system of the above probability model, in the probability inference module of the hardware part, the multiplication node operation module is used for the specific operation implementation of the multiplication node. It can be designed according to a tree / chain / discrete structure and perform operations on floating-point / integer / exponential domain data; the addition node operation module is used for the specific operation implementation of the addition node. It can also be designed according to a tree / chain / discrete structure. At the same time, the design concept of memory-computation integrated / near-memory computing is introduced, and the weights of weighted multiplication and addition are stored inside the operation module to further reduce the operation overhead. The operation data structure can also use floating-point / integer / exponential domain data; the node operation control module is used to control the specific operations of the multiplication node operation module or the addition node operation module. By receiving the data flow control signal from the top-level control module or storing the operation matrix converted from the small subgraph structure received externally locally, it controls the operation input and process of the multiplication node operation module or the addition node operation module. Each multiplication node operation module or addition node operation module has an independent node operation control module; the top-level control module distributes the control signals of the node operation control module and realizes parallel processing between multiple operation modules through address conversion. Each node operation control module has at least one separate or shared top-level control module; multiple probability inference modules are integrated in parallel under the probability storage module for operation scheduling. The probability storage module is independent of the probability inference module and is used for information interaction with the outside to send or receive the probabilities obtained by inference.

[0014] Further, in the inference acceleration system of the above probability model, the node reordering process of the algorithm part is a delay optimization step, which is selectively executed in different application scenarios: when the reuse rate of the probability model is not high, the delay optimization is skipped to enable the rapid deployment of the probability model; when the reuse rate of the probability model is high, the delay optimization is executed to provide better inference efficiency.

[0015] Further, in the inference acceleration system of the above probability model, the process of the algorithm part is implemented through software programming code; the process of the software part is jointly implemented through software programming code and the compiled codes of C#, C++, and Python; the hardware part is designed through the hardware programming code Verilog code.

[0016] The present invention also proposes a probability inference acceleration method implemented on the above inference acceleration system of the probability model, which is characterized in that it includes the preprocessing of the algorithm and software parts and the operation process of the hardware part; the operation process of the preprocessing is as follows: when using a certain probability circuit P for probability inference, first decompose it into m small subgraphs through the subgraph decomposition of the algorithm part. The number of multiplication nodes in each small subgraph is denoted as N p and the number of addition nodes is denoted as N s, the longest distance between the subgraph where the root node of the operation graph output is located and the input node is denoted as x. After the subgraph undergoes node reordering and then through the graph-matrix conversion in the software part, it becomes a matrix structure operation format, and the matrix structure operation format is then compiled by the compilation module into a language format recognizable by the hardware. Thus, the probability circuit P has the ability to be deployed on the hardware for probability inference after preprocessing;

[0017] In the operation process of the hardware part, a probability inference task of the probability circuit P requires n groups of inference processes with different input probabilities to obtain n groups of composite probability distributions, and each group of input probabilities is sequentially assigned a group number from 1 to n; the m subgraph data generated by the probability circuit P through the preprocessing of the algorithm and the software part, that is, the model of the probability circuit P, is deployed on m addition node operation modules and m multiplication node operation modules, and each multiplication / addition node operation module can process z groups of probabilities simultaneously; the top-level control module uses the operation pointers i and j to control the node operation control module to further control the operation modules below it. i represents the number of groups of probability inferences that have been performed, and j represents the number of subgraphs for which probability inference has been completed; the probability inference is completed through the following operation steps, specifically:

[0018] 1) Model loading; when this calculation step is the first step of probability inference, the node operation control module is initialized; during the initialization process, the probability storage module serves as a data transfer station, accepts the m subgraph data generated by external preprocessing, and stores each subgraph into the node operation control module of a multiplication node operation module or an addition node operation module respectively; according to the longest distance from the root node of the subgraph to the input node in the probability circuit P, each node operation control module is given a subgraph distance label x by the probability storage module, and the front and back connection order is recorded for use in subsequent calculations and probability assignments;

[0019] 2) Data transmission; the n groups of different input probabilities that need to be inferred are transmitted from the outside into the hardware part and are received and saved by the probability storage module;

[0020] 3) Module probability assignment; the probability storage module distributes the probabilities stored internally to different operation control modules according to the internal operation pointers i and j, specifically as follows:

[0021] If no operation module is activated at this time and the previous step is step 2), then the pointers i and j are initialized to 1, and the first group of input probabilities is assigned to the operation control module with a distance label of 0;

[0022] If there is already an operation module activated at this time, and i ≤ n / z, then i = i + 1; allocate the input probability of the (i - 1)-th group to the operation control module with a distance label of 0. If i < 2 at this time, skip; allocate the probability of the (i - 2)-th group from the operation module with a distance label of 0 to the operation control module with a distance label of 1 that is connected before and after it. If i < 3 at this time, skip; and so on for probability allocation until the input probability of the first group or all operation control modules have been allocated input probability;

[0023] If there is already an operation module activated at this time, and i > n / z, j ≤ m, then j = j + 1; after all the input probabilities have been allocated, allocate the probability of the (i - j)-th group from the operation module with a distance label of j - 1 to the operation control module with a distance label of j that is connected before and after it; allocate the probability of the (i - j - 1)-th group from the operation module with a distance label of j to the operation control module with a distance label of j + 1 that is connected before and after it; and so on for probability allocation until the input probability of the first group or all operation control modules have been allocated input probability;

[0024] The above three allocation methods indicate that not all input probabilities have completed probability inference. After the probability allocation is completed, the operation process jumps to step 4) to perform subsequent probability inference operations;

[0025] If there is no operation module activated at this time, and i > n / z, j > m, it means that the probability operation is completed, and the operation process jumps to step 7) to transfer the probability inference result;

[0026] 4) Activation / Deactivation of the operation module; through the above probability allocation method, the top-level control module determines whether the operation modules it manages have received the input probabilities that need to be processed; if the operation control module has received the input probability, the operation module controlled by this operation control module is activated to perform subsequent probability inference calculations; if the operation control module has not received the input probability, it means that the operation dependency of the sub-graph corresponding to the operation module controlled by this operation control module has not been achieved at this time, or the operation related to this operation module has been completed, and it is deactivated and enters the sleep state while maintaining the internal sub-graph structure data;

[0027] 5) The operation module performs probability inference; each independent sub-graph is assigned to a pair of coupled multiplication node operation modules and addition node operation modules. The operation module is internally composed of multiple multiplication operation units or multiply-add operation units, and can complete n p multiplication nodes and n s addition node operations respectively in each cycle; the probability inferences between the uncoupled operation modules are independent of each other; the coupled operation modules perform probability inferences for different node layers in the same sub-graph, which includes N p multiplication nodes and N sAdd nodes;

[0028] 6) Module probability reception: After all operation modules complete z sets of probability inferences, the probability inference results are sent to the probability storage module for the next module probability allocation; after the probability inference results of all activated operation modules are received, return to step 3);

[0029] 7) Inference completed. The probability inference results of n sets of different input probabilities are transmitted externally, and the overall circuit returns to the sleep state to wait for the next set of probability inference requests;

[0030] When the probability circuit P is used again for other probability inferences, repeat the operation steps 2)-7) of the hardware part to update the input probabilities until the probability inference is completed; when using a different probability circuit P2 for inference, after preprocessing P2 through the algorithm and software part, then through the operation steps 1)-7) of the hardware part to perform probability inferences of different input probabilities under different probability circuits.

[0031] Furthermore, in the probability inference acceleration method of the above probability model, when the operation module in step 5) performs probability inference, the coupled operation modules perform probability inferences on different node layers in the same small subgraph, specifically:

[0032] If the input probability is the input of the multiplication node layer, the multiplication node operation module directly starts the operation, performing probability operations on n multiplication nodes per cycle to complete the probability inferences of z sets of different input probabilities on N multiplication nodes; the addition node operation waits for the result of the multiplication node probability inference and directly starts the operation after the operation dependence of any node is satisfied to shorten the waiting delay. At the latest, after the operation of the first set of input probabilities in the multiplication node layer is completed, start the probability inferences of the first set of input probabilities on N addition nodes; p p s

[0033] Conversely, if the input probability is the input of the addition node layer, the addition node operation module directly starts the operation, performing probability operations on n addition nodes per cycle to complete the probability inferences of z sets of different input probabilities on N addition nodes; the multiplication node operation waits for the result of the addition node probability inference and directly starts the operation after the operation dependence of any node is satisfied to shorten the waiting delay. At the latest, after the operation of the first set of input probabilities in the addition node layer is completed, start the probability inferences of the first set of input probabilities on N multiplication nodes; for z sets of input probabilities, the operation delay is affected by the internal connection method of the small subgraph, and the theoretical upper limit of the operation delay is s s p cycles.

[0034] ​​​​​​The technical effects of the present invention are as follows:

[0035] The present invention proposes an efficient inference acceleration system and method for a probability model. For the specific task of probabilistic circuit inference, the present invention uses local storage technology, computing-in-memory technology, and operation concurrency technology, making full use of the computing and storage resources of different operation modules and the bandwidth resources of the system to achieve hardware acceleration of probabilistic inference, and capable of performing parallel and fast decomposition operations on complex probability distributions in hardware. The present invention adopts a modular and parameterized design, can be applied to various types of probabilistic inference environments, and can also be used in various types of probabilistic circuit algorithm structures, with strong versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A method for probabilistic inference of a probability model;

[0037] Figure 2 The basic definition of a probabilistic circuit;

[0038] Figure 3 Schematic diagram of the algorithm and software part of the present invention;

[0039] Figure 4 Schematic diagram of the hardware part of the system of the present invention;

[0040] Figure 5 Flowchart of the method of the present invention;

[0041] Figure 6 Schematic diagram of the "image recognition" probabilistic inference task in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] The present invention will be further clearly and completely described below with reference to the accompanying drawings and through specific embodiments.

[0043] The present invention proposes a design of an efficient inference acceleration system for a probability model, which consists of algorithm, software, and hardware parts. The algorithm part optimizes the probabilistic circuit operation flow at the algorithm structure level, mainly including sub-graph decomposition and node reordering technique processes, for accelerating the operation speed deployed on the hardware part; the software part transforms the graph structure into operation codes readable by the hardware, including a graph-matrix conversion module and a compilation module; the hardware part is the hardware implementation at the system level, including a multiplication node operation module, an addition node operation module, a node operation control module, a top-level control module, and a probability storage module, for efficiently implementing the operation of the probabilistic circuit.

[0044] The algorithm and software parts are as Figure 3As shown in the figure. The sub - graph decomposition process is used to disassemble the complete probability circuit operation directed acyclic graph for parallel operation on multiple hardware modules. The rules of sub - graph decomposition are as follows: Each node on the operation graph is first layered according to its longest distance to the input node. Then, based on the layering results, nodes in the same layer are planned into the same sub - graph. Subsequently, these sub - graphs are connected in sequence before and after pairwise according to the operation dependencies to form a large sub - graph. Under such a premise, the operation node types in each layer of the large sub - graph are the same, and the layers are alternately distributed directly. At the same time, each large sub - graph is divided into multiple small sub - graphs for hardware deployment to improve flexibility and configurability. The longest distance x from the root node of each small sub - graph to the input node is saved as a position marker. The node rearrangement process rearranges the nodes of the decomposed small sub - graphs to provide a more compact operation process. The rules of node rearrangement are as follows: Fix the input layer order of each small sub - graph. On this basis, re - sort the nodes in the next layer according to the satisfaction order of operation dependencies. After sorting a layer, re - sort the nodes in the next layer again until each node of the last small sub - graph is traversed. Node re - sorting is a delay optimization step, which can be selected whether to execute in different application scenarios: In the case where the probability model reuse rate is not high, the delay optimization is skipped for the rapid deployment of the probability model; In the case where the probability model reuse rate is high, the delay optimization is executed to provide better inference efficiency.

[0045] In the software part described above, the graph - matrix conversion module is used to implement the conversion from an arbitrary directed acyclic graph structure operation format to a matrix structure operation format. The processed directed acyclic graph structure can go through, or not go through, the sorting and optimization process described in the algorithm part. The rules of graph - matrix conversion are as follows: In an arbitrary directed acyclic graph structure, its internal connections are first layered according to the longest distance from each node to the input node. The connections between nodes in any two layers can be represented in the form of an adjacency matrix. At the same time, row compression or column compression is used for the adjacency matrix to reduce operation and storage overheads. The compilation module of the software part converts the completed operation structure into assembly language or machine code supported by the hardware, that is, converts high - level languages such as C# / C++ / Python into assembly language or machine code recognizable by the hardware. This process is the same as the normal compilation process after the operation structure format is converted by graph - matrix conversion.

[0046] The hardware part is as Figure 4As shown in the figure, it includes a probability storage module and a probability inference module. The probability inference module includes: a multiplication node operation module, an addition node operation module, a node operation control module, and a top-level control module. The multiplication node operation module is used for the specific operation implementation of the multiplication node. It can be designed according to a tree / chain / discrete structure and perform operations on floating-point / integer / exponential domain data. The addition node operation module is used for the specific operation implementation of the addition node. It can also be designed according to a tree / chain / discrete structure. At the same time, the design concept of memory-compute integrated / near-memory computing can be introduced to store the weights of weighted multiplication and addition inside the operation module to further reduce the operation overhead. The operation data structure can also use floating-point / integer / exponential domain data. The node operation control module is used to control the specific operations of the multiplication node operation module or the addition node operation module. By receiving the data stream control signal from the top-level control module, or storing the operation matrix converted from the small sub-graph structure received externally locally, it controls the operation input and process of the multiplication node operation module or the addition node operation module. Each multiplication node operation module or addition node operation module has an independent node operation control module. The top-level control module distributes the control signals of the node operation control module and realizes parallel processing between multiple operation modules through address conversion. Each node operation control module has at least one separate or shared top-level control module. Multiple probability inference modules are integrated in parallel under the probability storage module for operation scheduling. The probability storage module is independent of the probability inference module and is used for information interaction with the outside to send or receive the inferred probabilities.

[0047] In this embodiment, the modules in the hardware part are all implemented through Verilog code (a hardware programming code), which can be synthesized into specific hardware components through synthesis. The modular and parameterized design helps to adapt to different environmental tasks and different parallelisms.

[0048] Figure 5 It is a flowchart of the method of the present invention. Below, an efficient probability inference acceleration method implemented on the system of the present invention will be introduced with "small image inference" as an example.

[0049] "Small image inference" is a class of classic probability inference problems, and its tasks are as Figure 6 shown. In a 32x32x3 RGB image, each color channel of each pixel is regarded as having an independent probability distribution. When performing probability inference, the probability circuit believes that these 3072 color channels can form a complex distribution through node operations, and information such as the type of the picture is inferred from the marginal probabilities between different channels finally obtained.

[0050] When preparing to perform probabilistic inference using the probabilistic circuit P extracted from the "small image", the operation process first goes through preprocessing in the algorithm and software parts: a probabilistic circuit is first decomposed into 4 small subgraphs by the subgraph decomposition in the algorithm part. Each subgraph contains 512 multiplication nodes and 512 addition nodes. The longest distance from the subgraph where the output root node of the operation graph is located to the input node is 3. After node reordering and the graph-matrix conversion method in the software part, it is transformed into a matrix structure operation format, and then the matrix structure operation format is compiled by the compilation module into a language format recognizable by the hardware, such as: assembly language or machine code format. Through such preprocessing, the probabilistic circuit P has the ability to be deployed on the hardware for probabilistic inference. In the specific operation process of the hardware part, a probabilistic inference task requires 32 sets of inference processes with different input probabilities to obtain 32 sets of composite probability distributions. Each set of input probabilities is sequentially assigned a group number from 1 to 32; the 4 small subgraphs generated by the preprocessing of the probabilistic circuit P in the algorithm and software parts are deployed to 4 addition node operation modules and 4 multiplication node operation modules. Each multiplication node / addition node operation module can process 4 sets of probabilities simultaneously; the top-level control module uses the operation pointers i and j to control the node operation control module to further control the operation modules below it. i represents the number of groups of probabilistic inferences that have been performed, and j represents the number of subgraphs for which probabilistic inference has been completed; the probabilistic inference is completed through the following operation steps, specifically including:

[0051] 1) Model loading; when this calculation step is the first step of probabilistic inference, this step needs to initialize the node operation control module. During the initialization process, the probability storage module serves as a data transfer station, receives the 4 small subgraph data generated by external preprocessing, and stores each small subgraph into the node operation control module of a multiplication node operation module or an addition node operation module respectively. According to the longest distance from the root node of the small subgraph to the input node in the probabilistic circuit P, each node control module is given a small subgraph distance label x by the probability storage module, and the front and back connection order is recorded for use in subsequent calculations and probability assignments;

[0052] 2) Data transmission; 32 sets of different input probabilities to be inferred are transmitted from the outside into the hardware module and received and saved by the probability storage module;

[0053] 3) Module probability assignment; the probability storage module distributes the 32 sets of input probabilities stored internally to different operation control modules according to the internal operation pointers i and j. Specifically:

[0054] If no operation module is activated at this time and the previous step is step 2), then the pointers i and j are initialized to 1, and the first set of input probabilities is assigned to the operation control module with a distance label of 0;

[0055] If there is already an operation module activated at this time and i ≤ 8, then i = i + 1; allocate the (i - 1)-th group of input probabilities to the operation control module with a distance label of 0, and skip if i < 2 at this time; allocate the (i - 2)-th group of probabilities from the operation module with a distance label of 0 to the operation control modules with a distance label of 1 that are connected before and after it, and skip if i < 3 at this time; and so on for probability allocation until the 1st group of input probabilities are allocated or all operation control modules have been allocated input probabilities;

[0056] If there is already an operation module activated at this time and i > 8, j ≤ 4, then j = j + 1; after all input probabilities have been allocated, allocate the probabilities of the (i - j)-th group from the operation module with a distance label of j - 1 to the operation control modules with a distance label of j that are connected before and after it; allocate the probabilities of the (i - j - 1)-th group from the operation module with a distance label of j to the operation control modules with a distance label of j + 1 that are connected before and after it; and so on for probability allocation until the 1st group of input probabilities are allocated or all operation control modules have been allocated input probabilities;

[0057] The above three allocation methods indicate that not all input probabilities have completed probability inference. After probability allocation is completed, the operation process jumps to step 4) to perform subsequent probability inference operations.

[0058] If there is no operation module activated at this time and i > 8, j > 4, it means that probability calculation is completed, and the operation process jumps to step 7) to transfer the probability inference result;

[0059] 4) Activation / Deactivation of Operation Modules; Through the above probability allocation method, the top-level control module determines whether the operation modules it manages have received the input probabilities that need to be processed; if the operation control module receives the input probabilities, the operation module controlled by the operation control module is activated to perform subsequent probability inference calculations; if the operation control module does not receive the input probabilities, it means that the operation dependency of the sub-graph corresponding to the operation module controlled by the operation control module at this time has not been achieved, or the operations related to the operation module have been completed. In this case, it is deactivated and enters the sleep state while maintaining the internal sub-graph structure data;

[0060] 5) Probability Inference by Operation Modules; Each independent sub-graph is assigned to a pair of coupled multiplication node operation modules and addition node operation modules. The operation modules are internally composed of multiple multiplication operation units or multiply-accumulate operation units, and can complete 16 multiplication nodes and 8 addition nodes of operations respectively in each cycle. The probability inferences between uncoupled operation modules are independent of each other, and the coupled operation modules perform probability inferences for different node layers in the same sub-graph, which includes 512 multiplication nodes and 512 addition nodes;

[0061] If the input probability is the input of the multiplication node layer, the multiplication node operation module directly starts the operation, performing probability operations on 16 multiplication nodes per cycle to complete the probability inference of 4 different input probabilities on 512 multiplication nodes; the addition node operation waits for the result of the multiplication node probability inference and directly starts the operation after the operation dependencies of any node are satisfied to shorten the waiting delay. At the latest, after the multiplication node layer completes the probability operation on the first set of input probabilities, it starts the probability inference of the first set of input probabilities on 512 addition nodes.

[0062] Conversely, if the input probability is the input of the addition node layer, the addition node operation module directly starts the operation, performing probability operations on 8 addition nodes per cycle to complete the probability inference of 4 different input probabilities on 512 addition nodes; the multiplication node operation waits for the result of the addition node probability inference and directly starts the operation after the operation dependencies of any node are satisfied to shorten the waiting delay. At the latest, after the addition node layer completes the probability operation on the first set of input probabilities, it starts the probability inference of the first set of input probabilities on 512 multiplication nodes. For the 4 sets of input probabilities, the operation delay is affected by the internal connection method of the sub-graph, and the theoretical upper limit of the operation delay is 288 cycles.

[0063] 6) Module probability reception; after all operation modules complete the 4 sets of probability inferences, the probability inference results are all sent to the probability storage module for the next module probability allocation. After the probability inference results of all activated operation modules are received, return to step 3).

[0064] 7) Inference completed, the probability inference results of n different input probabilities are transmitted externally, and the overall circuit returns to the sleep state to wait for the next set of probability inference requests.

[0065] The above steps 1)-7) complete the probability inference of 32 different picture input probabilities using the probability circuit P. After this process is completed, the hardware part enters the sleep state. If it is necessary to use the probability circuit P again for other probability inferences, only need to repeat steps 2)-7) to update the input probabilities until the probability inference is completed. If it is necessary to use a different probability circuit P2 for inference, then P2 needs to go through the preprocessing process of algorithms and software, and then go through steps 1)-7) to perform probability inferences of different input probabilities under different probability circuits.

[0066] In the embodiments of the present invention, the task of probability inference for image recognition is taken as an example, but the present invention is not limited to a specific probability circuit inference task and is also applicable to the probability inference of other probability circuit structures. Although the present invention emphasizes the inference acceleration for probability circuits, it is also applicable to the forward propagation process in the extraction and training of probability circuits. For application scenarios involving probability circuits, it can be applied.

[0067] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention, enabling those skilled in the art to understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.

Claims

1. A probability model reasoning acceleration system, characterized in that: It consists of an algorithm part, a software part and a hardware part; the algorithm part optimizes the probability circuit operation flow at the algorithm structure level, including subgraph decomposition and node reordering processes; the subgraph decomposition process is used to disassemble the directed acyclic graph of the complete probability circuit operation, so as to deploy it on multiple hardware modules for parallel operation and improve the operation speed; the software part transforms the directed acyclic graph structure into hardware-readable operation code, including a graph-matrix conversion module and a compilation module; the hardware part is a system-level hardware implementation, which is used to realize the operation of the probability circuit, including a probability storage module and a probability reasoning module, and the probability reasoning module includes a multiplication node operation module, an addition node operation module, a node operation control module and a top-level control module.

2. The probability model reasoning acceleration system according to claim 1, characterized in that: The subgraph decomposition process of the algorithm part has the following rules for subgraph decomposition: each node on the operation graph is first layered according to the longest distance from the node to the input node, and then the nodes of the same layer are planned to the same subgraph according to the result of the layering, and then the subgraphs are connected in pairs according to the operation dependency to form a large subgraph; thus, the operation nodes in each layer of the large subgraph are of the same type, and the layers are directly alternately distributed. At the same time, each large subgraph is divided into multiple small subgraphs for hardware deployment to improve flexibility and configurability; the longest distance x from the root node to the input node of each small subgraph is saved as a position mark; the node rearrangement process of the algorithm part rearranges the nodes of the small subgraph decomposed by the subgraph decomposition process, and the rule for node rearrangement is: fix the input layer order of each small subgraph, and on this basis, reorder the nodes of the next layer according to the order of satisfaction of the operation dependency; after the sorting of one layer is completed, the nodes of the next layer are sorted again until each node of the last small subgraph is traversed.

3. The probability model reasoning acceleration system as claimed in claim 1, characterized in that: The graph-matrix conversion module of the software part is used to realize the conversion of any directed acyclic graph structure operation format to a matrix structure operation format; the rule of graph-matrix conversion is: in an arbitrary directed acyclic graph structure, its internal connections are first layered according to the longest distance from each node to the input node, and the connection between nodes between any two layers is represented in the form of an adjacency matrix, and the adjacency matrix is ​​compressed in row or column to reduce the operation and storage overhead; the compilation module of the software part converts the converted operation structure into an assembly language or machine code that can be supported by the hardware after the structural operation format is converted through the graph-matrix conversion.

4. The probability model reasoning acceleration system as claimed in claim 1, characterized in that: The probability reasoning module of the hardware part, wherein the multiplication node operation module is used for the specific operation implementation of the multiplication node, which can be designed according to the tree / chain / discrete structure, and operates on the floating point / integer / exponential domain data; the addition node operation module is used for the specific operation implementation of the addition node, which can also be designed according to the tree / chain / discrete structure, and at the same time introduces storage and calculation integration / near storage operation, and stores the weights of weighted multiplication and addition inside the operation module to further reduce the operation overhead, and the operation data structure can also use floating point / integer / exponential domain data; the node operation control module is used to control the specific operation of the multiplication node operation module or the addition node operation module, by receiving the data flow control signal from the top-level control module , or locally store the operation matrix converted from the small subgraph structure received from the outside, control the operation input and process of the multiplication node operation module or the addition node operation module, each multiplication node operation module or the addition node operation module has an independent node operation control module; the top-level control module distributes the control signal of the node operation control module, and realizes parallel processing between multiple operation modules through address conversion. Each node operation control module has at least one separate or shared top-level control module; multiple probability reasoning modules are integrated into the probability storage module in parallel for operation scheduling. The probability storage module is independent of the probability reasoning module and is used to interact with the outside to send or receive the probability obtained by reasoning.

5. The probability model reasoning acceleration system according to claim 1 or 2, characterized in that: The node reordering process of the algorithm part is a delay optimization step, which is selectively performed in different application scenarios: when the probability model reuse rate is not high, the delay optimization is skipped to allow the probability model to be deployed quickly; In cases where the reuse rate of probabilistic models is high, latency optimization is performed to provide better inference efficiency.

6. The probability model reasoning acceleration system according to claim 1 or 2, characterized in that: The process of the algorithm part is implemented through software programming code.

7. The probability model reasoning acceleration system according to claim 1 or 3, characterized in that: The process of the software part is implemented by the collaboration of software programming code and compiled code of C#, C++ and Python.

8. The probability model reasoning acceleration system according to claim 1 or 4, characterized in that: The hardware part is designed through hardware programming code Verilog code.

9. A method for accelerating the reasoning of a probability model, characterized in that: It includes the preprocessing of the algorithm and software parts and the operation flow of the hardware part; the operation flow of the preprocessing is: when using a certain probability circuit P for probability reasoning, it is first decomposed into m small subgraphs by the subgraph of the algorithm part, and the number of nodes in each small subgraph is multiplied by N p The number of nodes added is recorded as N s , the longest distance between the small subgraph where the root node of the operation graph output is located and the input node is recorded as x. The small subgraph is reordered after the nodes are reordered, and then converted into a matrix structure operation format through the graph-matrix conversion of the software part. The matrix structure operation format is then compiled by the compilation module into a language format that can be recognized by the hardware. Therefore, the probability circuit P has the ability to be deployed on the hardware for probabilistic reasoning after preprocessing; In the operation flow of the hardware part, a probability reasoning task of the probability circuit P needs to perform a reasoning process of n groups of different input probabilities to obtain n groups of composite probability distributions, and each group of input probabilities is assigned a group number of 1 to n in turn; the m small subgraph data generated by the probability circuit P after the preprocessing of the algorithm and the software part, that is, the model of the probability circuit P, are deployed to m adding node operation modules and m multiplying node operation modules, and each multiplying node / adding node operation module can process z groups of probabilities at the same time; the top-level control module uses the operation pointers i and j to control the node operation control module to further control the operation module below it, i represents the number of probability reasoning groups that have been performed, and j represents the number of subgraphs that have completed probability reasoning; the probability reasoning is completed through the following operation steps, specifically: 1) Model loading; when this calculation step is the first step of probabilistic reasoning, the node operation control module is initialized; during the initialization process, the probability storage module acts as a data transfer station, accepts m small subgraph data generated by external preprocessing, and stores each small subgraph in a node operation control module of a multiplication node operation module or an addition node operation module; according to the longest distance from the root node of the small subgraph to the input node in the probability circuit P, each node operation control module is assigned a small subgraph distance mark x by the probability storage module, and the front and back connection order is recorded at the same time, which is used in the subsequent calculation and probability allocation process; 2) Data transmission: n groups of different input probabilities that need to be inferred are transmitted from the outside to the hardware part, and are received and stored by the probability storage module; 3) Module probability allocation: The probability storage module allocates the internally stored probabilities to different operation control modules according to the internal operation pointers i and j, as follows: If no operation module is activated at this time, and the previous step is step 2), the pointers i, j are initialized to 1, and the first group of input probabilities are assigned to the operation control module with a distance mark of 0; If an operation module has been activated at this time, and i≤n / z, then i=i+1; the i-1th group of input probabilities is assigned to the operation control module with a distance mark of 0, and if i<2 at this time, it is skipped; the i-2th group of probabilities from the operation module with a distance mark of 0 is assigned to the operation control module with a distance mark of 1 connected to it before and after, and if i<3 at this time, it is skipped; and the probability assignment is performed in this way until the first group of input probabilities is assigned or all operation control modules are assigned input probabilities; If a computing module has been activated at this time, and i>n / z, j≤m, then j=j+1; After all the input probabilities have been allocated, the probability of the ijth group from the distance mark j-1 operation module is allocated to the operation control module with distance mark j connected to it before and after; the probability of the ij-1th group from the distance mark j operation module is allocated to the operation control module with distance mark j+1 connected to it before and after; and so on, the probability allocation is carried out until the first group of input probabilities is allocated or all operation control modules are allocated input probabilities; The above three allocation methods indicate that not all input probabilities have completed probability reasoning. After the probability allocation is completed, the operation flow jumps to step 4) for subsequent probability reasoning operations; If no operation module is activated at this time, and i>n / z, j>m, it means that the probability operation is completed, and the operation process jumps to step 7) to transmit the probability reasoning result; 4) Operation module activation / deactivation: Through the above-mentioned probability allocation method, the top-level control module determines whether the operation module under its jurisdiction has received the input probability to be processed; if the operation control module has received the input probability, the operation module controlled by the operation control module is activated to perform subsequent probability reasoning calculations; if the operation control module has not received the input probability, it means that the small subgraph operation dependency corresponding to the operation module controlled by the operation control module has not been achieved, or the operation related to the operation module has been completed, and it is closed and enters a dormant state while maintaining the internal small subgraph structure data; 5) Operation module performs probabilistic reasoning; each independent small subgraph is assigned to a pair of coupled multiplication node operation modules and addition node operation modules. The operation modules are composed of multiple multiplication operation units or multiplication and addition operation units, and each cycle can complete n p multiplication nodes and n s The probability reasoning between the uncoupled operation modules is independent of each other; the coupled operation modules perform probability reasoning on different node layers in the same small subgraph, which includes N p multiplication nodes and N s Add nodes; 6) Module probability reception; after all operation modules complete the z group probability reasoning, the probability reasoning results are sent to the probability storage module for the next module probability allocation; after the probability reasoning results of all activated operation modules are accepted, return to step 3); 7) After the reasoning is completed, the probability reasoning results of n groups of different input probabilities are transmitted to the outside, and the whole circuit returns to the sleep state to wait for the next group of probability reasoning requests; When the probability circuit P is used again for other probability reasoning, the operation steps 2)-7) of the hardware part are repeated to update the input probability until the probability reasoning is completed; when a different probability circuit P2 is used for reasoning, P2 is preprocessed by the algorithm and software part, and then passes through the operation steps 1)-7) of the hardware part to perform probability reasoning with different input probabilities under different probability circuits.

10. The method for accelerating the reasoning of a probability model as claimed in claim 9, characterized in that: In step 5), when the operation modules perform probability reasoning, the coupled operation modules perform probability reasoning on different node layers in the same small subgraph, specifically: If the input probability is the input of the multiplication node layer, the multiplication node operation module starts the operation directly, and performs n operations per cycle. p The probability operation of the multiplication nodes is performed to complete the probability of z groups of different inputs in N p The addition node operation waits for the result of the probability reasoning of the multiplication node, and starts the operation directly after the operation dependency of any node is satisfied, so as to shorten the waiting delay. At the latest, after the multiplication node layer completes the operation of the first group of input probabilities, the first group of input probabilities starts to be calculated in N s Probabilistic reasoning on the added nodes; On the contrary, if the input probability is the input of the node adding layer, the node adding operation module starts the operation directly, and performs n operations per cycle. s The probability operation of adding nodes is performed to complete the probability of z groups of different inputs in N s The multiplication node operation waits for the result of the probability reasoning of the adding node, and starts the operation directly after the operation dependency of any node is satisfied, so as to shorten the waiting delay. At the latest, after the operation of the first group of input probabilities is completed at the adding node layer, the first group of input probabilities is started in N p Probabilistic reasoning on multiplication nodes; for z groups of input probabilities, the operation delay is affected by the connection mode within the small subgraph, and the theoretical upper limit of the operation delay is cycle.