Data flow optimization method, system, medium and equipment for adapting unified probabilistic graph computation and hardware circuit thereof
By employing a data flow optimization method using iterative multiplication units and Gaussian approximation in probabilistic graph computation, the problems of excessive complexity and power consumption in probabilistic graph computation are solved, enabling low-power hardware circuit computation.
Patent Information
- Application Number
- CN202511675819.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-11-17
AI Technical Summary
In existing technologies, the computation of probability graphs and their hardware circuits suffers from excessive complexity and power consumption, especially when performing continuous multiplication operations directly in a processor.
By acquiring the probabilistic graphical model of the baseband signal processing task, message passing between variable nodes and verification nodes is implemented based on hardware circuits. A data flow optimization method using iterative multiplication units is adopted, including reuse of intermediate computation results and Gaussian approximation, to reduce computational complexity and power consumption.
It effectively reduces computational complexity, supports low-power hardware implementation of probabilistic graph computation, and reduces the energy consumption of hardware circuits.
Smart Images

Figure CN121144253B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of signal processing, and particularly relates to a data flow optimization method, system, medium and equipment for adaptive uniform probability graph calculation and hardware circuit thereof. BACKGROUND
[0002] In the prior art, when the probability graph and the hardware circuit thereof are calculated, the messages transmitted from the check node to the variable node and the joint probability distribution transmitted from the variable node to the check node in the probability graph calculation involve a large number of continuous multiplication operations, and if the operations are directly calculated in a processor, the complexity is too high, which makes the direct implementation of the probability graph and the hardware circuit thereof result in high power consumption. SUMMARY
[0003] In view of the above problems, the present application aims to provide a data flow optimization method, system, medium and equipment for adaptive uniform probability graph calculation and hardware circuit thereof, which can reduce the calculation complexity and support low-power hardware implementation of the probability graph calculation.
[0004] To achieve the above object, in a first aspect, the technical scheme adopted by the present application is as follows: a data flow optimization method for adaptive uniform probability graph calculation and hardware circuit thereof, comprising: obtaining a probability graph model of a baseband signal processing task, and realizing the message transmission between the variable nodes and the check nodes in the probability graph model based on the hardware circuit; wherein the probability graph model comprises variable nodes, check nodes and undirected edges, and the undirected edges are used to connect the variable nodes and the check nodes having a connection relationship; iteratively updating the message transmission to realize data flow rescheduling, and determining an iterative multiplication unit in the transmitted message, and performing data flow optimization on the iterative multiplication unit to reduce the power consumption of the hardware circuit.
[0005] Further, the iterative multiplication unit in the transmitted message is determined, and is as follows:
[0006] ;
[0007] wherein, represents the variable state contained in each state; is the check node i transmitted to the variable node j ; x j the message whose state value is ; is the iterative multiplication unit in the message transmitted from the check node i to the variable node j ; is the variable node j transmitted to the check node iThe iterative multiplication unit in the transmitted message, specifically the variable node. j To the verification node i The message is from the verification node. i Other than variable nodes j Connected verification nodes Provide message generation, For verification nodes To variable node j Passing information about variables The state value is The message; For variable nodes The information transmitted to the verification node regarding The message; for The probability configuration function; For the first i The verification node and the first j The candidate vector constructed from the undirected edges between the n variable nodes is the th... m The first line k One element, This represents the element value corresponding to variable node j in the candidate vector; Represents the variable nodes in the candidate vector The corresponding element value, express From the verification node i A set of indexes of variable nodes with connections Take the value from; Indicates and verifies nodes i Connected and variable nodes j Values The combination of all variable node states.
[0008] Furthermore, the iterative multiplication unit is optimized for data flow, including:
[0009] Select a row from the candidate vectors that has the highest similarity to elements in other rows. j -1 element Construct the first backbone unit, using the first backbone unit as the first row element of each subarray as the initial driving vector. The driving vector is used to represent the message being transmitted, and the driving vector is divided into... j There are M subarrays, each subarray being a column vector containing M elements, where M is the total number of candidate vectors. Among them, the element with the highest similarity is the element in the remaining rows that is the selected element. j -1 element The difference position does not exceed 2 difference positions;
[0010] Determine the difference bits between the elements of the remaining rows in each subarray and the corresponding backbone unit, and perform vertical propagation calculations based on the difference bits to calculate the elements of the remaining rows in each subarray, thereby obtaining the complete first column driving vector and outputting it.
[0011] During iterative updates, the calculation of the (j-1)th backbone unit in the (j-1)th column of the driving vector is determined based on the (j-2)th backbone unit in the adjacent preceding column of the driving vector. The calculation result of the (j-1)th backbone unit is obtained by horizontal propagation of the difference position of the backbone units in the two adjacent columns of the driving vector. Then, the (j-1)th backbone unit is vertically propagated to obtain the (j-1)th column of the driving vector and output.
[0012] Furthermore, the difference bits between the elements of each remaining row in each subarray and the corresponding backbone unit are determined, and vertical propagation is performed based on the difference bits, including:
[0013] When the difference bit is 1: compare the backbone unit with the elements in the second row, and take the extra element in the backbone unit as the first difference bit; compare the elements in the second row with the backbone unit, and take the extra element in the second row as the second difference bit;
[0014] When the difference bit is 2, the product of the two extra elements in the backbone unit is used as the first difference bit when comparing the backbone unit with the second row element; the product of the two extra elements in the second row element is used as the second difference bit when comparing the second row element with the backbone unit.
[0015] When calculating the second row of elements, only the ratio of the second difference position to the first difference position is calculated. The result of the multiplication of the backbone units is then multiplied by the ratio to obtain the second row of elements. This process is repeated to calculate the remaining alternating row elements based on the backbone units, thus completing the vertical propagation.
[0016] Furthermore, the calculation of the (j-1)th backbone unit in the (j-1)th column of the driving vector is determined based on the (j-2)th backbone unit in the adjacent preceding column of the driving vector. The (j-1)th backbone unit is obtained by lateral propagation of the difference position of the backbone units in the two adjacent columns of the driving vector, including:
[0017] In the first subarray, when the difference bit is 1: compared with the j-1 backbone unit, the missing element in the j-2 backbone unit is taken as the third difference bit; compared with the j-2 backbone unit, the missing element in the j-1 backbone unit is taken as the fourth difference bit.
[0018] When the difference bit is 2, the product of the two missing elements in the (j-2)th backbone unit is used as the third difference bit when comparing the (j-2)th backbone unit with the (j-1)th backbone unit; the product of the two missing elements in the (j-1)th backbone unit is used as the fourth difference bit when comparing the (j-1)th backbone unit with the (j-2)th backbone unit.
[0019] When calculating the (j-1)th backbone unit, only the ratio of the third difference position to the fourth difference position is calculated. The result of the multiplication of the (j-2)th backbone unit is then multiplied by the ratio to obtain the calculation result of the (j-1)th backbone unit.
[0020] By analogy, the calculation results of the (j-1)th backbone unit in each subarray are obtained, and the lateral propagation calculation is completed.
[0021] Furthermore, the iterative multiplication unit is optimized for data flow, including: adopting a probabilistic graphical message passing method with Gaussian approximation for the iterative multiplication unit, specifically: approximating the iterative multiplication unit as a Gaussian distribution, and replacing the multiplication operation with a Gaussian distribution.
[0022] Furthermore, the iterative multiplication unit is approximated as a Gaussian distribution:
[0023] Determine the elements in the association mapping matrix ,make:
[0024] ,
[0025] Therefore, the iterative multiplication unit can be approximated as:
[0026]
[0027] In the formula, Represents the elements in the association mapping matrix. Represents a candidate vector; Indicates the noise variance. express variance This represents the transpose of a vector. Represents the association mapping matrix of the first i row element, Indicates and verifies nodes i Connected divisor nodes j Other variable node indexes.
[0028] Secondly, the technical solution adopted by the present invention is as follows: a data flow optimization system adapted to unified probabilistic graph computation and its hardware circuit, comprising: a message passing module, which acquires the probabilistic graph model of the baseband signal processing task and realizes the message passing between variable nodes and verification nodes in the probabilistic graph model based on the hardware circuit; wherein, the probabilistic graph model includes variable nodes, verification nodes and undirected edges, and the undirected edges are used to connect variable nodes and verification nodes that have a connection relationship; and an iterative update optimization module, which iteratively updates the message passing to realize data flow rescheduling, and determines the iterative multiplication unit in the transmitted message, and optimizes the data flow of the iterative multiplication unit to reduce the power consumption of the hardware circuit.
[0029] Thirdly, the technical solution adopted by the present invention is: a computer-readable storage medium for storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0030] Fourthly, the technical solution adopted by the present invention is: a computing device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.
[0031] The present invention has the following advantages due to the adoption of the above technical solutions:
[0032] This invention reduces computational complexity by reusing intermediate results and using Gaussian approximation, thereby reducing the computational complexity of multiplication data stream processing from... Reduced to It supports low-power hardware implementation of probabilistic graph computation. Attached Figure Description
[0033] Figure 1 This is a flowchart of the data flow optimization method for adapting unified probability graph calculation and its hardware circuit in an embodiment of the present invention;
[0034] Figure 2 This is a schematic diagram of the connection relationship between the verification node and each variable node in the probabilistic graphical model of this invention.
[0035] Figure 3a This is a schematic diagram illustrating the selection of backbone units during parallel pipeline computation in the iterative multiplication unit optimization of this invention embodiment;
[0036] Figure 3b This is a schematic diagram of the first stage of vertical propagation calculation and horizontal propagation calculation during parallel pipeline calculation in the iterative multiplication unit optimization of this invention embodiment;
[0037] Figure 3c This is a schematic diagram of the second-stage vertical propagation calculation and horizontal propagation calculation during parallel pipeline calculation in the iterative multiplication unit optimization of this invention embodiment;
[0038] Figure 3d This is a schematic diagram of the final output of parallel pipelined computation in the iterative multiplication unit optimization in this embodiment of the invention. Detailed Implementation
[0039] To address the high complexity caused by numerous multiplication terms in probabilistic graph computation circuits, this invention provides a data flow optimization method and system adapted to unified probabilistic graph computation and its hardware circuits. By optimizing the message transmission from the verification node to the variable node and the multiplication part of the message transmission from the variable node to the verification node in the probabilistic graph computation, the computational complexity is reduced, and the high power consumption problem caused by direct implementation in the hardware circuit is avoided.
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0041] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0042] Example 1, in an embodiment of the present invention, as follows: Figure 1 As shown, a data flow optimization method adapted to unified probabilistic graph computation and its hardware circuitry is provided, which is implemented using ultra-low power signal processing technology. In this embodiment, the method includes the following steps:
[0043] 1) Obtain the probabilistic graphical model of the baseband signal processing task, and implement message passing between variable nodes and check nodes in the probabilistic graphical model based on hardware circuits; wherein, the probabilistic graphical model includes variable nodes, check nodes and undirected edges, and undirected edges are used to connect variable nodes and check nodes that have a connection relationship.
[0044] 2) Iteratively update message passing to achieve data flow rescheduling, and determine the iterative multiplication unit in the passed message. Optimize the data flow of the iterative multiplication unit to reduce the power consumption of the hardware circuit.
[0045] In step 1) above, the probabilistic graphical model is: based on the common characteristics of baseband signal processing tasks, a unified signal model is established, and the unified signal model is represented as a probabilistic graphical model. Baseband signal processing tasks include applications such as channel decoding, channel detection, channel estimation, multi-antenna detection, and image recognition; this is used as an example, but is not limited to these. This embodiment uses a multi-antenna detection task as an example for explanation.
[0046] Based on the common characteristics of baseband signal processing tasks, a unified signal model is established as follows: ,in, For unknown variables, For given and The correlation mapping matrix between , For observed variables, This is unknown disturbance noise. Each is a certain unknown variable. Each of these represents a specific observed variable.
[0047] When facing different tasks in baseband signal processing, the parameters of the unified signal model have different meanings. For example, in a multi-antenna detection task: H This is the characteristic matrix of a multi-antenna channel. x For the transmitted symbol to be estimated, y For the mixed signal observed by the receiver from multiple antennas, H continuous, y continuous, x For unknown discrete variables.
[0048] In this embodiment, the unified signal model is represented as a probabilistic graphical model, specifically including the following implementation process:
[0049] The probabilistic graphical model consists of multiple variable nodes, multiple check nodes, and multiple undirected edges. Undirected edges connect variable nodes and check nodes that are connected. Each variable node corresponds to an association mapping matrix. H The column vectors in the matrix, and the corresponding association mapping matrix for each verification node. H Row vectors in a matrix. For example, the associative mapping matrix. H The Middle i Row vector representation , for H No. i Line number j Column elements, representing validation nodes. i With variable nodes j The relationship between them, if the association mapping matrix H middle If the element is 1, then the variable nodej With verification node i If there is a connection relationship, such as an association mapping matrix H middle If the element is 0, then the variable node... j With verification node i No connection exists, for example: association mapping matrix of different tasks. H The correlation mapping matrix has different meanings in the LDPC decoding task. H For the codeword check matrix, the sending end sends the original bit information K Each codeword is encoded into N codes. x The receiver receives N symbols. y The decoding task received by the decoding end is to decode... y Restore to x Therefore, based on x Get the initial K Each code character.
[0050] Specifically, this embodiment characterizes the unified signal model based on unified probabilistic graphical theory as follows:
[0051] Variable nodes represent unknown variables. , of which each Belongs to the state set space Its cardinality (number of elements) is The signal state of a certain task is the state of the variable to be estimated. The set of values that can be denoted as . Specifically, in this embodiment, it is assumed that it has . K Each state is denoted as: , For example, in decoding scenarios, K Each state represents the value of the transmitted bit in LDPC decoding. K =2, u 1=1, u 2=0.
[0052] Verification nodes represent the coupling relationship between unknown variables, observed variables, and correlation mapping matrices.
[0053] Association patterns: Constructing probability configuration functions This characterizes the differences in probabilistic iterative messages from the verification node to the variable node across different tasks in baseband signal processing. Based on different tasks, a probabilistic configuration function is constructed. They are all different; the following describes the probability allocation function. This is an example, but not limited to:
[0054] Multi-antenna detection task Represented as:
[0055] ;
[0056] in, Indicates the noise variance. y i Indicates the received symbol. Representation matrix H The first in Row vectors.
[0057] The association patterns include:
[0058] (1) Message passing from the verification node to the variable node:
[0059] ;
[0060] in, Representing variables x j The included states; For verification nodes i To variable node j Passed variables x j The state value is The news, This represents a probability configuration function. Indicates the first The candidate vector constructed from the nth undirected edge Each element value Represents the variable nodes in the candidate vector The corresponding element value, Represents the variable nodes in the candidate vector The corresponding element value, Indicates and verifies nodes i Connected ( China satisfies of j The corresponding variable node is related to the verification node. i (with connections) and variable nodes j Values The combination of all variable node states, express From the verification node i A set of indexes of variable nodes with connections Take the value from; Represents variable nodes To the verification node The message conveyed about The news.
[0061] (2) Message passing from variable node to check node:
[0062] ,
[0063] in, Represents variable nodes j To the verification node Passing information about variable nodes between nodes The state value is The news.
[0064] Each signal processing problem in the baseband signal processing task is transformed into a problem of estimating unknown variables based on the observed data vector and correlation rules. The message passing process from each verification node to the variable node is calculated on the hardware circuit.
[0065] In this embodiment, the core of hardware-driven probabilistic graph computation lies in the iterative message update and belief propagation between variable nodes and verification nodes, which is executed through two coupled stages in each iteration:
[0066] (1) Posterior propagation (backward propagation): Each verification node uses prior messages from neighboring variable nodes. Compute posterior messages The result is then broadcast to the connected variable nodes.
[0067] (2) Prior propagation (forward propagation): each variable node Use the post-validation message from the neighboring validator node Update prior confidence messages And the update will be propagated to the verification nodes of all links.
[0068] In this embodiment, for any verification node in the probabilistic graphical model, each variable node connected to the verification node via undirected edges can be determined, as can the undirected edges between the verification node and each variable node. This is based on each undirected edge and the association mapping matrix. H Construct the node state matrix corresponding to each undirected edge, and the total number of node state matrices corresponds to the total number of undirected edges.
[0069] Furthermore, based on the probabilistic graphical model, the first... i The message passing between all variable nodes connected by undirected edges to a verifier node is represented as a unified signal processor architecture for vector-matrix multiplication operations. That is, the K messages from the verifier node to the variable node are represented as vector-matrix multiplication (VMM) operations.
[0070]
[0071] The unified signal processor architecture includes a node state matrix. driving vector and state value probability vector ,in, The specific implementation of the above formula is as follows:
[0072]
[0073] Based on the unified signal processor architecture described above, the operations involved in the probabilistic graph calculation are implemented using continuous physical operators (voltage and current), specifically:
[0074] The node state matrix is pre-stored into the hardware circuit;
[0075] The input to the hardware circuit is a voltage value vector that is proportional to the driving vector, for example: ;
[0076] The output of the hardware circuit is the current value. By naturally performing a weighted summation of probability products using Kirchhoff's laws, the state value probability vector is obtained. .
[0077] Specifically, the unified signal processor architecture based on hardware circuits will be described in this embodiment using only memristors as an example, and is not limited to this. A memristor is the fourth basic circuit element after resistors, capacitors, and inductors. Its resistance is determined by the excitation and changes continuously. It features high integration density, fast operation speed, low power consumption, and non-volatility. A memristor cross array can complete vector and matrix multiplication and accumulation operations within one cycle. The multiplication factors are directly stored in the memristor array, eliminating the need for separate storage units.
[0078] The current output by the memristor is iteratively updated to obtain the message transmission from each variable node to each verification node, which is then used as the input to the memristor cross array, thus completing the iterative process until the iteration ends.
[0079] In step 2) above, the iterative multiplication unit is determined in the transmitted message as follows:
[0080] ;
[0081] in, Representing variables x j The included states; For verification nodes i To variable node j Passed variables x j The state value is The message; All are iterative multiplication units. For verification nodes i To variable nodej Iterative multiplication units in the transmitted message; For variable nodes j To the verification node i The iterative multiplication unit in the transmitted message, specifically the variable node. j To the verification node i The message is from the verification node. i Other than variable nodes j Connected verification nodes Provide message generation, For verification nodes To variable node j Passing information about variables x j The state value is The message; For variable nodes To the verification node The message conveyed about The message; for The probability configuration function; For the first i The verification node and the first j The candidate vector constructed from the undirected edges between the n variable nodes is the th... m The first line k One element, This represents the element value corresponding to variable node j in the candidate vector; Represents the variable nodes in the candidate vector The corresponding element value, express From the verification node i A set of indexes of variable nodes with connections Take the value from; Indicates and verifies nodes i Connected and variable nodes j Values The combination of all variable node states.
[0082] Furthermore, to more clearly illustrate the content of this invention, the following explanation uses a decoding task to illustrate the content of the candidate vector. For example... Figure 2 As shown, in this embodiment, for any verification node The verification node can be determined in the probabilistic graphical model. Connecting each target variable node and the target undirected edge. When When the value is 3, the target variable nodes connected by the 3rd verification node are the 2nd, 3rd, and 4th variable nodes, respectively. At this time, the number of undirected edges in the target node is... The value is 3. Each target undirected edge is an undirected edge connecting the 3rd verification node to the 2nd, 3rd, and 4th variable nodes, respectively. The internal order of each target undirected edge among all target undirected edges is 1, 2, and 3.
[0083] The target character includes 0 and 1. At this point... The result is 3. Multiple candidate vectors are constructed by permuting and combining the three target characters. The number of candidate vectors is 2^3, or 8: (0,0,0), (0,0,1), (0,1,0), (0,1,1), (1,0,0), (1,0,1), (1,1,0), and (1,1,1). It's understood that the number of elements in each candidate vector is the same as the number of undirected edges in the target vector. The first, second, and third target characters in each candidate vector are then associated with the target variable nodes connected by the undirected edges in order 1, 2, and 3, respectively. In other words, the first, second, and third target characters in each candidate vector are associated with the second, third, and fourth variable nodes, respectively.
[0084] The internal order of the target undirected edge. The process iterates through the first, second, and third target undirected edges in sequence. Each time a target undirected edge is encountered, the eight candidate vectors are divided into first candidate vectors and second candidate vectors. For example, when iterating through the third target undirected edge... If the third element of a candidate vector is 1, it is determined as the first candidate vector. If the third element of a candidate vector is 0, it is determined as the second candidate vector. The first candidate vectors include (0, 0, 1), (0, 1, 1), (1, 0, 1), and (1, 1, 1), while the second candidate vectors include (0, 0, 0), (0, 1, 0), (1, 0, 0), and (1, 1, 0).
[0085] Each of the first and second candidate vectors mentioned above is taken as a candidate vector to be processed. For example, when the first candidate vector (0, 1, 1) is taken as a candidate vector to be processed, a zero vector (0, 0, 0, 0, 0, 0) is first constructed with a total number of elements equal to the total number of variable nodes (6). The first, second, and third target characters in the first candidate vector are associated with the second, third, and fourth variable nodes, respectively. Therefore, in this embodiment, the second, third, and fourth zero elements in the zero vector can be replaced with the first, second, and third target characters in the first candidate vector, respectively, to obtain the corresponding processed first candidate vector (0, 0, 1, 1, 0, 0).
[0086] At this point, when this embodiment traverses to the third target undirected edge, it can obtain a candidate vector set consisting of four processed first candidate vectors and four processed second candidate vectors based on the target undirected edge.
[0087] In step 2) above, the data flow of the iterative multiplication unit is optimized by using an optimization method that reuses intermediate computation results, such as... Figures 3a to 3d As shown, it includes the following steps:
[0088] 2.1) Select a row from the candidate vectors that has the highest similarity to the elements in the other rows. j -1 element Construct the first backbone unit, using the first backbone unit as the first row element of each subarray as the initial driving vector. The driving vector is used to represent the message being transmitted, and the driving vector is divided into... j There are M subarrays, each subarray being a column vector containing M elements, where M is the total number of candidate vectors. The element with the highest similarity is the element in the remaining rows that is selected. j -1 element The difference position does not exceed 2 difference positions;
[0089] For example, such as Figure 3a As shown, the selected j -1 element is Then the first backbone unit is .
[0090] 2.2) Determine the difference bits between the elements of each remaining row in each subarray and the corresponding backbone unit, and perform vertical propagation calculations based on the difference bits (e.g., Figure 3b , Figure 3c As shown in the figure, the elements of the remaining rows in each subarray are calculated, and then the complete first column driving vector is obtained and output.
[0091] 2.3) As Figure 3b , Figure 3c As shown, during iterative updates, the th j -1 column of driving vectors j The calculation of the -1 backbone unit is based on the first element in the driving vector of the adjacent preceding column. j -2 backbone units are determined by calculating the difference position of the backbone units in two adjacent driving vectors through lateral propagation. j The calculation results of the -1 backbone unit are then used by the first... j -1 backbone units undergo vertical propagation to obtain the first... j -1 column driving vectors and output, such as Figure 3d As shown.
[0092] In step 2.2) above, the difference bits between the elements of each remaining row in each subarray and the corresponding backbone unit are determined, and vertical propagation is performed based on the difference bits, including the following steps:
[0093] 2.2.1) When the difference bit is 1: compare the backbone unit with the second row of elements, and take the extra element in the backbone unit as the first difference bit; compare the second row of elements with the backbone unit, and take the extra element in the second row of elements as the second difference bit;
[0094] For example, j When = 5, the element selected from the candidate set If the result is 0, 0, 0, then the backbone unit... for second row element for Then the last element in the backbone unit is different from the last element in the second row (that is, the difference position). In other words, compared to the second row, the backbone unit has an extra... This is the second difference position.
[0095] 2.2.2) When the difference bit is 2, compare the backbone unit with the second row of elements, and the product of the two extra elements in the backbone unit is taken as the first difference bit; compare the second row of elements with the backbone unit, and the product of the two extra elements in the second row of elements is taken as the second difference bit.
[0096] 2.2.3) When calculating the second row of elements, only the ratio of the second difference position to the first difference position is calculated. The result of the multiplication of the backbone units is then multiplied by the ratio to obtain the second row of elements. Similarly, the remaining alternate row elements are calculated according to the backbone units to complete the vertical propagation.
[0097] In step 2.3 above, the first j -1 column of driving vectors j The calculation of the -1 backbone unit is based on the first element in the driving vector of the adjacent preceding column. j -2 backbone units are determined by calculating the difference position of the backbone units in two adjacent driving vectors through lateral propagation. j -1 backbone unit, including the following steps:
[0098] 2.3.1) In the first subarray, when the difference bit is 1: [The following text appears to be incomplete and requires further context:] j -2 backbone unit and the first j Compared to the -1 backbone unit, the first j -2 The missing element in the backbone unit is taken as the third difference position; the first j -1 Backbone Unit and the First j Compared to the -2 backbone unit, the first j -1 The missing element in the backbone unit is used as the fourth difference position;
[0099] 2.3.2) When the difference bit is 2, the first bit... j -2 backbone unit and the first j Compared to the -1 backbone unit, the first j -2 The product of the two missing elements in the backbone unit is taken as the third difference bit; the first j -1 backbone unit and j Compared to the -2 backbone unit, the first j -1 The product of the two missing elements in the backbone unit is used as the fourth difference position;
[0100] 2.3.3) Perform the first j When calculating the -1 backbone unit, only the ratio of the third difference position to the fourth difference position is calculated. The result of the product of the j-2th backbone units is then multiplied by the ratio to obtain the j-1th backbone unit. j -1 Calculation results of the backbone unit;
[0101] By analogy, the first subarray in each subarray is obtained. j -1 Calculation results of the backbone unit, completing the lateral propagation calculation.
[0102] In step 2) above, the data flow optimization of the iterative multiplication unit can also be achieved by using the Gaussian approximation optimization method: the iterative multiplication unit is approximated by the probabilistic graphical message passing method of Gaussian approximation, specifically: the iterative multiplication unit is approximated as a Gaussian distribution, and the multiplication operation is replaced by the Gaussian distribution.
[0103] In this embodiment, the iterative multiplication unit is approximated as a Gaussian distribution:
[0104] Determine the elements in the association mapping matrix ,make:
[0105] ,
[0106] Therefore, the iterative multiplication unit can be approximated as:
[0107]
[0108] In the formula, Represents the elements in the association mapping matrix. Represents a combination of candidate vectors; Indicates the noise variance. express variance This represents the transpose of a vector. Represents the association mapping matrix of the first i row element, Indicates and verifies nodes i Connected divisor nodes j Other variable node indexes.
[0109] Example 2: In an embodiment of the present invention, a data flow optimization system adapted to unified probabilistic graph computation and its hardware circuitry is provided, comprising:
[0110] The message passing module acquires the probabilistic graphical model and obtains the message passing between variable nodes and verification nodes in the probabilistic graphical model based on hardware circuitry. The probabilistic graphical model includes variable nodes, verification nodes, and undirected edges, which are used to connect variable nodes and verification nodes that have a connection relationship.
[0111] The iterative update optimization module iteratively updates message passing to achieve data flow rescheduling, and determines the iterative multiplication unit in the passed message, and optimizes the data flow of the iterative multiplication unit to reduce the power consumption of the probability graph and its hardware circuit.
[0112] In the above embodiments, the iterative multiplication unit is determined in the transmitted message as follows:
[0113]
[0114] in, Representing variables The included states; For verification nodes i To variable node j Passed variables x j The state value is The news, For verification nodes i To variable node j Iterative multiplication units in the transmitted message; For variable nodes j To the verification node i The iterative multiplication unit in the transmitted message, specifically the variable node. j To the verification node i The message is from the verification node. i Other than variable nodes j Connected verification nodes Provide message generation, For verification nodes To variable node j Passing information about variables x j The state value is The message; For variable nodes To the verification node The message conveyed about The message; for The probability configuration function; For the first i The verification node and the first j The candidate vector constructed from the undirected edges between the n variable nodes is the th... m The first line k One element, This represents the element value corresponding to variable node j in the candidate vector; Represents the variable nodes in the candidate vector The corresponding element value, express From the verification node i A set of indexes of variable nodes with connections Take the value from; Indicates and verifies the node i Connected and variable nodes j Values The combination of all variable node states.
[0115] In the above embodiments, data flow optimization of the iterative multiplication unit includes:
[0116] Select a row from the candidate vectors that has the highest similarity to elements in other rows. j -1 element Construct the first backbone unit, using the first backbone unit as the first row element of each subarray as the initial driving vector. The driving vector is used to represent the message being transmitted, and the driving vector is divided into... j There are M subarrays, each subarray being a column vector containing M elements, where M is the total number of candidate vectors. The element with the highest similarity is the element in the remaining rows that is selected. j -1 element The difference position does not exceed 2 difference positions;
[0117] Determine the difference bits between the elements of the remaining rows in each subarray and the corresponding backbone unit, and perform vertical propagation calculations based on the difference bits to calculate the elements of the remaining rows in each subarray, thereby obtaining the complete first column driving vector and outputting it.
[0118] During iterative updates, the first j -1 column of driving vectors j The calculation of the -1 backbone unit is based on the first element in the driving vector of the adjacent preceding column. j -2 backbone units are determined by calculating the difference position of the backbone units in two adjacent driving vectors through lateral propagation. j The calculation results of the -1 backbone unit are then used by the first... j -1 backbone units undergo vertical propagation to obtain the first... j-1 column driving vectors and output them.
[0119] In the above embodiments, the difference bits between the elements of each remaining row in each subarray and the corresponding backbone unit are determined, and vertical propagation is performed based on the difference bits, including:
[0120] When the difference bit is 1: compare the backbone unit with the elements in the second row, and take the extra element in the backbone unit as the first difference bit; compare the elements in the second row with the backbone unit, and take the extra element in the second row as the second difference bit;
[0121] When the difference bit is 2, the product of the two extra elements in the backbone unit is used as the first difference bit when comparing the backbone unit with the second row element; the product of the two extra elements in the second row element is used as the second difference bit when comparing the second row element with the backbone unit.
[0122] When calculating the second row of elements, only the ratio of the second difference position to the first difference position is calculated. The result of the multiplication of the backbone units is then multiplied by the ratio to obtain the second row of elements. This process is repeated to calculate the remaining alternating row elements based on the backbone units, thus completing the vertical propagation.
[0123] In the above embodiments, the first j -1 column of driving vectors j The calculation of the -1 backbone unit is based on the first element in the driving vector of the adjacent preceding column. j -2 backbone units are determined by calculating the difference position of the backbone units in two adjacent driving vectors through lateral propagation. j -1 backbone unit, including:
[0124] In the first subarray, when the difference bit is 1: the first... j -2 backbone unit and the first j Compared to the -1 backbone unit, the first j -2 The missing element in the backbone unit is taken as the third difference position; the first j -1 Backbone Unit and the First j Compared to the -2 backbone unit, the first j -1 The missing element in the backbone unit is used as the fourth difference position;
[0125] When the difference bit is 2, the first j -2 backbone unit and the first j Compared to the -1 backbone unit, the first j -2 The product of the two missing elements in the backbone unit is taken as the third difference bit; the first j -1 backbone unit and j Compared to the -2 backbone unit, the first j -1 The product of the two missing elements in the backbone unit is used as the fourth difference position;
[0126] Conduct the first j When calculating the -1 backbone unit, only the ratio of the third difference position to the fourth difference position is calculated. j -2 The product of the core units, multiplied by the ratio, yields the first... j -1 Calculation results of the backbone unit;
[0127] By analogy, the first subarray in each subarray is obtained. j -1 Calculation results of the backbone unit, completing the lateral propagation calculation.
[0128] In the above embodiments, the data flow optimization of the iterative multiplication unit includes: using a probabilistic graphical message passing method with Gaussian approximation for the iterative multiplication unit, specifically: approximating the iterative multiplication unit as a Gaussian distribution, and replacing the multiplication operation with a Gaussian distribution.
[0129] In the above embodiments, the iterative multiplication unit is approximated as a Gaussian distribution:
[0130] Determine the elements in the association mapping matrix ,make:
[0131] ,
[0132] Therefore, the iterative multiplication unit can be approximated as:
[0133]
[0134] In the formula, Represents the elements in the association mapping matrix. Represents a candidate vector; Indicates the noise variance. express variance This represents the transpose of a vector. Represents the association mapping matrix of the first... i row element, Indicates and verifies the node i Connected divisor nodes j Other variable node indexes.
[0135] The system provided in this embodiment is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0136] Example 3: This embodiment of the invention provides a computing device, which can be a terminal and may include: a processor, a communication interface, memory, a display screen, and an input device. The processor, communication interface, and memory communicate with each other via a communication bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs, which are executed by the processor to implement the methods described in the above embodiments. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, a management network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display or an e-ink display. The input device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse. The processor can call logical instructions stored in the memory.
[0137] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0138] Example 4: In an embodiment of the present invention, a computer program product is provided. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is able to execute the methods provided in the above-described method embodiments.
[0139] Example 5: In an embodiment of the present invention, a non-transitory computer-readable storage medium is provided, which stores server instructions that cause a computer to execute the methods provided in the above embodiments.
[0140] The computer-readable storage medium provided in the above embodiments has a similar implementation principle and technical effect to the above method embodiments, and will not be described again here.
[0141] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 Figure 1 The steps of the function specified in one or more boxes.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data flow optimization method adapted to unified probabilistic graph computation and its hardware circuit, characterized in that, include: Obtain the probabilistic graphical model of the baseband signal processing task, and implement message passing between variable nodes and check nodes in the probabilistic graphical model based on hardware circuits; wherein, the probabilistic graphical model includes variable nodes, check nodes and undirected edges, and undirected edges are used to connect variable nodes and check nodes that have a connection relationship. The message passing is iteratively updated to achieve data flow rescheduling, and the iterative multiplication unit is determined in the passed message. The iterative multiplication unit is then optimized for data flow to reduce the power consumption of the hardware circuit.
2. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 1, characterized in that, The iterative multiplication unit is determined in the transmitted message as follows: , ; in, Representing variables x j The included states; For verification nodes i To variable node j Passed variables x j The state value is The news, For verification nodes i To variable node j Iterative multiplication units in the transmitted message; For variable nodes j To the verification node i The iterative multiplication unit in the transmitted message, specifically the variable node. j To the verification node i The message is from the verification node. i Other than variable nodes j Connected verification nodes l Provide message generation, For verification nodes l To variable node j Passing information about variables x j The state value is The message; For variable nodes l To the verification node The message conveyed about The message; for The probability configuration function; For the first i The verification node and the first j The candidate vector constructed from the undirected edges between the n variable nodes is the th... m The first line k One element, Represents the variable nodes in the candidate vector j The corresponding element value; Represents the variable nodes in the candidate vector The corresponding element value, express From the verification node i A set of indexes of variable nodes with connections Take the value from; Indicates and verifies nodes i Connected and variable nodes j Values The combination of all variable node states; Represents variable nodes j To the verification node Passing information about variable nodes between nodes The state value is The news.
3. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 2, characterized in that, Data flow optimization of the iterative multiplication unit includes: Select a row from the candidate vectors that has the highest similarity to elements in other rows. j -1 element Construct the first backbone unit, using the first backbone unit as the first row element of each subarray as the initial driving vector. The driving vector is used to represent the message being transmitted, and the driving vector is divided into... j There are M subarrays, each subarray being a column vector containing M elements, where M is the total number of candidate vectors. Among them, the element with the highest similarity is the element in the remaining rows that is the selected element. j -1 element The difference position does not exceed 2 difference positions; Determine the difference bits between the elements of the remaining rows in each subarray and the corresponding backbone unit, and perform vertical propagation calculations based on the difference bits to calculate the elements of the remaining rows in each subarray, thereby obtaining the complete first column driving vector and outputting it. During iterative updates, the first j -1 column of driving vectors j- The calculation of the backbone unit is based on the first driving vector in the adjacent preceding column. j -2 backbone units are determined. The calculation result of the j-1th backbone unit is obtained by lateral propagation of the difference position of the backbone units in the two adjacent driving vector columns. Then, the j-1th backbone unit is vertically propagated to obtain the j-1th column driving vector and output it.
4. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 3, characterized in that, Determine the difference bits between the elements of each remaining row in each subarray and the corresponding backbone unit, and perform vertical propagation calculations based on the difference bits, including: When the difference bit is 1: compare the backbone unit with the elements in the second row, and take the extra element in the backbone unit as the first difference bit; compare the elements in the second row with the backbone unit, and take the extra element in the second row as the second difference bit; When the difference bit is 2, the product of the two extra elements in the backbone unit is used as the first difference bit when comparing the backbone unit with the second row element; the product of the two extra elements in the second row element is used as the second difference bit when comparing the second row element with the backbone unit. When calculating the second row of elements, only the ratio of the second difference position to the first difference position is calculated. The result of the multiplication of the backbone units is then multiplied by the ratio to obtain the second row of elements. This process is repeated to calculate the remaining alternating row elements based on the backbone units, thus completing the vertical propagation.
5. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 3, characterized in that, The calculation of the (j-1)th backbone unit in the (j-1)th column of the driving vector is determined based on the (j-2)th backbone unit in the adjacent preceding column of the driving vector. The calculation result of the (j-1)th backbone unit is obtained by lateral propagation of the difference position of the backbone units in the two adjacent columns of the driving vector, including: In the first subarray, when the difference bit is 1: compared with the j-1 backbone unit, the missing element in the j-2 backbone unit is taken as the third difference bit; compared with the j-2 backbone unit, the missing element in the j-1 backbone unit is taken as the fourth difference bit. When the difference bit is 2, the product of the two missing elements in the (j-2)th backbone unit is used as the third difference bit when comparing the (j-2)th backbone unit with the (j-1)th backbone unit; the product of the two missing elements in the (j-1)th backbone unit is used as the fourth difference bit when comparing the (j-1)th backbone unit with the (j-2)th backbone unit. When calculating the (j-1)th backbone unit, only the ratio of the third difference position to the fourth difference position is calculated. The result of the multiplication of the (j-2)th backbone unit is then multiplied by the ratio to obtain the calculation result of the (j-1)th backbone unit. By analogy, the calculation results of the (j-1)th backbone unit in each subarray are obtained, and the lateral propagation calculation is completed.
6. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 2, characterized in that, The data flow optimization of the iterative multiplication unit includes: using a probabilistic graphical message passing method with Gaussian approximation for the iterative multiplication unit, specifically: approximating the iterative multiplication unit as a Gaussian distribution and replacing the multiplication operation with a Gaussian distribution.
7. The data flow optimization method for adapting unified probabilistic graph computation and its hardware circuit as described in claim 6, characterized in that, Approximate the iterative multiplication unit as a Gaussian distribution: Determine the elements in the association mapping matrix ,make: , , Therefore, the iterative multiplication unit can be approximated as: In the formula, Represents the elements in the association mapping matrix. Represents a candidate vector; Indicates the noise variance. express variance This represents the transpose of a vector. Represents the association mapping matrix of the first i row element, Indicates and verifies nodes i Connected divisor nodes j Other variable node indexes.
8. A data flow optimization system adapted to unified probabilistic graph computation and its hardware circuitry, characterized in that, include: The message passing module acquires the probabilistic graphical model of the baseband signal processing task and implements message passing between variable nodes and check nodes in the probabilistic graphical model based on hardware circuitry. The probabilistic graphical model includes variable nodes, check nodes, and undirected edges, which are used to connect variable nodes and check nodes that have a connection relationship. The iterative update optimization module iteratively updates message transmission to achieve data flow rescheduling, and determines the iterative multiplication unit in the transmitted message, and optimizes the data flow of the iterative multiplication unit to reduce the power consumption of the hardware circuit.
9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 7.
10. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 7.
Citation Information
Patent Citations
Decoding method for simulating decoding circuit stop criterion based on probability calculation
CN114584151A
Fast robust sparse Bayesian image reconstruction method based on generalized approximate message passing
CN119090705A