A Coarse-Grained Reconfigurable Processor Architecture Based on Stochastic Computing
By introducing a random computing architecture into the coarse-grained reconfigurable processor, a random computing multiplier and an addition shifter are used, combined with Sobol sequence and precision control module, the high power consumption problem of traditional processors in low-precision scenarios is solved, and energy efficiency improvement and flexible precision control are achieved.
Patent Information
- Application Number
- CN202210219843.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Traditional coarse-grained reconfigurable processors consume a lot of power in scenarios where the calculation accuracy requirements are not high, and the delay and circuit area are increased after approximate calculation are introduced. The random calculation method has the problems of excessive hardware overhead and slow processing speed.
A coarse-grained reconfigurable processor architecture based on random calculations is introduced, using a random calculation multiplier and an addition shifter, combined with the Sobol sequence generator and a progressive precision control module, optimize the operator mapping through an integer linear programming model, and dynamically change the operation accuracy to reduce power consumption.
In scenarios where the accuracy is not required, the processor energy efficiency is improved, the hardware resource consumption is reduced, and the computing speed is accelerated, achieving flexible accuracy control.
Smart Images

Figure CN114722000B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of microprocessing technology, and in particular to a coarse-grained reconfigurable processor architecture based on random computing and an operator mapping method thereof. Background Art
[0002] Since the traditional coarse-grained reconfigurable processor (CGRA) uses precise computing units, although the computing accuracy is high, the power consumption is large. In some scenarios where the computing accuracy requirements are not high, the advantages are not obvious.
[0003] In order to better adapt to the characteristics and needs of intelligent computing and improve resource utilization and computing energy efficiency, the spatial computing architecture with hardware reconstruction capability - coarse-grained reconfigurable processor (CGRA) has become a research hotspot. The conference report of the 2019 International Solid-State Circuits Conference (ISSCC'19) pointed out that since 2016, more than half of the artificial intelligence chips have adopted a computing architecture with certain hardware reconstruction capabilities. CGRA takes "dynamic reconstruction and data-driven" as the basic design idea, supports dynamic hardware reconstruction at four levels / granularities, including core computing components, processing units PE, PE arrays, and computing modes. The entire architecture has great advantages in flexibility, high energy efficiency, and low power consumption.
[0004] In order to further improve the energy efficiency of CGRA, some improvements have been made to the traditional CGRA acceleration array, and the approximate computing mode has been introduced to integrate the approximate computing with the CGRA architecture. The traditional CGRA architecture is further improved, and a new X-CGRA architecture is proposed. Approximate computing is a computing mode that can improve hardware. The multipliers and adders under approximate computing have lower area, power consumption and delay. Its essence is to simplify the structure of the multiplier or adder, but it will bring a certain loss of accuracy. Based on this computing mode, X-CGRA introduces the QSPE structure. In order to ensure the accuracy during the operation, X-CGRA controls the computing mode of the operator through power-gating technology. Its approximate computing mode is divided into multiple levels. The article also proposes an accurate algorithm to measure the accuracy sensitivity of each operator, so as to specify the approximate mode. This problem is solved by the ILP solver. Because X-CGRA introduces power-gating technology, the processor delay in the operation process increases, and the hardware designed according to approximate computing increases the circuit area instead, without changing the traditional multiplication operation mode.
[0005] Stochastic computing (SC) is an unconventional computing method that treats data as probabilities, thus revolutionizing the operation mode of traditional multiplication. Typically, we can regard a decimal P as a string of N-bit random bit stream X composed of 0 or 1. The probability of 1 appearing in X is exactly equal to P, and X can be generated and processed by conventional logic circuits. Thus, a single AND gate can perform multiplication. The X value of SN is measured by the number of 1s in SN, which is also an information coding scheme found in biological nervous systems. SC has a wide range of applications in large-scale parallel systems and has strong tolerance to errors. Its disadvantages are low precision, slow processing speed, and complex design. It can effectively execute tasks such as communication decoding and neural network inference, which has rekindled people's interest in this field. Despite the above advantages of SC, the excessive hardware overhead brought by the stochastic number generator (SNG) is still the main obstacle to the widespread application of SC-based architectures. Traditional SC serially generates a random bit stream, converts the input data (binary format) into a long bit stream, and each data is compared with the data generated by a linear feedback shift register (LFSR) in a step-by-step manner, which is the main task of the SNG. Due to the necessity of the SNG, its internal LFSR will cause a large amount of resource consumption. To reduce the internal hardware consumption, an SNG design based on a deterministic generation method is proposed. Using a Sobol sequence generator, the input data is compared with the pre-generated Sobol sequence one by one, thus greatly reducing the hardware resource consumption and further improving the operation accuracy. A parallel SNG design based on the Sobol sequence is proposed, that is, increasing the number of comparators, so that a string of random bit streams can be directly generated within one clock cycle, further accelerating the execution speed of SC.
[0006] Since the arithmetic units in traditional reconfigurable processors still adopt traditional arithmetic paradigms, such as multipliers with large resource consumption, there is still room for improvement in their energy efficiency. Therefore, it is urgent to introduce a new computing mode to further improve its hardware architecture. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0008] For this purpose, the object of the present invention is to propose a coarse-grained reconfigurable processor architecture based on stochastic computing, which can apply stochastic computing to the design of CGRA and dynamically change the operation precision of CGRA, so that the computing task can be completed with lower power consumption in scenarios with low requirements for operation precision, greatly improving the energy efficiency of CGRA.
[0009] Another object of the present invention is to propose an operator mapping method for a coarse-grained reconfigurable processor architecture based on stochastic computing.
[0010] To achieve the above object, on the one hand, the present invention proposes a coarse-grained reconfigurable processor architecture based on stochastic computing, including:
[0011] A configuration memory for storing configuration information; a data memory for storing data; a central processing unit for performing computational processing on the data stored in the data memory; and a dynamically reconfigurable PE array for controlling data operations on the computational processing performed by the central processing unit by using a multiplier and an adder-shifter based on stochastic computing according to the configuration information stored in the configuration memory.
[0012] The coarse-grained reconfigurable processor architecture based on stochastic computing according to the embodiments of the present invention can apply stochastic computing in the design of CGRA, dynamically change the computing precision of CGRA, so that computing tasks can be completed with lower power consumption in scenarios where the computing precision requirement is not high, greatly improving the energy efficiency of CGRA.
[0013] In addition, the coarse-grained reconfigurable processor architecture based on stochastic computing according to the above embodiments of the present invention may further have the following additional technical features:
[0014] Further, in an embodiment of the present invention, the dynamically reconfigurable PE array includes an arithmetic logic unit SC-ALU based on stochastic computing, and the SC-ALU includes the multiplier ISC-MUL based on stochastic computing.
[0015] Further, in an embodiment of the present invention, the multiplier ISC-MUL based on stochastic computing includes: a primitive multiplier NSC-MUL based on stochastic computing and a leading zero shift module ZSAS.
[0016] Further, in an embodiment of the present invention, the primitive multiplier NSC-MUL based on stochastic computing includes: two parallel random number generators SNG, a multi-bit AND, an APC, and a left shifter; wherein, the random number generator SNG is implemented according to the Sobol sequence.
[0017] Further, in an embodiment of the present invention, the leading zero shift module ZSAS is used to improve the precision of the primitive multiplier NSC-MUL based on stochastic computing; wherein, the leading zero shift module ZSAS includes: two leading zero counters, two left shifters, a 4-bit adder, and a 4-bit subtractor.
[0018] Further, in an embodiment of the present invention, it further includes: a progressive precision control module, which is used to connect and fix the Sobol sequence in different PE arrays to obtain the Sobol sequence with a preset length.
[0019] Further, in an embodiment of the present invention, it further includes: a shift averaging module, which is used to determine the output result of the NSC-MUL by calculating the numbers in the AND gate output, add the output result of the NSC-MUL to multiple said adder shifters, and right-shift the sum of each said adder shifter for averaging.
[0020] To achieve the above object, on the other hand, the present invention proposes an operator mapping method for a coarse-grained reconfigurable processor architecture based on stochastic computing, including:
[0021] Using an integer linear programming model to formulate precision level allocation to determine the precision level of each operator in the data flow graph; and, presetting an objective function for the integer linear programming model, and obtaining the optimization result of each operator in the data flow graph by minimizing the value of the objective function.
[0022] The operator mapping method of the embodiment of the present invention can allow the CGRA to dynamically change its precision performance in the face of scenarios with different precision requirements.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1) By introducing stochastic computing, the architecture of CGRA is further improved in energy efficiency;
[0025] 2) The designed architecture allows dynamic precision control. Utilizing the flexible interconnection characteristics of CGRA, a precision expansion strategy is proposed;
[0026] 3) It is necessary to dynamically change the overall precision performance for different precision application scenarios and design a dynamic precision mapping algorithm, which allows the CGRA to dynamically change its precision performance in the face of scenarios with different precision requirements.
[0027] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Description of the Drawings
[0028] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0029] Figure 1 It is a schematic structural diagram of a coarse-grained reconfigurable processor architecture based on stochastic computing according to an embodiment of the present invention;
[0030] Figure 2 Schematic diagram of another coarse-grained reconfigurable processor architecture based on stochastic computing according to an embodiment of the present invention;
[0031] Figure 3 Schematic diagram for controlling the precision of the random bit stream according to an embodiment of the present invention;
[0032] Figure 4 Flowchart of an operator mapping method for a coarse-grained reconfigurable processor architecture based on stochastic computing according to an embodiment of the present invention. Detailed implementation manners
[0033] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] The coarse-grained reconfigurable processor architecture and operator mapping method based on stochastic computing according to an embodiment of the present invention will be described below with reference to the drawings.
[0036] Figure 1 Schematic diagram of a coarse-grained reconfigurable processor architecture based on stochastic computing according to an embodiment of the present invention.
[0037] As Figure 1 shown, the coarse-grained reconfigurable processor architecture 10 based on stochastic computing includes:
[0038] A configuration memory 100 for storing configuration information;
[0039] A data memory 200 for storing data;
[0040] A central processing unit 300 for performing computational processing on the data stored in the data memory 200;
[0041] A dynamically reconfigurable PE array 400 for controlling data operations on the computational processing performed by the central processing unit 300 by using a multiplier and an adder-shifter based on stochastic computing according to the configuration information stored in the configuration memory 100.
[0042] It can be understood that the present invention aims to improve the hardware structure of traditional reconfigurable architectures, introduce the new computing mode of stochastic computing into the design of traditional reconfigurable architectures, so as to further reduce the overall circuit area and energy consumption, and improve the energy efficiency ratio of traditional reconfigurable processors. The combination of stochastic computing and CGRA can not only bring out the advantages of both but also alleviate their respective disadvantages to a certain extent.
[0043] Due to the flexible on-chip interconnection of CGRA, the progressive precision characteristic of stochastic computing can be easily implemented on CGRA. Moreover, CGRA has the inherent coarse-grained operation characteristics and can perform operations with a wider bit width, so that stochastic computing can be conveniently parallelized. In addition, the low energy consumption benefit brought by stochastic computing can greatly enhance the energy efficiency of traditional CGRA. Therefore, the present invention proposes a CGRA based on stochastic computing, called SC-CGRA.
[0044] As an example, as Figure 2 (a) in shows the architecture design of a 4×4 SC-CGRA, which mainly consists of a central processor 300, a data memory 200, a configuration memory 100, and a dynamically reconfigurable PE array 400 with scalable quality. SC-CGRA supports spatial and temporal mapping and uses a 2D mesh plus interconnection mode. As Figure 2 shown, the ALU in SC-CGRA is different from the ALU in traditional CGRA in two aspects: 1) The precise multiplier is replaced by an improved stochastic computing-based multiplier (ISC-MUL). 2) The adder is replaced by an add-shifter, so that the sum of the adder can be right-shifted by 0 (for a normal adder) or 1 bit (for precision scaling), depending on the configuration information.
[0045] Furthermore, as Figure 2 shown by the yellow area in (c) of, the original MUL based on SC (NSC-MUL) consists of two parallel SNGs, a multi-bit AND, an APC, and a left shifter. To implement the parallel SNG, this embodiment introduces the Sobol sequence to implement SNG. To further reduce resource consumption, a fixed LD sequence pattern is also used, so that the Sobol sequence in each PE is fixed in hardware implementation. Therefore, the corresponding wires will be connected to the power supply or grounded, so the Sobol sequence does not require additional storage resources. To further improve the accuracy, each PE has an independent Sobol sequence with different cell values.
[0046] To further improve the precision of NSC-MUL, this embodiment designs a leading zero shift module for Figure 2The gray area shown in (c). This module includes two leading zero counters, two left shifters, a 4-bit adder, and a 4-bit subtractor (4 bits are used to calculate the number of 0s in a 16-bit operand). This module is called ZSAS. After adding the ZSAS module, the accuracy of the multiplier will be greatly improved immediately.
[0047] Figure 2 (c) also shows the entire process of ISC-MUL: First, the 4-bit input operands (a = 8, b = 6) will be shifted left by ZSAS. Since a = 8(1000)₂ and b = 6(0110)₂ have 0 and 1 leading zeros (Sa = 0 and Sb = 1) respectively, a and b are amplified to 8 and 12 respectively. At the same time, Sf = L - (Sa + Sb) = 3 can be easily calculated in ZSAS. Second, we use SNG to quickly convert a' and b' into two 16-bit random bitstreams (M = 4). After AND and APC process the two bitstreams, the number of 1s (N1s = 6) is obtained. Finally, shift 1 to the left by Sf = 4 - 1(3) to obtain the output of MUL (48).
[0048] Furthermore, it also includes: a progressive precision control module. Although the precision of ISC-MUL has been greatly improved, it is still limited by the length of the parallel fixed Sobol sequence in each PE and will not be too long in actual implementation. To further improve the accuracy of the MUL operation, using the progressive precision control module, the fixed Sobol sequences in different PEs can be connected to form a longer Sobol sequence.
[0049] Furthermore, it also includes: a shift averaging module. Since the SC-MUL result is determined by calculating the number of 1s in the AND gate output, through the shift averaging module, only need to add the SC-MUL output result to multiple adders and right-shift the sum of each adder by 1 bit for averaging. As Figure 3 shown, in the original DFG, to improve the precision of the blue MUL operator, an additional MUL operator and an add-shift operator are added (as Figure 3 shown in (b)), where the two MUL operators use the same operands (a = 6, b = 8) as input. As Figure 3 shown in (c), by executing the two MUL operators on PE1 and PE2, 47 and 49 are obtained as results. Then, taking the two results as input, the adder-shifter obtains a high-precision output (48).
[0050] Meanwhile, as an implementation method, the specific implementation of the coarse-grained reconfigurable processor architecture based on random calculation of the present invention is divided into the following steps:
[0051] Step A: Write a specific hardware implementation in Verilog HDL according to the hardware architecture;
[0052] Step B: Feed the hardware implementation into an EDA tool for power consumption, area, and timing evaluation and verification;
[0053] Step C: Burn the verified HDL on a reconfigurable array such as an FPGA or CPLD;
[0054] Step D: Write an implementation of the precision mapping algorithm;
[0055] Step E: Feed the program input by the user into the algorithm to obtain the DFG that needs to be mapped finally;
[0056] Step F: Obtain the configuration information of the CGRA through the time-domain mapping algorithm;
[0057] Step G: Feed the configuration information into the hardware to start the computing task.
[0058] Thus, according to the coarse-grained reconfigurable processor architecture based on stochastic computing of the embodiments of the present invention, through the introduction of stochastic computing, the CGRA architecture further obtains an improvement in energy efficiency. At the same time, the designed architecture allows dynamic precision control and can dynamically change the overall precision performance for different precision application scenarios. In addition, applying it in the CGRA enables a single PE in the CGRA to achieve high energy efficiency without causing too much precision loss, and through the flexible interconnection between PEs in the CGRA, adjacent PEs are connected to achieve flexible precision control.
[0059] To implement the above embodiments, as Figure 4 shown, this embodiment also provides an operator mapping method for a coarse-grained reconfigurable processor architecture based on stochastic computing, and the method includes:
[0060] Step S1, use an integer linear programming model to formulate a precision level assignment to determine the precision level of each operator in the data flow graph; and,
[0061] Step S2, preset an objective function for the integer linear programming model, and obtain the optimization result of each operator in the data flow graph by minimizing the value of the objective function.
[0062] Specifically, since splicing between PEs to improve the computing precision will cost more resources, it is necessary to carefully determine the precision level of each operator in the DFG data flow graph to meet the output quality requirements. Therefore, use integer linear programming (ILP) to formulate the precision level assignment, which has a formal expression and helps to find the optimal solution.
[0063] First is the definition of decision variables: Since different operators have different impacts on the output quality of the entire DFG data flow graph. In this embodiment, K binary decision variables will be defined For each MUL operator i in the DFG, where k ∈ {1, K} represents the k-th precision level. K is an empirical parameter, which specifically indicates how many basic random bitstreams will be concatenated to form a longer bitstream. The larger the value of K, the higher the precision of the operator
[0064] Secondly is the establishment of model constraints. In this embodiment, the establishment of the ILP model requires constructing two constraints
[0065] ILP model constraint 1 (C1): The model constrains that the precision level of each MUL is unique, that is, each MUL operator can only have one precision level. Therefore, the uniqueness constraint of the MUL operator decision variable is constructed as follows
[0066]
[0067] where G represents the number of MUL operators in the DFG
[0068] ILP model constraint 2 (C2): The precision constraint of the entire DFG. For a given precision requirement, the degradation of the output quality should not be greater than a given upper limit The present invention also calculates the loss degree of the output quality according to the error distance (ED) and error sensitivity (ES). ED is expressed as follows
[0069] ED = |O - O k |
[0070] where O and O k represent the exact output and approximate output with an exact level of k respectively
[0071] According to the definitions of ED and ES, ES is as follows
[0072]
[0073] where and represent the output ED of the entire DFG and the output ED of operator i respectively. At this time, operator i is a MUL operator with a precision level of k, and other operators are in the exact mode. Meanwhile, the error variance of DFG(v o ) can be calculated as
[0074]
[0075] where v i is the output variance of the i-th operator. After that, the output error constraint of the entire DFG is formulated as
[0076]
[0077] wherein, N and O n respectively represent the total number of samples and the exact output of the Nth DFG sample. As for the upper limit of precision can be freely set to various error representations, such as variance, average error distance, etc.
[0078] Finally, an objective function also needs to be set for the ILP model, and the optimization result is obtained by minimizing the value of the objective function. Since the precision extension strategy of the present invention needs to extend the original DFG, this may increase the minimum resource startup interval (II) and reduce the performance of software pipelining. Therefore, it is necessary to minimize the total number of added operators while satisfying constraints C1 and C2.
[0079] Given a MUL operator i, if its precision level is to be increased to k i ∈{1, K}, k i -1 additional MUL operators are required for the precision improvement of operator i, and k i -1 addition-shift operators are added to obtain the final result. To sum up, the objective function of the ILP model in this embodiment will represent the minimization of the newly added operators in the DFG:
[0080]
[0081] According to the operator mapping method of the coarse-grained reconfigurable processor architecture based on stochastic computing proposed in the embodiment of the present invention, the CGRA can be allowed to dynamically change its precision performance in the face of scenarios with different precision requirements.
[0082] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0083] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0084] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A coarse-grained reconfigurable processor architecture based on stochastic computing, characterized in that, Comprising: A configuration memory for storing configuration information; A data memory for storing data; A central processing unit for performing computational processing on the data stored in the data memory; A dynamically reconfigurable PE array for controlling data operations on the computational processing performed by the central processing unit by using a multiplier and an adder-shifter based on stochastic computing according to the configuration information stored in the configuration memory; The dynamically reconfigurable PE array includes an arithmetic logic unit SC-ALU based on stochastic computing, and the SC-ALU includes the multiplier ISC-MUL based on stochastic computing; The multiplier ISC-MUL based on stochastic computing includes a primitive multiplier NSC-MUL based on stochastic computing and a leading zero shifter module ZSAS; The primitive multiplier NSC-MUL based on stochastic computing includes two parallel random number generators SNG, a multi-bit AND, an APC, and a left shifter; wherein, the random number generator SNG is implemented according to the Sobol sequence; The leading zero shifter module ZSAS is used to improve the precision of the primitive multiplier NSC-MUL based on stochastic computing; wherein, the leading zero shifter module ZSAS includes two leading zero counters, two left shifters, a 4-bit adder, and a 4-bit subtractor; Further comprising: a progressive precision control module, which is used to connect the fixed Sobol sequences in different PE arrays to obtain the Sobol sequence of a preset length; Determine the output result of the NSC-MUL based on the numbers in the output of the AND gate calculation, add the output results of the NSC-MUL to multiple of the adder-shifters, and right-shift the sum of each adder-shifter for averaging.
2. An operator mapping method for the reconfigurable processor architecture according to claim 1, characterized in that Including the following steps: Formulate precision level allocation using an integer linear programming model to determine the precision level of each operator in the data flow graph; And, Preset an objective function for the integer linear programming model, and obtain the optimization result of each operator in the data flow graph by minimizing the value of the objective function.
3. The operator mapping method according to claim 2, wherein The formulating precision level allocation using an integer linear programming model includes: the definition of decision variables and the establishment of constraints of the integer linear programming model; wherein, The definition of the decision variables includes: defining K binary decision variables For each MUL operator i in the corresponding data flow graph DFG, where k ∈ {1, K} represents the k-th precision level; The constraints of the integer linear programming model include: the first integer linear programming model constraint and the second integer linear programming model constraint; wherein, The first integer linear programming model constraint includes: constructing a uniqueness constraint for the decision variables of the MUL operator; The second integer linear programming model constraint includes: the precision constraint for the entire data flow graph DFG.
4. The operator mapping method according to claim 3, characterized in that The constructing of the uniqueness constraint for the decision variables of the MUL operator is: Wherein, G represents the number of MUL operators in the DFG; The precision constraint for the entire data flow graph DFG is: Calculate the loss degree of the output quality according to the error distance ED and the error sensitivity ES, and the ED is expressed as: ED = |O - O k | where O and O k represent the exact output and the approximate output with the exactness level of k, respectively; According to the definitions of ED and ES: Among them, and represent the output ED of the entire DFG and the output ED of operator i respectively. Then, the error variance of the data flow graph v o is calculated as: where v i is the output variance of the i-th operator, and the output error constraint of the entire DFG is formulated as: where N and O n represent the total number of samples and the exact output of the Nth DFG sample, respectively.