Parallel processing appratus
The parallel processing device addresses inefficiencies by performing both multiplication and addition operations with integrated preprocessing units, enhancing efficiency through flexible operation mode switching and data exchange, thus optimizing hardware utilization.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- MORUMI CO LTD
- Filing Date
- 2022-07-05
- Publication Date
- 2026-07-29
AI Technical Summary
Parallel processing devices face inefficiencies due to separate adders and multipliers being underutilized based on operation type, and challenges in data exchange between processing units, leading to suboptimal performance.
A parallel processing device designed to perform both multiplication and addition operations, with integrated preprocessing units capable of shifting and selecting signals, facilitating efficient data exchange and operation mode switching.
Enhances parallel processing efficiency by enabling simultaneous performance of various operations without significantly increasing hardware complexity, optimizing utilization based on operational demands.
Smart Images

Figure 112022070101681-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The technology described below relates to a parallel processing device. Background Technology
[0002] Much research is being conducted on parallel processing devices to achieve high data processing performance. An example of a parallel processing device is the multi-core processor. A multi-core processor is a processor equipped with multiple cores (processing units), and the reason for using multi-core processors is to improve the overall processor performance by increasing the number of cores. However, for various reasons, the overall processor performance does not increase proportionally even when the number of cores is increased.
[0003] The inventor of the present invention has been carrying out continuous development to improve these problems, and based on this, has performed the inventions of Korean Patent Publication No. 10-2019-0132295, No. 10-2018-0057950, No. 10-2018-0058166, No. 10-2018-0058167, No. 10-2018-0007523, No. 10-2018-0007652 and Korean Patent Registration No. 10-1859294. Prior art literature
[0004] Korean Patent Publication No.: 10-2019-0132295, 10-2018-0057950, 10-2018-0058166, 10-2018-0058167, 10-2018-0007523, 10-2018-0007652 Korean Patent Registration No.: 10-1859294 The problem to be solved
[0005] Parallel processing devices according to the prior art are equipped with separate adders and multipliers. This reduces the efficiency of the parallel processing device. More specifically, when many multiplication operations are required, all multipliers of the parallel processing device are utilized, but some adders are not utilized. Also, when many addition operations are required, all adders of the parallel processing device are utilized, but some multipliers are not utilized. Furthermore, parallel processing devices according to the prior art have aspects where data exchange between processing units is not easy. This reduces the overall performance of the parallel processing device.
[0006] The present disclosure aims to solve the problems of the prior art by designing the processing unit of a parallel processing device to be capable of performing both multiplication and addition operations, thereby increasing the efficiency of the parallel processing device. Furthermore, the present disclosure aims to increase the overall efficiency of the parallel processing device by facilitating data exchange between processing units. Additionally, the present disclosure aims to increase the overall efficiency of the parallel processing device by enabling each processing unit to perform various operations over time. Furthermore, the present disclosure aims not to significantly increase the complexity of the overall hardware despite having the improvements described above. means of solving the problem
[0007] A parallel processing device according to one embodiment receives first to N inputs ({X1, Y1}, {X2, Y2}, ... {XN, YN}) and outputs first to N outputs (M1, M2, ... MN). The first to N inputs ({X1, Y1}, {X2, Y2}, ... {XN, YN}) include a plurality of first signals (X1, X2, ... XN) and a plurality of second signals (Y1, Y2, ... YN), and N is a natural number greater than or equal to 4. When the parallel processing device operates in summing mode, it transmits the i-th first signal (Xi) to summing units (SUM1, SUM2, ... SUMN) respectively according to the i-th bits (Y1[i], Y2[i], ... YN[i]) of the plurality of second signals (Y1, Y2, ... YN), and when operating in multiplication mode, it transmits signals ((X1<<(i-1)), (X2<<(i-1)), ... (XN<<(i-1)))) in which the first signals (X1, X2, ... XN) are shifted by (i-1) bits to the summing units (SUM1, SUM2, ... SUMN) respectively according to the i-th bits (Y1[i], Y2[i], ... YN[i]) of the second signals (Y1, Y2, ... YN), wherein i is a natural number greater than or equal to 1 and less than or equal to N; and includes summers (SUM1, SUM2, ... SUMN) that output the first to Nth outputs (M1, M2, ... MN), and each summer among the summers (SUM1, SUM2, ... SUMN) includes a main processing unit that sums the transmitted signals.
[0008] A parallel processing device according to one embodiment includes a preprocessing unit comprising preprocessing units; and a main processing unit comprising summing units, wherein each preprocessing unit among the preprocessing units comprises a selection operation unit and a shift operation unit, wherein the selection operation unit operates when a corresponding preprocessing unit among the preprocessing units operates in a summing mode and performs the function of transmitting first signals to a corresponding summing unit among the summing units according to bits of a corresponding second signal among the second signals, and wherein the shift operation unit operates when a corresponding preprocessing unit operates in a multiplication mode and transmits signals in which a corresponding first signal among the first signals is shifted to a corresponding summing unit according to bits.
[0009] A parallel processing device according to one embodiment includes a preprocessing unit comprising preprocessing units; and a main processing unit comprising summing units, wherein each preprocessing unit among the preprocessing units comprises a selection operation unit and a shift operation unit, wherein the selection operation unit operates when a corresponding preprocessing unit among the preprocessing units operates in a summing mode and performs the function of transmitting a corresponding portion of first signals among first signals to a corresponding summing unit among the summing units according to bits of a corresponding second signal among second signals, and wherein the shift operation unit operates when a corresponding preprocessing unit operates in a multiplication mode and transmits signals in which a corresponding first signal among the first signals has been shifted to a corresponding summing unit according to bits. Effects of the invention
[0010] The parallel processing device according to the present disclosure has high parallel processing efficiency because the unit can perform both multiplication and addition operations. In addition, the parallel processing device has the advantage of being able to additionally perform displacement and shift operations. Furthermore, the parallel processing device has the advantage of enabling easy data exchange between processing units. In addition, the parallel processing device has the advantage that each processing unit can perform various operations over time. Moreover, the parallel processing device has the advantage that the complexity of the hardware does not increase significantly despite the improvements described above. Brief explanation of the drawing
[0011] FIG. 1 is a drawing showing a parallel processing device according to a first embodiment. FIG. 2 is a drawing for explaining an example of the i-th preprocessing unit of the first embodiment. FIG. 3 is a drawing showing a parallel processing device according to a second embodiment. Specific details for implementing the invention
[0012] The technology described below is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the technology described below to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the technology described below.
[0013] Terms such as first, second, A, B, etc., may be used to describe various components, but such components are not limited by the said terms and are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of rights of the technology described below, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related described items or any of the multiple related described items.
[0014] In terms used in this specification, singular expressions should be understood to include plural expressions unless the context clearly indicates otherwise, and terms such as “includes” should be understood to mean that the described features, number, step, action, component, part, or combination thereof exist, and not to exclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0015] Before providing a detailed description of the drawings, it is to clarify that the classification of components in this specification is merely based on the primary function each component is responsible for. That is, two or more components described below may be combined into a single component, or a single component may be divided into two or more components based on more subdivided functions. Furthermore, each component described below may additionally perform some or all of the functions of other components in addition to its own primary function, and it goes without saying that some of the primary functions of each component may be exclusively performed by other components.
[0016] Furthermore, in performing the method or operation method, each process constituting the method may occur differently from the specified order unless a specific order is clearly indicated in the context. That is, each process may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.
[0017] FIG. 1 is a drawing illustrating a parallel processing device according to a first embodiment. Referring to FIG. 1, the parallel processing device receives first to Nth inputs ({X1, Y1}, {X2, Y2}, ... {XN, YN}) and outputs first to Nth outputs (M1, M2, ... MN). Here, N represents a natural number greater than or equal to 4, and for example, N may be 32. The first to Nth inputs ({X1, Y1}, {X2, Y2}, ... {XN, YN}) include first signals (X1, X2, ... XN) and second signals (Y1, Y2, ... YN). The parallel processing device includes a preprocessing unit (100) and a main processing unit (200). The parallel processing device may further include a delay unit (300) and a selection unit (400).
[0018] When the preprocessor (100) operates in summing mode, the first signal (Xi) of the i-th input ({Xi, Yi}) is transmitted to the summers (SUM1, SUM2, ... SUMN) of the main processing unit (200) respectively, according to the i-th bits (Y1[i], Y2[i], ... YN[i]) of the second signals (Y1, Y2, ... YN). Here, i is a natural number greater than or equal to 1 and less than or equal to N. Let the inputs of the first summer (SUM1) be S1_1, S1_2, ... S1_N, the inputs of the second summer (SUM2) be S2_1, S2_2, ... S2_N, and the inputs of the third summer (SUM3) be S3_1, S3_2, ... S3_N. At this time, the summing mode operation of the preprocessor (100) can be expressed in Verilog language as an example as follows.
[0019] (Y1[1] ? X1 : 0) => S1_1,
[0020] (Y1[2] ? X2 : 0) => S1_2,
[0021] ...
[0022] (Y1[N] ? XN : 0) => S1_N,
[0023] (Y2[1] ? X1 : 0) => S2_1,
[0024] (Y2[2] ? X2 : 0) => S2_2,
[0025] ...
[0026] (Y2[N] ? XN : 0) => S2_N,
[0027] ...
[0028] (YN[1] ? X1 : 0) => SN_1,
[0029] (YN[2] ? X2 : 0) => SN_2,
[0030] ...
[0031] (YN[N] ? XN : 0) => SN_N
[0032] When the preprocessor (100) operates in multiplication mode, the first signals (X1, X2, ... XN) are shifted by (i-1) bits ((X1<<(i-1)), (X2<<(i-1)), ... (XN<<(i-1))) and are transmitted to summers (SUM1, SUM2, ... SUMN) respectively according to the i-th bits (Y1[i], Y2[i], ... YN[i]) of the second signals (Y1, Y2, ... YN). The multiplication mode operation of the preprocessor (100) can be expressed in Verilog language as an example as follows.
[0033] (Y1[1] ? (X1 << 0) : 0)=> S1_1,
[0034] (Y1[2] ? (X1 << 1) : 0)=> S1_2,
[0035] ...
[0036] (Y1[N] ? (X1 << (N-1)) : 0)=> S1_N,
[0037] (Y2[1] ? (X2 << 0) : 0)=> S2_1,
[0038] (Y2[2] ? (X2 << 1) : 0)=> S2_2,
[0039] ...
[0040] (Y2[N] ? (X2 << (N-1)) : 0)=> S2_N,
[0041] ...
[0042] (YN[1] ? (XN << 0) : 0)=> SN_1,
[0043] (YN[2] ? (XN << 1) : 0)=> SN_2,
[0044] ...
[0045] (YN[N] ? (XN << (N-1)) : 0)=> SN_N
[0046] For example, the preprocessor (100) operates in either an aggregation mode or a multiplication mode according to the operation mode selection signals (SF1, SF2, ... SFN). For example, one operation mode selection signal may be assigned to the entire parallel processing unit. In this case, the entire parallel processing unit must operate in either an aggregation mode or a multiplication mode. As another example, N operation mode selection signals (SF1, SF2, ... SFN) may be assigned. In this case, some of the N outputs (M1, M2, ... MN) may be set to be results obtained according to the aggregation mode, and the rest may be results obtained according to the multiplication mode. For instance, when N is 4, the operation mode selection signals (SF1, SF2, SF3, SF4) may be set so that M1, M2, M3 operate in a multiplication mode and M4 operates in an aggregation mode.
[0047] For example, the preprocessing unit (100) includes a plurality of preprocessing units (150_1, 150_2, ... 150_N). The plurality of preprocessing units (150_1, 150_2, ... 150_N) include selection operation units (110_1, 110_2, ... 110_N) and shift operation units (120_1, 120_2, ... 120_N). The preprocessing unit (150_i) includes a selection operation unit (110_i) and a shift operation unit (120_i).
[0048] The selection operation unit (110_i) operates when the preprocessing unit (150_i) operates in summing mode. The selection operation unit (110_i) performs the function of transmitting the first signals (X1, X2, ... XN) to the summer (SUMi) according to the bits (Yi[1], Yi[2], ... Yi[N]) of the second signal (Yi). At this time, the operation of the selection operation unit (110_i) can be expressed in Verilog language as an example as follows.
[0049] (Yi[1] ? X1 : 0) => Si_1,
[0050] (Yi[2] ? X2 : 0) => Si_2,
[0051] ...
[0052] (Yi[N] ? XN : 0) => Si_N,
[0053] The shift operation unit (120_i) operates when the preprocessing unit (150_i) operates in multiplication mode. The shift operation unit (120_i) performs the function of transmitting signals ((Xi<<0), (Xi<<1), ... (Xi<<(N-1)))) in which the first signal (Xi) is shifted by 0, 1, ... (N-1) bits to the summer (SUMi) according to the bits (Yi[1], Yi[2], ... Yi[N]) of the second signal (Yi). At this time, the operation of the shift operation unit (120_i) can be expressed in Verilog language as an example as follows.
[0054] (Yi[1] ? (Xi << 0) : 0)=> Si_1,
[0055] (Yi[2] ? (Xi << 1) : 0)=> Si_2,
[0056] ...
[0057] (Yi[N] ? (Xi << (N-1)) : 0)=> Si_N,
[0058] The preprocessing unit (150_i) operates the selection operation unit (110_i) or the shift operation unit (120_i) according to the operation mode selection signal (SFi). For example, when SFi is 0, it means the operation of the selection operation unit (110_i), and when it is 1, it means the operation of the shift operation unit (120_i). In this case, SF1=0, SF2=0, and SFN=1 mean that the first preprocessing unit (150_1), the second preprocessing unit (150_2), and the Nth preprocessing unit (150_N) each operate the selection operation unit (120_1), the selection operation unit (110_2), and the shift operation unit (110_N).
[0059] The main processing unit (200) includes summers (SUM1, SUM2, ... SUMN). The i-th summer (Mi) sums the transmitted signals (Si_1, Si_2, ... Si_N) and outputs the summed result as the i-th output (Mi). The operation of the main processing unit (200) can be expressed in Verilog language as an example, as shown below.
[0060] S1_1 + S1_2 + ... S1_N => M1,
[0061] S2_1 + S2_2 + ... S2_N => M2,
[0062] ...
[0063] SN_1 + SN_2 + ... SN_N => MN,
[0064] The delay unit (300) delays and outputs the first to Nth outputs (M1, M2, ... MN) according to the clock signal (CLK). To this end, the delay unit (300) includes a plurality of delay units (DU1, DU2, ... DUN). The signals (D1, D2, ... DN) output from the delay unit (300) correspond to the first to Nth outputs (M1, M2, ... MN), respectively.
[0065] The selection unit (400) outputs signals selected according to input control signals (SI1, SI2, ... SIN) among signals (R1, R2, ... RN) transmitted from memory (not shown) and signals (D1, D2, ... DN) output from delay unit (300) as first signals (X1, X2, ... XN). For example, as shown in the drawing, the first signal (Xi) may be a signal selected according to input control signal (SIi) among a signal (Ri) transmitted from memory and a signal (Di) output from delay unit (300). As another example, the first signal (Xi) may be a signal selected according to input control signal (SIi) among a signal (Ri) transmitted from memory and two signals (D(i-1), Di) output from delay unit (300). That is, the first signal (Xi) may be a signal selected according to the input control signal (SIi) among the signal (Ri) transmitted from the memory, the delay part output signal (Di) corresponding to the i-th output (Mi), and the delay part output signal (D(i-1)) corresponding to the (i-1)-th output (M(i-1)). The memory (not shown) may, for example, have a plurality of banks. For example, the memory may have N banks, and the N banks may each be connected to N inputs ({X1, Y1}, {X2, Y2}, ... {XN, YN}). In addition, the memory may have 2*N banks, and among these, N banks may each be connected to N first signals (X1, X2, ... XN), and the remaining N banks may each be connected to N second signals (Y1, Y2, ... YN).
[0066] By having such a configuration, the parallel processing device can perform various operations with a single piece of hardware. For example, the parallel processing device can perform partial summation operations. Here, partial summation means summing all or part of the first signals (X1, X2, ... XN). In order to perform partial summation operations, the preprocessing unit (100) must operate in summation mode. At this time, the i-th output (Mi) corresponds to the summation of signals selected according to the bits (Yi[1], Yi[2], ... Yi[N]) of the second signal (Yi) among the first signals (X1, X2, ... XN). For example, if N is 4 and the second signals (Y1, Y2, Y3, Y4) are 1011, 1100, 0010, and 0111 in binary, the outputs (M1, M2, M3, M4) correspond to X4+X2+X1, X4+X3, X2, and X3+X2+X1, respectively. In this way, when the preprocessing unit (100) operates in summing mode, N partial summing operations can be performed simultaneously.
[0067] When the preprocessing unit (100) operates in summing mode, a displacement operation may also be performed. Here, a displacement operation means changing the position when transmitting the first signal (Xi) to the output signals (M1, M2, ... MN). For example, if N is 4 and the second signals (Y1, Y2, Y3, Y4) are 1000, 0100, 0010, 0001 in binary, the outputs (M1, M2, M3, M4) correspond to X4, X3, X2, X1, respectively. Therefore, the first signals (X1, X2, X3, X4) are transmitted to the outputs (M1, M2, M3, M4), but their positions are changed.
[0068] As described above, when a parallel processing unit performs partial summation and displacement operations, it facilitates data exchange between processing units. Here, the i-th processing unit is a concept that includes the i-th preprocessing unit and the i-th summer. For example, the first processing unit (150_1, SUM1) can receive not only the first first signal (X1) but also the second through Nth first signals (X2, ... XN) to perform partial summation. Additionally, the second processing unit (150_2, SUM2) can receive the Nth first signal (XN), which is a first signal other than the second first signal (X2).
[0069] For example, a parallel processing unit can perform multiplication operations. In order to perform multiplication operations, the preprocessing unit (100) must operate in multiplication mode. At this time, the i-th output (Mi) corresponds to the product (Xi*Yi) of the first signal (Xi) and the second signal (Yi). For instance, if N is 4, the outputs (M1, M2, M3, M4) correspond to X1*Y1, X2*Y2, X3*Y3, and X4*Y4, respectively. In this way, when the preprocessing unit (100) operates in multiplication mode, N multiplications can be performed simultaneously.
[0070] When the preprocessing unit (100) operates in multiplication mode, a shift operation may also be performed. For example, if N is 4 and the second signals (Y1, Y2, Y3, Y4) are 1000, 0100, 0010, 0001 in binary, the outputs (M1, M2, M3, M4) correspond to (Y1<<3), (Y2<<2), (Y3<<1), and (Y4<<0), respectively.
[0071] Parallel processing units can perform various operations simultaneously. For example, when N is 4, the following operations can be performed simultaneously.
[0072] M1 = X2 + X3 + X4 [Partial Sum Operation]
[0073] M2 = X1 [Displacement operation]
[0074] M3 = X3 * Y3 [Multiplication operation]
[0075] M4 = (X4 << 2) [Shift operation]
[0076] To this end, selection signals (SF1, SF2) must be set so that the first and second preprocessing units (150_1, 150_2) become summing mode, and selection signals (SF3, SF4) must be set so that the third and fourth preprocessing units (150_3, 150_4) become multiplying mode. Additionally, Y1 must be set to 1110 so that partial summing operations can be performed as described above, Y2 must be set to 0001 so that displacement operations can be performed as described above, and Y4 must be set to 0100 so that shift operations can be performed as described above.
[0077] In addition, the parallel processing device can independently change the operation of the parallel processing device in each turn by changing the selection signal (SFi) and the second signal (Yi). For example, in the first turn, after M1, M2, M3, and M4 each perform partial summation, displacement, multiplication, and shift operations as described above, in the second turn, multiplication, multiplication, partial summation, and partial summation operations can be performed as follows.
[0078] M1 = X1 * Y1 [Multiplication operation]
[0079] M2 = X2 * Y2 [Multiplication operation]
[0080] M3 = X1 + X2 + X3 + X4 [Partial Sum Operation]
[0081] M4 = X2 + X4 [Partial Sum Operation]
[0082] To this end, selection signals (SF1, SF2) must be set so that the first and second preprocessing units (150_1, 150_2) become multiplication mode, and selection signals (SF3, SF4) must be set so that the third and fourth preprocessing units (150_3, 150_4) become summing mode. Additionally, Y3 and Y4 must be set to 1111 and 1010, respectively, so that the partial summing operation can be performed as described above.
[0083] As such, the parallel processing device according to the first embodiment can perform N independent operations simultaneously, and the N operations can also be changed independently each time. This can maximize the efficiency of the parallel processing device.
[0084] If we assume there is a parallel processing unit equipped with a partial summation parallel processing unit that performs multiple partial summation operations, a displacement parallel processing unit that performs multiple displacement operations, a multiplication parallel processing unit that performs multiple multiplication operations, and a shift parallel processing unit that performs multiple shift operations, then at moments when many multiplication operations are required, the multiplication parallel processing unit can be utilized 100%, but the utilization of the partial summation parallel processing unit, the displacement parallel processing unit, and the shift parallel processing unit will be low. Also, at moments when many partial summation operations are required, the partial summation parallel processing unit can be utilized 100%, but the utilization of the displacement parallel processing unit, the multiplication parallel processing unit, and the shift parallel processing unit will be low.
[0085] In contrast, the parallel processing device according to the first embodiment can maximize the utilization of the parallel processing device by changing the operations performed by the preprocessing units (150_1, 150_2, ... 150_N) at every moment. For example, at a moment when many multiplication operations are required, most of the preprocessing units (150_1, 150_2, ... 150_N) can be utilized by setting a large portion of the preprocessing units (150_1, 150_2, ... 150_N) to perform multiplication operations and the remaining portions to perform other operations. In addition, at a moment when many partial sum operations are required, most of the preprocessing units (150_1, 150_2, ... 150_N) can be utilized by setting a large portion of the preprocessing units (150_1, 150_2, ... 150_N) to perform partial sum operations and the remaining portions to perform other operations.
[0086] FIG. 2 is a drawing for explaining an example of the i-th preprocessing unit of the first embodiment. Referring to FIG. 2, the preprocessing unit includes a selection operation unit (110_i) and a shift operation unit (120_i).
[0087] The selection operation unit (110_i) includes a plurality of demultiplexers (DM1, DM2, ... DMN). The plurality of demultiplexers (DM1, DM2, ... DMN) each output signals selected according to the bits (Yi[1], Yi[2], ... Yi[N]) of the second signal (Yi) among a plurality of first signals (X1, X2, ... XN) and 0. For example, the first demultiplexer (DM1) outputs a signal selected according to the first bit (Yi[1]) of the second signal (Yi) among the first signal (X1) and 0, the second demultiplexer (DM2) outputs a signal selected according to the second bit (Yi[2]) of the second signal (Yi) among the first signal (X2) and 0, and the Nth demultiplexer (DMN) outputs a signal selected according to the Nth bit (Yi[N]) of the second signal (Yi) among the first signal (XN) and 0.
[0088] The shift operation unit (120_i) includes a plurality of shift units (SH1, SH2, ... SHN). The plurality of shift units (SH1, SH2, ... SHN) each output signals selected according to the bits (Yi[1], Yi[2], ... Yi[N]) of the second signal (Yi) among the signals and 0 of the first signal (Xi). For example, the first shift unit (SH1) outputs a signal selected according to the first bit (Yi[1]) of the second signal (Yi) among the first signal (Xi) shifted by 0 bits (Xi<<0) and 0, the second shift unit (SH2) outputs a signal selected according to the second bit (Yi[2]) of the second signal (Yi) among the first signal (Xi) shifted by 1 bit (Xi<<1) and 0, and the Nth shift unit (SHN) outputs a signal selected according to the Nth bit (Yi[N]) of the second signal (Yi) among the first signal (Xi) shifted by (N-1) bits (Xi<<(N-1)) and 0.
[0089] Depending on the selection signal (SFi), either the selection operation unit (110_i) or the shift operation unit (120_i) operates. For example, when the selection signal (SFi) is 0, the selection operation unit (110_i) operates, and the shift operation unit (120_i) does not operate. At this time, the selection operation unit (110_i) transmits the signals output from the demultiplexers (DM1, DM2, ... DMN) to the summer (SUMi) as summer inputs (Si_1, Si_2, ... Si_N), and the shift operation unit (120_i) outputs high impedance signals. Also, when the selection signal (SFi) is 1, the selection operation unit (110_i) does not operate, and the shift operation unit (120_i) operates. At this time, the selection operation unit (110_i) outputs high impedance signals, and the shift operation unit (120_i) transmits the signals output from the shift units (SH1, SH2, ... SHN) to the summer (SUMi) as summer inputs (Si_1, Si_2, ... Si_N).
[0090] Unlike the drawing, separate demultiplexers can be added to select some of the outputs of the selection operation unit (110_i) and the shift operation unit (120_i) according to the selection signal (SFi). For example, if the selection signal (SFi) signifies the summing mode, the separate demultiplexers can transmit the outputs of the selection operation unit (110_i) to the summing unit (SUMi), and if it signifies the multiplication mode, the separate demultiplexers can transmit the outputs of the shift operation unit (120_i) to the summing unit (SUMi).
[0091] FIG. 3 is a diagram illustrating a parallel processing device according to a second embodiment. Referring to FIG. 3, the parallel processing device receives first to P inputs ({X1, Y1}, ... {X(p-1), Y(p-1)}, {Xp, Yp}, {X(p+1), Y(p+1)}, ... {XP, YP}) and outputs first to P outputs (M1, ... M(p-1), Mp, M(p+1) ... MP). Here, P represents a natural number greater than or equal to 4, and for example, P may be 1024. Also, p represents a natural number greater than or equal to 1 and less than or equal to P. The first to P inputs ({X1, Y1}, ... {X(p-1), Y(p-1)}, {Xp, Yp}, {X(p+1), Y(p+1)}, ... {XP, YP}) comprise first signals (X1, ... X(p-1), Xp, X(p+1), ... XP) and second signals (Y1, ... Y(p-1), Yp, Y(p+1) ... YP). The parallel processing device comprises a preprocessing unit (100A) and a main processing unit (200A). The parallel processing device may further comprise a delay unit (300A) and a selection unit (400A).
[0092] The preprocessing unit (100A) includes a plurality of preprocessing units (... 150_(p-1), 150_p, 150(p+1), ...). The plurality of preprocessing units (... 150_(p-1), 150_p, 150(p+1), ...) include selection operation units (... 110_(p-1), 110_p, 110_(p+1), ...) and shift operation units (... 120_(p-1), 120_p, 120_(p+1), ...). The preprocessing unit (150_p) includes a selection operation unit (110_p) and a shift operation unit (120_p).
[0093] The selection operation unit (110_p) operates when the preprocessing unit (150_p) operates in summing mode. The selection operation unit (110_p) performs the function of transmitting a first signal (Xp) corresponding to the preprocessing unit (150_p) and adjacent first signals (e.g., X(pQ / 2+1), ... X(p-1), X(p+1), ... X(p+Q / 2)) to a summer (SUMp) according to the bits (Yp[1], Yp[2], ... Yp[Q]) of the second signal (Yp). Here, Q means an even number greater than or equal to 4, and for example, Q can be 32. Also, q means a natural number greater than or equal to 1 and less than or equal to Q. At this time, the operation of the selection operation unit (110_p) can be expressed in Verilog language as, for example, as follows.
[0094] (Yp[1] ? X(pQ / 2+1) : 0) => Si_1,
[0095] (Yp[2] ? X(pQ / 2+2) : 0) => Si_2,
[0096] ...
[0097] (Yp[Q] ? X(p+Q / 2) : 0) => Si_Q,
[0098] The shift operation unit (120_p) operates when the preprocessing unit (150_p) operates in multiplication mode. The shift operation unit (120_p) performs the function of transmitting signals ((Xp<<0), (Xp<<1), ... (Xp<<(Q-1)))) in which the first signal (Xp) is shifted by 0, 1, ... (Q-1) bits to the summator (SUMp) according to the bits (Yp[1], Yp[2], ... Yp[Q]) of the second signal (Yp). At this time, the operation of the shift operation unit (120_p) can be expressed in Verilog language as an example as follows.
[0099] (Yp[1] ? (Xp << 0) : 0)=> Sp_1,
[0100] (Yp[2] ? (Xp << 1) : 0)=> Sp_2,
[0101] ...
[0102] (Yp[Q] ? (Xp << (Q-1)) : 0)=> Sp_Q,
[0103] The preprocessing unit (150_p) operates the selection operation unit (110_p) or the shift operation unit (120_p) according to the operation mode selection signal (SFp).
[0104] The main processing unit (200A) includes summers (... SUM(p-1), SUMp, SUM(p+1), ...). The p-th summer (Mp) sums the transmitted signals (Sp_1, Sp_2, ... Si_Q) and outputs the summed result as the p-th output (Mp). The operation of the main processing unit (200A) can be expressed in Verilog language as follows, for example.
[0105] ...
[0106] S(p-1)_1 + S(p-1)_2 + ... S(p-1)_Q => M(p-1),
[0107] Sp_1 + Sp_2 + ... Sp_Q => Mp, ...
[0108] S(p+1)_1 + S(p+1)_2 + ... S(p+1)_Q => M(p+1),
[0109] ...
[0110] The delay unit (300A) delays and outputs outputs (... M(p-1), Mp, M(p+1), ...) according to the clock signal (CLK). To this end, the delay unit (300A) includes a plurality of delay units (... DU(p-1), DUp, DU(p+1), ...). The signals output from the delay unit (300A) (... D(p-1), Dp, D(p+1), ...) correspond to the outputs (... M(p-1), Mp, M(p+1), ...) respectively.
[0111] The selection unit (400A) outputs signals selected according to input control signals (... SI(p-1), SIp, SI(p+1), ...) among signals transmitted from memory (not shown) (... R(p-1), Rp, R(p+1), ...) and signals output from delay unit (300A) (... D(p-1), Dp, D(p+1), ...) as first signals (... X(p-1), Xp, X(p+1), ...). For example, the memory has P banks, and the P banks can each be connected to P inputs ({X1, Y1}, {X2, Y2}, ... {XP, YP}). In addition, the memory has 2P banks, of which P banks are each connected to P first signals (X1, X2, ... XP), and the remaining P banks can be each connected to P second signals (Y1, Y2, ... YP).
Claims
Claim 1 A parallel processing device comprising N preprocessing units and N summing units, wherein N is a natural number greater than or equal to 4, wherein when the first preprocessing unit among the preprocessing units operates in a summing mode, N first signals (X1, X2, … XN) are selectively transmitted to the first summing unit among the summing units according to bits (Y1[1], Y1[2], … Y1[N]) of the first second signal (Y1), and when the first preprocessing unit operates in a multiplication mode, signals obtained by shifting the first first signal (X1) among the first signals (X1, X2, … XN) by different bits are selectively transmitted to the first summing unit according to bits (Y1[1], Y1[2], … Y1[N]) of the first second signal (Y1), and the first summing unit sums the signals transmitted from the first preprocessing unit. Claim 2 A parallel processing device according to claim 1, wherein when the second preprocessing unit among the preprocessing units operates in a summing mode, the first signals (X1, X2, … XN) are selectively transmitted to the second summing unit among the summing units according to the bits (Y2[1], Y2[2], … Y2[N]) of the second second signal (Y2), and when the second preprocessing unit operates in a multiplication mode, the signals obtained by shifting the second first signal (X2) among the first signals (X1, X2, … XN) by different bits are selectively transmitted to the second summing unit according to the bits (Y2[1], Y2[2], … Y2[N]) of the second second signal (Y2), and the second summing unit sums the signals transmitted from the second preprocessing unit. Claim 3 In claim 2, when the Nth preprocessing unit among the preprocessing units operates in a summing mode, the first signals (X1, X2, … XN) are selectively transmitted to the Nth summing unit among the summing units according to the bits (YN[1], YN[2], … YN[N]) of the Nth second signal (YN), and when the Nth preprocessing unit operates in a multiplication mode, the signals obtained by shifting the Nth first signal (XN) among the first signals (X1, X2, … XN) by different bits are selectively transmitted to the Nth summing unit according to the bits (YN[1], YN[2], … YN[N]) of the Nth second signal (YN), and the Nth summing unit sums the signals transmitted from the Nth preprocessing unit. Claim 4 A parallel processing device according to claim 3, wherein the first preprocessing unit operates in one mode selected according to the first selection signal among the summing mode and the multiplication mode, the second preprocessing unit operates in one mode selected according to the second selection signal among the summing mode and the multiplication mode, and the Nth preprocessing unit operates in one mode selected according to the Nth selection signal among the summing mode and the multiplication mode. Claim 5 A parallel processing device according to claim 4, wherein some of the preprocessing units operate in the summing mode, and simultaneously some or all of the remainder of the preprocessing units operate in the multiplication mode. Claim 6 A parallel processing device according to claim 4, wherein the value of the first selection signal, the value of the second selection signal, and the value of the Nth selection signal can be changed over time. Claim 7 In claim 3, the preprocessing units are parallel processing devices that operate in one mode selected according to a selection signal among the summing mode and the multiplication mode. Claim 8 A parallel processing device according to claim 3, further comprising N delay units, wherein the first delay unit among the delay units delays and outputs the output of the first summerator according to a clock signal, the second delay unit among the delay units delays and outputs the output of the second summerator according to the clock signal, and the Nth delay unit among the delay units delays and outputs the output of the Nth summerator according to the clock signal. Claim 9 A parallel processing device further comprising, in claim 8, a selection unit that outputs a signal selected according to a first input control signal among a first signal transmitted from memory and a signal output from the first delay unit as the first first signal (X1), outputs a signal selected according to a second input control signal among a second signal transmitted from memory and a signal output from the second delay unit as the second first signal (X2), and outputs a signal selected according to an Nth input control signal among an Nth signal transmitted from memory and a signal output from the Nth delay unit as the Nth first signal (XN).