High-Speed Prediction Processor
The hybrid analog-digital processing system with a photonic accelerator and digital equalization addresses speed and efficiency limitations in conventional computing by enabling high-throughput matrix-vector multiplication at clock frequencies beyond 10 GHz.
Patent Information
- Application Number
- JP2022580376
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-29
- Filing Date
- 2021-06-25
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2041-06-25
AI Technical Summary
Conventional computing systems face limitations in speed and efficiency due to parasitic capacitance in electrical interconnects, leading to significant delays and heat dissipation, which are particularly problematic in applications requiring high data throughput and rapid calculations.
A hybrid analog-digital processing system utilizing a photonic accelerator for matrix-vector multiplication, combined with digital equalization techniques to enhance bandwidth and reduce inter-calculation interference, allowing for clock frequencies exceeding 10 GHz.
The system achieves significantly faster data throughput and reduced interference, supporting clock frequencies up to 20 GHz by leveraging optical signals and digital equalization to overcome parasitic capacitance issues.
Smart Images

Figure 0007733682000005 
Figure 0007733682000006 
Figure 0007733682000007
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to high speed prediction processors. [Background technology]
[0002] Deep learning, machine learning, latent variable models, neural networks, and other matrix-based differentiable programs are used to solve a variety of problems, including natural language processing and object recognition in images. Solving these problems using deep neural networks generally requires long processing times to perform the required computations. When solving these problems, the most computationally intensive operations are often mathematical matrix operations, such as matrix multiplication. Summary of the Invention
[0003] In one embodiment, a hybrid analog-digital processing system includes a photonic accelerator configured to perform matrix-vector multiplication using light, the photonic accelerator exhibiting a frequency response having a first bandwidth; a plurality of analog-to-digital converters (ADCs) coupled to the photonic accelerator; and a plurality of digital equalizers coupled to the plurality of ADCs, the digital equalizers configured to set the frequency response of the hybrid analog-digital processing system to a second bandwidth greater than the first bandwidth.
[0004] In one embodiment, a method for performing a mathematical operation using a hybrid analog-digital processing system including a photonic accelerator includes: obtaining a plurality of parameters representing a first matrix; obtaining a first plurality of inputs representing a first input vector; and obtaining a second plurality of inputs representing a second input vector; at a first time, generating a first output vector by performing matrix-vector multiplication using the photonic accelerator based at least in part on the plurality of parameters and the first plurality of inputs; at a second time, subsequent to the first time, generating a second output vector by performing matrix-vector multiplication using the photonic accelerator based at least in part on the second plurality of inputs; and generating an equalized output vector by combining the first output vector with the second output vector.
[0005] Various aspects and embodiments of the present application are described with reference to the following figures. It should be appreciated that the figures are not necessarily drawn to scale. Items that appear in more than one figure are designated with the same reference numeral in the figures in which they appear. [Brief explanation of the drawings]
[0006] [Figure 1A] 1 illustrates an exemplary matrix-vector multiplication according to some embodiments. [Figure 1B] FIG. 1 is a block diagram illustrating a hybrid analog-digital processor configured to perform matrix-vector multiplication, according to some embodiments. [Figure 2] 1C is a block diagram illustrating a portion of the photonic accelerator of FIG. 1B in accordance with some embodiments. [Figure 3A] 1 is a plot illustrating a representative frequency response of a photonic accelerator, according to some embodiments. [Figure 3B] 1 is a plot illustrating a representative time-domain response of a photonic accelerator according to some embodiments. [Figure 3C]1 is a plot illustrating a representative time-domain response of a photonic accelerator according to some embodiments. [Figure 4] FIG. 1 is a block diagram illustrating a photonic accelerator including multiple digital equalizers, according to some embodiments. [Figure 5A] FIG. 1 is a block diagram illustrating a representative implementation of a digital equalizer, according to some embodiments. [Figure 5B] 1 is a plot illustrating a representative time-domain response of a digital equalizer according to some embodiments. [Figure 5C] 1 is a plot illustrating a representative frequency response of a photonic accelerator with and without equalization, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0007] I. Overview The inventors recognize and understand that conventional computing systems have not adequately kept up with the ever-increasing demand for data throughput in modern applications. Conventional electronic processors face severe limitations in speed and efficiency, primarily due to the presence of parasitic capacitance inherent in electrical interconnects. Every wire and transistor in an electrical processor's circuitry has resistance, inductance, and capacitance, which contribute to propagation delays and power dissipation for any electrical signal. For example, connecting multiple processor cores to memory and / or connecting a single processor core uses conductive traces with non-zero impedance. High impedance values limit the maximum speed at which data can be transferred through the traces with a negligible bit error rate. Most conventional processors cannot support clock frequencies greater than 2-3 GHz.
[0008] In applications where time delays are crucial, such as high-frequency stock trading, delays of even a few hundredths of a second can render an algorithm unusable. For processes requiring billions of operations by billions of transistors, these delays represent significant lost time. In addition to the inefficient speed of electrical circuits, the heat generated by energy dissipation due to circuit impedance also poses a barrier to developing electrical processors.
[0009] In digital computers, the output of a calculation must fully settle to its final one / zero value before it is sampled. Otherwise, sampling the output of a calculation before the output has settled can lead to errors. Typically, when the output of a calculation exceeds the linear switching threshold of a transistor, a sample is taken and the next calculation cannot begin until the previous calculation has fully settled and been sampled. This limits the throughput of a digital computer.
[0010] In contrast, analog computers do not operate in the saturation region of transistors. Instead, they operate on a continuum of values. Such values must be resolved to a level of precision that is defined for a particular window of time. Certain analog computers are designed to work in conjunction with digital systems (e.g., digital memory and processors). These systems are called hybrid analog-digital computers. The digital system multiplies the output of the analog computer by a sum of 2 bwhere b is the bit precision of the output. Analog computers, like digital processors, are also characterized by finite bandwidth due to the presence of parasitic capacitances. Finite bandwidth increases the time it takes for the analog computer's output to settle to the desired discretized level. For example, analog computers with a unipolar response are characterized by a time constant τ, which sets the time scale over which the output signal rises and falls. Typically, e-times (the time it takes for the signal to rise and fall by e times) is used to define this time constant τ. In such systems, the signal will not settle exactly at the desired discretized level, which is 1 / 2 the value of the desired discretized level. b It may take a multiple of τ to settle to 1 / 2. b The performance of such a system is limited if a subsequent calculation is not started until the level of has settled to its final value. However, starting a new calculation too early can cause interference between two overlapping calculations, which can result in sampling the wrong output (an effect referred to herein as "inter-calculation-interference").
[0011] As a result, conventional computers, whether digital or analog in nature, are throughput limited. The present inventors have developed techniques for increasing the data throughput of hybrid analog-digital computing systems. The techniques developed by the present inventors and described herein are used to improve the data throughput of hybrid analog-digital computing systems. bThis involves starting a new set of operands before settling to the final value of the level. In some embodiments, this can be achieved using digital equalization. Digital equalization involves inverting the channel characteristics of an analog processor to allow for faster settling of the received signal, thereby enabling inter-computation interference disambiguation. Digital equalization may be performed on the transmitter side of the computation, the receiver side of the computation (or both). Several types of digital equalization techniques may be used, including, but not limited to, pre-emphasis equalization, continuous time linear equalization (CTLE), and discrete feedback equalization (DFE). Processors utilizing the digital equalization techniques described herein can be fast enough to support clock frequencies greater than 10 GHz, 15 GHz, or even 20 GHz, representing a significant improvement over conventional processors.
[0012] II. Photonic Hybrid Processor The inventors have recognized and understood that the use of optical signals (in place of, or in combination with, electrical signals) overcomes some of the problems associated with electronic computing. Optical signals travel in a medium in which light travels at the speed of light. Because of this, the latency of optical signals is much less limiting than electrical propagation delays. Additionally, power is not dissipated as the optical signal travels greater distances, allowing for new topologies and processor layouts that were not possible using electrical signals. Thus, photonic processors offer much better speed and efficiency performance than traditional electronic processors.
[0013] Some embodiments relate to photonic processors designed to perform machine learning algorithms or other types of data-intensive computations. Certain machine learning algorithms (e.g., support vector machines, artificial neural networks, and probabilistic graphical model learning) rely heavily on linear transformations on multidimensional arrays / tensors. The simplest linear transformation is matrix-vector multiplication, which takes approximately O(N 2 ), where N is the number of dimensions of a square matrix being multiplied by a vector of the same dimension. General matrix-matrix (GEMM) operations are widespread in software algorithms, including those for graphical processing, artificial intelligence, neural networks, and deep learning. In modern computers, GEMM computations are typically performed using transistor-based systems such as GPUs or systolic array systems.
[0014] FIG. 1A illustrates matrix-vector multiplication according to some embodiments. Matrix A is referred to herein as the "input matrix" or simply "matrix," and individual elements of matrix A are referred to herein as "matrix values" or simply "matrix parameters." Vector X is referred to herein as the "input vector," and individual elements of vector X are referred to herein as "input values" or simply "inputs." Vector Y is referred to herein as the "output vector," and individual elements of vector Y are referred to herein as "output values" or simply "outputs." In this example, A is an N×N matrix, although embodiments of the present application are not limited to square matrices or to any particular dimensionality. In the context of artificial neural networks, matrix A can be a weight matrix, or a block of submatrices of weight tensors, or an activation (batch) matrix, or a block of submatrices of (batch) activation tensors, among other possible examples. Similarly, input vector X can be, for example, a vector of weight tensors, or a vector of activation tensors.
[0015] The matrix-vector multiplication in Figure 1A can be decomposed into scalar multiplications and scalar additions. For example, the output value y i (In this case, i=1,2…N) is the input value x1,x2…x N It can be calculated as a linear combination of y i To obtain , use scalar multiplication (e.g., A i1 Multiply by x1 and A i2 multiplying by x2) and scalar addition (e.g., A i1 x1 to A i2 In some embodiments, scalar multiplication, scalar addition, or both may be performed in the optical domain.
[0016] FIG. 1B illustrates a hybrid analog-digital processor 10 implemented using photonic circuits, according to some embodiments. The hybrid processor 10 may be configured to perform matrix-vector multiplication (e.g., of the type shown in FIG. 1A). The hybrid processor 10 includes a digital controller 100 and a photonic accelerator 150. The digital controller 100 operates in the digital domain, while the photonic accelerator 150 operates in the analog photonic domain. The digital controller 100 includes a digital processor 102 and a memory 104. The photonic accelerator 150 includes an optical encoder module 152, an optical computation module 154, and an optical receiver module 156. Digital-to-analog (DAC) modules 106 and 108 convert digital data to analog signals. An analog-to-digital (ADC) module 110 converts analog signals to digital values. Thus, the DAC / ADC modules provide an interface between the digital and analog domains. In this example, DAC module 106 generates N analog signals (one for each entry in the input vector), DAC module 108 generates N×N analog signals (one for each entry in the matrix), and ADC module 110 receives N analog signals (one for each entry in the output vector). Matrix A is square in this example, but in some embodiments the matrix may be rectangular, such that the size of the output vector is different from the size of the input vector.
[0017] Hybrid processor 10 receives an input vector represented by a set of input bit strings as input from an external processor (e.g., a CPU) and generates an output vector represented by a set of output bit strings. For example, if the input vector is an N-dimensional vector, the input vector can be represented by N separate bit strings, each representing a component of the vector. The input bit strings can be received as electrical signals from the external processor, and the output bit strings can be sent as electrical signals to the external processor. In some embodiments, digital processor 102 does not necessarily output an output bit string after each iteration of the process. Instead, digital processor 102 may use one or more output bit strings to determine a new input bit stream to feed through the components of hybrid processor 10. In some embodiments, the output bit string itself can be used as an input bit string for a subsequent iteration of the process implemented by hybrid processor 10. In other embodiments, multiple output bit streams are combined in various ways to determine a subsequent input bit string. For example, one or more output bit strings can be summed together as part of determining a subsequent input bit string.
[0018] The DAC module 106 is configured to convert the input bit string into an analog signal. The optical encoder module 152 is configured to convert the analog signal into optically encoded information that is processed by the optical computation module 154. The information may be encoded in the amplitude, phase, and / or frequency of the optical pulses. Accordingly, the optical encoder module 152 may include an optical amplitude modulator, an optical phase modulator, and / or an optical frequency modulator. In some embodiments, the optical signal represents the value and sign of the associated bit string as the amplitude and phase of the optical pulses. In some embodiments, the phase may be limited to a binary selection of either a zero phase shift or a π phase shift, representing positive and negative values, respectively. Embodiments are not limited to real input vector values. Complex vector components may be represented when encoding the optical signal, for example, by using more than two phase values.
[0019] Optical encoder module 152 outputs N separate optical pulses that are transmitted to optical computation module 154. Each output of optical encoder module 152 is coupled one-to-one to an input of optical computation module 154. In some embodiments, optical encoder module 152 may be located on the same substrate as optical computation module 154 (e.g., optical encoder module 152 and optical computation module 154 are on the same chip). In such embodiments, the optical signals may be transmitted from optical encoder module 152 to optical computation module 154 over a waveguide, such as a silicon photonic waveguide. In other embodiments, optical encoder module 152 may be located on a separate substrate from optical computation module 154. In such embodiments, the optical signals may be transmitted from optical encoder module 152 to optical computation module 154 using optical fiber.
[0020] Optical computation module 154 performs the multiplication of input vector X by matrix A. In some embodiments, optical computation module 154 includes multiple optical multipliers, each configured to perform a scalar multiplication in the optical domain between entries of the input vector and entries of matrix A. Optionally, optical computation module 154 may further include an optical adder for adding together the results of the scalar multiplications in the optical domain. Alternatively, this addition may be performed electrically. For example, optical receiver module 156 may generate a voltage obtained by integrating (over time) the photocurrent received from the photodetector.
[0021] Optical computation module 154 outputs N separate optical pulses that are transmitted to optical receiver module 156. Each output of optical computation module 154 is one-to-one coupled to an input of optical receiver module 156. In some embodiments, optical computation module 154 may be located on the same substrate as optical receiver module 156 (e.g., optical computation module 154 and optical receiver module 156 are on the same chip). In such embodiments, optical signals may be transmitted from optical computation module 154 to optical receiver module 156 using silicon photonic waveguides. In other embodiments, optical computation module 154 may be located on a separate substrate from optical receiver module 156. In such embodiments, optical signals may be transmitted from photonic processor 103 to optical receiver module 156 using optical fiber.
[0022] The optical receiver module 156 receives the N optical pulses from the optical computing module 154. Each optical pulse is then converted to an electrical analog signal. In some embodiments, the intensity and phase of each optical pulse is detected by a photodetector within the optical receiver module. The electrical signals representing these measurements are then converted to the digital domain using the ADC module 110 and sent back to the digital processor 102.
[0023] Digital processor 102 controls optical encoder module 152, optical computation module 154, and optical receiver module 156. Memory 104 can be used to store input and output bit strings from optical receiver module 156, as well as measurement results. Memory 104 also stores executable instructions that, when executed by digital processor 102, control optical encoder module 152, optical computation module 154, and optical receiver module 156. Memory 104 may also include executable instructions that cause digital processor 102 to determine a new input vector to send to the optical encoder based on a set of one or more output vectors determined by measurements performed by optical receiver module 156. In this manner, digital processor 102 can control the iterative process in which an input vector is multiplied by multiple matrices by adjusting settings of optical computation module 154 and feeding detected information from optical receiver module 156 back to optical encoder module 152. In this way, the output vector sent by hybrid processor 10 to an external processor can be the result of multiple matrix multiplications rather than just one.
[0024] Figure 2 shows a portion of photonic accelerator 150 in more detail, according to some embodiments. More specifically, Figure 2 shows circuitry for computing y1, the first entry of output vector Y. For simplicity, in this example, the input vector has only two entries, x1 and x2. However, the input vector can have any suitable size.
[0025] DAC module 106 includes DAC 206, DAC module 108 includes DAC 208, and ADC module 110 includes ADC 210. DAC 206 generates electrical analog signals (e.g., voltages or currents) based on the values they receive. For example, voltage Vx1 represents value x1, voltage Vx2 represents value x2, and voltage V A11 is the value A 11 represents the voltage V A12is the value A 12 The optical encoder module 152 includes an optical encoder 252, the optical computation module 154 includes an optical multiplier 154 and an optical adder 255, and the optical receiver module 156 includes an optical receiver 256.
[0026] The light source 402 generates light S0. The light source 402 may be implemented in any suitable manner. For example, the light source 402 may include a laser, such as a vertical cavity surface emitting laser (VCSEL) edge-emitting laser, examples of which are described in more detail below. In some embodiments, the light source 402 may be configured to generate light at multiple wavelengths, which enables light processing utilizing wavelength division multiplexing (WDM), as described in more detail below. For example, the light source 402 may include multiple laser cavities, each individually sized to generate a different wavelength.
[0027] The optical encoders 252 encode an input vector into multiple optical signals. For example, one optical encoder 252 encodes an input value x1 into an optical signal S(x1), and another optical encoder 252 encodes an input value x2 into an optical signal S(x2). The input values x1 and x2 provided by the digital processor 102 are digital signed real numbers (e.g., having floating-point or fixed-point digital representations). The optical encoders modulate the light S0 based on their respective input voltages. For example, the optical encoder 404 modulates the amplitude, phase, and / or frequency of light to generate the optical signal S(x1), and the optical encoder 406 modulates the amplitude, phase, and / or frequency of light to generate the optical signal S(x2). The optical encoders may be implemented using any suitable optical modulator, including, for example, an optical intensity modulator. Examples of such modulators include Mach-Zehnder modulators (MZM), Franz-Keldysh modulators (FKM), resonant modulators (e.g., ring-based or disk-based), nano-opto-electro-mechanical-system (NOEMS) modulators, and the like.
[0028] The optical multipliers are designed to generate a signal representing the product between an input value and a matrix value. For example, one optical multiplier 254 may be configured to multiply an input value x1 by a matrix value A 11 The signal S(A 11 x1), and another optical multiplier 254 multiplies the input value x2 by the matrix value A 12 The signal S(A 12x2). Examples of optical multipliers include Mach-Zehnder modulators (MZM), Franz-Keldysh modulators (FKM), resonant modulators (e.g., ring-based or disk-based), nano-opto-electro-mechanical systems (NOEMS) modulators, etc. In one example, the optical multiplier may be implemented using a modulatable detector. A modulatable detector is a photodetector that has a characteristic that can be modulated using an input voltage. For example, the modulatable detector can be a photodetector that has a responsivity that can be modulated using an input voltage. In this example, an input voltage (e.g., V A11 ) sets the responsivity of the photodetector. As a result, the output of the modulatable detector depends not only on the amplitude of the input optical signal, but also on the input voltage. When operating the modulatable detector in its linear region, the output of the modulatable detector depends on the product of the amplitude of the input optical signal and the input voltage (thereby achieving the desired multiplication function).
[0029] The optical adder 412 generates an electronic analog signal S(A 11 x1) and S(A 12 x2), as well as light S0′ (generated by light source 414), and A 11 x1, A 12 x2 and the optical signal S(A 11 x1+A 12 x2).
[0030] The optical receiver 256 receives the optical signal S(A 11 x1+A 12 x2) based on the sum A 11 x1+A 12 x2. In some embodiments, optical receiver 256 includes a coherent detector and a transimpedance amplifier. The coherent detector generates an output indicative of the phase difference between the waveguides of the interferometer. The phase difference is calculated by summing A 11 x1+A 12 Since the output of the coherent detector is a function of x2, the output of the coherent detector also represents the sum. The ADC converts the output of the coherent receiver to the output value y1 = A 11 x1+A 12The output value y1 may be provided as an input back to the digital processor 102 so that the output value can be used for further processing.
[0031] III. Digital Equalization The inventors recognize and understand that hybrid optical-based processors of the type described in the previous section, while substantially faster than conventional fully digital processors, are still bandwidth-limited. Hybrid optical-based processors of the type described herein are faster than fully digital processors because some of the conductive traces are replaced with optical waveguides, which do not suffer from parasitic capacitance issues. Nevertheless, these hybrid optical-based processors still include some conductive traces to accommodate electrical signals that control the operation of the photonic processor. Unfortunately, such conductive traces exhibit parasitic capacitance. The longer the conductive traces, the greater the parasitic capacitance, and the lower the bandwidth of the hybrid optical-based processor. For example, electrical paths longer than 1 cm may limit the processor's bandwidth to less than 3 GHz. Figure 3A shows a typical amplitude response of a hybrid processor as a function of frequency. In this example, the hybrid processor exhibits a single-pole response with a bandwidth between 2 GHz and 3 GHz.
[0032] The effect of having such a response is illustrated in Figure 3B, which shows the time-domain response of a hybrid processor, according to some embodiments. The plot shows the target settling level (the final desired level) and the actual response of the processor as a function of time. As a result of the processor's limited bandwidth, it takes several nanoseconds for the response to reach the desired level. A characteristic time constant τ, which is inversely proportional to the processor's bandwidth, sets the rate at which the output signal rises and falls. Typically, e-times (the time it takes for the signal to rise and fall by e times) is used to define this time constant τ. In such systems, it is difficult to ensure that the signal is exactly at the desired discretized level, i.e., 1 / 2 the value of the desired discretized level. bIt may take a multiple of τ (where b represents the bit precision) to settle to 1 / 2. b The performance of such a system is limited if a subsequent calculation is not started until the level of has settled to its final value. However, starting a new calculation too early can cause interference between two overlapping calculations, which can lead to inter-calculation interference.
[0033] 3C is a plot showing the analog response of a hybrid processor as a function of time and the digital data for a non-return-to-zero coded output, according to some embodiments. As shown, the analog response does not change quickly enough to accurately replicate the levels of the digital data, which can lead to errors.
[0034] The inventors have developed a technique that allows a new set of operands to be initiated before the previous set of operands has settled to its final value. The technique developed by the inventors involves determining the channel characteristics of a hybrid processor (e.g., the frequency response of a particular signal path in the hybrid processor) and equalizing the channel characteristics to extend the bandwidth of the hybrid processor. Channel equalization allows for faster settling of the received signal, thereby enabling disambiguation of inter-computation interference. In some embodiments, channel equalization may be performed using a digital filter. Several types of digital equalization techniques may be used, including, but not limited to, pre-emphasis equalization, continuous-time linear equalization (CTLE), and discrete feedback equalization (DFE). Processors utilizing the digital equalization techniques described herein can be fast enough to support clock frequencies exceeding 10 GHz, 15 GHz, or even 20 GHz, representing a significant improvement over conventional processors.
[0035] Figure 4 is a block diagram of a portion of a photonic accelerator including multiple channels, according to some embodiments. In this example, each channel is arranged in the manner described in connection with Figure 2. As such, each channel includes two or more optical encoders 252, two or more optical multipliers 254, an optical adder 255, an optical receiver 256, and an ADC 210. In some embodiments, the photonic accelerator may be designed to perform matrix-vector multiplication on very large matrices, for example, on the order of 256 x 256, 512 x 512, or 1024 x 1024 (although the matrices need not be square). In these examples, the photonic accelerator may include 256, 512, or 1024 channels, one channel for each row of the matrix.
[0036] Each channel includes a digital equalizer 400 coupled to the output of the ADC. The input to the equalization channel is specified as y[n] (where n=1, 2...N is a discretized time variable) and the output is specified as w[n]. In some embodiments, the equalization channel 400 generates the output w[n] by computing a linear combination of the previous state samples y[n], y[n-1], y[n-2], etc. The linear combination can be expressed as:
[0037]
number
[0038] Coefficient c i may be determined in any of a number of ways based on the characteristics of the channel of the hybrid processor. In some embodiments, the coefficient c ican be obtained by exciting the hybrid processor with a known excitation and sampling the output at a desired rate. In some such embodiments, the coefficients can be calculated based on the following equation (where the last coefficient is chosen such that the sum of all coefficients equals 1):
[0039]
number
[0040] In some embodiments, the digital equalizer can suppress low-frequency components of the output signal and amplify high-frequency components. This provides greater time-domain separation of potentially interfering calculations and, therefore, the ability to perform calculations significantly faster. FIG. 5C is a plot showing the amplitude response of a photonic accelerator with and without equalization, according to some embodiments. As shown in this figure, the response of the unequalized photonic accelerator exhibits a bandwidth between 2 GHz and 3 GHz (e.g., a 3 dB bandwidth). The limited bandwidth may be due, for example, to parasitic capacitances present in the circuitry controlling the photonic accelerator. The digital equalizer suppresses low frequencies and amplifies high frequencies (in this example, the peak is between 15 GHz and 20 GHz). As a result, the equalized photonic accelerator exhibits a larger bandwidth than the unequalized one. For example, the bandwidth of the equalized photonic accelerator may be between 10 GHz and 30 GHz. By leveraging the bandwidth expansion available thanks to digital equalization, the hybrid processor can be timed using a clock having a frequency greater than the bandwidth of the unequalized photonic accelerator while still allowing for disambiguation of inter-computation interference. Figure 5C further illustrates a clock having a frequency between the bandwidth of the unequalized photonic accelerator and the bandwidth of the equalized photonic accelerator.
[0041] It should be appreciated that the equalization function that generates w[n] need not be a linear function. In some embodiments, the equalization function can be nonlinear. In some embodiments, the digital equalizer may be designed to implement a discrete linear time-varying system and to convert the transfer function to a linear constant coefficient difference equation using a z-transform. Consider the following transfer function:
[0042]
number
[0043]
number
[0044] VI. Additional Comments Having thus described several aspects and embodiments of the technology of the present application, it should be recognized that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. Accordingly, it should be understood that the foregoing embodiments have been presented by way of example only, and that, within the scope of the appended claims and their equivalents, embodiments of the invention may be practiced otherwise than as specifically described. Additionally, any combination of two or more features, systems, articles, materials, and / or methods described herein is within the scope of the present disclosure, provided such features, systems, articles, materials, and / or methods are not mutually inconsistent.
[0045] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of a method may be ordered in any suitable manner. Accordingly, embodiments may be constructed in which acts are performed in an order different from that illustrated, which may include performing some acts simultaneously even though they are shown as sequential acts in an exemplary embodiment.
[0046] Definitions defined and used herein should be understood to take precedence over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms. As used in this specification and the claims, the indefinite articles "a" and "an" should be understood to mean "at least one," unless expressly indicated otherwise.
[0047] As used in this specification and claims, the term "and / or" should be understood to mean "either or both" of the conjoined elements, i.e., elements that may be present conjunctively or disjunctively.
[0048] As used in this specification and claims, the phrase "at least one," when referring to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but does not necessarily include at least one of each and every element specifically listed in the list of elements, nor does it exclude any combinations of elements in the list of elements. This definition also allows for elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related or unrelated to those specifically identified elements, may optionally be present.
[0049] The terms "approximately," "substantially," and "about" may be used in some embodiments to mean within ±20% of a target value. The terms "approximately," "substantially," and "about" may include the target value.
[0050] In the claims, the use of order terms such as "first," "second," "third," etc. to modify claim elements does not, in itself, imply a precedence, priority, or order of one element in a claim over another element, or a temporal order in which method actions are performed, but is merely used as a marker to distinguish an element in a claim having a particular name from an element in another claim having the same name (apart from the use of order terms) to distinguish between elements in the claims.
[0051] Also, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," and variations thereof herein is meant to encompass not only the items listed thereafter and equivalents thereof, but also additional items. The technical ideas that can be understood from the above embodiment will be described below. [Appendix 1] 1. A hybrid analog-digital processing system comprising: an analog processing unit configured to perform matrix-vector multiplication, the analog processing unit exhibiting a frequency response having a first bandwidth; a plurality of analog-to-digital converters (ADCs) coupled to the analog processing unit; a plurality of digital equalizers coupled to the plurality of ADCs, the digital equalizers configured to set a frequency response of the hybrid analog-digital processing system to a second bandwidth greater than the first bandwidth. [Appendix 2] 10. The hybrid analog-digital processing system of claim 1, further comprising clock circuitry configured to time the digital equalizer using a clock having a frequency between the first bandwidth and the second bandwidth. [Appendix 3] 10. The hybrid analog-digital processing system of claim 1, wherein at least one of the plurality of digital equalizers includes a plurality of registers configured to store a sample history and to provide an output based at least in part on the sample history. [Appendix 4] 2. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer is configured to perform continuous-time linear equalization (CTLE). [Appendix 5] 2. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer includes a finite impulse response (FIR) filter, the FIR filter including a plurality of respective coefficients configured to equalize the frequency response of the analog processing unit. [Appendix 6] 6. The hybrid analog-digital processing system of claim 5, wherein each of the plurality of coefficients is determined by passing one or more known inputs through the analog processing unit. [Appendix 7] 7. The hybrid analog-digital processing system of claim 6, wherein the known input comprises a step input. [Appendix 8] 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes at least one electrical path having a length greater than 1 cm, and the first bandwidth is less than 3 GHz. [Appendix 9] 2. The hybrid analog-digital processing system of claim 1, wherein the second bandwidth is between 10 GHz and 30 GHz. [Appendix 10] 2. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer is configured to perform discrete feedback equalization (DFE). [Appendix 11] 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes a plurality of modulatable detectors, and the analog processing unit is configured to perform matrix-vector multiplication using the plurality of modulatable detectors. [Appendix 12] 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes a plurality of optical adders, and the analog processing unit is configured to perform matrix-vector multiplication using the plurality of optical adders. [Appendix 13] 1. A method for performing mathematical operations using a hybrid analog-digital processing system including an analog processing unit, comprising: obtaining a plurality of parameters representing a first matrix; obtaining a first plurality of inputs representing a first input vector and a second plurality of inputs representing a second input vector; At a first time, generating a first output vector by performing matrix-vector multiplication using the analog processing unit based at least in part on the plurality of parameters and the first plurality of inputs; At a second time subsequent to the first time, generating a second output vector by performing a matrix-vector multiplication using the analog processing unit based at least in part on the second plurality of inputs; generating an equalized output vector by combining the first output vector with the second output vector. [Appendix 14] 14. The method of claim 13, wherein combining the first output vector with the second output vector comprises linearly combining the first output vector with the second output vector. [Appendix 15] 15. The method of claim 14, further comprising determining a plurality of coefficients by passing one or more known inputs through the analog processing unit, and wherein linearly combining the first output vector with the second output vector comprises linearly combining the first output vector with the second output vector using the plurality of coefficients. [Appendix 16] 16. The method of claim 15, wherein passing one or more known inputs through the analog processing unit comprises passing one or more step inputs through the analog processing unit. [Appendix 17] 14. The method of claim 13, further comprising clocking the hybrid analog-digital processing system using a clock having a frequency greater than a bandwidth of the analog processing unit. [Appendix 18] 18. The method of claim 17, wherein the bandwidth is less than 3 GHz and the frequency of the clock is between 10 GHz and 30 GHz. [Appendix 19] 14. The method of claim 13, wherein combining the first output vector with the second output vector includes performing continuous-time linear equalization (CTLE). [Appendix 20] 14. The method of claim 13, wherein combining the first output vector with the second output vector includes performing discrete feedback equalization (DFE). [Appendix 21] 14. The method of claim 13, wherein performing matrix-vector multiplication using an analog processing unit comprises performing matrix-vector multiplication using a plurality of modulatable detectors.
Claims
1. 1. A hybrid analog-digital processing system comprising: an analog processing unit configured to perform matrix-vector multiplication, the analog processing unit exhibiting a frequency response having a first bandwidth; a plurality of analog-to-digital converters (ADCs) coupled to the analog processing unit; a plurality of digital equalizers coupled to the plurality of ADCs, the digital equalizers configured to equalize channel characteristics of the hybrid analog-digital processing system to set a frequency response of the hybrid analog-digital processing system to a second bandwidth greater than the first bandwidth.
2. 10. The hybrid analog-digital processing system of claim 1, further comprising clock circuitry configured to time said digital equalizer using a clock having a frequency between said first bandwidth and said second bandwidth.
3. 10. The hybrid analog-digital processing system of claim 1, wherein at least one of the plurality of digital equalizers includes a plurality of registers configured to store a history of samples and to provide an output based at least in part on the history of samples.
4. 10. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer is configured to perform continuous-time linear equalization (CTLE).
5. 2. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer includes a finite impulse response (FIR) filter, the FIR filter including a respective plurality of coefficients configured to equalize the frequency response of the analog processing unit.
6. 6. The hybrid analog-digital processing system of claim 5, wherein the respective plurality of coefficients are determined by passing one or more known inputs through the analog processing unit.
7. 7. The hybrid analog-digital processing system of claim 6, wherein the known input comprises a step input.
8. 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes at least one electrical path having a length greater than 1 cm, and the first bandwidth is less than 3 GHz.
9. 10. The hybrid analog-digital processing system of claim 1, wherein the second bandwidth is between 10 GHz and 30 GHz.
10. 10. The hybrid analog-digital processing system of claim 1, wherein the digital equalizer is configured to perform discrete feedback equalization (DFE).
11. 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes a plurality of modulatable detectors, and wherein the analog processing unit is configured to perform matrix-vector multiplication using the plurality of modulatable detectors.
12. 2. The hybrid analog-digital processing system of claim 1, wherein the analog processing unit includes a plurality of optical adders, and wherein the analog processing unit is configured to perform matrix-vector multiplication using the plurality of optical adders.
Citation Information
Patent Citations
Optoelectronic computing systems
US20190370652A1