Optical circuit and optical computing device method for executing artificial intelligence accelerator
By using optical circuits and time multiplexing methods for optical computing devices, the problem of high power consumption of electronic circuits in artificial intelligence computing is solved, and low-energy consumption and efficient computing-intensive operations are achieved.
Patent Information
- Application Number
- CN202510130322.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-02-05
- Publication Date
- 2025-08-01
AI Technical Summary
Existing electronic circuits face high power consumption problems in computing-intensive applications such as artificial intelligence and deep learning. Traditional electronic-based MAC unit computing frameworks are difficult to effectively reduce energy consumption and improve computing efficiency.
Optical calculation devices are used to realize multiplication and accumulation operations through optical circuits, and optical components such as source pixels, modulator pixels and detector pixels are used for optical signal processing. Combined with the time multiplexing method, it replaces the traditional spatial multiplexing method and reduces the demand for large-size pixel matrix.
It significantly reduces the energy consumption of computing-intensive devices, improves computing speed, reduces time delay, and is compatible with existing manufacturing processes to achieve efficient optical MAC computing.
Smart Images

Figure CN120405990A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to optical circuits and optical computing methods for performing artificial intelligence accelerators. Background Art
[0002] Artificial intelligence (AI) is a technology that has emerged in recent years and has become a powerful tool to simulate human intelligence by programming machines to think and act like humans. Artificial intelligence has attracted wide attention because its application scenarios are more prevalent than any other previous high-tech and can be used in various applications and industries. An AI accelerator is a hardware device built from processors, memories, and interface elements for efficiently processing AI workloads such as neural networks. Summary of the Invention
[0003] According to one aspect of embodiments of the present application, there is provided an optical circuit, including: a source pixel configured to generate a plurality of input optical pulses having time intervals; a modulator pixel optically coupled to the source pixel and configured to modulate the plurality of input optical pulses to generate a plurality of modulated optical pulses; a detector pixel optically coupled to the modulator pixel and configured to generate charges in response to the plurality of modulated optical pulses; and a controller configured to electrically control the intensity of each of the plurality of input optical pulses and the modulation level of the modulator pixel to perform a multiply-accumulate operation.
[0004] According to another aspect of embodiments of the present application, there is provided an optical circuit, including: an M-by-N matrix of source pixels, each source pixel configured to generate a plurality of input optical pulses having time intervals, where M and N are natural numbers; an M-by-N matrix of modulator pixels optically coupled to the source pixel matrix, each modulator pixel configured to modulate the plurality of input optical pulses with a separate transmittance or reflectance to generate a plurality of modulated optical pulses; an M-by-N matrix of detector pixels optically coupled to the modulator pixel matrix, each detector pixel configured to generate charges in response to the plurality of modulated optical pulses from a corresponding one of the modulator pixels; and a controller configured to electrically control the intensity of each of the plurality of input optical pulses of each source pixel and the separate transmittance or reflectance of each modulator pixel to perform a multiply-accumulate operation.
[0005] In yet another aspect of the embodiments of the present application, a method for performing optical computing of an artificial intelligence accelerator is provided. The method includes: generating, by a first source pixel, K first input optical pulses having a time interval, where K is a natural number; guiding the K first input optical pulses to a first modulator pixel, thereby receiving, at the time interval, K first modulated optical pulses, where the first modulator pixel is configured to modulate in correspondence with K first weights and K first output optical pulses; receiving, by a first detector pixel, the K first modulated optical pulses; and accumulating, by an integrator in response to the K first modulated optical pulses, the charge of the first detector pixel to generate a result of a multiply-accumulate operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various aspects of the present invention are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, in accordance with standard practice in the industry, the various components are not drawn to scale and are for illustrative purposes only. In fact, for clarity of discussion, the dimensions of the various components may be arbitrarily increased or decreased.
[0007] Figure 1 A block diagram of a neural network according to some embodiments is shown.
[0008] Figure 2 A block diagram of an optical computing device for optical computing according to some comparative embodiments of the present disclosure is shown.
[0009] Figure 3A and Figure 3B A block diagram of a spatial light multiplexer (SLM) according to some embodiments of the present disclosure is shown.
[0010] Figure 4A A schematic diagram of an optical computing device according to some comparative embodiments of the present disclosure is shown.
[0011] Figure 4B A schematic diagram of an optical computing device according to some comparative embodiments of the present disclosure is shown.
[0012] Figure 5A A schematic diagram of an optical computing device according to some embodiments of the present disclosure is shown.
[0013] Figure 5B A schematic diagram of an optical computing device according to some embodiments of the present disclosure is shown.
[0014] Figure 6 A schematic diagram of an optical computing device according to some embodiments of the present disclosure is shown.
[0015] Figure 7A and Figure 7B A block diagram of a semiconductor optical computing device according to some embodiments of the present disclosure is shown.
[0016] Figure 8 A schematic flowchart showing a method of operating an optical computing device according to some embodiments of the present disclosure.
[0017] In the following detailed description, for purposes of illustration, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. However, it will be apparent that one or more embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown schematically in order to simplify the drawings. In addition, the same reference numerals in different drawings indicate similar features, so that when these features are first introduced in the present disclosure, a detailed explanation of the similar features may be provided, and subsequent repetitions may not be made. Detailed Description
[0018] The following disclosure provides many different embodiments or examples for implementing different features of the present invention. Specific embodiments or examples of components and arrangements are described below to simplify the present invention. Of course, these are merely examples and are not intended to be limiting. For example, in the following description, forming a first component above or on a second component may include embodiments where the first component and the second component are in direct contact, and may also include embodiments where additional components may be formed between the first component and the second component, such that the first component and the second component may not be in direct contact. In addition, the present invention may repeat reference numerals and / or letters in various examples. This repetition is for the purpose of simplicity and clarity, and does not in itself indicate a relationship between the various embodiments and / or configurations discussed.
[0019] In addition, for ease of description, spatially relative terms such as "below", "beneath", "lower", "above", "upper", etc. may be used herein to describe the relationship of one element or component to another as shown in the figures. In addition to the orientation shown in the figures, the spatially relative terms are intended to include different orientations of the device in use or operation. The device may be otherwise oriented (rotated 90 degrees or in other orientations), and the spatially relative descriptors used herein may be interpreted accordingly.
[0020] As used herein, although terms such as "first", "second", and "third" describe various elements, components, regions, layers, and / or parts, these elements, components, regions, layers, or parts should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or part from another. Unless the context clearly indicates otherwise, the terms "first", "second", and "third" as used herein do not imply an order or sequence.
[0021] The electronics industry has an increasing demand for smaller and faster electronic devices that can support more and more complex and sophisticated functions simultaneously. To meet these demands, the integrated circuit (IC) industry has a continuous trend of manufacturing low-cost, high-performance, and low-power ICs. To achieve these goals, device performance is mainly improved and related costs are reduced by shrinking the IC size (e.g., the minimum IC component size). However, when approaching its physical limits using electronic engineering methods, improvements may not be as rapid as before.
[0022] One of the major problems commonly faced by most existing electronic circuits is the increasing power consumption in computationally intensive applications such as artificial intelligence, deep learning, and machine learning, which require performing a large amount of data calculations in a short time. Such computational frameworks are typically implemented as simulating a neural network constructed from multiple computational units, such as a convolutional neural network (CNN), a deep neural network (DNN), which are referred to as multiply-accumulate (MAC) units in this paper. The more MAC units that a computationally intensive device can utilize, the faster or more computational tasks it can achieve. However, the highly increased power consumption of computationally intensive devices formed by MAC units will hinder the application and popularization of computationally intensive semiconductor devices.
[0023] In this regard, the present disclosure proposes implementing optical computing in optical / photon devices to replace the electronics-based MAC units, where the input activations of the neural network are represented by multiple light beams or pulses, the weights of the neural network are implemented by the filtering operations of one or more spatial light modulators (SLMs), and the output activations of the neural network are represented by multiple photodetectors. Thus, the energy consumption for performing MAC operations by optical computing is greatly reduced. In addition, the proposed optical MAC architecture employs a time-division multiplexing method, where the multiplication operations that are performed in a space-division multiplexing manner in existing MAC units are replaced by a time-division multiplexing manner, thereby greatly reducing the need for a large-size pixel matrix for accommodating input activations and super-large matrix weights. Thus, computationally intensive devices can process large neural network models with almost unlimited input lengths at the cost of an insignificant time delay. In addition, the hardware framework for implementing such computationally intensive devices can be compatible with the current manufacturing processes of photon devices, making energy-efficient optical-based MAC computing possible.
[0024] Figure 1 [[ID=lo]]A schematic diagram of a neural network (model) 100 according to some embodiments is shown. As Figure 1 shown, the neural network 100 is formed by a reticular interconnect structure formed by multiple neuron inner layers (or hidden layers). In Figure 1 it, for simplicity, only one neuron 101 and the weights of the input connections (w1, w2, w3…w n)。Neurons in adjacent inner layers are connected by connections with weights, and the weights are set according to the influence or effect of previous neurons in the previous inner layer on subsequent neurons in the next inner layer. The output value or output activation of a previous neuron is multiplied by its connection weight to the subsequent neuron to determine the specific stimulus exerted by the previous neuron on the subsequent neuron.
[0025] The total input stimulus of a neuron includes the stimuli from all its weighted input connections from previous neurons in the previous inner layer. According to various configurations, if the total input stimulus of a neuron exceeds a certain threshold, the neuron is triggered to provide an output or output activation that follows a linear or non-linear function based on its input stimulus. The above process is repeated for each neuron in each inner layer until all neurons in the last inner layer provide their respective outputs.
[0026] Based on the characteristics of neural network 100, the more connections there are between neurons, the more neurons there are in each inner layer. In addition, the more neuron layers there are, the stronger the artificial intelligence that neural network 100 can simulate. Therefore, neural network 100 for practical, real-world artificial intelligence applications typically features a large number of neurons and a large number of connections between neurons. Therefore, an extremely large amount of input and computation (including neuron output functions and weighted connections) is involved in operating neural network 100.
[0027] Neural network 100 can be implemented in software or hardware. However, although neural network 100 can be fully implemented in software as program code executed on the cores of one or more general-purpose central processing units (CPUs) or graphics processing units (GPUs), without being restricted by the wiring limitations that hardware computing may encounter, the read / write activities between the CPU / GPU cores and the system memory required to perform computations (such as MAC operations) are very intensive. The overhead and energy associated with repeatedly moving data between the processing unit and the system memory to complete a large number of computations in large neural network 100 are not entirely satisfactory in many aspects.
[0028] Reference Figure 1 , neural network 100 can perform general MAC operations using a 1-by-K vector of input activations X, a K-by-N matrix of weights W, and a 1-by-N vector of output activations Y, where K and N are natural numbers. Therefore, the MAC operation of the vector matrix multiplication operation can be represented by the matrix expression of Equation (1) shown below.
[0029]
[0030] In Equation (1), the vector of input activations X is X = [x 1,1 , x 1,2 , … x 1,k , … x 1,K , and the weights W are represented by the matrix [wk,n structure, the output activation Y is from the array [Y 1,1 , Y 1,2 , …Y 1,k , …Y 1,N structure, where the index k represents the row index and the index n represents the column index.
[0031] Figure 2 FIG. shows a block diagram of an optical computing device 200 for optical computing according to some comparative embodiments of the present disclosure. The optical computing device 200 can be used to simulate Figure 1 the neural network 100 shown in. The optical computing device 200 operates in the optical domain and performs at least a part of the MAC operation based on optical signal processing. For example, an algebraic accumulation operation of multiple input activations is performed by optically converging multiple optical signals representing the input activations and outputting a converging or fan-in optical signal representing the output activation. In addition, an algebraic multiplication operation of the multiplier on the input activation can be achieved by optically guiding the optical signal representing the input activation to an optical element having a controllable modulation level in terms of transmittance or reflectance value, and outputting the obtained output optical signal as the output activation through the optical element.
[0032] The optical computing device 200 includes an array of source pixels 302, a first optical element 304, a spatial light modulator (SLM) 306, a second optical element 308, detector pixels 309, and a controller (CTRL) 310. The SLM 306 includes an array of modulator pixels 307 formed thereon. Although Figure 2 only the above elements of the optical computing device 200 are shown, the present disclosure is not limited thereto. In other embodiments, more or fewer elements may be included in the optical computing device 200.
[0033] According to some embodiments, the array of source pixels 302 is configured to receive the input activation values of the neural network 100 and convert these values into the intensities of an input beam / pulse array. According to some embodiments, the source pixels 302 are formed by a liquid crystal display (LCD) having a light-emitting pixel array. According to some embodiments, the source pixels 302 include light-emitting diodes (LEDs), such as organic OLEDs, mini-OLEDs, micro-LEDs, etc. According to some embodiments, the source pixels 302 are formed by vertical cavity surface emitting lasers (VCSELs).
[0034] According to some embodiments, the first optical element 304 is configured to adjust its optical properties before an input beam / pulse transmitted by the source pixel 302 impinges on a corresponding modulator pixel 307 of the SLM 306. The first optical element 304 may include a lens, a mirror, a meta-lens, a combination thereof, etc. According to some embodiments, the first optical element 304 includes a diffractive lens, a concave lens, or a convex lens, which is configured to direct the input beam / pulse and align the input beam / pulse with the corresponding modulator pixel 307 on the SLM 306.
[0035] The SLM 306 may include a liquid crystal display (LCD) panel that includes an array of modulator pixels 307. Each modulator pixel 307 may be a transmissive optical medium having an adjustable transmittance as a modulation level with respect to the beam / pulse incident thereon. The transmissive optical medium of the modulator pixel 307 allows the modulator pixel 307 to filter or modulate the input beam / pulse based on a modulation value of a transmittance value according to a corresponding weight to effect multiplication on the input beam or pulse. The modulation operation performed by the modulator pixel 307 is equivalent to the multiplication step of each input activation and its corresponding weight in a MAC operation.
[0036] In some other embodiments, the SLM 306 alternatively employs a reflective optical medium ( Figure 2 not shown separately herein, but shown and labeled as 306B in Figure 3B ), which is configured to reflect the incident input beam / pulse and modulate the intensity of the input beam / pulse according to its reflectivity. The transmissive optical medium and the reflective optical medium are regarded as various types of SLM 306 having one or more modulator pixels 307. The modulator pixels 307 of the SLM 306 are configured to provide a modulated beam / pulse in response to the input beam / pulse, based on a modulation value in terms of its transmittance or reflectivity according to a corresponding weight factor based on its modulation level, to exhibit a multiplicative effect on the input beam / pulse. The modulation operation performed by the modulator pixel 307 is equivalent to the multiplication step of each input activation and its corresponding weight in a MAC operation.
[0037] According to some embodiments, the second optical element 308 is configured to adjust its optical properties before the modulated beam / pulse transmitted by the modulator pixel 307 impinges on a detector pixel 309 in a receiver of the optical computing device 200. The second optical element 308 may include a lens, a mirror, a meta-lens, a combination thereof, etc. According to some embodiments, the second optical element 308 further includes a converging or concave optical element, which is configured to direct or fan the modulated beam / pulse into the detector pixel 309 in the receiver.
[0038] According to some embodiments, detector pixel 309 includes a photodetector or a photodiode, which is configured to receive a modulated light beam / pulse transmitted via the second optical element 308. According to some embodiments, detector pixel 309 is configured to accumulate the modulated light beam / pulse from each modulation pixel 307 and convert the photons of the accumulated light beam / pulse into current or charge. The value of the current or charge corresponds to the value of the output activation of the neural network 100. The above light accumulation and charge conversion steps are equivalent to the addition step in the MAC operation.
[0039] According to some embodiments, controller 310 is configured to receive the parameters of the input activation and weights and send control signals to source pixel 302 and modulator pixel 307 to determine the intensity of the light beam / pulse of each source pixel 302 and the transmittance of each modulator pixel 307. According to some embodiments, controller 310 is configured to receive the current or charge generated by detector pixel 309 and convert it into the value of the output activation. According to some embodiments, controller 310 includes a microcontroller, a field programmable gate array (FPGA), a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), etc.
[0040] Referring to Figure 2 , during operation, a training source S1 is received or provided for training Figure 1 the neural network 100 shown. The training source S1 can be in the form of an image, video, audio, text, data, or other types of information suitable for training an artificial intelligence model. The training source S1 is divided into multiple segments or parts as the input activation of the neural network, denoted as vectors [x1, x2,..., x K . Each segment x1, x2,..., x K of the training source S1 is associated with each other to form complete information, such as a picture of a dog or a segment of human speech.
[0041] The optical computing device 200 is configured to perform a MAC operation for output activation y1 given an input activation vector [x1, x2,..., x K and a weight array [w1, w2,..., w K . Each of the K source pixels 302 is configured to emit an input light beam / pulse with a separate intensity based on the value of the input activation [x1, x2,..., x K , and each of the K modulator pixels 307 is configured to have a transmittance based on the weight vector [w1, w2,..., w KThe transmittance or reflectance of the value. After the input beam / pulse propagates through the first optical element 304, the modulator pixel 307, and the second optical element 308 to reach the detector pixel 309, the modulated beam / pulse is accumulated and converted into an equivalent current or equivalent charge corresponding to the output activation value. Thus, a MAC operation based on optical computing is achieved.
[0042] Figure 3A FIG. shows a block diagram of a spatial light multiplexer (SLM) 306A for optical computing according to some embodiments of the present disclosure. Referring to Figure 3A , the input beam / pulse 322 is transmitted, for example, from the source pixel 302 and incident on the SLM 306A. According to some embodiments, the SLM 306A is a transmissive SLM having an optical transmission medium configured to modulate or filter the input beam / pulse 322 such that only a portion of the input beam / pulse 322 is allowed to pass through and become the modulated beam / pulse 324.
[0043] According to some embodiments, the SLM 306A is formed by a layer stack including a pair of polarizing layers or polarizers 311, a pair of substrates 312, a pair of conductive layers 314, a pair of alignment layers 316, and a liquid crystal layer 318. Although Figure 3A only the above elements of the SLM 306A are shown, the present disclosure is not limited thereto. More or fewer elements may be included in the SLM 306A.
[0044] According to some embodiments, the pair of polarizing layers 311 are disposed on the two outermost sides of the SLM 306A and configured to provide different polarization directions for the input beam / pulse and the modulated beam / pulse. The pair of polarizing layers 311 may be configured with different polarizations, such as vertical polarization or other polarization angles. As a result, a portion of the input beam / pulse may be blocked from propagating through the SLM 306A. According to some embodiments, the pair of polarizing layers 311 are formed of polyvinyl alcohol (PVA) or other suitable materials. According to some embodiments, the pair of polarizing layers 311 are omitted in the SLM 306A.
[0045] According to some embodiments, a pair of substrates 312 are disposed between the pair of polarizing layers 311 and the pair of conductive layers 314. The substrates 312 may provide physical support for other layers such that the other layers of the SLM 306A can be formed on the pair of substrates 312. According to some embodiments, the pair of substrates 312 are formed of a transparent material, such as glass, quartz, or other suitable materials.
[0046] Each conductive layer 314 is disposed between the substrate 312 and the alignment layer 316 and is configured to receive power from an external power source (e.g., Figure 2The controller 310 shown therein receives a bias voltage and generates an electric field between the pair of conductive layers 314. According to some embodiments, the conductive layers 314 are formed of a transparent material, such as indium tin oxide (ITO) or other suitable transparent conductive oxide materials, conductive polymers, metal grids and random metal networks, carbon nanotubes, graphene, nanowire networks, and the like.
[0047] According to some embodiments, each alignment layer 316 is disposed between the liquid crystal layer 318 and the conductive layer 314. The alignment layer 316 is configured to align liquid crystal molecules in a helical twist manner between the alignment layers 316, where the first layer of molecules and the last layer of molecules are formed perpendicular to each other. According to some embodiments, the pair of alignment layers 316 is formed of polyvinyl alcohol (PVA) or other suitable materials. According to some embodiments, the liquid crystal layer 318 is located in the innermost layer of the SLM 306A and is configured to control the transmittance of the SLM 306 according to the rotation direction of the liquid crystal molecules in the SLM 306A.
[0048] Although not shown separately, the SLM 306A can be divided into a plurality of modulator pixels 307 and modulated according to different weights provided by the controller 310 shown Figure 2 therein. During operation, when modulating the SLM 306A according to the voltage difference, due to the different voltage differences between the two conductive layers 314, each modulator pixel 307 in the liquid crystal layer 318 can experience a separate electric field across its two ends. This causes the liquid crystal molecules to align themselves in the direction of the separate electric field. The relative orientation of the molecules and their birefringence properties result in phase modulation of the SLM 306A. According to some embodiments, the amplitude modulation of the SLM 306A is achieved by using a pair of polarizing layers 311.
[0049] Figure 3B A block diagram of an SLM 306B for optical computing according to some embodiments of the present disclosure is shown. Referring to Figure 3B , an input beam / pulse 322 is transmitted, for example, from a source pixel 302 and incident on the SLM 306B. According to some embodiments, the SLM 306 is a reflective SLM having a reflective optical medium configured to modulate or filter the input beam / pulse 322 such that only a portion of the input beam / pulse 322 is allowed to be reflected back according to the drive beam 323 and becomes a modulated beam / pulse 324.
[0050] According to some embodiments, the SLM 306B is formed by a stack of layers including a polarizer layer 311, a pair of substrates 312, a pair of antireflection layers 313, a pair of conductive layers 314, a reflector 315, a pair of alignment layers 316, a silicon layer 317, and a liquid crystal layer 318. According to some embodiments, the SLM 306B is an optically addressed LCOS (Liquid Crystal on Silicon) SLM, wherein the liquid crystal layer 318 includes nematic liquid crystal molecules aligned in parallel. Although Figure 3B only the above-described elements of the SLM 306B are shown, the present disclosure is not limited thereto. More or fewer elements may be included in the SLM 306B. Some components of the SLM 306B have been discussed for the SLM 306A and will not be repeated for the sake of brevity.
[0051] According to some embodiments, the polarizer layer 311 is disposed between one of the antireflection layers 313 and the substrate 312, on the side close to where the input beam / pulse is incident. According to some embodiments, the pair of antireflection layers 313 are disposed on the outermost layer of the SLM 306B. The antireflection layer 313 is configured to prevent or mitigate light reflection at the interface between the SLM 306B and the environment, to ensure that most of the input beam / pulse 322 is allowed to propagate into the SLM 306B and is reflected according to a determined reflectivity.
[0052] According to some embodiments, the reflector 315 is disposed between the liquid crystal layer 318 and the silicon layer 317 and is configured to reflect the input beam / pulse 322 into a reflected beam / pulse or a modulated beam / pulse 324 according to a determined reflectivity. According to some embodiments, the reflector 315 may be formed of a dielectric mirror or other suitable material.
[0053] According to some embodiments, the silicon layer 317 is disposed between the reflector 315 and one of the conductive layers 314 adjacent to the side where the drive beam is incident and serves as a configurable resistor layer. By changing the resistivity distribution of the silicon layer, the voltage difference across the liquid crystal molecules in the liquid crystal layer 318 and the resulting electric field can be different.
[0054] During operation, when the SLM 306B is modulated according to the voltage difference and the drive beam 323, the silicon layer 317 is configured to generate a pattern with a variable resistance distribution, thereby adjusting the overall voltage difference at the positions of different modulator pixels 307 in the SLM 306B. This causes the liquid crystal molecules to align themselves in the direction of the individual electric fields. The relative orientation of the molecules and their birefringence characteristics result in different phase modulations of different modulator pixels 307 on the SLM 306B. According to some embodiments, the amplitude modulation of the SLM 306B is achieved by using the polarizer layer 311.
[0055] Figure 4AFIG. 0 shows a schematic diagram of an optical computing device 400 according to some comparative embodiments of the present disclosure. The optical computing device 400 includes an array of source pixels 302, a first optical element 332, an SLM 306, and an array of modulator pixels 307, an array of second optical elements 308, and an array of detector pixels 309. Although Figure 4A only the above-described elements of the optical computing device 400 are shown, the present disclosure is not limited thereto. More or fewer elements may be included in the optical computing device 400. Some components of the optical computing device 400 have been discussed with respect to the optical computing device 200 and will not be repeated for the sake of brevity.
[0056] Referring Figure 4A , a training source S1 is received or provided and divided into a plurality of segments or portions as input activations for a neural network, represented as vectors [x1, x2, …, x K . The input activations [x1, x2, …, x K may be sent to a controller 310, which is configured to correspond the values of the received input activations [x1, x2, …, x K to the intensities of the input beam / pulse array. According to some embodiments, the controller 310 is further configured to control the array of source pixels 302 to transmit the input beam / pulse array based on the input activations [x1, x2, …, x K . Although Figure 4A the source pixels 302 are shown arranged in rows or columns, the present disclosure is not limited thereto. The source pixels 302 may also be arranged in an array shape having adjustable numbers of rows and columns.
[0057] According to some embodiments, K input beams / pulses are fan - out by a first optical element 332 into N sets of copies of the K input beams / pulses. Since the fan - out operation is implemented by the first optical element 332, the optical computing device 400 is referred to as a “optical fan - out” optical computing device. According to some embodiments, the first optical element 332 includes a lens, a mirror, a meta - lens, a combination thereof, etc., and is configured to generate N copies of each of the K input beams / pulses. The first optical element 332 may also include a diffractive lens, a concave lens, a convex lens, etc. Although Figure 4A the fan - out input beams / pulses are shown arranged in rows or columns, the present disclosure is not limited thereto. The fan - out input beams / pulses may be configured with a combination of multiple sub - arrays, each sub - array including a set of K input beams / signals, representing the input activations [x1, x2, …, x K . Thus, the fan - out input beams / pulses may be configured with an input activation matrix [x k,nis represented, where the row index k is a natural number from 1 to k, and the column index n is a natural number from 1 to N. Here, k is the length of the input activation, and N is the number of columns of the weights of the neural network 100, as shown in Equation (1).
[0058] According to some embodiments, the controller 310 is further configured to determine the transmittance or reflectance of the modulator pixels 307 on the SLM 306 according to the weight matrix [w k,n . Each fanned-out input beam / pulse (corresponding to the input activation [x k,n ) is imaged onto the corresponding modulator pixel 307 (corresponding to the weight [w k,n ) to achieve optical coupling between the source pixel 302 and the corresponding modulator pixel 307. The above modulation step is equivalent to the multiplication step in the MAC operation.
[0059] Subsequently, each group of K modulated beams / pulses is converged or fanned into the corresponding detector pixels 309 (e.g., detector pixels 309_1, 309_2, …, 309_N) through the corresponding second optical element 308 (e.g., second optical elements 308_1, 308_2, …, 308_N). According to some embodiments, the second optical element 308 is configured to adjust its optical characteristics before the modulated beams / pulses transmitted by the modulator pixels 307 are incident on the corresponding detector pixels 309 in the receiver of the optical computing device 400. The second optical element 308 may include a lens, a mirror, a meta-lens, a combination thereof, etc. According to some embodiments, the second optical element 308 includes a converging or concave lens, which is configured to direct the modulated beams / pulses to the corresponding detector pixels 309 in the receiver.
[0060] According to some embodiments, each detector pixel 309 includes a photodetector or a photodiode, which is configured to receive the modulated beams / pulses transmitted through the corresponding second optical element 308. According to some embodiments, the detector pixel 309 is configured to accumulate the modulated beams / pulses from each modulator pixel 307 in the same group and convert the photons of the accumulated beams / pulses into current or charge. The value of the current or charge corresponds to the value of the corresponding output activation of the neural network 100 (i.e., [y1, y2, … y N ). For example, the multiply-accumulate operation of the first output activation y1 represented by the current of the detector pixel 309_1 is represented by the following Equation (2):
[0061] y1 = x1 w11 + x1 w 12 + … + x K w 1K , …………………………………(2)
[0062] Similarly, the final output activation y represented by the current of detector pixel 309_N N for the multiply-accumulate operation is represented by the following equation (3):
[0063] y N = x1 w N1 + x1 w N2 + … + x K w NK ……………………………………(3)
[0064] The above optical accumulation and charge conversion steps are equivalent to the accumulation step in the MAC operation.
[0065] Figure 4B FIG. shows a schematic diagram of an optical computing device 401 according to some comparative embodiments of the present disclosure. The optical computing device 401 includes a source panel 412 having source pixels 402 of a K-by-N matrix, a first optical element 304, a modulator panel 416 having modulator pixels 406 of a K-by-N array serving as an SLM, a plurality of second optical elements 308, and a detector panel 419 having detector pixels 409 of an N-by-1 array. The optical computing device 401 can perform the same MAC operation using a structure similar to that of the optical computing device 400, where the source pixels 402, the modulator pixels 406, and the detector pixels 409 correspond to the matrix of the source pixels 302, the matrix of the modulator pixels 307, and the array of the detector pixels 309, respectively. The first optical element 304 and the second optical elements 308 have been discussed for the optical computing device 400 and will not be repeated for brevity.
[0066] According to some embodiments, the source panel 412 is formed of a liquid crystal display (LCD) panel on which a matrix of source pixels 302 is formed, where the source pixels 302 are formed of LCD pixels. The source panel 412 can also be formed of a light emitting diode (LED) display panel on which a matrix of source pixels 302 is formed, where the source pixels 302 are formed of VCSELs or LEDs, such as OLEDs, mini-LEDs, micro-LEDs. Similarly, the modulator panel 416 is formed of a liquid crystal display (LCD) panel on which a matrix of modulator pixels 307 is formed, where the modulator pixels 307 are formed of LCD pixels. In addition, the detector panel 419 can also be formed with an array of detector pixels 309 formed thereon, where the detector pixels 309 are formed of photodetectors or photodiodes.
[0067] According to some embodiments, the controller 310 is configured to copy the first column of the source panel 412 to the other N - 1 columns in the source panel 412. Thus, the N columns of source pixels 302 on the source panel 412 will be configured with the same input activation X = [x1, x2, …, x KEmit an input optical beam / pulse. This is equivalent to the optical fan-out step performed by the first optical element 332 of the optical computing device 400, but is implemented in the pixel matrix on the panel. In other words, the source pixels 302 in the same row but different columns of the source panel 412 are electrically interconnected. As a result, since all columns have been electrically connected together row by row, the controller 310 only needs to send a control signal to any one of the first column or N columns to complete the configuration of all the source pixels 302. This can also be seen from the repeated labels of X1 at the bottom of each column near the source panel 522. This will further simplify the control complexity.
[0068] According to some embodiments, the modulator pixels 307 of the modulator panel 416 are configured according to a matrix of weights [w n,k . The source pixels 302 are imaged onto the corresponding matrix of the modulator pixels 307 on the modulator panel 416. According to some embodiments, if the alignment and focusing of the source pixels 302 are well controlled such that the direct imaging of the source pixels 302 onto the modulator pixels 307 is considered successful, the first optical element 304 can be omitted.
[0069] According to some embodiments, the modulated optical beam / pulse is fanned into the corresponding detector pixels 309 on the detector panel 419 through the corresponding second optical element 308.
[0070] Compared with the electrical computing device, the optical computing devices 400 and 401 have the advantages of optical computing in implementing MAC operations, with fast processing time and low energy consumption. According to some embodiments, the length K of the input activation [x k is determined by the size of the source panel 412 and the number of columns N of the weights [w k,N or the length N of the output activation [y n . For example, if the total number of modulator pixels 307 of the modulator panel 416 is M2*N2 (the parameters M2, N2 are natural numbers), the number K of the input activation [x k is constrained by the upper limit M2*N2 / N. Therefore, the length or type of the training source S1 is correspondingly limited to relatively small training data.
[0071] Figure 5A FIG. shows a schematic diagram of an optical computing device 500 according to some embodiments of the present disclosure. The optical computing device 500 can be used to implement Figure 1The neural network 100 shown. The optical computing device 500 operates in the optical domain and performs at least a part of the MAC operation based on optical signal processing. For example, an algebraic accumulation operation of multiple input activations is performed by optically converging multiple optical signals representing the input activations and outputting a converged or fan-in optical signal representing the output activation. In addition, an algebraic multiplication operation of the multiplier on the input activation can be achieved by optically guiding the optical signal representing the input activation to an optical element with a controllable transmittance or reflectance value and outputting the obtained output optical signal as the output activation through the optical element.
[0072] The optical computing device 500 includes a source pixel 502, a first optical element 504, an SLM 506, a second optical element 508, a detector pixel 509, a controller 510, a clock generator 511, and an integrator 505. Modulator pixels 507 are formed on the SLM 506. Although Figure 5A only the above elements of the optical computing device 500 are shown, the present disclosure is not limited thereto. In other embodiments, more or fewer elements may be included in the optical computing device 500. Some components of the optical computing device 500, such as the source pixel 502, the first optical element 504, the SLM 506, the modulator pixels 507, the second optical element 508, and the detector pixel 509, are similar to their respective corresponding elements in the optical computing device 200, for example, the source pixel 302, the first optical element 304, the SLM 306, the modulator pixels 307, the second optical element 308, and the detector pixel 309, and thus their descriptions are not repeated for the sake of brevity.
[0073] Refer to Figure 5A , receive or provide the training source S1 and divide it into multiple segments or parts as the input activations of the neural network 100, represented as an array [x1, x2,..., x K . The input activations [x1, x2,..., x K can be sent to the controller 510, which is configured to convert the received input activations [x1, x2,..., x K into an input optical pulse array. According to some embodiments, the controller 510 is further configured to control the source pixel 502 to determine the intensity of the input optical pulse array based on the values of the input activations [x1, x2,..., x K .
[0074] The source pixel 502 is a single source pixel 502 configured to transmit K input optical pulses in a time-division multiplexing manner at a predetermined time interval T. The clock generator 511 is configured to generate a clock signal with a time interval T to manage the different transmission times of the input optical pulses by the source pixel 502. The clock generator 511 can also be configured to synchronize the transmission times of multiple input optical pulses of the source pixel 502 with the modulation times of the modulator pixels 507 on the SLM 506. According to some embodiments, the controller 510 is configured to transmit control signals to manage the clock generator 511. The controller 510 can be configured to determine or receive the value of the time interval T and send it to the clock generator 511.
[0075] According to some embodiments, the K input optical pulses are transmitted through the first optical element 504 and guided to the modulator pixels 507 of the SLM 506. According to some embodiments, the modulator pixel 507 is a single modulator pixel 507 configured to modulate the input optical pulses transmitted by the source pixel 502 according to the corresponding weights [w1, w2, …, w K . The modulation step of the K input optical pulses is performed in a time-division multiplexing manner at a time interval T. According to some embodiments, the clock generator 511 is configured to generate a clock signal with a time interval T to control the modulation time and interval of the modulator pixel 507. The above modulation step is equivalent to the multiplication step in the MAC operation. According to some embodiments, the clock generator 511 is integrated into the controller 510 such that the controller 510 can generate clock signals and control signals for the source pixel 502, the SLM 506, the detector pixel 509, and the integrator 505.
[0076] The modulated optical pulses are transmitted through the second optical element 508 and guided to the detector pixel 509. The detector pixel 509 is configured to collect and accumulate the K modulated optical pulses in a time-division multiplexing manner at a time interval T through the photodetector / photodiode of the detector pixel 509. The detector pixel 509 can also be configured to convert the K modulated optical pulses into K corresponding currents or charges and transmit the K currents or charges to the integrator 505.
[0077] According to some embodiments, the integrator 505 is configured to integrate the current or current charge transmitted by the detector pixel 509. The integrated charge or current is converted into a corresponding value and sent back to the controller 510 and converted into an output activation y1. The integrator 505 can include a capacitor or other integrator circuits including operational amplifiers. The above accumulation of the modulated optical pulses and the integration of the converted charge / current are equivalent to the accumulation step in the MAC operation. As a result, the output activation y1 can be represented by Equation (4) shown below.
[0078] y1 = x1w1 + x2w2 + … + xK w K …………………………………………(4)
[0079] For any output activation yn, if the input activations [x1, x2, …, x K , weights [w1, w2, … w 1,k , … w K , and output activations [y1, y2, … y n …, y N are respectively extended to the corresponding matrix forms [x 1,1 , x 1,2 , … x 1,k …, x 1,K , [w k,n |1 ≤ k ≤ K, 1 ≤ n ≤ N], and [y 1,1 , y 1,2 , … y 1,n …, y 1,N , then the n-th output activation y1,n can be represented by Equation (5) shown below:
[0080]
[0081] The optical computing device 500 provides advantages over the optical computing devices 200, 400, or 401. Different from the multiplication steps implemented in a space-division multiplexing manner by the source pixels 302 and modulator pixels 307 of the SLM 306, the multiplication steps implemented by a single source pixel 502 and a single modulator pixel 507 are performed in a time-division multiplexing manner. Therefore, compared with the optical computing devices 200 or 400, the hardware cost and device footprint of the optical computing device 500 can be greatly reduced and are independent of the length K of the input activation.
[0082] In addition, according to some embodiments, since the input optical pulse is transmitted by a single source pixel 502, modulated by a single modulator pixel 507, and accumulated by a single detector pixel 509, the optical path of the optical computing device 500 has only a single incident angle, which is perpendicular to the modulator pixel 507 or the detector pixel 509, different from the various incident angles of the first optical element 332 for splitting out the input optical pulse or the second optical element 308 for splitting in the modulated optical pulse. In the optical computing device 500, the energy loss or noise generated by the splitting-out or splitting-in operation performed by the first optical element 332 or the second optical element 308 is greatly reduced or eliminated. Therefore, the optical coupling efficiency and accuracy can be significantly improved.
[0083] In addition, the processing time for generating a single output activation y1 for the optical computing device 500 is greater than that of the optical computing devices 400 or 401 and depends on the time interval T of the input activation and [x1, x2, …, xK The length K of []. However, in practical usage scenarios, the time delay is not obvious. For example, under the operation of a clock generator with a gigahertz sampling rate, given a relatively long input activation [x1, x2, …, x K sequence, where K is equal to approximately 8 million, the time interval T is basically equal to 1e(-9) seconds, and the average time delay for generating the output activation is approximately 0.008 seconds. Based on the above, the time delay of the optical computing device 500 can be ignored.
[0084] Figure 5B FIG. shows a schematic diagram of an optical computing device 501 according to some embodiments of the present disclosure. The optical computing device 501 includes a source panel 512 having an array of 1-by-N source pixels 502, an array of 1-by-N first optical elements 504, a modulator panel 516 having an array of 1-by-N modulator pixels 507 serving as an SLM, an array of 1-by-N second optical elements 508, a detector panel 519 having an array of 1-by-N detector pixels 509, and an integrator panel 517 having an array of 1-by-N integrators 505 formed thereon. The optical computing device 501 can perform the same MAC operation using a similar structure to the optical computing device 500, where the first optical element 504 and the second optical element 508 have been discussed for the optical computing device 500 and will not be repeated for brevity.
[0085] According to some embodiments, the source panel 512 is formed by a liquid crystal display (LCD) panel having an array of source pixels 502 formed thereon, where the source pixels 502 are formed by LCD pixels. The source panel 512 can also be formed by a light-emitting diode (LED) display panel having an array of source pixels 502 formed thereon, where the source pixels 502 are formed by VCSELs or LEDs, such as OLEDs, mini-LEDs, micro-LEDs, etc. Similarly, according to some embodiments, the modulator panel 516 is formed by a liquid crystal display (LCD) panel having an array of modulator pixels 507 formed thereon, where the modulator pixels 507 are formed by LCD pixels. In addition, according to some embodiments, the detector panel 519 can also be formed by a panel having an array of detector pixels 509 formed thereon, where the detector pixels 509 are formed by photodetectors or photodiodes. According to some embodiments, the integrator panel 51 is also formed by a panel having an array of integrators 505 formed thereon, where the integrators 505 are formed by capacitors or integrator circuits.
[0086] According to some embodiments, the controller 510 is configured to copy the first source pixel 502 of the source panel 512 to the remaining N - 1 source pixels 502 of the source panel 512. The N source pixels 502 can be electrically interconnected. Thus, the N source pixels 502 on the source panel 512 will be configured with the same input activation X = [x1, x2, …, x at the same time interval T.K Emit an input optical pulse. This is equivalent to the optical fan - out step performed by the first optical element 332 of the optical computing device 400, but implemented in a time - multiplexed manner on the panel. By controlling the electrical clock signal, the input optical pulse is fanned out in the time domain. Thus, the optical computing device 501 is also referred to as an “electrical fan - out” optical computing device. As a result, since all source pixels 502 are already electrically connected together, as shown by the solid lines of the entries of the source pixels 502 connecting the source panel 522, the controller 510 only needs to send a control signal to the first source pixel 502 or any one of the source pixels 502 to complete the configuration of all source pixels 502. The control complexity will be further simplified.
[0087] According to some embodiments, the modulator pixels 507 of the modulator panel 516 are configured according to the weight matrix [w n,k , where the nth modulator pixel 507 is configured to be modulated at time intervals T successively according to K entries in the nth column of the weight matrix (i.e., the weights represented in the form of the entries in the nth column [w n =[w n,1 ; w n,2 ; …; w n,k ). Then, through the nth first optical element 504, the nth source pixel 502 is imaged onto the corresponding nth modulator pixel 507 on the modulator panel 516 K times separated by the time interval T, to achieve optical coupling between the input optical pulse for the input activation X = [x1, x2, …, x K and the modulator pixels 507 for the nth row of weights [w n =[w n,1 ,; wn,2; …; wn,K,] in a time - multiplexed manner. According to some embodiments, the nth first optical element 504 is omitted because the alignment and focusing of the source pixels 502 with respect to the respective modulator pixels 507 are achieved through a simple one - to - one pixel mapping. Thus, their optical coupling performance is well managed without the nth first optical element 504.
[0088] According to some embodiments, each group of K modulated optical pulses of the nth modulator pixel 507 is transmitted through the nth second optical element 508 and imaged onto the nth detector pixel 509 on the detector panel 519 K times separated by the time interval T. According to some embodiments, the nth second optical element 508 is omitted because the alignment and focusing of the modulator pixels 507 with respect to the respective detector pixels 509 are achieved through a simple one - to - one pixel mapping. Thus, their optical coupling performance is well managed without the nth second optical element 508.
[0089] According to some embodiments, each set of K currents or K sets of charges in the n-th detector pixel 509 is accumulated or integrated by the n-th integrator 505 to provide a value corresponding to the output activation y n The N integrators 505 can obtain the total output activation, such as [y1, y2, …, y N .
[0090] Figure 6 FIG. shows a schematic diagram of an optical computing device 600 according to some embodiments of the present disclosure. The optical computing device 600 is regarded as an extended version of the optical computing device 501 because the optical computing device 600 can simultaneously process multiple training data streams based on different training sources S1, S2, … S m … S M where the parameters M, m are natural numbers and m ranges from 1 to M. Each training stream for performing MAC operations on the training source S m is similar to the training stream performed by the optical computing device 501 on the training source S1.
[0091] Referring to Figure 6 , the optical computing device 600 includes a source panel 522 with an M-by-N matrix of source pixels 502, an M-by-N matrix of first optical elements 504, a modulator panel 526 with an M-by-N matrix of modulator pixels 507 as an SLM, an M-by-N matrix of second optical elements 508, a detector panel 529 with an M-by-N matrix of detector pixels 509, and an integrator panel 517 with an M-by-N matrix of integrators 505. The source panel 522, the modulator panel 526, the detector panel 529, and the integrator panel 527 can respectively correspond to the source panel 512, the modulator panel 516, the detector panel 519, and the integrator panel 517 of the optical computing device 501 with extended dimensions. The matrices of the first optical elements 504 and the second optical elements 508 have been discussed for the optical computing device 500 or 501 and will not be repeated for simplicity.
[0092] Receives or provides multiple training sources S1 to S M for training different neural networks, each neural network being similar to Figure 1 the neural network 100 shown. The training sources S1 to S M [[ID=!28]]can be independent of each other.
[0093] According to some embodiments, the N source pixels 502 in the m-th row of the source panel 522 are used to receive the m-th training source S m . In other words, M training streams running on the M rows of the source panel 522 are performed simultaneously in a space-division multiplexing manner. For the training source S mFor each training stream of , the source pixels 502 in the mth row of the source panel 522 are configured to generate or transmit K input light pulses according to the K input activations of the mth training source, where the input activations are represented as [X m,K ] or simply [X m ], where index m represents the training source S m , index k represents the kth input activation of the mth training source to be transmitted at the kth time instant. In a similar arrangement to the source panel 512, the N source pixels 502 in the same row of the source panel 522 are electrically interconnected, as shown by the solid lines connecting the entries of the source pixels 502 on the same row in the first three rows of the source panel 522. This means that the input activations [X m ] are the same, and the N source pixels 502 in the same row are configured to transmit the same input light pulse for each column of the source panel 522 at the same time interval T.
[0094] According to some embodiments, the modulator pixels 507 of the modulator panel 526 are trained according to the mth training source S m The weight matrix of the [n,k]th weight Configuration. The [m, n]th modulator pixel 507 in the mth row and nth column of the modulator panel 526 is configured according to the weight matrix (ie, ) are modulated continuously at a time interval T. Then, the n-th source pixel 502 on the m-th row of the source panel 522 is imaged onto the corresponding n-th modulator pixel 507 on the m-th row of the modulator panel 516 by the [n,m]-th first optical element 504 at a time interval T for K times to realize the m-th training source S in a time multiplexing manner. m The input activation X = [x m,1 ,x m,2 ,…,x m,K ] input optical pulses and weights for the nth column According to some embodiments, the [n,m]th first optical element 504 is omitted because the alignment and focusing of the [n,m]th source pixel 502 with respect to the corresponding [n,m]th modulator pixel 507 is achieved by a simple one-to-one pixel mapping, and therefore, such optical coupling performance is well managed without the [n,m]th first optical element 504.
[0095] According to some embodiments, the training sources S1 to S m The same neural network 100 is trained so that the weights remain constant for different indices m. Therefore, for all m, the weights can be simplified to In this case, the M modulator pixels 507 on the same column will be modulated with M identical weights [w n,k at intervals of the same time interval T. According to some embodiments, as Figure 6 shown, the modulator pixels 507 in the same column of the modulator panel 516 are electrically interconnected, as shown by the solid lines, which connect the entries of the modulator pixels 507 in the same column of the first three columns of the modulator panel 526. The control complexity will be further simplified.
[0096] According to some embodiments, each group of K modulated light pulses of the [m, n]th modulator pixel 507 for the mth training source S m is transmitted through the [m, n]th second optical element 508 and imaged onto the [m, n]th detector pixel 509 on the detector panel 529 at a time interval T. According to some embodiments, each group of K modulated light pulses of the [m, n]th modulator pixel 507 is converted by the [m, n]th detector pixel 509 into a corresponding group of currents or charges. According to some embodiments, the [m, n]th second optical element 508 is omitted because the alignment and focusing of the [m, n]th modulator pixel 507 relative to the [m, n]th detector pixel 509 are achieved by simple pixel-by-pixel mapping. Therefore, their optical coupling performance is well managed in the absence of the [m, n]th second optical element 508.
[0097] According to some embodiments, each group of K currents in the [m, n]th detector pixel 509 is accumulated or integrated by the [m, n]th integrator 505 of the integrator panel 527 to provide a value corresponding to the [m, n]th output activation For the mth training source S m , the N integrators 505 can obtain a total output activation of
[0098] Figure 7A FIG. shows a semiconductor optical computing device 700A according to some embodiments of the present disclosure. According to some embodiments, the semiconductor optical computing device 700A is used to implement the optical computing devices 500, 501 or 600. The semiconductor optical computing device 700A includes a source substrate 702, a first optical substrate 704, a modulator substrate 706, a second optical substrate 708, and a detector substrate 710 that are parallel to each other and spaced apart from each other. The semiconductor optical computing device 700A further includes spacers 712 between the above substrates 702-710. Although Figure 7A only the above elements of the semiconductor optical computing device 700A are shown, the present disclosure is not limited thereto. In other embodiments, more or fewer elements may be included in the semiconductor optical computing device 700A.
[0099] According to some embodiments, the source panel 512 or 522, the modulator panel 516 or 526, and the detector panel 519 or 529 are respectively formed on the source substrate 702, the modulator substrate 706, and the detector substrate 710. The detector substrate 710 may further include an integrator panel 517 or 527 on which an integrator 505 is formed. According to some embodiments, the first optical element 504 and the second optical element 508 are respectively formed on the first optical substrate 704 and the second optical substrate 708. According to some embodiments, in order to ensure that light pulses can propagate smoothly in the semiconductor optical computing device 700A, each of the source substrate 702, the first optical substrate 704, the modulator substrate 706, and the detector substrate 710 is formed of a transparent substrate formed of glass, quartz, or other suitable transparent materials. The source substrate 702 and the detector substrate 710 may also be formed of non-transparent semiconductor substrates, such as silicon, germanium, or other suitable semiconductor materials, provided that these substrates do not pose an obstacle to the propagation of light pulses.
[0100] Each adjacent pair of the source substrate 702, the first optical substrate 704, the modulator substrate 706, the second optical substrate 708, and the detector substrate 710 may be spaced apart from each other by a spacing D. The spacing D may be in the range between millimeters and micrometers or sub-micrometers, depending on the design requirements of the first optical substrate 704 or the second optical substrate 708. Between different adjacent pairs of the above substrates 702 - 710, the spacing D may be substantially equal or unequal.
[0101] The spacer 712 is disposed between each adjacent pair of the source substrate 702, the first optical substrate 704, the modulator substrate 706, the second optical substrate 708, and the detector substrate 710. The spacer 712 is used to provide physical support and spacing fixation for the semiconductor optical computing device 700A. From a top view perspective, the spacer 712 may include a closed-loop shape. According to some other embodiments, the spacer 712 includes a plurality of columns distributed around the periphery of the above substrates 702 - 710. The spacer may be formed of a molded material, a packaging material, or other suitable materials. According to some embodiments, the gap between adjacent pairs of substrates 702 - 710 is maintained under low-pressure conditions and filled with air or filled with a transparent material, such as plastic or polymer.
[0102] Figure 7BShows a semiconductor optical computing device 700B according to some embodiments of the present disclosure. The semiconductor optical computing device 700B is similar to the semiconductor optical computing device 700A in many aspects, and thus for the sake of brevity, these similar features will not be repeated. The main difference between the semiconductor optical computing device 700B and the semiconductor optical computing device 700A lies in the arrangement of replacing the spacer 712 with the spacer 714. The spacer 714 laterally surrounds the source substrate 702, the first optical substrate 704, the modulator substrate 706, the second optical substrate 708, and the detector substrate 710. According to some embodiments, the source substrate 702, the first optical substrate 704, the modulator substrate 706, the second optical substrate 708, and the detector substrate 710 are clamped by the spacer 714 at the periphery of the above-mentioned substrates 702 - 710.
[0103] Figure 8 Shows a schematic flowchart of a method 800 for operating a semiconductor optical computing device according to some embodiments of the present disclosure. It should be understood that additional steps may be provided before, during, and after the steps in the method 800, and some of the steps described below may be replaced or eliminated with other embodiments. Figure 8 The order of the steps shown may be interchanged. Some steps may be performed simultaneously or independently.
[0104] In step 802, a first source pixel generates K first input optical pulses at a certain time interval, where K is a natural number.
[0105] In step 804, the K first input optical pulses are directed to a first modulator pixel, thereby receiving K first modulated optical pulses at that time interval, where the first modulator pixel is configured to be modulated corresponding to K first weights with respect to the K first output optical pulses.
[0106] In step 806, the K first modulated optical pulses are received by a first detector pixel.
[0107] In step 808, an integrator accumulates the charge generated in response to the K first modulated optical pulses to produce the result of a first multiply-accumulate operation.
[0108] According to an embodiment of the present invention, an optical circuit includes: a source pixel configured to generate a plurality of input optical pulses having a time interval; a modulator pixel optically coupled to the source pixel and configured to modulate the plurality of input optical pulses to generate a plurality of modulated optical pulses; a detector pixel optically coupled to the modulator pixel and configured to generate charge in response to the plurality of modulated optical pulses; and a controller configured to electrically control the intensity of each of the plurality of input optical pulses and the modulation level of the modulator pixel to perform a multiply-accumulate operation.
[0109] In some embodiments, the optical circuit further includes a clock generator for synchronizing each of the plurality of input optical pulses with the modulation time of the modulator pixels.
[0110] In some embodiments, the modulator pixels include an optical transmission medium having a variable transmittance, wherein modulating the plurality of input optical pulses includes modulating the plurality of input optical pulses with the individual transmittances of the modulator pixels.
[0111] In some embodiments, the modulator pixels include an optical reflection medium having a variable reflectance, wherein modulating the plurality of input optical pulses includes modulating the plurality of input optical pulses with the individual reflectances of the modulator pixels.
[0112] In some embodiments, the optical circuit further includes an integrator configured to accumulate the charge generated by the detector pixels to provide a value corresponding to the result of a multiply-accumulate operation.
[0113] In some embodiments, the optical circuit further includes a first optical element located between the source pixels and the modulator pixels, and the first optical element is configured to direct the plurality of input optical pulses to the modulator pixels.
[0114] In some embodiments, the optical circuit further includes a second optical element located between the modulator pixels and the detector pixels and configured to direct the plurality of modulated optical pulses to the detector pixels.
[0115] In some embodiments, the source pixels include at least one of a light emitting diode (LED), an organic LED, a mini-LED, a micro-LED, and a vertical cavity surface emitting laser (VCSEL).
[0116] In some embodiments, the optical circuit further includes a spatial light modulator (SLM), wherein the SLM includes a liquid crystal display (LCD) panel, the LCD panel includes a pixel array, and the pixel array includes modulator pixels.
[0117] In some embodiments, the detector pixels include photodiodes.
[0118] According to one embodiment of the present disclosure, an optical circuit includes: an M-by-N matrix of source pixels, each source pixel being configured to generate a plurality of input optical pulses having time intervals, where M and N are natural numbers; an M-by-N matrix of modulator pixels optically coupled to the source pixel matrix, each modulator pixel being configured to modulate the plurality of input optical pulses with a separate transmittance or reflectance to produce a plurality of modulated optical pulses; an M-by-N matrix of detector pixels optically coupled to the modulator pixel matrix, each detector pixel being configured to generate charge in response to the plurality of modulated optical pulses from a corresponding one of the modulator pixels; and a controller configured to electrically control the intensity of each of the plurality of input optical pulses of each source pixel and the separate transmittance or reflectance of each modulator pixel to perform a multiply-accumulate operation.
[0119] In some embodiments, the source pixels in the same row are configured to transmit the same input optical pulses separated by an input activation array transmission time interval.
[0120] In some embodiments, the source pixels in the same row of the source pixel matrix are electrically interconnected.
[0121] In some embodiments, the modulator pixels in the same column of the modulator pixel matrix are configured with equal transmittances or reflectances separated by time intervals according to a weight matrix.
[0122] In some embodiments, the input optical pulses in different rows of the source pixels are associated with different training sources.
[0123] In some embodiments, the optical circuit further includes a source substrate, a modulator substrate, and a detector substrate that are parallel to each other, wherein the source pixel matrix, the modulator pixel matrix, and the detector pixel matrix are respectively formed on the source substrate, the modulator substrate, and the detector substrate.
[0124] According to one embodiment of the present invention, an optical computing method for an artificial intelligence accelerator includes: generating, by a first source pixel, K first input optical pulses having time intervals, where K is a natural number; guiding the K first input optical pulses to a first modulator pixel, thereby receiving K first modulated optical pulses at time intervals, wherein the first modulator pixel is configured to modulate in correspondence with K first weights and the K first input optical pulses; receiving the K first modulated optical pulses by a first detector pixel; and accumulating the charge of the first detector pixel in response to the K first modulated optical pulses by an integrator to produce a result of a multiply-accumulate operation.
[0125] In some embodiments, the method further includes synchronizing the generation time of the K first input optical pulses with the K modulation times of the first modulator pixel.
[0126] In some embodiments, the method further includes sending a clock signal to the first source pixel and the first modulator pixel to achieve synchronization.
[0127] In some embodiments, the method further includes generating, by a second source pixel, K second input optical pulses having a time interval, wherein the K first input optical pulses and the K second input optical pulses are generated based on a first training source and a second training source, respectively, and the first source pixel and the second source pixel are disposed on the same panel.
[0128] The foregoing outlines the features of several embodiments so that those skilled in the art may better understand the various aspects of the present disclosure. Those skilled in the art should understand that they can readily use the present disclosure as a basis for designing or modifying other processes and structures for achieving the same purposes and / or achieving the same advantages as those introduced in the embodiments herein. Those skilled in the art should also recognize that such equivalent structures do not depart from the spirit and scope of the present invention, and that they can make various changes, substitutions, and alterations in the present invention without departing from the spirit and scope of the present invention.
Claims
1. An optical circuit, comprising: A source pixel configured to generate a plurality of input optical pulses having time intervals; A modulator pixel optically coupled to the source pixel and configured to modulate the plurality of input optical pulses to generate a plurality of modulated optical pulses; A detector pixel optically coupled to the modulator pixel and configured to generate charge in response to the plurality of modulated optical pulses; And A controller configured to electrically control the intensity of each of the plurality of input optical pulses and the modulation level of the modulator pixel to perform a multiply-accumulate operation.
2. The optical circuit according to claim 1, further comprising a clock generator to synchronize each of the plurality of input optical pulses with the modulation time of the modulator pixel.
3. The optical circuit according to claim 1, wherein, The modulator pixel includes an optical transmission medium having a variable transmittance, wherein the modulation of the plurality of input optical pulses includes modulating the plurality of input optical pulses with the individual transmittance of the modulator pixel.
4. The optical circuit according to claim 1, wherein, The modulator pixel includes an optical reflection medium having a variable reflectance, wherein the modulation of the plurality of input optical pulses includes modulating the plurality of input optical pulses with the individual reflectance of the modulator pixel.
5. The optical circuit according to claim 1, further comprising an integrator configured to accumulate the charge generated by the detector pixel to provide a value corresponding to the result of the multiply-accumulate operation.
6. An optical circuit, comprising: An M-by-N matrix of source pixels, each source pixel configured to generate a plurality of input optical pulses having time intervals, where M and N are natural numbers; An M-by-N matrix of modulator pixels optically coupled to the source pixel matrix, each modulator pixel configured to modulate the plurality of input optical pulses with an individual transmittance or reflectance to generate a plurality of modulated optical pulses; An M-by-N matrix of detector pixels optically coupled to the modulator pixel matrix, each detector pixel configured to generate charge in response to a plurality of modulated optical pulses from a corresponding one of the modulator pixels; And A controller configured to electrically control the intensity of each input optical pulse of the plurality of input optical pulses of each source pixel and the individual transmittance or reflectance of each modulator pixel to perform a multiply-accumulate operation.
7. The optical circuit according to claim 6, wherein, The source pixels in the same row are configured to transmit the same input optical pulses separated by time intervals according to an input activation array.
8. The optical circuit according to claim 7, wherein, The source pixels in the same row of the source pixel matrix are electrically interconnected.
9. A method for performing optical computing of an artificial intelligence accelerator, the method comprising: Generating, by a first source pixel, K first input optical pulses having time intervals, where K is a natural number; Directing the K first input optical pulses to a first modulator pixel, thereby receiving K first modulated optical pulses at the time intervals, wherein the first modulator pixel is configured to modulate in correspondence with K first weights and the K first input optical pulses; Receiving the K first modulated optical pulses by a first detector pixel; and Accumulating, by an integrator in response to the K first modulated optical pulses, the charge of the first detector pixel to generate a result of a multiply-accumulate operation.
10. The method according to claim 9 further includes generating, by a second source pixel, K second input optical pulses having the time interval, wherein, The K first input optical pulses and the K second input optical pulses are respectively generated based on a first training source and a second training source, and the first source pixels and the second source pixels are arranged on the same panel.