Photonic Neural Network System
The photonic neural network system addresses computational and power challenges in convolutional neural networks by utilizing optical Fourier transforms for faster and more efficient data processing, achieving high accuracy with reduced power consumption.
Patent Information
- Application Number
- JP2023138980
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-02
- Filing Date
- 2023-08-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2038-09-20
AI Technical Summary
Convolutional neural networks face challenges with high computational demands and power consumption due to the processing of large amounts of data, which can overwhelm traditional digital computing capabilities.
A photonic neural network system that performs convolution using optical Fourier transforms, enabling faster and more power-efficient processing by converting frames of data into their Fourier equivalents for modulation and inverse Fourier transforms, allowing for full-frame parallelism and almost 100% efficiency.
The photonic neural network system achieves significantly faster and more power-efficient convolution operations compared to traditional digital methods, maintaining accuracy while reducing computational and power requirements.
Smart Images

Figure 0007708448000001 
Figure 0007708448000002 
Figure 0007708448000003
Abstract
Description
Technical Field
[0001] The present invention relates to neural networks, and more particularly to convolutional neural networks involving optical processing. Status of the Prior Art
[0002] Neural networks are well-known as computing systems with a large number of simple and highly interconnected processing elements that process information by dynamic state responses to external inputs. Neural networks are useful for pattern recognition and for clustering and classifying data. A computer can perform machine learning using a neural network to learn to perform some task by analyzing training samples. Usually, these examples are pre-labeled by a user. For example, a neural network set up as an object or image recognition system is supplied with thousands of exemplary images labeled as "cat" or "no cat", and then the results are used to identify cats in other images or, in some cases, to indicate the absence of cats in other images. Alternatively, such a neural network set up as an object recognition system is supplied with thousands of examples of images having various objects such as cats, cows, horses, pigs, sheep, cars, trucks, boats, and airplanes, labeled as such, and then the results are used to identify whether other images have cats, cows, horses, pigs, sheep, cars, trucks, boats, or airplanes therein.
[0003] A CNN (Convolutional Neural Network) is a type of neural network that uses multiple identical copies of the same neuron, enabling the network to have many neurons while significantly reducing the number of real-valued (or actual values) that describe how the neurons that require learning behave, making it possible to represent computationally large models. Convolution is a method of combining two signals to form a third signal. A CNN is typically implemented in software or programmable digital hardware.
[0004] Deep learning is a term used in stack neural networks, i.e., networks that include several layers. Layers are composed of nodes. A node (or knot) is a place where calculations are performed, and is loosely patterned after neurons in the human brain and fires when it encounters sufficient stimuli. A node combines the input from data with a set of coefficients or weights that amplify or attenuate that input, thereby assigning significance to the input for the task that the algorithm, such as which input is most useful for classifying the data without error, is trying to learn. These input × weight products are summed, and that sum passes through the activation function of the node to determine whether and to what extent that signal further progresses through the network to affect the final result, such as the act of classification. The node layer is a row of switches such as neurons that turn on or off when input is supplied through the network. The output of each layer is simultaneously the input of the subsequent layer, starting from the initial input layer that receives the data. Three or more node layers are considered "deep" learning. In a deep learning network, each layer of nodes trains a separate set of features based on the output of the previous layer, and thus, the more layers the data (e.g., pictures, images, sounds, etc.) passes through, the more complex the features the nodes can recognize. During training, a process called backpropagation is used for adjustment to increase the network's ability to predict the same type of image the next time. Such data processing and backpropagation are performed repeatedly until the prediction is moderately accurate and no improvement is seen. The neural network is then utilized in inference mode to classify new input data and predict the results inferred from that training.
[0005] A typical convolutional neural network has, in addition to an input layer and an output layer, four essential neuron layers, namely convolution, activation (or activation), pooling, and fully connected. In the first one or more convolutional layers, thousands of neurons act as the first set of filters that search for patterns and scrutinize all parts and pixels within an image. As more and more images are processed, each neuron gradually learns to filter specific features, which improves accuracy. In fact, one or more convolutional layers decompose an image into different features. Next, the activation layer emphasizes prominent features, such as those that are likely to have value or importance in the final classification result. For example, eyes are more likely to indicate a face rather than a frying pan.
[0006] All of the convolution and activation across the entire image can generate a large amount of data, which may overwhelm the computer's computing capacity. Therefore, pooling is used to compress the data into a more manageable form. Pooling is a process of selecting the best data and discarding the rest, resulting in a dataset with a lower resolution. Several types of pooling can be used, and some of the more common types are "max pooling" and "average pooling".
[0007] Finally, in the fully connected layer, each reduced or "pooled" feature map (feature map) or data is connected to output nodes (neurons) that represent items that the neural network has learned to recognize or can distinguish, for example, for cats, cows, horses, pigs, sheep, cars, trucks, boats, and airplanes. the above When the feature map or data is run through these output nodes, each node feature Vote on the map or data. The final output of the network for the image data that has passed through the network is based on the votes of the individual nodes. At the beginning of the training of the network, the voting may produce more incorrect outputs, but as the number of images and backpropagation increase to adjust the weights and refine the training, the accuracy improves, and thus, ultimately, the prediction or inference of the result from the input data can be very accurate.
[0008] The foregoing examples of the related art and the limitations related thereto are illustrative of the present subject matter but are not intended to be exclusive or exhaustive. Other aspects and limitations of the related art will become apparent to those of ordinary skill in the art upon reading this specification and examining the drawings.
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate some exemplary embodiments and / or features, but do not represent a sole or exclusive set of embodiments and / or features. The embodiments and drawings disclosed herein are intended to be considered illustrative rather than limiting. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 is a functional block diagram of an exemplary photonic neural network system.
[0011] FIG. 2 is an isometric view of an exemplary photonic convolution assembly for optically processing and convolving an image for the photonic neural network system of FIG. 1, with a portion of a second Fourier transform lens broken away to reveal an array of optical sensor-display components of a second sensor-display device.
[0012] FIG. 3 is a front view of an exemplary radial modulator in the exemplary photonic neural network of FIG. 1.
[0013] Figure 4 is an isometric view of the central portion of the radial modulator of the example of FIG. 3, together with an enlarged view of a modulator segment of an example of a radial modulator.
[0014] Figure 5 is an isometric view showing an example of a convolution function of a photonic convolution assembly of the photonic neural network system of the embodiment.
[0015] Figure 6 is a schematic top view of an example of the photonic convolution assembly of FIG. 2, showing a first sensor-display device for displaying a frame of data (image) and a second sensor-display device for detecting a convolution frame of the data.
[0016] Figure 7 is a schematic top view of an example of the photonic convolution assembly of FIG. 2, showing a second sensor-display device for displaying a frame of data (image) and a first sensor-display device for sensing a convolution frame of the data.
[0017] Figure 8 is a functional block diagram of an array of transmit-receive modules (or transceiver modules) within a first sensor-display device. External Interface is the external interface, Column Controls are column controls, and Row Controls are row controls.
[0018] Figure 9 is an enlarged isometric view of a part of an array of transceiver modules.
[0019] Figure 10 is an enlarged isometric view of an exemplary transceiver module.
[0020] Figure 11 is a perspective view of an example of an optical transmitter element of the transceiver module examples of FIGS. 9 and 10.
[0021] Figure 12 is a functional block diagram of an exemplary system interface to the external interface of a sensor-display device. To All RedFives is to all RedFives. Output is the output, and Analog in is the analog input.
[0022] Figure 13 is a functional block diagram of an exemplary external interface of a sensor-display device. Digital Interface is a digital interface, Analog Interface is an analog interface, Row & Column Controls are row and column controls, Global Controls are global controls, Analog Bus Buffers are analog bus buffers, and Trixel Array is a Trixel array.
[0023] Figure 14 is a schematic diagram of row and column control line registers (or row and column control line registers) for an array of transceiver modules.
[0024] Figure 15 is a schematic diagram of analog data lines to transceiver modules.
[0025] Figure 16 is a schematic diagram of some of the transceiver modules (trixels) in an array interconnected by a pooling chain.
[0026] Figure 17 is an enlarged schematic diagram of the interconnection between the pooling boundary line (or pooling boundary line) of a transceiver module (pixel) and adjacent transceiver modules (trixels).
[0027] Figure 18 is an exemplary memory shift driver for a memory bank in an exemplary transceiver module of a photonic neural network system 10. MEM SHIFT DRIVER is a memory shift driver.
[0028] Figure 19 is a schematic diagram of an exemplary analog memory read interface for a memory bank. MEM READ PAD is a memory read pad, APPLY ReLU is to apply ReLU, READ AMP is a read amplifier, and Pooling Chain is a pooling chain.
[0029] Figure 20 is a schematic diagram showing the analog memory read average of the transceiver module (trixel) to the pooling chain. MEM READ PAD is the memory read pad, APPLY ReLU applies ReLU, READ AMP is the read amplifier, and Pooling Chain is the pooling chain.
[0030] Figure 21 is a schematic diagram showing the analog memory read maximum value of the transceiver module (trixel) to the pooling chain. MEM READ PAD is the memory read pad, APPLY ReLU applies ReLU, READ AMP is the read amplifier, and Pooling Chain is the pooling chain.
[0031] Figure 22 is a schematic diagram showing the analog memory of the transceiver module (trixel) read out to the external data line. MEM READ PAD is the memory read pad, APPLY ReLU applies ReLU, READ AMP is the read amplifier, and Pooling Chain is the pooling chain.
[0032] Figure 23 is a schematic diagram showing the peak value storage of the analog memory of the transceiver module (trixel). MEM READ PAD is the memory read pad, APPLY ReLU applies ReLU, READ AMP is the read amplifier, and Pooling Chain is the pooling chain.
[0033] Figure 24 shows the peak value reset of the analog memory of the transceiver module (trixel). MEM READ PAD is the memory read pad, APPLY ReLU applies ReLU, READ AMP is the read amplifier, and Pooling Chain is the pooling chain.
[0034] FIG. 25 shows a graphical representation of an example of ReLU (rectified linear unit) response. Max Voltage represents the maximum voltage, Min Voltage represents the minimum voltage, Linear (Non-ReLU) Response represents the linear (non-ReLU) response, Classical ReLU Response represents the classical ReLU response, and Soft ReLU Response represents the soft ReLU response.
[0035] FIG. 26 is a schematic diagram showing the writing to the analog memory of the transceiver module (trixel). MEM DRIVER is the MEM (memory) driver, MEM WRITE PAD is the MEM write pad, and Pooling Chain is the pooling chain.
[0036] FIG. 27 is a schematic diagram showing the loading of the analog memory from the external data line. MEM DRIVER is the MEM (memory) driver, MEM WRITE PAD is the MEM write pad, and Pooling Chain is the pooling chain.
[0037] FIG. 28 is a schematic diagram showing the flag memory write circuit. MEM DRIVER is the MEM (memory) driver, MEM WRITE PAD is the MEM write pad, and Pooling Chain is the pooling chain.
[0038] FIG. 29 is a schematic diagram showing the flag memory read circuit. MEM DRIVER is the MEM (memory) driver, MEM WRITE PAD is the MEM write pad, and Pooling Chain is the pooling chain.
[0039] FIG. 30 is a schematic diagram showing the optical control line setting for reading the transceiver module (trixel) sensor into the pooling chain. SNSR READ PAD is the SNSR (sensor) read pad, SNSR AMP is the SNSR (sensor) read pad, and Pooling Chain is the pooling chain.
[0040] Figure 31 is a schematic diagram showing an optical control line for resetting a transceiver module (trixel) sensor. SNSR READ PAD is the SNSR (sensor) read pad, SNSR AMP is the SNSR (sensor) read pad, and Pooling Chain is the pooling chain.
[0041] Figure 32 is a schematic diagram showing an optical control line setting for writing an optical transmitter element (modulator) from a pooling chain. MOD READ PAD is the MOD read pad, MOD DRIVER is the MOD driver, and Pooling Chain is the pooling chain.
[0042] Figures 33A - B show a schematic diagram of a transceiver module (trixel) circuit. MEM READ PAD is the memory read pad, READ AMP is the read amplifier, APPLY ReLU is to apply ReLU, and MEM SHIFT DRIVER is the MEM shift driver.
[0043] Figure 34 shows an exemplary photonic convolution assembly having a Fourier optical sensor device for Fourier - transforming a frame of correction data in a training mode.
[0044] [FIG. 35] Figure 35 is a graphical isometric view of an example Fourier optical sensor device
[0045] Figure 36 shows an example of a photonic convolution assembly with a camera lens embodiment for introducing a real - world frame of data (image) into the photonic convolution assembly.
DETAILED DESCRIPTION OF THE INVENTION
[0046] An exemplary photonic neural network system (or, a photon neural network system) 10 is shown in a functional block diagram in FIG. 1, and an isometric view of an exemplary photonic convolution assembly 12 for optically processing and convolving an image for the photonic neural network system 10 is shown in FIG. 2. The convolution by this photonic neural network system 10 is performed by an optical Fourier transform, which significantly increases speed, resolution, and power efficiency compared to digital spatial convolution. Thus, generating and using neural networks can be done much faster and with much less power consumption than processing and computing convolution by typical computer algorithms. Since all of convolution and summation are completely analog, full-frame photonic calculations, the power consumption is very low. As will be explained below, addition (or, summing) is achieved by building charges in a capacitive optical sensor, which is an analog process. The sensor has very low noise and no clocking or other transient noise sources, so summing is a very low-noise process. The photonic neural network 10 can accept and process any data such as images, videos, audio, speech patterns, or anything that is normally processed by a convolutional neural network, and supports all existing convolutional neural network architectures and training methods. The photonic neural network 10 also provides full-frame image parallelism using an architecture where all data elements are in their ideal positions for the next stage at full resolution processed at the speed of light, and thus is almost 100 percent efficient. Other advantages can be understood from this description.
[0047] Referring to both FIGS. 1 and 2, for example, the optical processing of an image for the photonic neural network system 10 is performed by the photonic convolution assembly 12. Essentially, as will be described in more detail below, the first sensor display device 14 projects a frame of data (e.g., an image of other data such as sound, speech pattern, video, etc. or an optical representation) as a modulated light field 15 through a first Fourier transform lens 16 and a polarizer 18 onto a radial (or radial direction) modulation device 20 disposed on the focal plane of the lens 16. The frame of data projected by the first sensor display device 14 is formed by the first sensor display device 14 based on values or signals supplied to the first sensor display device 14 by support electronics (described in more detail below) via an electronic data interface 22. The Fourier transform lens 16 can use a diffractive lens, a solid convex lens, or any other form of Fourier transform lens. Also, a fiber faceplate (not shown) can be disposed in front of the lens 16 to collimate the light before it enters the lens 16.
[0048] The lens 16 converts a frame of data (e.g., an image) to its Fourier equivalent at the focal plane (also called the Fourier transform plane), and thus, at the surface of the radial modulation device 20. The radial modulation device 20 modulates the light field 15 containing the Fourier equivalent of the frame of data based on a pattern (also called a "filter") loaded into the radial modulation device 20 by support electronics (described in more detail later) via an electronic data interface 24, and reflects the modulated frame of data to a second sensor display device 26, which detects the result. The reflected light field containing the modulated frame of the data inverse Fourier transform is inverse-transformed in the spatial domain at the distance from the radial modulation device 20 to the second sensor display device 26, and thus, the modulated data frame incident on the second sensor display device 26 has passed through, i.e., is the spatial domain feature of the data frame not removed by the filter by the radial modulation device 20. The result is for each pixel of the second sensor display on the Fourier transform plane device. on the Fourier transform plane The result is the spatial region feature of the data frame that was not removed by the filter by the radial modulation device 20. The result is for each pixel of the second sensor display device Detected by 26, where the light incident on each pixel generates charge in proportion to the intensity of the light and the time the light is incident on the pixel. The first sensor-display device 14 Each frame of data output from can be modulated by one or more filters (patterns) by the radial modulation device 20. Also, the second sensor-display device 26 can receive one or more data frames modulated by one or more filters applied to the radial modulation device 20 from the first sensor-display device 14. Therefore, the charge accumulation of each pixel in the second sensor-display device 26 can be the sum of one or more modulated (i.e., filtered) patterns of one or more frames of data, as will be described in more detail below, thereby constituting the convolution of one or more frames of data projected by the first sensor-display device 14.
[0049] For example, one frame of data is sequentially projected by the first sensor-display device 14, first projected in red, then in green, then in blue, and the radial modulation device (or, radial modulator / radial modulator) 20 can apply the same or different filters (pattern modulation) to each of the red, green, and blue projections. All of these modulated frames of data are added to the charge of each pixel by the light from each of the sequentially modulated frames of data and can be sequentially detected by the second sensor-display device 26. Then, of the second sensor-display device 26 for each pixel, those charges are of the second sensor-display device 26 transferred to their respective memory cells, and those memory cells each pixel of the second sensor-display device 26 store these total results for each pixel, thereby including the stored pixel values of the convolution of the frames of data in the spatial region projected by the first sensor-display device 14 and convolved in the Fourier transform region by the filters in the radial modulation device 20. of the second sensor-display device 26
[0050] The process is for the red, green, and blue projections of the same frame of data from the first sensor-display device 14, but with different filters within the radial modulation device 20, and thus can be repeated with different modulation patterns from the Fourier transform region reflected by the radial modulation device 20 to the second sensor-display device 26, thereby resulting in another set of accumulated pixel values and another addition result for another convolution frame of data within the memory bank of the second sensor-display device 26. What accumulates the convolution frames of data convolved within the second sensor-display device 26 forms a 3D convolution block from the frame of data projected by the first sensor-display device 14 for all of these different filter applications by the radial modulation device 20. In summary, the frame of data from the first sensor-display device 14 is multiplied by the radial modulation device 20 by a series of filters in the Fourier plane and added by the second sensor-display device 26 in a sequence that constructs a 3D convolution block in the memory of the second sensor-display device 26. Assuming sufficient memory capacity to store all of the pixel values for all of the convolution frames of data within the 3D convolution block, any number of such convolution frames of data can be accumulated within the 3D convolution block. That 3D convolution block can be considered as the first level in a neural network.
[0051] For the next convolution block or level, the first sensor-display device 14 and the second sensor-display device 26 swap functions. The 3D convolution block in the memory of the second sensor-display device 26 becomes a frame of data for the next convolution sequence. For example, the data of each convolution frame stored in the 3D convolution block in the memory of the second sensor-display device 26 is projected by the second sensor-display device 26 through the second Fourier transform lens 28 onto the radial modulation device 20, where it is multiplied by a filter and reflected to the first sensor-display device 14. The first sensor-display device 14 detects and adds a series of such convolution and added data frames to construct the next 3D convolution block in the memory of the first sensor-display device 14.
[0052] This process cycle is schematically shown in FIG. 5.
[0053] These convolution processes can repeat the frame of data projected back and forth (or mutually) between the first sensor-display device 14 and the second sensor-display device 26 as many times as required for any convolutional neural network architecture. When more filters are applied in subsequent cycles, as will be described in more detail below, instead of feeding the accumulated charge from each pixel detection to individual memory cells, the accumulated charge from multiple pixel detections can be fed to one memory cell to pool the convolution. Thus, a convolutional neural network with many levels of abstraction can be developed using the exemplary photonic neural network 10.
[0054] A front view of an exemplary radial modulation device 20 is shown in FIG. 3, and an exemplary segment light modulator of the exemplary radial modulation device 20 element 40 A perspective view of the central portion of an exemplary radial modulator device 20 having an enlarged view is shown in FIG. 4. The radial modulator device 20 has an optically active region 30 comprising a plurality of optical modulation wedge segments 32 (wedge segments), each of which is independently operable to modulate light incident on the respective wedge segment 32. In the example radial modulation device 20 shown in FIGS. 2, 3, and 4, the wedge segments 32 are grouped into a plurality of wedge sectors 34, each wedge sector extending radially outward from a central component 36, and together they form the optically active region 30 of the radial modulation device 20. In FIGS. 3 and 4, to avoid confusion in the drawings, only some of the wedge segments 32 and sectors 34 are labeled with these reference numerals, but those skilled in the art will understand where all of the wedge segments 32 and wedge sectors 34 are located within the exemplary radial modulator device 20 in this figure. In the example radial modulation device 20 shown in FIGS. 3 and 4, the wedge segments 32 are arranged to form a circular optically active region 30, but other shapes can also be used.
[0055] As described above, each of the wedge segments 32 is optically active in the sense that each wedge segment 32 can be activated to transmit light, block light, or modulate the transmission of light between full transmission and blocking. Thus, a beam or field of light incident on the optically active region 30 can be modulated by any combination of one or more wedge segments 32. Spatial light modulators can be designed and manufactured to modulate light in many ways. For example, U.S. Patent No. 7,103,223, issued September 5, 2006 to Rikk Crill, illustrates the use of a birefringent liquid crystal material for modulating wedge segments in a radial spatial light modulator similar to the radial modulation device 20 of FIGS. 2 and 3. The paper by Zhang et al., "Active meta surface modulator with electro-optic polymer using bimodal plasmonic resonance", Optics Express, Vol. 25, No. 24, November 17, 2017) describes an electrically tunable metal grating having an electro-optic polymer suitable for ultra-thin surface normal applications that modulates light. The optically active wedge segments 32 in the exemplary radial modulator device 20 Such a metasurface optical modulator element is shaped for use as an exemplary segment optical modulator 40 having a metallic grating structure 42 for element 40 and is shown in FIG. 4. The grating structure 42 is sandwiched between a bottom metal (e.g., Au) layer 46 and an interdigitated top thin film metal (e.g., Au) grating layer 48, all constructed on the substrate 50 interdigitated and includes an electro-optic polymer 44. The diffraction grating 42 has a period such that diffraction is prohibited incident for light LShorter than the wavelength of . The thickness of the upper metal layer 48 is greater than the skin depth in order to eliminate the direct coupling from the incident light L to the electro-optic polymer 44. The bottom metal layer 46 is also of the same thickness so as to operate as an almost perfect reflective mirror. Essentially, the light L enters at the top of the metasurface optical modulator element 40, is phase-shifted within the electro-optic polymer 44 that is periodically poled by the application of the poling voltage 45, is reflected from the bottom metal layer 46, is further phase-shifted during its second (i.e., the reflected) pass, and exits from the top surface together with the polarization of the light that has been rotated by 90 degrees. The other wedge segments 32 in the radial modulator 20 of the embodiment have the same type of optical modulator element 40 can is sized and shaped to fit and substantially fill each particular wedge segment 32. The central component 36 can also have an optical modulator element 40.
[0056] As shown in FIGS. 2 to 7 and as described above, for example, the radial modulator 20 is a reflection device, where the incident light is modulated and reflected by the wedge segments 32. However, the radial modulator can alternatively be a transmission device where the incident light is modulated and transmitted through the radial modulator. Of course, the positions of the optical components, such as sensor display devices, lenses, and polarizers, must be rearranged in an appropriate sequence for each optical component to route the optical field, but those skilled in the art will understand how to make such rearrangements if they are familiar with the exemplary photonic neural network 10 described above.
[0057] As shown in FIGS. 3 and 4 and as briefly described above, the optically active wedge segments 32 are grouped into a plurality of wedge sectors 34 that extend radially from the round central component 36 to the periphery of the optically active region 30. Also, the wedge segments 32 are arranged in a concentric ring shape around the central component 36. Each concentric ring of the wedge segments 32 other than the innermost concentric ring has an outer radius that is twice the outer radius of the immediately adjacent inner ring, which corresponds to the scale distribution in the Fourier transform. Thus, each of the wedge segments 32 that follows radially outward within the wedge sector 34 is twice as long as the immediately preceding wedge segment 32. A detailed description of the radial modulator that functions as a filter on the Fourier transform plane of an image can be found, for example, in U.S. Patent No. 7,103,223 issued to Rikk Crill on September 5, 2006. Here, the optical energy from the higher spatial frequency shape content in the spatial domain is dispersed more radially outward than the light energy from lower on the Fourier transform plane spatial frequency content, while it is sufficient to say that the angular orientation and intensity of the optical energy from both the higher spatial frequency shape content and the lower spatial frequency content are preserved in the Fourier transform of the image. Thus, the optical energy transmitted (or, passed through) by a particular wedge segment 32 arranged at a particular angular orientation and at a particular radial distance from the center (optical axis) of the Fourier transform image within the Fourier transform plane is inverse Fourier transformed in the projection to display only the shape content (shape content) (features) from the original image having the same angular orientation as the particular wedge segment 32, and the shape content (features) of that angular orientation having a spatial frequency within the range corresponding to the radial spread by which such optical energy is dispersed within the Fourier transform plane. returned to the spatial region The light intensity (brightness) of these inverse Fourier-transformed features (shape contents) corresponds to the light intensity (brightness) that these features (shape contents) had in the original image, and they are in the same positions as those in the original image. Of course, the shape contents (features) included in the light energy of the original image that are blocked and not transmitted by a specific wedge segment 32 in the Fourier transform plane will be missing when returning to the spatial domain in the inverse Fourier-transformed image. Also, the shape contents (features) composed of light energy that are only partially blocked and thus partially transmitted by a specific wedge segment 32 in the Fourier transform plane are inverse Fourier-transformed into a spatial domain having the same angular orientation and a specific spatial frequency as described above, but the intensity (luminance) decreases. (Or, also, the shape contents (features) composed of light energy that are only partially blocked and thus partially transmitted by a specific wedge segment 32 in the Fourier transform plane are inverse Fourier-transformed into a spatial domain at the same angular orientation and a specific spatial frequency as described above, but the intensity (luminance) decreases.) Therefore, as explained above and described in more detail below, some of the shape contents (features) of the original image are preserved in the inverse Fourier-transformed image with complete or partial intensity (brightness), and some of the shape contents (features) are partially or The inverse Fourier-transformed image returned to the spatial domain, where it is completely deleted, is the convolutional image detected and used in the construction of the 3D convolutional block of the neural network, as shown in FIG. 5.
[0058] Accordingly, referring to FIG. 5, a first filter 54 is loaded into the radial modulation device 20 via the data interface 24, whereby the wedge segment 32 is set to block light in a pattern set by the first filter 54, or to transmit light completely or partially. For example, a first frame of data including an image of the mountain of the LEGO® toy building block 52 is loaded into the first sensor-display device 14 via the data interface 22, and thus the display component (not shown in FIG. 5) within the first sensor-display device 14 is set to display a frame of data including an image of the LEGO® toy building block 52 as seen in FIG. 5. The laser irradiation 13 is essentially directed onto the first sensor-display device 14 that irradiates a frame of data 50 through the first Fourier transform lens 16 and through the polarizer 18 onto the radial modulator device 20, which is disposed at the focal distance Fl from the first Fourier transform lens 16, i.e., within the focal plane of the first Fourier transform lens 16 as also shown in FIG. 6. The Fourier transform lens 16 focuses the light field 15 including the image 50 onto the surface of the radial modulation device 20 at the focal point. A frame of data 50 including an image of the LEGO® toy building block 52 is in the Fourier transform region of the radial modulator 20 Filtered by the wedge segment 32, which, as described above, either fully or partially reflects some light or blocks some light that constitutes an image. The wedge segment 32 phase-shifts the reflected light and thus rotates the polarization, so that the light reflected by the radial modulation device 20 is reflected by the polarizer 18 to the second sensor-display device 26 as shown by the reflected light field 56. Thus, as described above, some shape contents (features) of the original frame of the data 50 of the image of the LEGO (registered trademark) toy construction block 52 are missing or not as strong, i.e., filtered out, in the convolutional image incident on the second sensor-display device 26 as shown in FIG. 5. The frame of the convolutional data (image) is detected by the second sensor-display device 26 and added to several subsequent convolutional images by the second sensor-display device 26 to form the first convolutional addition data frame (image) 58. The first convolutional added data frame (image) 58 is schematically shown in FIG. 5 as the 3D convolutional block 6 5 For constructing, it can be transferred to a memory bank for accumulation with subsequent convolutional added data frames (images).
[0059] Next, the second sensor display device 26 and the first sensor display device 14 swap roles as described above, with the second sensor display device 26 entering the display mode and the first sensor display device 14 entering the sensor mode. When the second sensor display device 26 is set to the display mode, then the first convolution and addition data frame (image) 58 is projected by the second sensor display device 26 onto the radial modulation device 20 as schematically shown in FIG. 7, where it is convolved with an additional filter and then reflected by the radial modulation device 20 back to the first sensor display device 14. This role swap is shown schematically in FIG. 7, where the second sensor display device 26 switches to the display mode and the first sensor display device 14 switches to the sensor mode. In the display mode, the display components of the second sensor display device 26 are programmed to display its first convolution and addition data frame (image) 58. Thus, the laser irradiation 60 on the second sensor display device 26 irradiates the first convolution and addition data frame (image) 58 along the second optical axis 62 through the second Fourier transform lens 28 onto the polarizer 18, which reflects the light field 64 along the first optical axis 61 to the radial modulation device 20. The optical distance between the second Fourier transform lens 28 and the radial modulation device 20 along the second optical axis 62 and the first optical axis 61 is equal to the focal length of the second Fourier transform lens 28. Therefore, the light field 64 at the radial modulation device 20 on the Fourier transform plane is the Fourier transform of the first convolution and addition data frame (image) 58. The radial modulation device 20 applies a filter to the Fourier transform of the first convolution and addition data frame to provide a second convolution to the data frame, and as described above, reflects it with a phase shift and propagates it along the first optical axis 61 to the first sensor display device 14. The first sensor display device 14 then detects the frame (image) of the data convolved by the filter applied by the radial modulation device 20 in the now-exchanged role of the detector as described above.Next, the convolution frames of the data (images) detected by the first sensor-display device 14 are added by the first sensor-display device 14 with several other convolution frames of the data (images) subsequently detected by the first sensor-detection device 14. The frames of such convolution-added data (images) are transferred to a memory bank and used to construct the second 3D convolution block 66, which is shown diagrammatically in FIG. 5 as shown diagrammatically in FIG.
[0060] Next, the roles of the first and second sensor-display devices 14, 26 are exchanged again. The convolution and sum frames of the data sensed and summed by the first sensor-display device 14 are back-projected through the system in the same manner as described above, convolved by the radial modulator device 20, and then detected and summed by the second sensor-display 26 to continue constructing the first 3D convolution block 65 and sent back through the system, convolved with additional filters and sums to continue constructing the second 3D convolution bank 66. This process is repeated the desired number of times to build deeper and deeper convolutions or until the inference neural network is complete.
[0061] The first sensor-display device 14 and the second sensor-display device 26 each construct 3D convolution blocks, as will be described in more detail below. 65, to construct blocks 66 and subsequent convolutional blocks, it is possible to have a memory bank for receiving and summing the data (images) of the subsequently received convolutional frames and storing the convolutional frames of the data (images). Thus, except for the first frame of the data (images) loaded into the system, the input frames of the data can always reside in one of the memory banks of the sensor-display devices 14, 26 from the previous convolutional cycle. A series of filters 68 are loaded into the radial modulation device 20 in synchronization with the display of the frames of the data (images) by the respective first and second sensor-display devices 14, 26 for convolution of the frames of the data (images) by the filters.
[0062] Except for being optically computed in the Fourier transform domain, the convolution in this photonic neural network system 10 is the same as the convolution computed by traditional digital methods. However, as will be described in more detail below, the all-frame parallelism at any resolution processed at the speed of light in an architecture where all data elements are in the ideal position for the next convolutional stage of the cycle is almost 100% efficient. Thus, as described above and as will be described in more detail below, constructing the convolutional blocks in the photonic neural network system 10 of the embodiments provides much more power and speed than the convolution computed by traditional digital methods.
[0063] As described above, each of the first and second sensor-display devices 14, 26 has the ability to both detect light and display images on a pixel-by-pixel basis. In this example of the photonic neural network system 10, since the first sensor-display device 14 and the second sensor-display device 26 have essentially the same components and structure, the details of those devices will be described mainly with reference to the first sensor-display device 14, but it is understood that such details also apply to the second sensor-display device 26. Therefore, in the following description, the first sensor-display device 14 may sometimes be simply referred to as the sensor-display device 14. A functional block diagram of an example sensor-display device 14 is shown in FIG. 8 and includes an array 80 of transceiver modules (transmission / reception modules) 82, each of which has a light transmission and light detection element and a memory bank, as will be described in more detail below. Row and column control for the transceiver modules 82 within the array 80, as well as a mixed analog and digital interface 24 to an external control circuit (not shown in FIG. 8), are provided for input / output data, which will be described in more detail below. An enlarged portion of the array 80, shown schematically in FIG. 9, illustrates an example of a transceiver module 82 within the array 80, and a further enlarged graphical representation of the exemplary transceiver module 82 is shown in FIG. 10. Each of the exemplary transceiver modules 82 includes both a micro light transmitter element 84 and a micro light detector (sensor) element 86, which are the transceiver modules 82Neural network results that are at least as useful as those from typical computational convolutions, and, similar to processing by computer algorithms, are small enough and close enough to each other to effectively function as a light transmitter and a light receiver at substantially the same pixel positions of an image or a frame of data with sufficient resolution. For example, for a neural network using a representative photonic neural network system 10 to produce results as useful as those from typical computational convolutions and processing by computer algorithms, the micro optical transmission element 84 and the micro optical detection element 86 can be offset from each other by 40 micrometers or less and both can be fitted within a transceiver module 82 having an area of 160 square micrometers or less.
[0064] As best shown in FIG. 10, in addition to the optical transmitter element 84 and the optical sensor or optical receiver detector element 86, an exemplary transceiver module 82 includes a modulator driver 88, a memory bank 90, a memory interface 92, analog and digital control elements 94, a pooling connection 96 for making a pooling connection with an adjacent transceiver module 82 in the array 80, a pooling control element 98, and a sense amplifier 100. In the display mode, for example, a first sensor-display device 14 (FIGS. 2, 5, and 6) As described above, when one frame of data (image) is projected onto the radial modulator 20, as shown in FIGS. 2, 5, and 6, the laser irradiation is directed towards the back of the transceiver module 82. The first frame of the data (image) is composed of pixel values for each pixel of the frame of the data (image). These pixel values are 13 (see FIGS. 2, 5, and 6) of the optical field 15 to generate the first frame of the data (image) in a pattern in the array 80 (FIGS. 8 to 10) 10, for each one of the transmit / receive modules 82, the pixel value is provided to an analog and digital control element 94, which shifts the pixel value to a modulator driver 88. The modulator driver 88 shifts the pixel value to the optical transmitter element 84 according to the pixel value. (FIG. 10) By modulating the voltage above, of the other transceiver modules 82 of the array 80 (FIGS. 8 to 10) The other phototransmitter elements 84 transmit light for each pixel while the laser is illuminated. 13 (FIGS. 2, 5, and 6) 3, modulating the laser illumination incident on the optical transmitter elements 84 in a manner that transmits a pixel of 15 (FIGS. 2, 5, and 6) After the first frame of data (image) is transmitted by the first sensor and display device 14 and the convolved frame of data (image) is returned to the first sensor and display device 14, the light field containing the frame of convolved data (image) is incident on the sensors 86 of all the transceiver modules 82 in the array 80 of the first sensor and display device 14. Thus, the array 80 The sensors 86 on each transceiver module 82 in the sensor 86 detect pixels of the incident light field and therefore pixels of a frame (image) of data (or date) constituted by (or contained in) the incident light field. Those skilled in the art will understand how light sensors, such as charge-coupled devices (CCDs), are constructed and function, and such light sensors or similar light sensors can be used for the sensors 86. Essentially, each light sensor has a light-sensitive photodiode or capacitive component that responds to incident photons by absorbing most of the energy in the photon, generating a charge proportional to the incident light intensity, and storing that charge in the capacitive component. The longer the light is incident on the sensor, the more charge accumulates in the capacitive component. Thus, each pixel of light energy incident on each sensor 86 causes a charge to build up in that sensor 86, the magnitude of the charge being proportional to the intensity of the incident light in that pixel and the time that the light in that pixel is incident on the sensor 86.
[0065] As discussed above, when a series of convoluted frames of data (images) are transmitted to and received by the sensor and display device 14, the light energy (photons) of the successive light fields comprising the successive (sequence) frames of data (images) will cause a charge to be generated in the sensor 86 such that successive pixels of the light field energy from the successive light fields sensed by each individual sensor 86 can be accumulated, i.e., added to the capacitive component of that individual sensor 86, thereby resulting in a stored charge in the sensor 86 that is the sum of the light energies from the sequence of light fields at that particular pixel location. Thus, the sequence of convoluted frames of data (images) received by the sensor and display device 14 are sensed and summed pixel by pixel by the array 80 of transceiver modules 82 of the sensor and display device 14. Then, as discussed above, once a predetermined number of individual convoluted frames of data (images) have been received and summed, the accumulated (summed) charge in the sensor 86 of each individual transceiver module 82 is shifted to the memory bank 90 of that individual transceiver module 82. The same action of shifting the accumulated (summed) charge in the sensor 86 to the memory bank 90 also 80 Therefore, when the shift operation is performed, the transmission / reception modules 82 Array of 80 a module for transmitting and receiving a complete convolved and added frame of data (image) resulting from that series or successive convolutions and additions of input frames of data (images) 82 10, the pixel values for that particular transmit / receive module 82's pixel location for that initial (or first) convolved and summed frame of data (image) are shifted from the sensor 86 to a first memory cell 102. Thus, the combination of all the first memory cells 102 in the memory banks 90 of all transmit / receive modules 82 in the array 80 comprises a pixel-by-pixel convolved and summed frame of data (image).
[0066] Next, as described above, when a subsequent second series or sequence of frames of data (images) are convolutionally added, the charge accumulated in sensor 86 for that pixel of the resulting second convolutionally added data (image) frame is shifted into the first memory cell 102 as the charge from that pixel of the first convolutionally added data (image) is simultaneously shifted into the second memory cell 104 of memory 90. A detailed description is not necessary for the purposes of this explanation as one of ordinary skill in the art will understand how such shift register memory is made and operated. This same process occurs simultaneously in other transceiver modules 82 within array 80. Thus, the combination of all first and second memory cells 102, 104 in memory banks 90 of all transceiver modules 82 of array 80 includes the first and second convolutionally added frames of data (images) on a pixel-by-pixel basis.
[0067] As more subsequent series or sequences of frames of data (images) are convolved and summed up as described above, the summed pixel values of the sequentially (or in a sequence) convolved and summed frames of such data (images) are sequentially shifted to the first memory cell 102, while each preceding pixel value is further shifted along the memory cells of the memory bank 90. This process is performed simultaneously in all of the transceiver modules 82 of the array 80 as described above. Thus, all of such convolved and summed frames of data (images) from all of the series or sequences of convolution and addition are stored pixel by pixel in the memory cells of the memory bank 90 within the array 80 of the transceiver module 82. Each of such convolved and summed frames of data (images) may be called a convolution. Thus, the array 80 of the transceiver module 82 can hold the same number of convolutions pixel by pixel as there are individual memory cells present within the individual memory banks 90 of the transceiver module 82. For example, the exemplary transceiver module 82 schematically shown in FIGS. 9 and 10 each has a memory bank 90 composed of 64 individual memory cells 102, 104,..., n. Thus, the exemplary array 80 of the transceiver module 82 can hold data (images) of 64 frames pixel by pixel at full resolution. When the optical transmission element 84 and the optical sensor element 86 within the transceiver module 82 (see FIG. 1) are pooled with the optical transmission element 84 and the optical sensor element 86 of the adjacent transceiver module 82 as will be described in more detail below, all of the optical transmission elements 84 and the optical sensor elements 86 within the pooled group display the same luminance for a coarser representation of the frame of data (images). Under such pooling conditions, the memory banks 90 of the transceiver modules 82 within the pooled group can be sequentially used to sense and store the summed results for all of the transceiver modules 82 of the entire pooled group, thereby increasing the effective memory capacity and depth.For example, when transmission / reception modules 82 each having a memory bank 90 including 64 memory cells are pooled into a 5×5 group, i.e., when 25 transmission / reception modules 82 per group are pooled, the effective memory capacity or depth of each group is 1,600 memory cells (64×25 = 1,600). Thus, sequential convolution addition frames of data (images) are first supplied to one of the transmission / reception modules 82 within the group until the memory bank 90 of that transmission / reception module 82 is filled, and then further sequential convolution added data (image) frames are supplied to a second transmission / reception module 82 within the group until the memory bank 90 of that second transmission / reception device 82 is filled, and then can continue to sequentially fill the respective memory banks 90 of the remaining transmission / reception modules 82 within the group. When the memory banks 90 of all transmission modules 82 within the group are filled, that block of convolution within the memory becomes 1,600 in depth. The aggregation of convolutions within the memory 90 of the transmission / reception module 82 within the array 80 is, for example, the 3D convolution block 65 schematically shown in FIG. 5. When the desired number of such convolutions is accumulated in the array for the last 3D convolution block, it can be read out from the memory bank 90 pixel by pixel for transmission by the sensor-display device 14 and returned through the electronic data interface 22 at the end of the process to output the neural network result.
[0068] However, during the deep learning process of repeatedly further convolutional and adding the data frames using the exemplary photonic neural network system 10, it is important to reiterate that the pixel values of the most recently formed convolutional block still exist within the memory cells of the individual memory banks 90 within the individual transceiver modules 82. Thus, when the sensor-display device 14 switches from the sensor mode where the convolutional block is stored in the memory 90 of the transceiver module 82 to the display mode where the convolutional block is transmitted back and forth through the optical components of the system 10 for deeper convolutional processing, the pixel values for each of the convolutional frame and the sum frame of the data (image) having (or constituting) the convolutional block can be directly read out (shifted) to the modulator driver 88 without further processing and transfer of the data to and from external computer processing, memory, and other components or functions from the memory cells 102, 104,..., n of the memory 90. Instead, when switching from the sensor mode to the display mode, the pixel values of the individual frames of the data (image) having (or constituting) the convolutional block are sequentially read out (shifted) from the memory 90 directly to the modulator driver 88, whereby the optical transmitter element 84 is driven to modulate the laser light incident on the transceiver module 82 in a manner that writes (imposes) the pixel values of the frame of the data (image) to be further convolved in that convolutional cycle onto the optical field. Thus, all of the transceiver modules 82 within the array 80 switch to the display mode simultaneously and the pixel values in each of them are written (imposed) into the laser light field, so that the synthesis of those pixel values in the optical field transmitted by the sensor-display device 14 replicates the frames previously convolved with the data (image) added and stored in the memory banks 90 of the transceiver modules 82 within the array 80. An optical field having a frame (image) of pre-convolved and summed data is projected through a Fourier transform lens 16 onto a radial modulator 20 for further convolution with a filter in the Fourier transform plane and then detected by another (e.g., second) sensor-display device 26 as described above.
[0069] Also, as described above, these convolution and addition processes are repeated many times through many cycles with many filters. Also, the first sensor-display device 14 and the second sensor-display device are aligned on respective optical axes 61, 62 such that the transceiver module 82 of the first sensor-display device 14 is optically aligned with the corresponding transceiver module 82 of the second sensor-display device 26 (see FIGS. 2, 6, and 7), and as a result, there is complete optical alignment between the respective arrays 80 of the first and second sensor-display devices 14, 26 including between the corresponding transceiver modules. Thus, the exemplary photonic neural network 10 performs full-frame full-resolution full-parallel convolution at the speed of light. Other effects such as gain, threshold (ReLU), max or average pooling, and other functions are simultaneously executed in dedicated circuitry as will be described in more detail below, and these effects exhibit no additional time delay. For example, virtually any convolutional neural network architecture including VGG16 or Inception-Resnet-v2 can be accommodated. All processing is done completely on the sensor-display devices 14, 26 without relocating the frames of data (images) to or from these devices. In the inference operation, the user application only needs to load the image and receive the result a few microseconds later.
[0070] The micro-optical transmitter element 84 within the transceiver module 82 can be any optical modulator device that emits or modulates light. The above-described illustrative photonic neural network system 10 includes an optical transmitter element 84 that modulates the laser light incident on the back surface of the optical transmitter element by allowing light to pass through the optical transmitter element or suppressing it. However, the optical transmitter element 84 can be replaced with a reflective light modulator device that modulates the incident light and reflects it, which requires laser irradiation to be incident on the same surface of the optical transmitter element where the light is reflected and will require rearrangement of the optical elements, as will be understood by those skilled in the art after understanding the above-described example of the photonic neural network. As another alternative, the optical transmitter element 84 can be replaced with a light emitter, thereby eliminating the need for the laser light field to be incident on the back surface and pass through the modulator.
[0071] An example of the optical transmitter element 84 is shown in FIG. 11, which performs the same phase modulation of incident light as the metasurface optical modulator element 40 shown in FIG. 4 and described above. However, since the optical transmitter element 84 in this example of FIG. 11 is an optical transmission element instead of the optical reflection element in FIG. 4, gaps are formed between the lattice structures so that the bottom metal layer 46 intersects like the top electrode 48. As a result, the incident light L is phase-modulated by the electro-optic polymer 44 and can be blocked by the electro-optic polymer 44 or pass through the lattice structure 42. The substrate 50 is transparent to the light L. The poling voltage 45 is driven by the modulator driver 88 of the transceiver module 82 according to the pixel value imposed by the light field L as described above. Further details of such a transmissive optical modulator can be found in the paper "Surface-normal electro-optic-polymer modulator with silicon subwavelength grating" by Kosugi et al., IEICE Electronics Express, Vol. 13, No. 17, pp. 1-9, September 10, 2016.
[0072] The back surface of the transceiver module 82 has an opaque cover or mask (not shown) covering the back surface to prevent laser irradiation on the back surface of the transceiver module 82 from passing through the transceiver module 82, except for an opening that allows light to reach and pass through the optical transmission element 84. Optical components including the Fourier transform lens 16 and the fiber faceplate that collimates the light in front of the Fourier transform lens 16 can be bonded to the front surface of the first sensor-display device 14. Similarly, the Fourier transform lens 28 and the fiber faceplate can be adhered to the front surface of the second sensor-display device 26.
[0073] Referring now to FIG. 1, in addition to the first and second sensor-display devices 14, 26 and the radial modulator device 20 of the photonic convolution assembly 12, an exemplary photonic neural network system 10 includes, for example, (i) a circuit block 110 that implements a pulse output for driving the radial modulator device 20, (ii) a high-speed analog-to-digital circuit block 112 to which digital data is loaded and received from the first and second sensor-display devices 14, 26, a high-bandwidth memory (HBM2) 114, and a field-programmable gate array (FPGA) 116, which is a basic control and interface device for other system components. The HBM2 114 provides storage for filters, state machine steps, and image data. The circuit block 110, the HBM2 114, and the FPGA 116 are on a multi-chip module (MCM) 118, and the user interface to the system 10 nominally goes through a PCI-Express bus 120.
[0074] A functional block diagram of an exemplary system interface 122 between a field programmable gate array (FPGA) 116 and a first sensor display device 14 is shown in FIG. 12, and the functional block diagram of FIG. 12 is also representative of the system interface between the FPGA 116 and a second sensor display device 26. For purposes of convenience and brevity in the drawings and the related descriptions for the circuit blocks 110 output circuitry 111 (FIG. 1) in, the sensor display devices 14, 26 (FIGS. 1, 2, 5-7, 10) For, any term "Sensay" (an abbreviation of sensor and display) or any term "RedFive" is sometimes used. For purposes of convenience and brevity, the transceiver module 82 (FIGS. 8 to 10) may be referred to as "Trixel". (Where "Trixel" is an abbreviation of "transmit-receive pixel".)
[0075] RedFives 111 Some of are responsible for generating analog data to load the sensor memory bank 90. These RedFive s 111 Since are the state machine sources managed by the FPGA 116 for the HBM2 114, they are interfaced via the memory module (HBM2) 114. The analog and digital input / output (I / O) are interfaced via the FPGA 116 because they are used for the control of the feedback loop. Some of the unused bits are wrapped back to the FPGA 116 as status flags for synchronization. The Sensay Digital 14, 26 I / O use the same memory lines as some of the RedFives 111 but since they are not accessed simultaneously, this dual use of the memory lines is not a conflict. Also, some of the output analog lines from the RedFives 111 are shared as input analog lines to the ADC 112 . The number of ADCs used to read the data and pass it to the FPGA 116 112 depends on the implementation.
[0076] sensor-display device ( Sensay) 14, 26 of External interface (see FIG. 8) The functional block diagram of is shown in FIG. 13. In FIG. 13, "Sx" is attached before "Sensay A" 14 or "Sensay B" 26 (When discussing the system to distinguish signals related to either Sensay 14 or 26, the digital input lines in FIG. 8 can be grouped into three general categories. Row and column control loads a set of latches within Sensay (see FIG. 14). The global control lines (or, rows) have various functions, each of which is explained in the context of use. The global rows can be routed along rows or columns. The global control lines are routed to all transceiver modules (trixels) 82 and are not specific to a particular column or row.)
[0077] SxPeakreset resets the analog peak hold circuit used for external gain control. This signal is asynchronous but should only be asserted when SxFreeze is asserted to avoid data contention (1).
[0078] of FIG. 13 SxSnsreset resets the sensor 86 (FIG. 10) to the level of the analog SxLevel line. Since the sensor 86 is designed to accumulate charge, this mechanism is required to drop to a preset level. This reset can be used as a global bias to preset the charge level of sensor 86 (and thus, the modulator level for the next pass).
[0079] SxPoolmode determines the average (1) or maximum (0) operation in pooling (1).
[0080] of FIG. 13 SxFreeze enables or disables global memory of the bank 90 (FIG. 10) access. When asserted (1), all trixels memory drives 92 are set to a safe state and the memoryof the bank 90 Access and shift are not permitted. SxFreeze is used when configuring other control lines to prevent data contamination before the line is calibrated. In the following description, the function of SxFreeze is not always mentioned, and its operation is not always a rule.
[0081] of FIG. 13 SxRWDir determines read / write for valid memory bank 90 (FIG. 10) When set to "1", data is written to the memory bank 90 When set to "0", data is read from the memory bank 90 This also gate-controls the operation of the sensor (photodetector) 86 and the modulator (optical transmitter) 84. This represents the modulator mode (0) or the sensor mode (1).
[0082] SxFlagRD, SxFlagWR, and SxFlagRST control the digital flag memory used for semantic labeling. SxFlagRST is the global address reset of all flag memories. SxFlagRD and SxFlagWR control memory access.
[0083] SxShift0,1,2 is externally driven in a three-phase sequence to move the charge of the shift register memory 90 either clockwise or counterclockwise only in the addressed trixels (transmitter-receiver modules) 82. If the trixel is not addressed, its memory drive line is forced to a safe state and does not affect the charging of the memory.
[0084] SxExtemal determines whether the SxAnalog and SxData lines are active (1) or if the data movement and access are only internal (0).
[0085] Consider the four combinations of these signals:
[0086] Image Load: SxFreeze = 0, SxRWDir = 1, SxExternal = 1. This means that the memory cells of the addressed trixels 82 take in data from the external SxAnalog lines and apply voltage to the trixels memory bank 90 via the internal SxData lines. Since there are 120 SxAnalog lines, this action can be up to 120-wide. In implementations where a 120-width DAC set is not appropriate, the lines can be externally connected in groups and the MEMR lines can be enabled in sequence to accommodate narrower access. Regardless of the implemented external line width, normally only one MEMC line is enabled at a time to avoid contention (although, if necessary, the same DAC value can be sent across an entire row at once).
[0087] Result Save: SxFreeze = 0, SxRWDir = 0, SxExternal = 1. This means that the memory cells of the addressed trixels 82 send data to the external SxAnalog lines for conversion with the external ADC. Again, this can be up to 128-wide, but narrower implementations can be adapted without design changes to Sensay. Regardless of the implemented external line width, only one MEMC is enabled at a time to avoid contention (this is not optional for reads to avoid data contention).
[0088] Sensor Mode: SxFreeze = 0, SxRWDir = 1, SxExternal = 0. This means that any memory cell in the addressed trixels 82 acquires data from the sensor 86 (via the polling chain described below and also in relation to SxShift 0,1,2, while shifting the existing voltage as a shift register set of memory values and saving the voltage as a new memory charge.
[0089] Modulator mode: SxFreeze = 0, SxRWDir = 0, SxExternal = 0. This means that any memory cell of the addressed trixels 82 sends data to the modulator (optical transmitter element) 84 (via the pooling chain), and in relation to SxShiftO, 1, 2, shifts the existing voltage as a shift register set of memory charges. Memory readout is non-destructive.
[0090] Exemplary row and column control line registers for the trixels (transceiver modules) 82 are schematically shown in FIG. 14. In this example, the row and column control line registers comprise 235 individually addressed static 64-bit latches arranged as 5 rows and 5 column lines per trixel. These outputs are always active and are set to zero at power-on. These row and column control lines are used by each trixel 82 to configure its function in relation to its adjacent devices. Each of the latches is individually addressed by asserting data with SxControl, setting an 8-bit address with SxAddr, and pulsing SxLatch.
[0091] The memory 90 of a trixel is said to be "addressed" when both its MEMR and MEMC are asserted. Similarly, when both its OPTC and OPTR are asserted, its optical sensor is said to be "addressed". The other functions of the trixels 82 are disabled when their ENBR and ENBC are de-asserted. To completely disable the trixels 82, their MEMR, MEMC, OPTR, OPTC, FLAGR, and FLAGC are also de-asserted.
[0092] The pool boundary lines 86 (POOLC and POOLR) affect the entire columns and rows of trixels 82 and define the boundaries of the super trixels, as will be explained in more detail below. Since the rightmost and bottommost lines are always enabled, only the 1079POOLR and 1919POOLC lines exist. The unused lines of the 64-bit latch are not connected.
[0093] The *_SL and *_SR lines shift their respective registers left or right on the rising edge.
[0094] SxReLUl and SxReLU2 (Figure 13) are driven by an external DAC. They are global for all trixels 82 and are applied to the sensor 86 readings to remove weak information. SxLevel (Figure 13) is also driven by an external DAC. It is used by all trixels - sensors 86 as a preset level and is also added to the modulator driver 88 level, where it is used as a phase offset. The sensay (sensor - display device) 14 or 26 is always in the sensor or modulator (transmit) mode as described above, so there is no contention. The SxPeak (Figure 13) analog output signal is the signal from all trixels (transceiver module) 82. As will be explained in more detail below, each trixels - memory cell passes its highest value to a common trace. The value of this trace represents the highest global value seen by the entire trixels array since SxPeakreset was last asserted. This is used by external circuitry for system gain and normalization.
[0095] An example of the analog interface is schematically shown in Fig. 15. The SxAnalog line is 120 traces that connect nine adjacent SxData lines respectively. In other words, internally, the SxData0000 to SxData0008 line traces are all connected to the output pin SxAnalog00. The SxData0009 to SxData0017 line traces are all connected to the output pins such as SxAnalog001. All SxAnalog pins are hardwired to nine internal SxData traces. Only one trixels memory bank 90 at a time can drive or sense its local trace (forced by an external controller). When TMS is asserted, all SxAnalog and SxData lines are connected together.
[0096] As described above, since the control lines are individually controllable, input or output schemes of any size from 1 to 120 widths can be implemented by simply connecting these lines together outside the Sensay (sensor - display device) and matching only the appropriate trixels 82 to the architecture. Note that the wider the interface, the faster the load and unload operations, but more external circuitry is required. This allows for advanced customization without changing the design.
[0097] In the exemplary photonic neural network system 10, the Sensay (sensor - display device) 14, 26 architecture is constructed around a pooling chain. As shown in FIGS. 9, 10, and 16, each of the transceiver modules (trixels) 82 in the array 80 has a pooling boundary line 96 along two of its edges, for example, along the right and bottom edges when FIGS. 9, 10, and 16 are oriented on the paper. All sensor, modulator, memory read or memory write accesses use the pooling chain to pass analog data within and between the trixels (sensor - display devices) 82. The function of the pooling boundary line 96 is to connect or disconnect adjacent transceiver modules (trixels) 82 from the pooling chain, creating super - trixels or "islands". The pooling - chain circuit connections to the boundary lines 96 of each adjacent trixel 82 are shown in the enlarged schematic of the connections in FIG. 17 at the virtual positions nnnn, mmmm within the array 80 of trixels 82. When POOLC = 0, all of the east - west trixel - pooling - chain connections for the entire column are opened. When POOLR = 0, all of the north - south trixel pooling - chain connections for the entire row are opened. All other trixel pooling - chain connections remain closed. The effect of this pooling structure is to create islands of connected pooling - chain lines. All trixels on a super - trixel island share this chain, which is essentially a single low - impedance "trace". When POOLR is asserted, the transistors connecting the pooling chain of these trixels conduct, connecting the pooling chain to the trixels 82 to its south in the next row. When POOLC is asserted, POOLC connects eastward to the pooling chain for the trixels 82.
[0098] As described above, the memory banks 90 in each of the transceiver modules (trixels) 82 are essentially shift registers, and the design and technology of shift registers are well understood and readily available to those skilled in the art. FIG. 18 shows an analog memory shift driver scheme. When addressed (both MEMC and MEMR are asserted) and unfrozen (SxFreeze is de-asserted), any combination of SxShiftO,1,2 simply propagates to the outputs (MemShiftO,1,2) that actually drive the analog memory cell shift plates. If either MEMC or MEMR is de-asserted for a Trixel, or if SxFreeze is asserted, the analog memory driver is automatically placed in a safe state (MemShiftO,1,2 = 010).
[0099] FIG. 19 is a schematic diagram of an exemplary analog memory read interface for the memory bank 90 (FIG. 10). The memory can be read, and analog data can be sent to the external SxAnalog interface via the internal SxData lines, or can be sent to the pooling chain 126 (set to zero if greater than SxReLU, otherwise) via either the maximum (diode) or average (resistor) circuit paths. The unchanged value read from the analog memory is also used to charge the diode isolation capacitor (sample & hold circuit), and ultimately drives the SxPeak value for the entire sensor display device (sensay) 14, 26 (used externally for system gain control). Examples of these modes are shown schematically in FIGS. 20 - 24. FIG. 20 shows the trixel analog memory read average to the pooling chain. FIG. 21 shows the trixels analog memory read maximum to the pooling chain. FIG. 22 shows the analog memory read to the external data lines. FIG. 23 shows the trixels analog memory peak value storage. FIG. 24 shows the analog memory peak value reset.
[0100] The Rectified Linear Unit (ReLU) is often applied to data to suppress weak responses. The first sensor-display device 14 and the second sensor-display device 26 (sensay, 14, 26) each have a flexible dual-slope ReLU implementation that can produce the various responses shown in Fig 25 and the responses range from ineffective (Example A) to conventional cutoff (Example B) to variable slope cutoff (Example C). Two external analog voltages driven by a DAC control the transfer function. Since sensors 14, 26 are of a unipolar design, the “zero” position is nominally at the center of the voltage range of memory bank 90.
[0101] Writing to analog memory 90 is simpler than reading. When the analog memory 90 of transceiver modules (trixels) 82 is addressed (both MEMC and MEMR are asserted, SxRWDir = 1), whatever the value on the local pooling chain is, it is placed on the write pad as shown in Fig 26. To actually store the value in the analog memory cell, the shift line is cycled. Loading analog memory 90 from the external data line is shown in Fig 27.
[0102] The flag memory is a 640-bit Last-In-First-Out (LIFO) device (i.e., a “stack”) in each transceiver module (trixels) 82 used for the implementation of semantic labeling. When SxFlagRST = l, for all transceiver modules (trixels) 82, the internal address pointer is unconditionally set to zero. There is no need to set the value to zero. Except for reset, the memory is active only when FLAGRmmmm = 1 and FLAGCNNnn = 1 for the virtual trixel position nnnnnnn,mmmm. If either FLAGRmmmm = 0 or FLAGCNNnn = 0, the signal does not affect the memory. See Fig 14 for FLAGR and FLAGC.
[0103] Schematic diagrams of flag memory writing and flag memory reading are shown in Figures 28 and 29 respectively. When SxFlagWR = 1, the output of the comparator is valid at the "D" memory input. At the falling edge when SxFlagWR changes from "1" to "0", if SxFlagRD = 0, the current flag bit determined by the state of the current read value of the trixel compared with the value on the pooling chain is pushed onto the stack. That is, when the analog memory read voltage matches the pooling chain voltage, this trixel 82 is the "master" and "1" is stored; otherwise, "0" is stored. Refer to Figure 19 for FlagVAL.
[0104] Since the hysteresis is very small, if multiple trixels 82 have very similar voltage levels, it is possible to view itself as the "master". In such a case, finally read out will be the average voltage of the enabled (or, enabled) trixels 82 within this pooling group during the extended path. Since the "competing" voltages were almost the same, this will have little practical impact.
[0105] At the rising edge when SxFlagRD = 1, while SxFlagWR = 0, the last bit written (i.e., the top of the stack) is read out and applied to the Trixel memory read circuit as the enabled FlagEN = 1 (refer to Figure 19). The output is enabled as long as SxFlagRD = 1.
[0106] When SxFlagWR = 0 and SxFlagRD = 0, FlagEN = 1. This is applied. SxFlagWR = 1 and SxFlagRD = 1 are invalid, and the external controller should not apply it. To avoid competition between the memory output and the comparator output, the flag EN is tri-stated in such cases.
[0107] Examples of optical control line settings for loading the sensor 86 of the transceiver module (trixel) 82 into the pooling chain, resetting the sensor 86, and writing the modulator (optical transmitter element) 84 from the pooling chain are shown in FIGS. 30, 31, and 32, respectively. The function of the optical control line is to connect their optical elements (modulator 84 or sensor 86) to the pooling chain for the trixels 82 at the intersection of the enabled OPTR and OPTC lines. When SxRWDir = 0 and SxExternal = 0, data is read from the pooling chain to drive this trixel modulator 84. When SxRWDir = 1 and SxExternal = 0, data is buffered from the sensor 86 of this trixel and placed on the pooling chain. When SxExternal = l, both the modulator 84 and the sensor 86 are disconnected. Multiple sensors 86 can be enabled (or, enabled) simultaneously, and the average of their values will appear on the pooling chain due to lower noise. Also, when the sensor 86 is adding an optical signal (data frame) as described above, note that there is no other activity on the sensors 14, 26 (without a clock, etc.), resulting in a very low-noise measurement result.
[0108] In the case of the modulator mode (SxRWDir = 0) and internal drive (SxExternal = 0), the outputs of all addressed trixel memory banks 90 are automatically pooled, and all optical transmitter elements (modulators) 84 within the same super trixel (connected to the same pooling chain) "shine" at the same brightness. This constitutes resampling by replication.
[0109] Individual optical transmitter elements (modulators) 84 can be disabled by the local ENB (ENBRmmmm = l and ENBCNNn = l).
[0110] The drive level DL of the optical transmission element (modulator) 84 is obtained by adding SxLevel to the product of the total of the pooling chain PC and the calibration sensor value CS + 1, and the formula is DL = (PC * (CS + 1)) + SxLevel. When Sxlnvert = 1, the drive is inverted, that is, the 100% level becomes 0% modulation, 90% becomes 10%, and so on.
[0111] The schematic diagrams of FIGS. 33A to 33B show an overview of the transceiver (trixel) circuit.
[0112] The above description is based on inference mode, for example, in the case where a trained neural network is used to recognize images, sounds, voices, etc., based on photonic neural network processing. Training a neural network using a photonic neural network, such as the photonic neural network system 10 described above, has several differences compared to a digital convolutional network system. As described above, during the training of a typical digital convolutional neural network system, adjustments are made using a process called backpropagation to increase the likelihood that the network will predict the same type of image the next time. In a typical digital convolutional neural network, such data processing and backpropagation are performed many times until the prediction becomes moderately accurate and no longer improves. The neural network is then utilized in inference mode to classify new input data and predict the results inferred from its training. In a digital convolutional neural network, since all backpropagation terms (or, periods / terms) and filters are in the spatial domain, training is relatively straightforward. Going back through the structure to take the "correct answer" and calculate the correction terms is slow but still does not require a change in domain. Training in a photonic neural network is not as direct because the terms that need to be trained are in the frequency domain and the convolution result is in the spatial domain. Although spatial domain data can be used and correction terms can be calculated using the fast Fourier transform (FFT) algorithm and applied to the Fourier filters used in the radial modulator device 20, such calculations are very computationally intensive.
[0113] Instead, the example of the photonic neural network system 10 described above is adapted such that the correction term for training is converted into a Fourier transform term that can then be added to the filter applied to the convolution in the iterative training process by the radial modulation device 20. An adaptation example for optically performing such a conversion instead of digital calculation includes adding a Fourier optical sensor device 130 specialized for the photonic convolution assembly 12, as shown in FIG. 34. The Fourier optical sensor device 130 is arranged axially aligned with the second sensor display device 26 from the second sensor display device 26 on the optical axis 62 on the opposite side of the polarizer 18. Also, the Fourier optical sensor device 130 is second equal to the focal length F2 of the 28 Fourier transform lens second Fourier transform lens 28 and is located in the Fourier transform plane at a distance from. Therefore, second the Fourier optical sensor device 130 is second located on the Fourier transform plane of the 28 Fourier transform lens. In its Fourier transform plane, second the Fourier optical sensor device 130 can detect the Fourier transform of the data or image frame of the light emitted from the second sensor display device 26. Therefore, the correction term required to train the photonic neural network system 10 (FIG. 1) can be supplied to the second sensor display device 26 in the spatial region frame of the correction data, and then the frame of the correction data in the optical field 132 is displayed (projected) on the Fourier optical sensor device 130. Therefore, the frame of the correction data in the optical field 132 is Fourier transformed when it reaches the Fourier optical sensor device 130 by the 28 Fourier transform lens, that is, the frame of the correction data in the spatial region is Fourier transformed into the Fourier region at the speed of light in the Fourier optical sensor device 130. Next, the frame of the correction data in the Fourier transform region is detected by the Fourier optical sensor device 130 and used to adjust the filter of the radial modulation device 20.
[0114] Typically, in the inference mode, the 3D convolution block is shifted out of memory and sent back through the photonic convolution assembly 12 for further levels of convolution and addition cycles to the memory bank 90 of the transceiver module (trixel) 82 (FIG. 10) The data frame present in a particular iterative convolution cycle within is lost, and the memory bank 90 is refilled with subsequent 3D convolution blocks, all of which are done very quickly as described above. However, in the training mode, these intermediate data frames are from the memory banks 90 of the first and second sensor-display devices 14, 26 of the transceiver module 82 Extracted, transferred to external memory, used to perform backpropagation digital calculations, and write correction terms in the spatial domain. These correction terms are then shown in FIG. 34 and, as described above, are supplied as a frame of corrected data in the spatial domain to the second sensor-display device 26 for projection onto and Fourier transform by the Fourier optical sensor device 130, so that the Fourier-transformed frame of corrected data can be detected by the Fourier optical sensor device 130 in the Fourier domain for use as a filter for the radial modulator device 20 for further convolution cycles. This training mode extraction of the intermediate correlation data, backpropagation digital calculations, and writing of the correction terms takes some time and thus slows down the iterative convolution addition cycles compared to the inference operation mode, but is still much faster than digital convolutional neural network training.
[0115] To incorporate the Fourier light sensor device 130 into the photonic convolution assembly 12, for example, as shown in FIG. 34, a half-wave variable polarizer 134 is placed between the second sensor display device 26 and the polarizer 18 to rotate the plane of polarization 90 degrees when frames of correction data are being projected by the second sensor display device 26 onto the Fourier light sensor 130. For example, in a normal inference mode of operation, the second sensor display device 26 displays P-polarized light reflected from the polarizer 18 onto the radial modulator device 20, and then, to display or project frames of correction data onto the Fourier light sensor 130 for training, the half-wave variable polarizer 134 is activated to rotate the plane of polarization of the projected light field 90 degrees to S-polarized, such that the resulting light field 132 passes through the polarizer 18 to the Fourier light sensor 130.
[0116] The frames of corrected data are then filtered to train the neural network. radial modulator 20 (see FIG. 3) has a value that needs to be provided to a particular wedge segment 32. Thus, of FIG. 34 The frames of correction data provided to the second sensor and display unit 26 for projection onto the Fourier light sensor unit 130 are provided in a format corresponding to the wedge segments 32 of the radial modulator device 20 (see FIG. 3) that need to be modulated in a corrected manner to train the neural network, so that these correction data are ultimately placed into a filter that drives the appropriate wedge segments 32 in a corrected manner. Thus, the Fourier light sensor unit 130 detects light 132 from the second sensor and display unit 26 in the same pattern as the wedge segments 32 of the radial modulator unit 20, so that the correction data for the light 132 is detected, processed and provided to the appropriate wedge segments 32 of the radial modulator unit 20.
[0117] The light projected from the second sensor display device 26 is of the radial modulator 20 (FIG. 3) To facilitate detection according to the same pattern as the wedge segment 32, as described above, for example, the Fourier optical sensor device 130 as an example has a radial modulation device 20 as shown, for example, in FIG. 35 (FIG. 3) and has an optical sensor board 135 having a plurality of optical sensor elements 136 arranged in an optical sensor array 138 corresponding to the patterns of the wedge segment 32 and the wedge sector 34 of (FIG. 3) . As shown in FIG. 35, The radial array lens plate 140 is arranged in front of the optical sensor array 138 and has a plurality of individual lens elements 142 arranged in a radial pattern of wedges and sectors corresponding to the wedge segment 32 and the sector 34 of the radial modulation device 20. (FIG. 3) These lens elements 142 capture the incident light 132 from the second sensor display device 26 in a radial pattern corresponding to the radial patterns of the wedge segment 32 and the wedge sector 34 of the radial modulation device 20, thereby capturing a frame of correction data of the incident light 132 when formulated and programmed in the second sensor display device 26. The segments of light captured by each lens element 142 are focused onto the respective optical sensor elements 136 as individual sub-beams 138 by the lens elements 142 and converted into electrical signals corresponding to the intensity of the light incident on the sensor elements 136, thus converting the frame of correction data in the incident light 132 into an electrical data signal corresponding to the frame of correction data. As shown in FIG. 34, These analog electrical data signals can be converted into digital signals for processing by a correction filter by the FPGA 116 and can then be supplied by the circuit block 110 for connection to the radial modulator device 20 via the interface 24. Also in this case, the frame of data in the incident light 132 is a Fourier transform lens 28is Fourier-transformed and detected by the sensor element 136 in the Fourier-transform domain, and thus the correction data in the signal sent from the Fourier optical sensor device 130 to the FPGA 116 or other electrical processing components is in the Fourier domain as required to drive the wedge segment 32 of the radial modulator device 20. Due to the arrangement of the optical components, the frame of the correction data supplied to the second sensor-display device 26 may have to be inverted so that the segments of light captured by the sensor element 136 and the corresponding signals generated match the appropriate wedge segment 32 of the radial modulator device 20. However, as described above, since the correction terms are calculated in the spatial domain, no algorithmic constraints are imposed on the training. Once the normal training backpropagation calculation is performed, the optical systems described above and shown in FIGS. 34 and 35 convert the spatial domain correction terms into their radial Fourier domain equivalents.
[0118] In another embodiment illustrated in FIG. 36, the camera lens 150 is, as described above, for example, a photonic neural network 10 (FIG. 1) for processing, the camera lens 150 is mounted on the photonic convolution assembly 12 in such a way that it irradiates the real-world scene 152 as a frame of data (image) into the photonic convolution assembly 12 in the spatial domain. For example, as shown in FIG. 35, the camera lens 150 is mounted on the optical axis 62 so as to be axially aligned with the second sensor-display device 26 on the opposite side of the polarizer 18 from the second sensor-display device 26. The polarizing plate 154 is disposed between the camera lens 150 and the polarizer 18 to polarize the light field 156 transmitted by the camera lens 150 in the polarization plane that reflects from the polarizer 18. Thus, the light field 156 is reflected by the polarizer 18 to the first sensor-display device 14 as shown in FIG. 36. The optical sensor element 86 in the transceiver module 82 (see FIGS. 9 and 10) of the first sensor-display device 14 detects and captures the frame of data (image) in the light field 156, and as described above, the memory bank 90 in the first sensor-display device 14(FIG. 10) It processes the frame of data (image) therein. Then, the shutter device 158 on the camera lens 150 closes on the camera lens 150, terminating the light transmission through the camera lens 150. Next, the first sensor-display device 14 can start processing the frame of data (image) passing through the photonic convolution assembly 12 in either the above-described inference operation or training operation.
[0119] In order to enable only a specific spectral frequency of light to be transmitted to the photonic convolution assembly 12 as needed, a band-pass filter 160 can also be provided together with the camera lens 150. The band-pass filter 160 can be a variable band-pass filter as needed. As a result, various spectral frequency bands of light from the real-world scene 152 can be sequentially transmitted from the camera lens 150 to the photonic convolution assembly, while the frames of data (image) within each frequency band are sequentially captured, whereby a hyperspectral image set for convolution is sequentially provided through the photonic convolution assembly. Such a variable band-pass filter is well known. For example, a variable half-wave retarder can be used in combination with a fixed polarizing plate as a variable band-pass filter. Such a variable half-wave retarder combined with a fixed polarizing plate can also be used as a shutter.
[0120] The foregoing description is considered to explain the principles of the present invention. Further, since numerous modifications and changes will readily occur to those skilled in the art, it is not desirable to limit the present invention to the exact structures and processes shown and described above. Therefore, reliance can be placed on all suitable modifications and equivalents within the scope of the present invention. When the terms "comprising," "including," "containing," and "having" are used in this specification, they are intended to specify the presence of the stated features, integers, components, or steps, but do not preclude the presence or addition of one or more other features, integers, components, steps, or groups thereof. The following is the invention described at the initial filing of the present application. <Claim 1> A first sensor-display device comprising an array of transceiver modules, each transceiver module comprising an optical sensor element, an optical transmitter element, and a memory bank having a plurality of memory cells. A second sensor-display device comprising an array of transceiver modules, each transceiver module comprising an optical sensor element, an optical transmitter element, and a memory bank having a plurality of memory cells. A radial modulator device having a plurality of modulation elements arranged in a plurality of radial distances and angular directions with respect to the optical axis. A first Fourier transform lens disposed between the optical transmitter element of the first sensor-display device and the optical transmitter element of the radial modulator device. A second Fourier transform lens disposed between the optical transmitter element of the first sensor-display device and the radial modulator device. And having A system for convolving and adding data frames, wherein the radial modulator device is disposed at the focal distances from the first Fourier transform lens and the second Fourier transform lens such that the radial modulator device is disposed on the Fourier transform surfaces of both the first Fourier transform lens and the second Fourier transform lens. <Claim 2> A system control component for forming and supplying a filter to the radial modulator device, controlling the order of transmission of an optical field having frames of data from the first and second sensor-display devices, convolving the frames of data with the filter of the radial modulator device, and detecting an optical field having the convolved frames of data from the radial modulator device. The system according to claim 1. <Claim 3> The system according to claim 1, wherein the optical sensor element is a capacitive optical sensor that accumulates charge from the sensed light. <Claim 4> A method of convolving and adding frames of data for a convolutional neural network, sequentially projecting the frames of data as optical fields in a spatial region along a first optical axis; sequentially Fourier-transforming the optical fields in a Fourier transform plane; sequentially convolving the optical fields in the Fourier transform plane with an optical modulator having optical modulation segments spaced apart at various radial distances and angular directions with respect to the optical axis; inverse Fourier-transforming the sequence of convolved optical fields to a spatial region at a first sensor-display position; at the first sensor-display position, detecting each of the convolved optical fields pixel-by-pixel with a capacitive optical sensor at a pixel position capable of charge storage in a spatial domain; accumulating charges in the capacitive optical sensor resulting from sequentially detecting the convolved optical fields at the first sensor-display position The method comprising. <Claim 5> After sensing a plurality of convolved optical fields such that the memory cell has accumulated charges resulting from the light detected at a specific pixel position for the sequentially sensed optical fields, shifting the charges accumulated in each sensor to the memory cell in the memory bank, the method according to claim 4. <Claim 6> applying different filters to convolve an additional sequence of optical fields having frames of data with the optical modulator; detecting an additional sequence of convolved optical fields pixel-by-pixel with the capacitive sensor and accumulating the charges resulting from the detection at each pixel position; after detecting a plurality of convolved optical fields, shifting the charges accumulated in each sensor to a memory cell having pre-accumulated charges while shifting the pre-accumulated charges to another memory cell in the memory bank; Repeating the above process to construct a 3D convolution block of a frame of data that is convolved and added within a memory bank at each pixel position at the first sensor-display position The method according to claim 5, having <Claim 7> Performing a Fourier transform on a frame of convolved and added data that forms a 3D convolution block in a sequential optical field from a pixel position at the first sensor-display position and sending it back to a modulator in the Fourier transform plane; Sequentially convolving the optical field in the Fourier transform plane by an optical modulator having optical modulation segments arranged at various radial distances and angular directions with respect to the optical axis; Performing an inverse Fourier transform on the sequence of convolved optical fields in a spatial region at a second sensor-display position; Detecting each of the convolved optical fields in the spatial region for each pixel with a capacitive optical sensor at a pixel position where charge can be accumulated at the second sensor-display position; Accumulating the charge generated from sequentially detecting the convolved optical field at the second sensor-display position in the capacitive optical sensor; Applying different filters and convolving an additional sequence of optical fields having a frame of data with the optical modulator; Detecting an additional sequence of convolved optical fields for each pixel at the second sensor-display position with a capacitive sensor and accumulating the charge generated from the detection at each pixel position; After detecting a plurality of convolved optical fields, shifting the charge accumulated in each sensor at a second sensor-receiver position to a memory cell having a previously accumulated charge while shifting the previously accumulated charge to another memory cell in a memory bank; Repeating the above process to construct a 3D convolution block of the frame of data folded and added in the memory bank at each pixel position at the second sensor - display position The method according to claim 6, having <Claim 8> The method according to claim 7, having the step of repeating the process in an additional cycle. <Claim 9> The method according to claim 8, having the step of pooling a plurality of the sensors and memory banks together in the repeating cycle of the process. <Claim 10> The method according to claim 8, including maximum pooling of the plurality of sensors and memory banks. <Claim 11> The method according to claim 7, having the step of transmitting, pixel - by - pixel with the optical transmitter element at the pixel position at the first sensor - display position, the frame of data folded and added.
Claims
1. An array of modules arranged in a plurality of rows and a plurality of columns, each module in the array having a read / write memory, the array, A pooling chain having a first pooling boundary line extending to a first side of each module in each column of the array and a second pooling boundary line extending to a second side of each module of the array, wherein the first pooling boundary line extending to the first side of each module intersects and connects with the second pooling boundary line extending to the second side of that module to form a boundary pooling circuit for that module, and the boundary pooling circuit of each module is connected to an adjacent boundary pooling circuit of an adjacent module in the same column by a first switchable switch and to an adjacent pooling circuit of an adjacent module in the same row by a second switchable switch, the pooling chain; A memory pooling interface within each module, connecting the read / write memory within that module to the boundary pooling circuit of that module, enabling data to be read from the read / write memory of that module to the boundary pooling circuit of that module and enabling data to be written from the boundary pooling circuit to the read / write memory, the memory pooling interface; A data processing array having.
2. The data processing array according to claim 1, further comprising a pooling row switching control line connected to each first switch in the column of the modules for simultaneously opening and closing all the first switches in the column, and a pooling column switching control line connected to each second switch in the row of the modules for simultaneously opening and closing all the second switches in the row.
3. The data processing array according to claim 2, wherein each module has an optical modulator, a modulation driver connected to the optical modulator, and a modulation pooling interface connecting the modulation driver to the boundary pooling circuit of that module to enable data from the boundary pooling circuit to be written to drive the optical modulator.
4. The data processing array according to claim 3, wherein each module has an optical sensor and a sensor interface for arranging data from the optical sensor in the boundary pooling circuit.
5. The data processing array according to claim 1, wherein each module has a data interface for connecting an analog data line to the read / write memory of the module to read data from the read / write memory of the module to an analog-to-digital converter and write data from a digital-to-analog converter to the read / write memory of the module.
Citation Information
Patent Citations
Multiplexed optical system and feature vector transformer using the system, feature vector detection and transmission device, and recognition classifier using the above
JP1997258287A
Semiconductor Memory Device and System Conducting Parity Check and Operating Method of Semiconductor Memory Device
KR102142589B1
Semiconductor memory device and system conducting parity check and operating method of semiconductor memory device
US20140250353A1