Optical processing system

By using a spatial light modulator and a 4f optical correlator to perform optical convolution through an optical processing system, the problem of time-consuming training and inference of convolutional neural networks is solved, and efficient acceleration of convolution operations is achieved.

CN112400175BActive Publication Date: 2025-10-28OPTRIS GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980043334.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-27
Filing Date
2019-04-26
Publication Date
2025-10-28
Estimated Expiration
2039-04-26

AI Technical Summary

Technical Problem

The training and inference processes of existing convolutional neural networks are time-consuming, especially when using hardware such as GPUs. Furthermore, the computational load of convolutional layers is significant, and increasing the resolution will increase the computational burden. Existing algorithms have limited speed-up effects.

Method used

An optical processing system is employed, utilizing a spatial light modulator and a 4f optical correlator for optical convolution. Through block input and parallel processing of kernel modes, combined with digital computation to generate a data focusing pattern, the convolution operation is realized.

Benefits of technology

It significantly accelerates the inference process of convolutional neural networks, reduces computation time, improves computational efficiency, and is suitable for accelerating the inference stage of trained networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112400175B_ABST
    Figure CN112400175B_ABST
Patent Text Reader

Abstract

An optical processing system includes at least one spatial light modulator (SLM) configured to simultaneously display a first input data pattern (a) and at least one data focus pattern, the at least one data focus pattern being a Fourier domain representation (B) of a second input data pattern (b). The optical processing system further includes a detector for detecting light that has been optically processed sequentially by the input data pattern and the focus data pattern, thereby generating an optical convolution of the first and second input data patterns, the optical convolution being used in a neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to convolutional neural networks for machine learning. In particular, this invention relates to accelerating convolutional neural networks using optical correlation-based processing systems. Background Technology

[0002] Convolutional Neural Networks (CNNs or ConvNets) are well-known and have become a leading machine learning technique in image analysis. They are deep feedforward artificial neural networks that achieve state-of-the-art performance in image recognition and classification. The lifecycle of a ConvNet is typically divided into training and inference phases. However, training large convolutional networks is very time-consuming—even with state-of-the-art graphics processing units (GPUs), it can take weeks. During the training and inference phases, some more complex ConvNets may require even longer to run.

[0003] In ConvNet, convolutional layers represent a very significant portion—often the majority—of the computational load. Furthermore, increasing the resolution of convolutions (increasing the input size or kernel size) incurs even greater computational burden. This drives network configurations to avoid building a large number of high-resolution convolutional layers, or prompts network configurations to modify convolutional layers to reduce computational load.

[0004] Various algorithms have been explored, implemented on GPU or FPGA architectures, to digitally accelerate ConvNet training and inference. However, further acceleration is urgently needed.

[0005] The embodiments of the present invention attempt to provide a solution to this problem. Summary of the Invention

[0006] In a first independent aspect, an optical processing system is provided, comprising at least one spatial light modulator (SLM) configured to simultaneously display at least one data focusing mode as a first input data mode (a) and a Fourier domain representation (B) of a second input data mode (b). The optical processing system further comprises a detector for detecting light that has been optically processed sequentially by the input data mode and the focusing data mode, thereby generating an optical convolution of the first and second input data modes, the optical convolution being used in a neural network.

[0007] The optical processing system includes a 4f optical correlator, wherein the input data mode (A) is in the input plane of the correlator, and the data focusing mode is in the Fourier plane. The data focusing mode may include a convolution kernel or a filter. The SLM has a dynamic modulation effect on the incident light. As described in PCT / GB2013 / 051778, the SLM may be, for example, parallel layers, or they may be in the same plane. In such a 4f optical correlator, light from the optical input is incident on the displayed mode and is optically processed continuously, so that each input data mode and focusing mode forms a continuous optical path before the light is captured by the detector.

[0008] The data focusing pattern is selected as the Fourier domain representation (B) of the second input data pattern (b). That is, filters are computed from the second input data pattern to produce the convolutions required by the ConvNet. For example, the data focusing pattern can be computed digitally.

[0009] The neural network can be a convolutional neural network (ConvNet), where optical convolution is suitable for the convolutional layers of the neural network. Typically, the second input data pattern is called the "kernel," and the first input data pattern is convolved with this kernel.

[0010] Importantly, optical correlators are a better platform for evaluating (2D) convolutions compared to known digital implementations. Optical methods offer improved performance compared to such methods.

[0011] In some embodiments, the first input data pattern (a) includes a plurality of (N) block input data patterns—or “feature maps”—each of which corresponds to a member in a “batch” of images being processed, and wherein multiple convolutions are generated in parallel for each of the plurality of block input data patterns. In some embodiments, the second input data pattern includes a plurality of (M) block kernel data patterns, each of which corresponds to a different member in a set of filter kernels (b), and wherein multiple convolutions are generated in parallel for each pair (N×M) formed by the block input data patterns and the block kernel data patterns.

[0012] Therefore, the input 2D image and / or 2D kernel can be segmented. By segmenting, we mean that the patches do not overlap with each other and are within the detector's region. Thus, the resolution of each "patch" is less than the detector's resolution. This leads to different ways of performing multiple lower-resolution convolutions using a given hardware resolution. Therefore, advantageously, by appropriately selecting the segmented input and kernel, the full resolution of the detector (e.g., a camera) can be utilized. To fully utilize the convolution resolution, segmenting both the input and kernel produces results across the detector (camera) plane. Advantageously, for each pair of input and kernel data (each pair can be a "patch"), multiple convolutions can be performed in parallel. Parallel convolutions can be "batched." A batch represents multiple smaller Fourier transforms.

[0013] Note that in the case of kernel partitioning, these blocks are appropriately padded within the direct representation (b). They are then converted to a Fourier domain filter representation (B). Due to the characteristics of filter generation, this conversion is not perfect and may lead to crosstalk between different operations.

[0014] Preferably, the optical processing system further includes a processor for digitally obtaining the data focusing pattern (B) from the second input data pattern (b). Therefore, the filter of the 4f optical correlator is digitally calculated from the kernel. This is a relatively fast process due to the low filter resolution (compared to the optical detector resolution). Furthermore, these patterns can be reused once the network has been trained; they do not need to be recalculated.

[0015] The processor can be configured to obtain the data focusing pattern using the minimum Euclidean distance (MED) method. This method assumes that the appropriate available SLM modulation value used to represent a given complex value is the value that is closest to the distance measured on the complex plane.

[0016] Advantageously, a 4f optical correlator can be used to accelerate the inference process of ConvNet, which has already been trained on a conventional digital platform with knowledge of the characteristics of the optical system.

[0017] Alternatively, a 4f optical correlator can be incorporated into the network and used directly during training. The optical correlator is used during forward propagation through the network, while regular digital convolutions are used during backpropagation.

[0018] It should be understood that the constraints imposed on the filtered image by the properties of the spatial light modulator limit the data focusing patterns (B) that can be displayed, and thus limit the corresponding convolutional kernels (b) that can be implemented. However, in some embodiments, it is important to limit the maximum spatial frequency components of the data focusing pattern (B), as this determines the effective size of the corresponding kernel (b). In retrospect, the data focusing pattern (B) and the kernel (b) are linked via a Discrete Fourier Transform (DFT). This fact must be considered during the design of the data focusing pattern. Therefore, the system controls the maximum spatial frequency content of the filter to limit the effective kernel width; that is, the size of the kernel and the distance in the input beyond these points do not affect the corresponding points in the output. This prevents undesirable connectivity between different convolutional operations. Controlling this connectivity is crucial for successful operation.

[0019] In some embodiments, the second input data mode (b) is a signed kernel with positive and negative components.

[0020] In some embodiments, the kernel is decomposed into 1) a positive kernel p and 2) a uniform bias kernel b1, such that f*k = f*(p-bl) can be reconstructed from the two positive convolutions. This overcomes the fact that the detector (camera) cannot measure negative amplitudes.

[0021] In some embodiments, the SLM is a binary SLM instead of a multi-level SLM. Binary SLMs offer high bandwidth and high performance, but require applications to adapt to device limitations.

[0022] In some embodiments, the system further includes lenses or lens selection for adjusting amplification and for optical pooling neural networks.

[0023] In another aspect, a method for generating optical convolutions using the optical processing system described above is provided. Such a method can be used for layers of convolutional or convolutional-pooling neural networks. For example, this method can be used to train or infer convolutional neural networks.

[0024] In another aspect, the use of the system described above for image analysis in deep machine learning is provided.

[0025] In a second independent aspect, a method for configuring a neural network using a 4f optical correlator is provided, the method comprising the following steps:

[0026] Display the first input data mode (a);

[0027] Determine at least one data focusing pattern (B) of the Fourier domain representation (B) as the second input data pattern (b);

[0028] Display at least one data focus mode (B);

[0029] Detecting light that has been sequentially optically processed by the input data pattern and the focus data pattern, thereby generating optical convolutions of the first and second input data patterns; and

[0030] Optical convolution is used as part of a neural network.

[0031] In some embodiments, the method further includes the steps of: dividing at least one of the first input data pattern and the second input data pattern into blocks, and generating multiple optical convolutions in parallel.

[0032] In some embodiments, different system components operate at different speeds. A camera operating at a slower speed than the input SLM is exposed across multiple frames to produce a result that is the sum of multiple optical convolutions.

[0033] It will be understood that all alternative embodiments described in the reference system aspect are applicable to these methods. Attached Figure Description

[0034] Figure 1 is an optical path diagram derived from the applicant's own existing technology;

[0035] Figure 2 is another optical path diagram showing a known "4f" optical system;

[0036] Figure 3 (a) through 3(c) show examples of the optical correlation between the two functions;

[0037] Figure 4 It is the phase vector diagram (Argand diagram) of the complex plane;

[0038] Figure 5 This is a schematic representation of 2D convolution (ConvNet) based on existing technology; further examples can be found at http: / / deeplearning.net / tutorial / lenet.html.

[0039] Figure 6 This is an example of Fourier domain convolution;

[0040] Figures 7(a)-7(d) This demonstrates different ways to perform multiple low-resolution convolutions using fixed-input hardware.

[0041] Figure 8 It is a graph showing the "deceleration factor" S relative to the roofline performance when the 10MPx input is divided into P batches, compared with the digital Fourier transform.

[0042] Figure 9 The illustration shows an example of performing convolution using an optical system; and

[0043] Figure 10Two examples of signed convolutions and corresponding optical convolutions obtained by digital computation according to embodiments of the present invention are shown. Detailed Implementation

[0044] The inventors have recognized that optical correlators can be used in a specific manner to evaluate convolutions at the core of convolutional neural networks (CNNs) or ConvNets. In other words, optical processing is being applied to deep learning processes, particularly for evaluating convolutions required for ConvNets. However, this application is not insignificant due to the specific characteristics that need to be considered in optical convolutions. Specifically, as will be described in detail below, the problem relates to the limitations of available electro-optical and optoelectronic hardware.

[0045] Optical Fourier Transform and Optical Correlator

[0046] Coherent optical processing systems are known. In coherent processing systems such as optical correlators, one or more spatial light modulator (SLM) devices are typically used to modulate a laser or other coherent source in phase or amplitude, or a combination of both. An SLM is a device that dynamically modulates incident light. These typically include liquid crystal devices, but can also be micromirror microelectromechanical (MEMS) devices. Optical correlator devices are commonly used in optical pattern recognition systems. In 4f matched filter or joint transform correlator (JTC) systems, SLM devices utilize functions for addressing, typically representing the input or reference pattern (which can be an image) and / or filter pattern based on the Fourier transform representation of a reference function / pattern to be "matched" to the input function.

[0047] Cameras, such as complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) sensors, are typically located in the output plane of an optical system to capture the resulting optical intensity distribution, which, in the case of an optical correlator system, may contain locally correlated intensities indicating the similarity and relative alignment of the input and reference functions.

[0048] Coherent optical information processing utilizes the fact that simple lenses exhibit Fourier transforms, and the most common function used in the types of coherent optical systems involved in the prior art and this invention is the Optical Fourier Transform (OFT)—a time decomposition, or in this case, a spatial distribution of its frequency components. This is analogous to the pure form of the two-dimensional Fourier transform, as indicated by the following equation:

[0049]

[0050] Where (x, y) represents spatial / temporal variables, and (u, v) is a frequency variable.

[0051] OFT can be implemented using the optical system shown in Figure 1, where collimated coherent light (typically a laser) 1 with wavelength λ = 1 is phase- or amplitude-modulated by a spatial light modulator 2 (SLM, typically a liquid crystal or electromechanical MEM array). The modulated beam is then passed through a positive converging lens 3 with focal length f and focused in the lens's back focal plane, where a detector, such as a CMOS array 4, is positioned to capture the intensity of the resulting Fourier transform.

[0052] The front focal plane (which, for clarity, may be called the downstream focal plane) contains precisely the Fourier transform of the complex domain existing at the rear focal plane (which, for clarity, may be called the upstream focal plane)—both amplitude and phase. This is consistent with the fact that for a perfectly flat beam over an infinite spatial range, a single spot of light is obtained (the “DC” term in Fourier theory).

[0053] In optical processing systems, OFT can be used as a direct replacement for electronic / software-based Fast Fourier Transform (FFT) algorithms, offering significant advantages in processing time and resolution. This process can also serve as the basis for various functions.

[0054] The correlation between two or more functions can be achieved in an optical system in two main ways. The first way is through matched filtering, denoted as follows:

[0055] r(x,y)*g(x,y)=FT[R(u,v) * ×G(u, v)]

[0056] The uppercase function represents the Fourier transform of its lowercase equivalent; "*" indicates the complex conjugate of adjacent functions, and "*" indicates the related function.

[0057] The second way to achieve association is to use the joint transformation association process, such as the 1 / f JTC described in EP1546838 (W02004 / 029746).

[0058] In each case, the correlation is formed as the inverse Fourier transform of the product of two functions, which have already been Fourier transformed themselves. In the case of matched filtering, one function is optically Fourier transformed via a lens; and the other function is electronically Fourier transformed during the filter design process. In the case of JTC, both are optically Fourier transformed.

[0059] Figure 2 illustrates a "4f" optical system that can be used to implement matched filters or other convolution-based systems. Figure 2 shows a collimated coherent light source 5 with wavelength λ, modulated by an SLM pixel array 6 ("input data" a), and then transmitted and focused through a lens 7 ("lens 1") onto a second SLM pixel array 8, thereby forming a function displayed on the first SLM at the pixels of the second SLM 8. (Filter B). The resulting optical pointwise matrix multiplication is then subjected to an inverse Fourier transform via lens 9 (“lens 2”), and the result is captured at detector array 10 (which may or may include camera c).

[0060] For the matched filtering process, the pattern displayed by the pixels of the first SLM 6 will be “input scene, a”g(x,y), and the pattern displayed on the second SLM 8 will represent the version of the Fourier transform of the reference function r(x,y).

[0061] As shown in Figure 2, SLMs 6 and 8 are used to input data into the optical correlator system at planes a and B. For example, if a beam with complex amplitude A(x, y) passes through an SLM with spatially variable transmission function t(x, y), the resulting beam has amplitude A(x, y)·t(x, y).

[0062] Input data 'a' is placed in front of the system. Filter 8(B) effectively multiplies the light field by a 2D function B. Camera sensor 10 pairs the output beam 'c'. 2 Imaging. The system performs mathematical operations. Where c is a complex amplitude function. (In fact, the second shot performs a forward Fourier transform, not an inverse Fourier transform, but the net effect is a coordinate inversion compensated for by the camera sensor orientation.) The camera sensor measures the intensity of the field, I = |C|. 2 .

[0063] Convolution theory uses Fourier transform to influence the convolution of two functions f and g through simple multiplication (*):

[0064]

[0065] The inventors have recognized that the role of the optical system is to evaluate convolution.

[0066] One of the correlation inputs, 'a', is directly input to the optical system. The second correlation input, 'b', is processed digitally. This is done offline, using a digital discrete Fourier transform to generate 'B', followed by post-processing to map it to an available filter stage (proper use of an optical correlator introduces a small overhead when generating filter 'B' from the target 'b').

[0067] Therefore, the inventors recognized that correlation is closely related to convolution processing, corresponding to the inverted coordinates in the convolved function. Note that in a 2D image, the inversion of each coordinate is a rotation of the image. For clarity, these functions are defined in 1D, but they naturally extend to 2D.

[0068] The convolution (*) of two functions f(x) and g(x) is expressed in both discrete and continuous forms, and is defined as follows:

[0069] f*g(x)=∑ i f(i)g(xi),

[0070] f*g(x)=∫f(χ)g(x-χ)dχ.

[0071] The convolution (o) of two identical functions f(x) and g(x) is defined as:

[0072]

[0073]

[0074] As can be seen from these definitions, the operations are interchangeable under coordinate inversion of one of the functions. This inversion is performed in the optical correlator by rotating the filter. Therefore, for symmetric functions, correlation and convolution are equivalent. Importantly, the inventors have recognized that an optical correlator can also be considered as an optical convolutioner.

[0075] Among other things, correlation is very useful for pattern matching applications. The process can be viewed as dragging one function onto another and taking the dot product (projection) between them at the set of all displacements. The two functions will have a large dot product and optically produce "correlation points" corresponding to the locations where their displacement versions are matched.

[0076] Figure 3 Images (a) through 3(c) illustrate examples of the correlation between the two functions. The test image (a) is correlated with the target (b). The peak in the output (c) indicates the location where a match exists—this is in Figure 3 (c) circled.

[0077] Information is encoded into the light beam using a liquid crystal display (SLM). An SLM is essentially a very small display (in fact, some devices used in optical processor systems originated from display projectors). These devices use liquid crystal technology—combined with optical polarizers—to modulate the light beam. Typically, the amplitude and relative phase of the beam can be modulated. Different devices offer different capabilities. Generally, devices can be divided into two categories: multi-level SLMs and binary SLMs.

[0078] For multi-level SLMs, each pixel on the device can be set to one of several different levels (typically 8 bits or 256 levels). Depending on the optical polarization, the device can modulate the amplitude or phase of the light field in different ways. SLMs cannot independently modulate the amplitude and phase of the light field. Typically, some coupling modulation exists. This modulation can be expressed as an "operation curve"; that is, a path on the complex plane describing the accessible modulation. Figure 4 This is an Argento plot of the complex plane, showing an example of the curves representing the 256 modulation levels of an SLM. Multi-level SLM devices typically operate around 100Hz.

[0079] Binary SLMs typically operate much faster than multi-level SLMs, at around 10 kHz. However, binary SLMs only offer two different modulation levels, which can be amplitude modulation, phase modulation, or a combination of both. Despite offering only two levels, the significantly higher speed means that binary SLMs are the highest bandwidth devices available.

[0080] Note that if a binary SLM is used as a filter SLM, it is not limited to representing a binary convolution kernel, because the Fourier transform of a binary function is not necessarily binary. However, a binary filter does constrain the ability to control the spatial frequency content of the filter.

[0081] Convolutional Neural Network (ConvNet)

[0082] Figure 5 An example schematic of a 2D ConvNet is outlined. Specifically, the input image 11 is divided into patches, and 2D convolutions are performed on these patches. This is followed by processing at different layers. These layers include "pooling" layers to reduce the sample resolution; additional convolutional layers; the application of non-linear functions; and classic "fully connected" neural network layers.

[0083] ConvNet typically consists of several different layers linked together, including:

[0084] • Convolutional layers, where (multiple) inputs are convolved with a given kernel.

[0085] • Pooling layer, in which spatial downsampling is performed on the layer.

[0086] • Activation layer, in which a (non-linear) activation function is applied.

[0087] • Fully connected layers, which are classic fully connected neural networks, have fully specified weights between each node. These layers are used to generate the final classification node values.

[0088] While it will be understood that ConvNets have many nuances, this is only a high-level overview of the canonical architecture. These layers can be combined in a wide variety of ways. Properly configuring neural networks—with good performance and without excessive computational requirements—is one of the key challenges in this field.

[0089] At each layer (except for fully connected layers), the forward-propagated network state is a 3D object (or 4D when batch dimension is included), which includes a set of x, y "feature maps" due to the application of different convolutional kernels.

[0090] Within each layer, there are additional configuration options. For example, convolutional layers can combine different feature maps from the previous layer in many ways, and various non-linear activation functions exist.

[0091] The lifecycle of ConvNet can be divided into training and inference.

[0092] Once the network configuration is defined, it must be trained. The network contains a large number of parameters that must be determined empirically, including the weights of convolutional kernels and fully connected layers.

[0093] Initially, these parameters are set randomly. Training them requires a large pre-classified dataset. This training dataset is fed through the network, and the error between the correct classification and the reported classification is then fed back through the network (backpropagation). The error associated with each point in the network is used with gradient descent to optimize the variables at that point (i.e., the weights in the kernel). Convolutions are performed during backpropagation. Due to the high accuracy requirements, it is unlikely that these convolutions can be implemented optically.

[0094] Once training is complete, the network can be deployed for inference, where data is simply presented to the network and propagated until the final classification layer.

[0095] Applying optical processing to ConvNet / Implementing optical convolution

[0096] Optical 4f correlators can be used to evaluate convolutions and, in particular, accelerate ConvNet inference applications. Advantageously, inference can be made to work with relatively low accuracy (compared to training, which requires high accuracy, especially during backpropagation, to make numerical gradient descent methods robust).

[0097] The optical implementation of convolution is not a symmetric operation. Referring back to the "Optical Fourier Transform and Optical Correlator" section above, one independent variable 'a' is directly input into the optical system, while for a second independent variable 'b', its Fourier domain representation 'B' is computed locally in the optical correlator and displayed as a "filter". This is the opposite of a pure convolution in which 'a' and 'b' are interchangeable.

[0098] Therefore, the process of creating filters and inputs leads to performance asymmetry. Illustratively, this Fourier domain convolution overview is... Figure 6 middle. Figure 6 The diagram illustrates the principle of optical evaluation of convolution with reference to a 4f correlator. An optical Fourier transform is performed on the input, which is then optically multiplied by a filter, and then an inverse optical Fourier transform is performed to produce the convolution. The filter is digitally computed from the kernel. (This can be a fast process due to the low filter resolution).

[0099] This architecture requires digital computation of kernel-based filters. This is not a huge overhead. Once training is complete, the pre-computed filters can be stored, thus eliminating the overhead during interference. Second, the relatively low resolution of the kernel simplifies filter computation. The technical aspects of kernel computation will be discussed below.

[0100] Considering the hardware, optical convolution is an O(1) process. Convolution is performed at the resolution of the SLM and the camera during the system's "cycle time" (i.e., the time the system spends updating itself). Different aspects of the system (input SLM, filter SLM, and camera) can have different cycle times. The effective throughput of the system is determined by the speed of the fastest component; slower components limit the actual operations the system can perform.

[0101] These hardware considerations lead to at least four different technical problems in evaluating convolutions used in neural networks using optical processors:

[0102] 1. O(1) operations are fixed-resolution convolutions. On the other hand, applications that require flexibility in terms of convolution resolution should be accommodated.

[0103] 2. The limited operating range of the SLM means that arbitrary data cannot be input into the system. This affects both the input and the filter.

[0104] 3. By using a camera sensor to measure the results, what is measured is intensity, not complex amplitude. Even if the output is a real number, this not only prevents the measurement of the phase of the output, but also makes it impossible to determine whether the result is positive or negative (signed).

[0105] 4. Hardware limitations mean that different components of the system (input, filter, camera) may operate at different speeds. Specifically, the camera may operate slower than the input and / or the SLM filter. Therefore, the system output is integrated across several different frames.

[0106] Now we will address these problems and their solutions in turn.

[0107] 1. Fixed-resolution convolution

[0108] As discussed, the convolutions performed depend on the resolution of the system hardware. While ConvNets may have relatively high-resolution inputs, the pooling stage means that the resolution of the convolutions decreases as the network progresses. Even the first convolutional layer is unlikely to utilize the resolution of the correlator. This problem can be addressed by making the convolutions optically parallel. By arranging several inputs on the input SLM, parallel convolutions with respect to the same kernel can exist.

[0109] Referring back to Figure 2, a set of inputs is divided into blocks to form function a. The desired convolution kernel b is transformed into an appropriate filter B. The resulting optical convolution is arranged across the output plane c.

[0110] The input must be properly separated. When evaluating discrete convolutions, a "full" convolution will cause the kernel width to extend beyond the input coverage area. Therefore, the different input blocks must be separated sufficiently to allow the data to be extracted without crosstalk between different results. To obtain only "effective" convolutions, the blocks can be made more compact by ignoring the boundaries of the results. Preferably, when making the blocks more compact, the system can be configured to avoid overlap between full convolution regions and identical filled regions of adjacent images.

[0111] Technically, this can be called "same-padding" convolution. In CNN terminology, an "effective" convolution refers to one in which no zero padding is applied (i.e., the kernel must remain completely within the "effective" region defined by the image size): this results in an output image that is smaller than the kernel width of the input image. For small kernels, this is a small difference, but in most CNN cases, "same-padding" convolution is preferred over full or effective convolution.

[0112] However, input partitioning is not the only available form of partitioning. The kernel can also be partitioned before being converted into an optical filter. The corresponding convolution is then partitioned. Sufficient space should be allowed between the partitioned kernels so that the resulting convolutions do not collide. Implementing this partitioned form is more challenging because it places higher demands on the optical filter function and can lead to performance degradation as more kernels are partitioned together.

[0113] Figures 7(a) to 7(d) This demonstrates different ways to perform multiple lower-resolution convolutions using fixed-input hardware convolutions. By appropriately combining the input and kernel, it is possible to ensure efficient use of the hardware convolution resolution; the goal is to ensure that the output (camera) resolution is fully utilized.

[0114] Evaluating a batch of lower-resolution Fourier transforms is not as efficient as evaluating a single high-resolution Fourier transform using the system. This parallelization impacts the competitiveness of optical methods compared to digital methods. Advantageously, optical methods offer Fourier transforms as O(1) processes, compared to the O(N log N) process provided by digital Fourier transforms. The magnitude of the Fourier transform is determined by the system resolution. The maximum performance gain can be achieved when using this full resolution.

[0115] When we don't utilize the full-resolution Fourier transform, although block-based processing somewhat compensates for this problem, we are not taking advantage of the full system performance. While we are still utilizing the system's full resolution, we are not leveraging the high resolution of the corresponding Fourier transform. Instead, we use it to implement a "batch" convolution process. This batch approach is equivalent to several smaller Fourier transforms.

[0116] When the inherent full-resolution Fourier transform is not used directly, but instead a batch of lower-resolution transforms are performed, some of the potential performance is "lost." The following comparison between these two applications is made using a simple calculation of the scaling independent variable.

[0117] The comparison calculation ratio is considered to be determined by the ratio of the Fourier transform (which is "cheaper" in terms of element-wise multiplication). Therefore, our performance can be compared as follows:

[0118] FFT: O(N log N)

[0119] OFT: O(1),

[0120] There are a total of N pixels. Instead of evaluating a single transformation of size N, we evaluate a batch of P transformations of size M, where N = M·P. The same O(1) scaling applies to OFT, but FFT now has scaling:

[0121] Balch-FFT: O(P·M log M).

[0122] A “deceleration factor” S is defined, which describes the achievable performance of the roofline by following these simple proportional independent variables:

[0123]

[0124]

[0125] The logarithmic base in this formula is irrelevant. This formula is part of the roofline performance advantage of the optical system relative to the computer expected to achieve when performing batch operations.

[0126] Figure 8This is a graph showing the "deceleration factor" S relative to the roofline performance when the 10MPx input is divided into P batches compared to the digital Fourier transform. As can be seen, the behavior is relatively favorable—even when the input is divided into one million 10-pixel batches, the performance loss is only one-tenth.

[0127] 2. Accommodates SLM operating range

[0128] As discussed, both the input and the filter SLM have finite operating ranges. This means that arbitrary convolutions cannot be performed. Simply implementing the convolution kernel optically is not necessarily the direct approach.

[0129] This can be illustrated by considering a filter SLM. Consider that we have a target kernel and want to determine the corresponding filter. Given the known operating range of the SLM, we can find the best representation of that filter. However, it will not exactly represent the target kernel, but rather a different kernel. In CNN applications, we can call this actual kernel the "hidden kernel." This process is further elaborated in... Figure 9 The diagram is shown in the image.

[0130] like Figure 9 As shown, kernel 20 and input 21 are digitally convolved together within the neural network to produce an optical convolution 22. Given the range of SLM operations, an optimal filter 23 corresponding to a given kernel is generated. Given this filter and inversely computed against the true "hidden" kernel, a different convolution is expected, as shown at 24. The measured optical convolution is shown at 25—a good agreement with the theoretical optical convolution 24 can be observed.

[0131] The new nucleus retains the same properties as the original nucleus, but it is not entirely identical. However, as... Figure 9 As shown, the optical system performs as expected. To avoid any doubt, the system does perform optical convolutions with high fidelity; it just cannot perform all convolutions—including… Figure 9 The examples shown are arbitrary. However, the hidden kernels implemented in practice are well-defined and well-known. There are several different ways to derive the filter from the target kernel and thus the corresponding hidden kernel.

[0132] The general workflow for solving this problem involves training the network on a conventional digital computer and then deploying it on an optical platform to accelerate inference applications. Training must be based on an understanding of the performance of the optics. The underlying principle is that the minimization pattern explored during ConvNet training is not particularly compelling. Appropriate solutions should exist within the region accessible by optical convolutions. There are two fundamental approaches to solving the finite range of SLMs, which derive filters from the kernel in different ways:

[0133] 1. First, a strategy can be used that employs the optimal optical filter for a given target kernel, as determined by some metric. When training the network, it can be assumed that the optimal implementation is sufficiently close to a regular digital convolution, or the properties of this optical convolution can be incorporated into the network's training.

[0134] 2. Second, the network can be trained without knowledge of optical convolution properties and without ignoring any complex filter design strategies. This can be achieved, for example, by directly parameterizing the kernel. Instead of directly specifying the kernel using 25 parameters of a 5x5 kernel, the filters can be parameterized using these parameters through a suitable set of fundamental functions. The suitable set of fundamental functions is Fourier basis; the kernel will undergo a Fourier transform, and the result will be appropriately scaled to directly map to the driving values ​​of the SLM. This results in hidden kernels that are arbitrarily correlated with the driving convolution kernel values. For backpropagation to still allow for successful training, the network's connectivity must be the same as that of a regular convolution (identical elements of the kernel link the same pixel in the input and output feature maps together), and accurate convolutions must be used during backpropagation.

[0135] In one implementation, training is performed entirely digitally using conventional convolution operations or high-fidelity simulations of correlators. The trained neural network is then potentially deployed on optics for inference, potentially after intermediate quantization or calibration steps. In a second implementation, optical convolutions are used directly during training. The latter is expected to work most efficiently, taking into account optical aberrations, device defects, etc.

[0136] Note that constraints on the kernel representation have an impact during the training step. Backpropagation involves setting the convolution kernel to be error-prone and needs to be performed with appropriate fidelity. This is not a problem if the model is trained in a computer and then deployed optically (an important first step in development). Several techniques can be used to enable optical implementation of error-prone kernels by improving filter performance, such as: improved SLM characteristics and optimization; spatiotemporal jitter in filtering; and combining multiple SLMs to extend the SLM modulation range.

[0137] When attempting to implement a given kernel, the appropriate approach is the minimum Euclidean distance (MED) method, commonly used in generating optical filters for correlators. The principle is to find an accessible point in the complex plane that is closest to the given complex value required for the Fourier transform of the filter.

[0138] This method is most successful if the complex values ​​(“operation curves”) available for the SLM are in mixed mode, because they modulate both amplitude and phase.

[0139] Importantly, the filter does not contain spatial frequency components that would make the kernel effectively larger than it should be. In most embodiments, subsequent layers of the same input cannot be computed simultaneously, thus filter-bleed crosstalk occurs between different training examples in image batches rather than between different layers of the network.

[0140] One technique to avoid this is to ensure that the core design process does not introduce any high spatial frequency components. For example, if a high core resolution is required for display on an SLM, padding with zeros and then performing a Fourier transform, or performing a Fourier transform and then performing sine interpolation, are both effective methods (other interpolation schemes—though suboptimal—may provide sufficient performance).

[0141] 3. Bipolar results

[0142] Measuring optical intensity using a camera means that the sign of the output function cannot be directly determined. While some network configurations can tolerate this problem, it remains potentially significant because, for example, commonly used nonlinearities (e.g., ReLU) depend on the sign of the measured function. However, the inventors have developed a method to directly determine the sign of the resulting convolution by utilizing a bias function.

[0143] Considering the objective is to optically evaluate the convolution f*k, where f is positive and k is a bipolar kernel (2D matrices are indicated in bold here), this convolution can be treated as the difference between two separate convolutions:

[0144]

[0145] Where I is an identity matrix of the same size as the kernel. To produce a positive bias kernel p, a bias b is applied to the kernel. Then, we have a second kernel, which is simply a matrix of multiple 1s; the application of this kernel is the same as that of the boxcar average. Now, both convolutions exclusively have positive inputs and therefore positive outputs. The camera can take measurements without losing amplitude information because it is known a priori. The difference between these two functions then produces the desired convolution.

[0146] Figure 10 Two examples of digital computation of bipolar convolution and corresponding optical convolution obtained using the method described above are shown. Figure 10 As shown, this method is effective in obtaining bipolar convolutions. The presented convolutions are not directly from camera imaging, but rather the result of this difference process.

[0147] This method has minimal performance overhead. The filter corresponding to kernel I is negligible, and the differential overhead of the two functions is negligible. These two operations can be performed in parallel.

[0148] Furthermore, if the input f is also bipolar, it can be divided into positive and negative components, processed separately, and the result can be reconstructed by combining the resulting convolutions.

[0149] 4. The camera is slower than the SLM.

[0150] Due to hardware limitations, camera operation is typically slower than SLM. Therefore, while the system will be able to process data at a rate determined by the SLM's throughput, it cannot independently capture the resulting frames. A particular implementation could involve a fast input SLM, a slower filter SLM, and a medium-speed camera.

[0151] However, the system can still be used usefully. Camera exposure can span multiple input SLM frames. The camera then captures multiple convolutions, which are optically integrated; the frames are added together.

[0152] This facilitates the typical convolution operation present within neural networks, which actively implement this addition. For example, consider the convolutional layers in a Theano neural network operating on a 4D tensor. The operation inputs (A, B) and output (C) are defined as:

[0153] A=A(batch, input_channel, x, y)%Feature maps

[0154] B=B(output_channel, input_channel, x, y)%Kernels

[0155] C=C(batch,output_channel,x,y)%Outputs

[0156] =∑ i conv2[A(batch,i,:,:),B(output_channel,i,:,:)]

[0157] As you can see, although the basic operation is 2D convolution, this is immediately incorporated into the summation of other 2D convolutions. This can be optically simulated directly by exposing the camera across multiple input frames (but note that we will sum the intensity amplitudes, not the amplitudes representing the actual convolution result).

[0158] Other aspects

[0159] In some embodiments, the pooling scheme is implemented optically. The purpose of pooling is to reduce the resolution of a given feature map using some downsampling scheme. By engineering the optical scaling of the system, a given input resolution can be rendered at a given (smaller) output resolution. This effect naturally achieves L2 pooling (sum of intensities).

[0160] The highest bandwidth SLM is binary and can be used to represent filters. A binary filter B corresponding to a much more general kernel b can be made. Due to binary filters—for example, symmetry—there are some constraints on the kernel, but there are several ways to address this (e.g., combining a fixed random phase mask with the filter SLM). The constraint that the filter is binary does not correspond to the constraint that the kernel is binary. However, it does have a significant impact on limiting the spectral content of the filter, and controlling the kernel width with a binary filter is challenging. For this reason, it is preferable not to use a binary SLM as the filter SLM.

[0161] Another—more useful—way is to use a high-bandwidth binary SLM instead of the input, which represents multi-level inputs through time or space jitter, or to implement a network configuration designed to use only binary inputs.

[0162] A key requirement for driver-side integration is the rapid transfer of data into and out of the optical domain with minimal latency. ConvNet represents the high computational load of different basic operations—a requirement for high-performance systems to handle such computational bandwidth. Furthermore, to achieve high system utilization, several jobs must be batched together. However, this will impact latency.

[0163] The use of FPGA-based driver systems offers powerful capabilities and flexibility. Beyond simply driving I / O, it can also be used to leverage the capabilities of optical systems. Functions such as tissue batching, filter generation, and determination of bipolar convolutions can be implemented in hardware.

[0164] Furthermore, in some embodiments, the FPGA allows for the development of integrated solutions outside of convolutional layers. Nonlinear activation functions can be implemented in hardware. Additionally, pooling schemes can be implemented optically, in hardware within the FPGA, or in software.

[0165] In another implementation, an integrated solution can be provided, where the driving electronics also implement other layers of the neural network. This means that data can be digitized quickly and then transferred back to the optical domain between the convolutional layers with minimal latency.

Claims

1. An optical processing system for optically processing layers of a convolutional neural network (CNN), the optical processing system comprising at least one spatial light modulator (SLM) configured to display a first input data pattern (a); the optical processing system further comprising a second input data pattern (b) having one or more kernels, the one or more kernels being computed via a fast Fourier transform (FFT) as a Fourier domain representation (B) of the second input data pattern (b), the Fourier domain representation (B) being displayed as one or more optical filters; performing an optical Fourier transform (OFT) on the first input data pattern (a), optically multiplying it by the one or more optical filters, and then performing an inverse optical Fourier transform to produce an optical convolution; the optical processing system further comprising a detector for detecting light that has been optically processed, wherein, The first input data pattern (a) includes multiple (N) block input data patterns, each of which corresponds to a different segment of the first input data pattern, and the second input data pattern (b) includes a single block kernel data pattern, wherein multiple convolutions are generated in parallel for each pair (N×1) formed by the block input data patterns and the single block kernel data pattern; or The first input data pattern (a) includes a single block input data pattern, and the second input data pattern (b) includes a plurality of (M) block kernel data patterns, each of the plurality of block kernel data patterns corresponding to a different segment of the second input data pattern (b), and wherein, for each pair (1×M) formed by the single block input data pattern and the block kernel data patterns of the plurality of block kernel data patterns, a plurality of convolutions are generated in parallel.

2. The optical processing system according to claim 1, wherein, The detector and the at least one SLM operate at different speeds; the detector operates at a slower speed than the at least one SLM.

3. The optical processing system according to claim 1 or claim 2, wherein, The system has outputs integrated across several different frames.

4. The optical processing system according to claim 1 or claim 2 further includes a processor for digitally obtaining a data focusing mode from the second input data mode.

5. The optical processing system according to claim 4, wherein, The processor is also configured to use the minimum Euclidean distance (MED) to obtain the data focusing mode.

6. The optical processing system according to claim 4, wherein, The plurality of kernels consists of several kernels with appropriately padded blocks, such that a convolution of an input with the plurality of kernels is performed in parallel; the second input data pattern is converted for display on the filter SLM.

7. The optical processing system according to claim 1 or claim 2, wherein, The Fourier domain representation (B) contains only spatial frequency components up to a finite maximum frequency, thereby limiting the effective width of the kernel corresponding to the second input data pattern (b).

8. The optical processing system according to claim 1 or claim 2, wherein, The second input data pattern (b) is a bipolar kernel with positive and negative components, and wherein the optical convolution is the difference between a first convolution of the first data input and the positive representation of the kernel and a second convolution of the first data input and the kernel with applied bias.

9. The optical processing system according to claim 1 or claim 2, wherein, The SLM is a binary SLM.

10. The optical processing system of claim 1 or claim 2, further comprising a lens for controlling amplification and for optically pooling the neural network.

11. The optical processing system according to claim 1 or claim 2, wherein, The system includes an input, a filter, and a camera that operate at different speeds; the camera is slower than both the input and the filter and is used to integrate multiple frames to synthesize a higher level of convolutional operations than those used in existing CNN frameworks.

12. The optical processing system according to claim 11, wherein, The system includes at least one input SLM and at least one filter SLM; the input SLM and the filter SLM operate at the same speed; and wherein the camera operates at a slower speed than both of the SLMs.

13. The optical processing system according to claim 1 or claim 2, wherein, The detector has exposures spanning multiple input SLM frames.

14. The optical processing system according to claim 1 or claim 2, wherein, The detector is configured to capture multiple optically integrated convolutions.

15. The optical processing system according to claim 1 or claim 2, wherein, The optical convolution comprises several different layers, wherein the input is convolved with a given kernel.

16. The optical processing system according to claim 1 or claim 2, wherein, The optical convolution includes a pooling layer, wherein the layer is spatially downsampled.

17. The optical processing system according to claim 1 or claim 2, wherein, The optical convolution includes an activation layer in which a nonlinear activation function is applied.

18. The optical processing system according to claim 1 or claim 2, wherein, The optical convolution includes a fully connected layer with fully specified weights between each node.

19. A method for generating optical convolution using an optical processing system according to any one of the preceding claims.

20. The method of claim 19, for use in a layer of a convolutional or pooling neural network.

21. The method of claim 19, for training or inference of a convolutional layer of a neural network.

22. The use of the system according to any one of claims 1 to 18 for image analysis in deep learning or machine learning.

23. A method for configuring a neural network and layers of an optically processed convolutional neural network (CNN), the method comprising the following steps: Display the first input data mode (a); Provide one or more kernels and compute the Fourier domain representation (B) of the one or more kernels as a second input data pattern (b) by means of Fast Fourier Transform (FFT); The Fourier domain representation (B) is displayed as one or more optical filters; The first input data pattern (a) is subjected to an optical Fourier transform (OFT) and optically multiplied by the one or more optical filters, and then an optical Fourier transform is performed in reverse to produce an optical convolution. as well as Detecting optically processed light, in which... The first input data pattern (a) includes multiple (N) block input data patterns, each of which corresponds to a different segment of the first input data pattern, and the second input data pattern (b) includes a single block kernel data pattern, wherein multiple convolutions are generated in parallel for each pair (N×1) formed by the block input data patterns and the single block kernel data pattern; or The first input data pattern (a) includes a single block input data pattern, and the second input data pattern (b) includes a plurality of (M) block kernel data patterns, each of the plurality of block kernel data patterns corresponding to a different segment of the second input data pattern (b), and wherein, for each pair (1×M) formed by the single block input data pattern and the block kernel data patterns of the plurality of block kernel data patterns, a plurality of convolutions are generated in parallel.

24. The method according to claim 23, wherein, The method includes the following additional steps: providing at least one spatial light modulator (SLM) and detector for display modes; and operating the at least one SLM and the detector at different speeds; the detector operating at a slower speed than the at least one SLM.

25. The method of claim 24, further comprising the step of integrating the output on multiple different frames.

26. The method according to any one of claims 23 to 25, further comprising the following step: At least one block of the first input data pattern and the second input data pattern is divided, and multiple optical convolutions are generated in parallel.

Citation Information

Patent Citations

  • Optical correlator

    EP1546838A1

  • Advanced miniature processing handware for ATR applications

    US6529614B1