Neural network quantization method in optical calculation, optical calculation method and artificial intelligence equipment
By mapping tensor data elements to numerical values in preset sets that meet specific distribution characteristics in optical calculations, the accuracy problem caused by mismatch between data distribution and hardware requirements is solved, and higher computational accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510275128.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In optical computing, the data distribution does not match the hardware requirements, resulting in the accuracy problem of the computing model on the optical computing chip.
By receiving tensor data and mapping its elements to the closest numerical value in a preset set, the numerical distribution in the preset set satisfies the dense and sparse distribution characteristics in some intervals, adapting to the characteristics of optical computing hardware.
It improves the accuracy and efficiency of the computing model on the optical computing chip, reduces the data complexity, and adapts to the characteristics of optical computing hardware.
Smart Images

Figure CN120218151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for neural network quantization in optical computing, an optical computing method, and an artificial intelligence device. Background Art
[0002] In optical computing, optical computing hardware implements computing based on optical devices, so there are specific requirements for the type and distribution of input data. The inventor of the present invention has found that if the data distribution does not match the hardware requirements in optical computing, it may affect the accuracy of the computing model on the optical computing chip. Summary of the Invention
[0003] The present invention provides a method for neural network quantization in optical computing, an optical computing method, and an artificial intelligence device, aiming to solve the accuracy problem caused by insufficient adaptation of the hardware characteristics of the computing model on the optical computing chip.
[0004] In a first aspect, the present invention provides a quantization method in optical computing, including:
[0005] Receiving tensor data, where the tensor data includes a plurality of elements, and each element is a numerical value;
[0006] Through a mapping operation, mapping the elements of the tensor data to a certain numerical value in a preset set, where the preset set includes a group of preset numerical values, and the mapping operation includes: for each of the elements, selecting the numerical value closest to it in the set for mapping;
[0007] Wherein, the value distribution of the mapped numerical values satisfies: having a dense distribution in the first interval [a1, b1] and a sparse distribution in the second interval [a2, b2], where 0 < a1 < b1 <= a2 < b2.
[0008] In a second aspect, the present invention further provides a method for neural network quantization in optical computing, including:
[0009] Receiving tensor data, where the tensor data includes a plurality of elements, and each element is a numerical value;
[0010] Through a mapping operation, mapping the elements of the tensor data to a certain numerical value in a preset set, where the preset set includes a group of preset numerical values, and the mapping operation includes: for each of the elements, selecting the numerical value closest to it in the set for mapping, where the preset numerical values can be expressed as: Where a i ∈{0, -1, 1}; h i is selected from integers; b0 is a constant.
[0011] In a third aspect, the present invention further provides an optical computing method, which quantizes tensor data using the method described in the first aspect. The optical computing method further includes:
[0012] Representing the elements in the quantized tensor data with optical signals,
[0013] Performing multiplication calculations on the elements.
[0014] In a fourth aspect, the present invention further provides an artificial intelligence device, which performs quantization processing using any one of the methods described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a flowchart of the method for neural network quantization in optical computing provided by the present invention;
[0017] Figure 2 is a flowchart of neural network quantization in optical computing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0019] The terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here.
[0020] Optical computing can be used in, for example, neural networks. The inventor of the present invention found that if the data distribution does not match the hardware requirements in optical computing, it may affect the accuracy of the computing model on the optical computing chip. Therefore, a neural network quantization method that adapts to the characteristics of optical computing hardware is needed to improve the accuracy and efficiency of the computing model.
[0021] The following is combined with Figure 1 - Figure 2Describe the method of neural network quantization in optical computing, the optical computing method, and the artificial intelligence device of the present invention.
[0022] Please refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the method of neural network quantization in optical computing provided by the present invention. A method of neural network quantization in optical computing includes:
[0023] S110, receiving tensor data, where the tensor data includes multiple elements, and each element is a numerical value.
[0024] Specifically, the system receives the input tensor data. The tensor data is organized in the form of a multi-dimensional array and includes multiple elements, and each element is a specific numerical value. The tensor data is the core data structure in the computing model and is used to represent input features, weight parameters, or intermediate calculation results. By receiving and processing the tensor data, the system can provide a basis for subsequent quantization operations, thereby optimizing the data distribution and improving the computing accuracy of the optical computing hardware.
[0025] S120, through a mapping operation, map the elements of the tensor data to a certain numerical value in a preset set, where the preset set includes a group of preset numerical values, and the mapping operation includes: for each of the elements, select the numerical value in the set that is closest for mapping; wherein, the value distribution after the mapping processing satisfies: having a dense distribution in the first interval [a1, b1] and a sparse distribution in the second interval [a2, b2], where 0 < a1 < b1 <= a2 < b2.
[0026] Specifically, the system maps each element of the tensor data to a certain numerical value in a preset set through a mapping operation. The preset set contains a group of predefined numerical values. The mapping operation method is to select the numerical value in the preset set that is closest to each element as the mapping result. After the mapping processing, the distribution characteristics of the numerical values satisfy: being densely distributed in the first interval [a1, b1], while being sparsely distributed in the second interval [a2, b2], where 0 < a1 < b1 <= a2 < b2. This non-linear distribution makes the expression of small numerical values more refined and the expression of large numerical values more simplified, thereby optimizing the quantized data distribution and making it more adaptable to the characteristics of the optical computing hardware, and finally improving the computing accuracy.
[0027] In some embodiments, the elements in the preset set can be expressed as
[0028]
[0029] where, a i ∈ {0, -1, 1}, that is, a i can take the values of 0, -1, or 1; h iis an integer; b0 is a constant. By selecting a i , and different values of h i , multiple different elements p can be constructed to form the preset set.
[0030] The following specifically describes the above steps S110 to S120.
[0031] In some embodiments, in step S110, the element p in the preset set can be expressed as:
[0032]
[0033] is an integer, and m, n are integers greater than 0.
[0034] Specifically, a certain element in the preset set where each a i takes values from the set These values are expressed in the form of negative exponents of 2, and the exponent part is related to the parameters i, m, n. The number of available values in the preset set is the quantization level, and the quantization level is 2 mn . Such a design makes the numerical distribution in the preset set have non-linearity and scalability, can flexibly adapt to the requirements of different quantization precisions, thereby optimizing the quantization effect and improving the performance of the optical computing system.
[0035] Exemplarily, the value of q i is: q0 = {0, 2 0 , 2 -2 , 2 -4}, q1 = {0, 2 -1 , 2 -3 , 2 -5}, and the preset set is the combination of the sum of q0 and q1, with a total of 2 mn = 2 4 = 16 values, that is
[0036] In some embodiments, before performing the mapping operation, the tensor is scaled to a preset range by dividing the tensor by a scaling factor; and, the mapping operation is performed on the scaled tensor.
[0037] For example, the tensor T = [0.3, 0.6, 0.9, 1.2], and the scaling factor is s = 1.2. Divide the tensor T by the scaling factor s to scale the values to the preset range [0, 1], that is:
[0038] T_scaled = T / s = [0.3 / 1.2, 0.6 / 1.2, 0.9 / 1.2, 1.2 / 1.2] = [0.25, 0.5, 0.75, 1.0].
[0039] Regarding the mapping operation, for example, for a preset set Taking the element 0.75 in T_scaled as an example, the value 3 / 4 in the preset set that is closest to it is selected. Therefore, it is exactly mapped to 0.75. For another example, if the value before mapping is 0.99, it is closest to 1 in the preset set. Therefore, it is mapped to 1.
[0040] In some embodiments, the data after the mapping operation is multiplied by a scaling factor.
[0041] Therefore, through the scaling and mapping operations, the original data can be quantized to the values in the preset set, which can significantly reduce the data complexity, adapt to the characteristics of the optical computing hardware, and improve the computing accuracy and performance.
[0042] In some embodiments, the tensor data is activation value data, and the corresponding preset set is a set of non-negative values.
[0043] Specifically, the tensor data can be the activation value data of the computing model.
[0044] In some embodiments, the tensor data is a weight value, and the corresponding preset set is a set containing symmetric values.
[0045] Specifically, the tensor data can be the weight value in a computing model (such as a neural network model), that is, the parameter connecting different neurons, which is used to adjust the influence of the input data on the output.
[0046] In some embodiments, the method further includes performing a clipping operation on the tensor data before executing the mapping operation.
[0047] Specifically, in the quantization process, the mapping operation is used to map the tensor data to the preset set. Before performing the mapping operation, a clipping operation needs to be performed on the tensor data. The clipping operation is used to make the data value after clipping within the preset range. For example, when processing the weight value, the data is restricted to the interval [-1, 1] through clamp(data, -1, 1); when processing the activation value, the data is restricted to the interval [0, 1] through clamp(data, 0, 1). The purpose of the clipping operation is to prevent the data value from exceeding the preset range, so that the mapping operation can efficiently and accurately map the data to the preset set.
[0048] Exemplarily, the method for neural network quantization in the optical computing includes:
[0049] Step 1: Perform a clipping operation. This includes processing at least one of the weight values or the activations:
[0050] ① Process the weight values:
[0051] Use clamp(data, -1, 1) to limit the scaled weight values within the interval [-1, 1]: For example, T_weights_clipped = clamp(T_weights, -1, 1) = [1.0, -0.3, 0.8, -1.0].
[0052] ② Process the activation values:
[0053] Use clamp(data, 0, 1) to limit the activation values within the interval [0, 1]: For example,
[0054] T_activations_clipped = clamp(T_activations, 0, 1) = [0.7, 1.0, 0.0, 0.9].
[0055] Step 2: Perform a mapping operation. Specifically, map the clipped tensor data to the numerically closest value in a preset set. This includes:
[0056] ① Weight value mapping; and / or
[0057] ② Activation value mapping.
[0058] Optionally, if it is necessary to restore the quantized data to the original range, the quantization result can be multiplied by a scaling factor s, for example, to restore the weight values, and / or to restore the activation values.
[0059] The above clipping operation can limit the weight values and activation values within the intervals [-1, 1] and [0, 1] respectively through the clipping operation, so that the data values are within the preset range and prevent quantization errors. The mapping operation is to map the clipped data to the numerically closest value in the preset set to map the data into the preset set and improve the quantization accuracy. The restoration operation is to restore the quantization result to the original range by multiplying by the scaling factor s.
[0060] Please refer to Figure 2 , Figure 2 which is the flowchart of neural network quantization in optical computing provided by an embodiment of the present invention. A method for neural network quantization in optical computing includes:
[0061] S210, receive an input tensor and perform a data scaling operation.
[0062] Receive the input tensor t (as tensor data), where the input tensor t can be a weight tensor or an activation value tensor. Divide the input tensor t by the scaling factor α to obtain the scaled data d, with the formula d = t / α. The scaling factor α is used to narrow the numerical range of the input tensor t to make it more suitable for subsequent quantization operations. The scaled data d is within a smaller range (such as [-1, 1] or [0, 1]), which is convenient for quantization.
[0063] S220, perform a clipping operation.
[0064] Clip the scaled data d according to whether the input tensor is a weight.
[0065] If the input tensor is a weight, clip the scaled data d through d clamped = clamp(d, -1, 1) to clip d to the range [-1, 1].
[0066] If the input tensor is an activation, clip the scaled data d through d clamped = clamp(d, 0, 1) to clip d to the range [0, 1].
[0067] Among them, clamp(x, a, b) = max(a, min(x, b)). The purpose of the clipping operation is to make the data within a preset range before quantization to avoid interference from extreme values on the quantization accuracy.
[0068] S230, perform a mapping operation.
[0069] Map the clipped data d clamped to the closest value in the preset set S = {s1, s2,..., s n}.
[0070] Therefore, by performing the mapping operation, the clipped data is mapped to the preset set S, thereby reducing the computational and storage overhead while maintaining a high accuracy.
[0071] S240, restore the scaling operation.
[0072] The purpose of restoring the scaling step is to restore the quantized data d q to the original numerical range so that the quantized tensor is very close to the input tensor numerically. Specifically, multiply the data d after the mapping operation q by the scaling factor α to obtain the final quantized tensor t q , with the formula:
[0073] t q = d q × α.
[0074] In some embodiments, the present invention further provides an optical computing method, which quantifies tensor data by using the method of neural network quantization in optical computing as described above. The optical computing method further includes:
[0075] Representing the elements in the quantized tensor data with optical signals,
[0076] Performing multiplication calculations on the elements.
[0077] That is to say, using the method of neural network quantization in the aforementioned optical computing to quantify tensor data specifically includes the following steps: First, convert the elements in the quantized tensor data into optical signal representations, taking advantage of the efficient transmission characteristics of optical signals; Second, perform multiplication calculations on the elements represented by optical signals, giving full play to the advantages of optical computing chips in parallel computing and high-speed processing. It reduces the complexity of data processing and significantly improves the computing efficiency and accuracy, and is applicable to application scenarios that require high-performance computing, such as deep learning and large-scale data processing.
[0078] In some embodiments, the present invention further provides a method for neural network quantization in optical computing, which is characterized by including:
[0079] Receiving tensor data, where the tensor data contains multiple elements, and each element is a numerical value;
[0080] Through a mapping operation, map the elements of the tensor data to a certain numerical value in a preset set, where the preset set includes a group of preset numerical values, and the mapping operation includes: for each of the elements, select the closest numerical value in the set for mapping, where the preset numerical value can be expressed as: Where a i ∈{0, -1, 1}; h i is selected from integers; b0 is a constant.
[0081] This embodiment solves the accuracy problem caused by insufficient adaptation of the computing model to the hardware characteristics on the optical computing chip. Specifically, this method maps the received tensor data and its elements to the closest numerical value in a preset set, where the preset numerical value adopts a specific mathematical expression (such as Where a i ∈{0, -1, 1}, h i is an integer, and b0 is a constant), thereby realizing precise quantization of the data. This mapping operation not only reduces the computational complexity but also improves the hardware adaptability, ensuring that the quantized data can be efficiently and accurately executed on the optical computing chip.
[0082] In some embodiments, the present invention further provides an artificial intelligence device, which performs quantization processing by using the method for neural network quantization in optical computing described in any of the above embodiments.
[0083] That is to say, the artificial intelligence device performs quantization processing on data by using the method of neural network quantization in optical computing described in any of the foregoing embodiments. Specifically, the artificial intelligence device receives tensor data and maps its elements to the closest numerical value in a preset set, and uses a specific mathematical expression (such as ) to achieve efficient and accurate quantization. This method not only reduces the computational complexity but also improves the hardware adaptability, and is applicable to high-performance computing platforms such as optical computing chips.
[0084] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for quantizing neural networks in optical computing, characterized in that: include: Receive tensor data, wherein the tensor data includes a plurality of elements, each element being a value; By means of a mapping operation, the element of the tensor data is mapped to a certain value in a preset set, wherein the preset set includes a group of preset values, and the mapping operation includes: for each of the elements, selecting the closest value in the set for mapping; The distribution of the numerical values after the mapping process satisfies: a dense distribution in the first interval [a1, b1] and a sparse distribution in the second interval [a2, b2], wherein 0 <a1<b1<=a2<b2。 2. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: The element p in the preset set is expressed as: in, The number of possible values in the preset set is the quantization level; wherein i is an integer greater than or equal to 0, and m and n are integers greater than 0.
3. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: Before performing the mapping operation, the tensor is scaled to a preset range by dividing the tensor by a scaling factor; and the mapping operation is performed on the scaled tensor.
4. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: The data after the mapping operation is multiplied by a scaling factor to restore the quantized result to the original range.
5. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: The tensor data is activation value data, and the corresponding preset set is a set of non-negative values.
6. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: The tensor data is a weight value, and the corresponding preset set is a set containing symmetric values.
7. The method for quantizing a neural network in optical computing according to claim 1, characterized in that: The method further includes, before performing the mapping operation, performing a clipping operation on the tensor data.
8. A method for quantizing neural networks in optical computing, characterized in that: include: Receive tensor data, wherein the tensor data includes a plurality of elements, each element being a value; Through a mapping operation, the element of the tensor data is mapped to a certain value in a preset set, wherein the preset set includes a group of preset values, and the mapping operation includes: for each of the elements, selecting the closest value in the set for mapping, wherein the preset value is represented as: Among them, a i ∈{0,-1,1}; h i is selected from integers; b0 is a constant.
9. A method of optical computing, characterized in that: The tensor data is quantized using the method according to claim 1, wherein the optical computing method further comprises: Use optical signals to represent the elements in the quantized tensor data. Performs multiplication on the elements.
10. An artificial intelligence device, which performs quantization processing using the method described in any one of claims 1 to 8.