Neural network system, floating-point number processing method and device

Through fault-tolerant training (FTT) and fault-tolerant floating point number (FTF) formats, the impact of memory errors on artificial intelligence applications is solved, computing efficiency and accuracy are improved, and high-capacity and efficient computing needs of large language models are met.

CN120597943APending Publication Date: 2025-09-05MACRONIX INTERNATIONAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410294501.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2024-03-14
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing memory solutions have problems such as large memory errors, high costs and increased latency in artificial intelligence applications. Especially in new processors such as GPUs and HBMs, it is difficult to meet the high capacity and efficient computing needs of large language models.

Method used

The fault-tolerant training (FTT) method and fault-tolerant floating point number (FTF) format are used to identify and handle exception weights during the training process, set them as reference values, and combine custom floating point number formats to reduce the impact of bit errors and improve calculation stability and accuracy.

Benefits of technology

Effectively reduce the impact of errors in memory computing, improve the computing efficiency and accuracy of artificial intelligence applications, and reduce the delay and cost in model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597943A_ABST
    Figure CN120597943A_ABST
Patent Text Reader

Abstract

The invention provides a neural network system and a floating-point number processing method and device. The neural network system includes one or more memory devices storing one or more neural network models, the one or more memory devices when training the one or more neural network models performing: a training iteration of a start weight; judging whether an abnormal weight is found or not; when the abnormal weight is found, setting the abnormal weight as a reference value; and when no abnormal weight is found, carrying out the next training iteration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a neural network system, a floating-point number processing method and a device, and more particularly to a neural network system, a floating-point number processing method and a device that are compatible with a custom floating-point number format. Background Art

[0002] With the rise of artificial intelligence (AI) technology, especially neural network-based models, computing demands have shifted. This shift has shifted computing demands from being primarily concentrated on central processing units (CPUs) to new processors such as graphics processing units (GPUs), tensor processing units (TPUs), neural network processing units (NPUs), or field programmable gate arrays (FPGAs).

[0003] These new processors have significant differences from traditional CPUs in terms of memory requirements, mainly reflected in the following aspects.

[0004] (1) Less focus on latency: This means that these new processors are relatively less sensitive to latency during the computing process and may focus more on overall computing performance.

[0005] (2) Paying high attention to memory bandwidth: This means that these processors pay more attention to memory transfer rate to ensure efficient data access.

[0006] (3) Requires larger capacity: New processors have a larger demand for memory capacity, probably because processing large amounts of data or complex models requires more memory space.

[0007] (4) Higher requirements for energy efficiency: This means that these processors focus more on completing computing tasks at high efficiency to improve energy utilization efficiency.

[0008] Compared to the traditional CPU and dynamic random access memory (DRAM) combination, the typical memory solution for new processors is typically a GPU paired with graphics double data rate memory (GDDR) or a GPU paired with high bandwidth memory (HBM). These solutions feature both on the graphics card to reduce data transmission distance and energy consumption. Specifically, GDDR offers high bandwidth but limited capacity, while HBM offers higher bandwidth and greater capacity than GDDR.

[0009] The aforementioned memory solutions (GPU with GDDR or GPU with HBM) are quite expensive. Furthermore, the size of large language models (LLMs) is growing exponentially, meaning they are becoming larger and require processing more data and parameters. Therefore, the demand for higher-capacity memory in the AI ​​field is becoming more urgent. Simply put, as technology advances and model size grows, more advanced and larger memories are needed to meet the demands of AI applications.

[0010] In recent years, several new memory concepts have been developed to meet specific application requirements, including the following: (1) DRAM with ultra-wide I / O directly integrated with logic chips, providing extremely high bandwidth without the high cost of through-silicon vias (TSVs); (2) combining logic, DRAM, and volatile NAND to provide extremely high capacity; and (3) continued scaling to push the limits of memory to higher levels.

[0011] However, the unreliability of memory will affect these solutions, specifically: (1) DRAM and logic chips occupy a large area, making it difficult to ensure that there are no errors; (2) the NAND manufacturing process is rarely "naturally good"; and, (3) miniaturization will also increase the error rate.

[0012] While reliability issues can be addressed by adding additional controllers, this increases both cost and significantly increases latency.

[0013] Therefore, an object of the present disclosure is to provide a method to reduce the impact of memory errors and enable computing in memory (CIM) for artificial intelligence applications. Summary of the Invention

[0014] According to a first aspect of the present invention, a neural network system is proposed, comprising one or more memory devices, wherein the one or more memory devices store one or more neural network models. When the one or more neural network models are trained, the one or more memory devices execute: starting a training iteration of weights; determining whether an abnormal weight is found; when the abnormal weight is found, setting the abnormal weight to a reference value; and when no abnormal weight is found, proceeding to the next training iteration.

[0015] According to a second aspect of the present invention, a method for processing floating-point numbers is provided, which is applied to an electronic device. The method comprises: obtaining a custom floating-point number, the custom floating-point number comprising a sign field (Sign), an exponent field (Exponent), and a mantissa field (Mantissa), wherein a numerical value of the custom floating-point number is determined by bits of the sign field, bits of the exponent field, bits of the mantissa field, and a bias value (Bias), the bias value being determined by the total number of bits of the exponent field; and applying the custom floating-point number to numerical calculations.

[0016] According to a third aspect of the present invention, a floating-point number processing device is provided, comprising a processor, wherein the processor executes: obtaining a custom floating-point number, the custom floating-point number comprising a sign field, an exponent field, and a mantissa field, wherein a value of the custom floating-point number is determined by bits of the sign field, bits of the exponent field, bits of the mantissa field, and a bias value, wherein the bias value is determined by the total number of bits of the exponent field; and applying the custom floating-point number to a numerical calculation.

[0017] In order to better understand the above and other aspects of the present invention, the following embodiments are specifically described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1A A schematic diagram of a neural network (NN) is shown.

[0019] Figure 1B Shows the accuracy of the neural network during the training phase.

[0020] Figures 2A to 2D Shows the weight distribution after successful training using different model sizes and databases.

[0021] Figure 3 The training process of Fault-Tolerant Training (FTT) in an embodiment of the present disclosure is shown.

[0022] Figure 4 The error simulation experiment results of an embodiment of the present disclosure are shown.

[0023] Figure 5 A Fault-Tolerant Floating Point (FTF) format according to another embodiment of the present disclosure is shown.

[0024] Figure 6 FIG. 4 shows a comparison between the FTF16 format of an embodiment of the present disclosure and a conventional custom floating-point format under an exponent bit error.

[0025] Figure 7 The figure shows an error-tolerant floating point (FTF) format according to another embodiment of the present disclosure.

[0026] Figure 8 The figure shows an error-tolerant floating point (FTF) format according to another embodiment of the present disclosure.

[0027] Figure 9 The figure shows an error-tolerant floating point (FTF) format according to another embodiment of the present disclosure.

[0028] Figure 10 An example of a neural network processing system according to an embodiment of the present disclosure is shown.

[0029] Figure 11 This is a schematic diagram of a system architecture provided by an embodiment of the present disclosure.

[0030] Figure 12 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure.

[0031] Figure 13 A floating-point number processing method according to an embodiment of the present disclosure is shown.

[0032] Description of Reference Numerals

[0033] 310-340: Steps

[0034] 1000: Neural Network System

[0035] 1005: Memory device

[0036] 1010: Neural Networks

[0037] 1110, 1120, 1130: Electronic equipment

[0038] 1111: Decoder 1112: Encoder

[0039] 1113: Memory 1114: Computing unit

[0040] 1200: Electronic equipment 1201: Processor

[0041] 1202: Input device 1203: Output device

[0042] 1204: Computer-readable storage medium 1205: Database

[0043] 1206: Storage device

[0044] 1310-1320: Steps DETAILED DESCRIPTION

[0045] The technical terms used in this specification refer to the customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanations or definitions in this specification. Each embodiment of the present disclosure has one or more technical features. Under the premise of possible implementation, technicians in this technical field may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.

[0046] Figure 1A A schematic diagram of a neural network (NN) is shown. Figure 1B Shows the accuracy of the neural network during the training phase.

[0047] The basic structure and training of a neural network (NN) will now be described. A NN is a machine learning model that uses one or more layers of nonlinear units to perform a series of operations on its inputs, ultimately generating predictions based on the final output. In addition to the input and output layers, a neural network also includes one or more hidden layers. The output of each hidden layer serves as the input for the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output based on its input values ​​according to its current parameters (weights).

[0048] A neural network is a function, represented by f w (x i )≡y i The neural network function has a set of trainable weights w that will be used to train the input vector x. i Mapped to the output vector y i .

[0049] As for the neural network structure, the neural network has a layer-by-layer structure, such as Figure 1A As shown. The first layer is the input layer, which contains the input vector x i The various components The last layer is the output layer, which contains the output vector of the neural network. The hidden layer between the input layer and the output layer is the input vector x iThe connecting lines between the input layer and the hidden layer, the connecting lines between the hidden layer and the hidden layer, and the connecting lines between the hidden layer and the output layer represent all the trainable weights w.

[0050] As for the training accuracy of the neural network, Figure 1B As shown in Figure 2. The initial accuracy of the neural network before training is very poor. The goal of training is to find a set of optimal weights w o , so that all outputs y of the neural network i best matches the corresponding input x i Expected answer During the training process, the accuracy of the neural network is gradually improved by adjusting the weights, enabling it to effectively perform the given task.

[0051] This article describes the rapid increase in the cost of training artificial intelligence (AI) models and the development of custom floating-point number formats to accelerate AI applications.

[0052] Because operations in AI models rely on floating-point numbers, a variety of custom floating-point formats have emerged to accelerate AI applications. These custom floating-point formats may be optimized for specific application scenarios to improve computational efficiency. For example, a specialized memory designed specifically for operations in a specific custom floating-point format has been proposed, with a built-in accelerator for executing operations in that specific format. This is intended to more effectively support AI applications.

[0053] However, these methods are not designed for operations with errors. In other words, they may lack the ability to handle errors during operations. Therefore, these methods are not suitable for the context of computing in memory (CIM).

[0054] The basis of the present invention will now be described. Figures 2A to 2D The weight distribution shown is derived from the observation of the center. Figures 2A to 2D Shown are the weight distributions after successful training using different model sizes and databases, where each distribution is normalized for fair comparison.

[0055] Regarding the weight range, such as Figures 2A to 2D As shown, each weight distribution is in the range of (-1, 1), which means that all weights will fall within the range of (-1, 1).

[0056] Regarding the center of the weight distribution, such as Figures 2A to 2DAs shown in Figure 2, the center of each weight distribution is approximately near zero, which means that most weights are concentrated near zero, that is, close to zero. This may imply that most weights are zero, which corresponds to sparsity, which is a common feature in machine learning and neural network models.

[0057] The following describes a new training method inspired by observations of weight distribution in an embodiment of the present disclosure, as well as a proposed solution to errors that may occur in the computation-in-memory (CIM) process.

[0058] According to two observations on the weight distribution, after training, the weights are usually in the range (-1, 1), and most of the weights are almost zero. This suggests that the weights in the model have a certain sparsity and are concentrated in the range close to zero.

[0059] An embodiment of the present disclosure proposes a fault-tolerant training (FTT) method. When using the CIM architecture to further accelerate the training process, errors may occur. To address such errors, an embodiment of the present disclosure proposes a new training procedure called fault-tolerant training (FTT). In the fault-tolerant training of an embodiment of the present disclosure, weights that exceed the normal range are regarded as abnormal weights, and these abnormal weights are set to zero or other reference values ​​throughout the training process.

[0060] Experimental results show that in some cases, setting the outlier weights to zero does not harm the accuracy, and sometimes even improves it.

[0061] In addition, another embodiment of the present disclosure proposes a new custom floating point format: Fault-Tolerant Floating point (FTF) format.

[0062] Although existing custom floating-point formats can accelerate AI operations, these custom floating-point formats are usually designed to cover a wide range of values. If errors occur in the exponent bits, they may have a serious impact on the values.

[0063] Therefore, an embodiment of the present disclosure further proposes a new custom floating-point format, called the FTF format. The FTF format has a smaller numerical range (possibly between (-1, 1) or other small numerical ranges), thus making more efficient use of each bit and significantly reducing the impact of bit errors.

[0064] Fault-Tolerant Training (FTT) Example:

[0065] Figure 3The training process of the fault-tolerant training (FTT) of the embodiment of the present disclosure is shown. The fault-tolerant training (FTT) of the embodiment of the present disclosure involves modification of the training process, and adjustments are made to the existing training process to handle possible abnormal situations. Figure 3 The training process of the FTT can be executed by hardware or software. For example, the training process of the FTT can be executed by a computer system or a floating point computing chip.

[0066] In step 310 , weight training iterations are started.

[0067] In step 320, a determination is made as to whether any abnormal weights are found. The purpose of step 320 is to identify errors or abnormal weights that may occur during the CIM framework's weight training process (e.g., but not limited to, weights of a neural network model). When using the CIM framework, hardware defects (e.g., memory chip defects) may cause errors (e.g., bit errors) during the neural network weight training process.

[0068] If any abnormal weights are found in step 320, they are set as reference values ​​in step 330. This process is intended to eliminate abnormal weights that may have a negative impact on the model to maintain the stability and accuracy of the training.

[0069] If no abnormal weights are found in step 320 , the next training iteration is performed in step 340 .

[0070] During a training iteration, at least one of a pre-update weight or a post-update weight may be checked for anomalies in step 320. Step 320 helps identify whether anomalies were introduced during the weight update process.

[0071] As for the definition of abnormal weights, in one embodiment of the present disclosure, there can be a variety of abnormal weight definitions. For example, but not limited to, if the weight exceeds the range of the normal weight median plus or minus 1 or more standard deviations (standard deviation), the weight is judged to be an abnormal weight. For example, but not limited to, the normal weight median is -0.0072, and the standard deviation of the normal weight is 0.14748, then the range of the normal weight median (-0.0072) plus or minus 3 normal weight standard deviations (0.1478) is: -0.0072±3*0.1478, that is, between -0.4506 and 0.4362. Therefore, if the weight before or after the update exceeds -0.4506 to 0.4362, it will be judged as an abnormal weight. The normal weight median is the median of all normal weights.

[0072] Alternatively, if the weight exceeds the numerical range, the weight is determined to be an abnormal weight. The numerical range may be, but is not limited to, (-1, 1). In one embodiment of the present disclosure, the numerical range depends on a custom floating-point format. The relationship between the numerical range and the custom floating-point format will be described below.

[0073] In step 330, there are various ways to set these abnormal weights as reference values. For example, but not limited to, the median of the normal weights may be used as the reference value (i.e., the abnormal weights are set to the "median of the normal weights"). Alternatively, the abnormal weights may be set to zero (i.e., 0 is used as the reference value).

[0074] In summary, in one embodiment of the present disclosure, these methods and options provide flexibility in handling abnormal weights during training to ensure the stability and accuracy of the model.

[0075] In one embodiment of the present disclosure, an experiment was conducted to determine whether fault-tolerant training could provide assistance. The experiment used the RestNet50 model and the standard Cifar10 database for training. During this process, weight errors were manually introduced, and these errors were defined using two metrics: error severity and error count.

[0076] Amount of Error: Fix 10%, 1%, or 0.1% of the weights to a certain value. This simulates introducing different amounts of weight error during training.

[0077] Error severity: Error severity is defined as the error-free, standard deviation of all weights (σ) = 0.14 after training, with the values ​​of the selected weights fixed at 0, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, and 1024. The error severity setting measures the impact of an introduced error on the model and the corresponding weight change.

[0078] Figure 4 The error simulation experiment results of an embodiment of the present disclosure are shown. Figure 4 In order to present the horizontal axis in a logarithmic scale, the error severity on the graph needs to be subtracted by 1 to get the actual error severity, but this does not affect the inductive results. Figure 4 The error simulation experiment results can be summarized as follows.

[0079] For data of a given accuracy (points on the horizontal line): data with higher severity errors also have lower error amounts. In short, the fewer errors a data has, the more severe errors it can tolerate, while maintaining the same accuracy.

[0080] For data of a given error severity (points on the vertical line): data with higher accuracy also have lower error amounts. In short, under the premise of maintaining the same error severity, more errors mean lower accuracy.

[0081] For data with a given amount of error (such as the points on the dashed line labeled "0.1% error"): as the severity of the error decreases, the accuracy tends to increase.

[0082] In other words, when an error occurs, by setting the error severity to the minimum value (setting the weight to zero or a reference value), the accuracy can be improved and the impact of the error on the model can be repaired.

[0083] In general, these summaries demonstrate the correlation between error amount, severity, and accuracy under different conditions, and indicate that when errors occur, the performance of the model can be maintained through fault-tolerant training (FTT) according to an embodiment of the present disclosure.

[0084] As can be seen from the above, the fault-tolerant training (FTT) of an embodiment of the present disclosure can effectively reduce errors and repair the unreliable problem of computing in memory, thereby enabling computing to be performed in memory (CIM) and accelerating artificial intelligence computing.

[0085] Custom floating point format: Fault-tolerant floating point (FTF) format

[0086] A new custom floating point format disclosed in another embodiment of the present disclosure, called the Fault Tolerant Floating Point (FTF) format, will be described below.

[0087] Figure 5 Another embodiment of the present disclosure is a fault-tolerant floating-point number (FTF) format. A 16-bit version of the FTF format (referred to as FTF16) is used as an example for illustration, but the present disclosure is not limited thereto.

[0088] The 16-bit version of the fault-tolerant floating-point (FTF) format includes: a sign field (including 1 sign bit s), an exponent field (including 8 exponent bits e8-e1) and a mantissa field (including 7 mantissa bits f7-f1).

[0089] The 8 exponent bits are expressed in decimal as E (also called the value of the exponent field), then E is as follows (1):

[0090]

[0091] The 7 mantissa digits are expressed in decimal as M (also called the value of the mantissa field), then M is as follows (2):

[0092]

[0093] The actual value of the FTF16 format can be calculated in the following manner.

[0094] When E = 0 and M = 0, FTF16 = (-1) s 0, where when the sign s = 0, FTF16 = (-1) 0 0 = +0, and when the sign s = 1, FTF16 = (-1) 1 0 = -0. Therefore, when E = 0 and M = 0, FTF16 = (-1) s 0 is called FTF16 = plus or minus 0 (±0). E = 0 means that all exponent bits in the exponent field are 0. M = 0 means that all mantissa bits in the mantissa field are 0.

[0095] When E = 0 and M > 0, FTF16 is called a sub - normal number.

[0096] When 0 < E < 255 (regardless of the value of M), FTF16 is called a normal number.

[0097] When E = 255 and M = 0, FTF16 = (-1) s ∞, which is called FTF16 = plus or minus infinity (±infinity).

[0098] When E = 255 and M > 0, it is called FTF16 = not a number (NaN).

[0099] Summarize the above as:

[0100]

[0101] Therefore, Figure 5 the value range of FTF16 in

[0102] (generally referring to the range of normal numbers) is [-0.996, 0.996].

[0103] Next, the advantages of the FTF16 - format floating - point number of an embodiment of the present disclosure over the conventional custom floating - point number format will be described. The advantages mainly focus on the limitation of the value range and the impact on the bit error in the exponent field.

[0104] Regarding the impact of bit errors in the exponent field, Figure 6 The comparison between the FTF16 format floating point number of an embodiment of the present disclosure and the custom floating point number format of the prior art is shown when the exponent bit is wrong. When the exponent bit is correct, the FTF16 format of an embodiment of the present disclosure is 3.84×10 -34 , while the custom floating-point number format of the existing technology is 130560.

[0105] When the penultimate bit of the exponent is wrong, the FTF16 format value of an embodiment of the present disclosure becomes 1.13×10 -72 , while the value of the custom floating point format of the existing technology becomes 3.84×10 -34 In addition, when the second to last bit of the exponent is wrong, the FTF16 format value of an embodiment of the present disclosure becomes 7.08×10 -15 , while the value of the custom floating point format of the existing technology becomes 2.41×10 24 .

[0106] Therefore, by Figure 6 As can be seen from the above, regarding the impact of bit errors in the exponent, in the custom floating-point number format of the prior art, when an exponent bit error occurs, the value of the custom floating-point number format of the prior art changes greatly (130560 becomes 3.84×10 -34 and 2.41×10 24 In contrast, when an exponent bit error occurs, the value change of the FTF16 format in one embodiment of the present disclosure is very small (3.84×10 -34 becomes 1.13×10 -72 and 7.08×10 -15 ).

[0107] In general, since the numerical range of the FTF16 format of an embodiment of the present disclosure is smaller, the sensitivity to exponent bit errors is also lower, which makes the FTF16 format of an embodiment of the present disclosure more reliable in AI application scenarios, especially when higher accuracy and stability are required.

[0108] In one embodiment of the present disclosure, the weights of the neural network model may be in FTF format to reduce the impact of errors, enable calculation in memory (CIM), and thereby increase the calculation speed.

[0109] In another embodiment of the present disclosure, when the total number of digits of the exponent and the mantissa is fixed, the number of digits of the exponent and the mantissa can be adjusted, that is, 1+X+Y=16, which is also within the scope of the present disclosure.

[0110] Furthermore, in other possible embodiments of the present disclosure, FTF16 can be extended to other total number of bits, for example, 1+X+Y=8 can form an FTF8 format.

[0111] Figure 7 The figure shows an error-tolerant floating point (FTF) format according to another embodiment of the present disclosure.

[0112] The multiple bits of the FTF format include: sign field (including 1 sign bit s), exponent field (including X exponent bits e X -e1) and the mantissa field (including Y mantissa digits f Y -f1).

[0113] The X exponent bits are expressed as E in decimal, then E is as follows (4):

[0114]

[0115] The Y mantissa digits are expressed as M in decimal, and M is as follows (5):

[0116]

[0117] Therefore, the FTF values ​​are as follows.

[0118] When E=0 and M=0, FTF=(-1) s 0, where, when symbol s=0, FTF16=(-1) 0 0 = +0, and when symbol s = 1, FTF = (-1) 1 0 = -0. Therefore, when E = 0 and M = 0, FTF = (-1) s 0 is called FTF = plus or minus 0 (±0).

[0119] When E=0 and M>0, The FTF is called a sub-normal number. The bias value b = 2 X -1.

[0120] When 0 <E<2 X -1 (regardless of the value of M), The FTF is called a normal number.

[0121] When E=2 X -1 and M = 0, FTF = (-1) s ∞, called FTF is positive and negative infinity (±infinity).

[0122] When E=2 XWhen -1 and M>0, the FTF is called not a number (NaN).

[0123] To summarize the above:

[0124]

[0125] In one embodiment of the present disclosure, FTT and FTF are used to train neural network models on memory devices with unavoidable errors (e.g., NAND flash memory). FTT and FTF in one embodiment of the present disclosure can also be applied to situations with other error sources, such as errors in high / low temperature environments and errors caused by manufacturing defects.

[0126] Furthermore, although in the above embodiment, FTF16 is designed for 16-bit numbers within the range (-1, 1), such as weights in a neural network model, in other possible embodiments of the present disclosure, the bias value can be changed in the general formula to represent values ​​within different ranges. As for the range of the bias value, as long as the ratio of the total number of numbers with an absolute value less than 1 to the total number of numbers with an absolute value greater than 1 remains below 2, it can be considered a good range of the bias value.

[0127] Figure 8 The error-tolerant floating point format (FTF) of another embodiment of the present disclosure is shown. The multiple bits of the error-tolerant floating point format (FTF) include: a sign field (including 1 sign bit s), an exponent field (including X exponent bits e X -e1) and the mantissa field (including Y mantissa digits f Y -f1). The range of the deviation value b is: Among them, Round is the rounding function. That is, the range of the deviation value b is 2 X -1 and The rounded result of Determined.

[0128] so, Figure 8 The FTF values ​​are shown below.

[0129] When E=0 and M=0, FTF=(-1) s 0, where, when symbol s=0, FTF16=(-1) 0 0 = +0, and when symbol s = 1, FTF = (-1) 1 0 = -0. Therefore, when E = 0 and M = 0, FTF = (-1) s 0 is called FTF = plus or minus 0 (±0).

[0130] When E=0 and M>0, It is called that FTF is a sub-normal number.

[0131] When 0 < E < 2 X -1 (regardless of the value of M), It is called that FTF is a normal number.

[0132] When E = 2 X -1 and M = 0, FTF = (-1) s ∞, it is called that FTF is ±infinity.

[0133] When E = 2 X -1 and M > 0, it is called that FTF is a non-a-number (NaN).

[0134] Summarize the above as:

[0135]

[0136] To illustrate the influence of the setting of the bias value b on the numerical range of FTF, please refer to Figure 9 . Figure 9 Displays the fault-tolerant floating-point (FTF) format of another embodiment of the present disclosure. The multiple bits of the fault-tolerant floating-point (FTF) format include: a sign field (including 1 sign bit s), an exponent field (including 8 exponent bits e8 - e1), and a mantissa field (including 7 mantissa bits f7 - f1). The range of the bias value b is:

[0137]

[0138] Therefore, Figure 9 The numerical value of FTF is as follows.

[0139] When E = 0 and M = 0, FTF = (-1) s 0, where, when the sign s = 0, FTF16 = (-1) 0 0 = +0, and when the sign s = 1, FTF = (-1) 1 0 = -0. Therefore, when E = 0 and M = 0, FTF = (-1) s 0 is called FTF = ±0.

[0140] When E = 0 and M > 0, It is called that FTF is a sub-normal number.

[0141] When 0 < E < 255 (regardless of the value of M), It is called that FTF is a normal number.

[0142] When E=255 and M=0, FTF=(-1) s ∞, called FTF is positive and negative infinity (±infinity).

[0143] When E=255 and M>0, the FTF is called not a number (NaN).

[0144] To summarize the above:

[0145]

[0146]

[0147] The FTF value range is ±2 254-b (1+127 / 128).

[0148] When the bias value b = 255, the FTF value range is [-0.996, 0.996], which is a small range. When the bias value b = 254, the FTF value range is [-1.992, 1.992], which is a medium range. When the bias value b = 253, the FTF value range is [-3.984, 3.984], which is a large range. In other words, the larger the bias value, the smaller the FTF value range, and vice versa.

[0149] Figure 10 An example of a neural network processing system 1000 according to an embodiment of the present disclosure is shown. The neural network processing system 1000 is an example of a system implemented as a computer program on one or more computers in one or more locations.

[0150] Neural network system 1000 includes one or more memory devices 1005, which store neural network 1010. Neural network 1010 includes one or more neural network models. When training the one or more neural network models, the one or more memory devices 1005 execute the aforementioned FTT (Fault Tolerant Training) method. Furthermore, neural network 1010 may also be compatible with the FTF (Fault Tolerant Floating Point Format) of another embodiment of the present disclosure. For example, the input values, weight values, and output values ​​of neural network 1010 are represented in the FTF (Fault Tolerant Floating Point Format).

[0151] The neural network processing system 1000 is a processing system that uses floating point calculations to perform neural network operations.

[0152] Floating point calculations refer to calculations performed using floating point data types. Neural network 1010 is an example of a neural network that can be configured to receive any type of numeric data input and produce any type of score or classification output based on the numeric data input.

[0153] Neural network 1010 includes multiple neural network layers, including one or more input layers, one output layer, and one or more hidden layers between the input and output layers. Each neural network layer includes one or more neural network nodes. Each neural network node has one or more weight values. Each node processes a series of input values ​​using corresponding weight values ​​and performs an operation on the processed results to generate an output value.

[0154] In some implementations, each node in each input layer of neural network 1010 receives a set of floating-point input values. Output values ​​are numerical values ​​generated by the output layer nodes of neural network 1010 when processing the network inputs.

[0155] The neural network processing system 1000 may store the generated neural network output in an output data repository, or provide the neural network output for other purposes, such as display on a user device or further processing by another system.

[0156] See also Figure 11 , is a schematic diagram of a system architecture provided by an embodiment of the present disclosure. The technical method of the above embodiment of the present disclosure can be Figure 11 The system architecture shown in the example or a similar system architecture is specifically implemented. Figure 11 As shown, the system architecture may include multiple electronic devices, such as electronic device 1110, electronic device 1120, and electronic device 1130. Electronic devices 1110, 1120, and 1130 may establish communication connections via wired or wireless networks (such as WiFi, Bluetooth, and mobile networks) to perform floating-point data storage, calculation, and transmission in various fields (finance, engineering, scientific research, aerospace, etc.).

[0157] like Figure 11As shown, taking electronic device 1110 as an example, electronic device 1110 may include a decoder 1111 and an encoder 1112 for floating-point number processing, a memory 1113, and corresponding multiple computing units 1114 (such as computing unit 1, computing unit 2, computing unit 3 ... computing unit N), etc. When electronic device 1110 performs general computing, high-performance computing or AI training, a large amount of floating-point type data is needed. At this time, the electronic device can obtain the corresponding floating-point number through decoder 1111 based on a floating-point number processing method provided in an embodiment of the present disclosure (the floating-point number can be obtained from the local memory 1113, or from the electronic device 1120 or the electronic device 1130 through a wired or wireless network), and transmit the floating-point number data to the computing unit, and complete the corresponding calculation through the computing unit. Accordingly, the operation result finally obtained by the computing unit can also be encoded into a floating-point number through encoder 1112, and the floating-point number can be used for data storage and data transfer. The embodiments of the present disclosure can flexibly meet different requirements for the numerical range and numerical precision of floating-point numbers in various scenarios (such as general computing, high-performance computing, or AI training, etc.) without increasing the total number of bits, that is, without increasing the cost of data storage or data transfer, thereby improving the use effect of floating-point numbers.

[0158] Figure 11 The structures and functions of the electronic device 1120 and the electronic device 1130 may specifically refer to the electronic device 1110. In some possible embodiments, the electronic device 1110, the electronic device 1120 and the electronic device 1130 may include the following: Figure 11 There are more or fewer elements shown, and the embodiments of the present disclosure do not specifically limit this.

[0159] In summary, electronic device 1110, electronic device 1120 and electronic device 1130 can be smart wearable devices, smart phones, smart home appliances, tablet computers, notebook computers, desktop computers, vehicle-mounted computers or servers with the above functions, etc., among which it can be a server, a server cluster composed of multiple servers, or a cloud computing service center, etc., and the embodiments of the present disclosure do not make specific limitations on this.

[0160] Based on the description of the above method and device embodiments, an embodiment of the present disclosure further provides an electronic device. Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 12As shown, the electronic device 1200 includes at least a processor 1201, an input device 1202, an output device 1203, and a storage device 1206. The storage device 1206 includes a computer-readable storage medium 1204 and a database 1205. The electronic device 1200 may also include other common components, which will not be described in detail here. The processor 1201, input device 1202, output device 1203, and computer-readable storage medium 1204 in the electronic device 1200 may be connected via a bus or other means. Figure 12 The electronic device 1200 can be used to implement Figure 11 electronic device 1110, electronic device 1120 and electronic device 1130.

[0161] The processor 1201 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the above embodiments. The processor 1201 may be used to execute the fault-tolerant training (FTT) method of the above embodiments of the present disclosure. In addition, the processor 1201 may be compatible with the fault-tolerant floating-point number (FTF) of the above embodiments of the present disclosure. In addition, the processor 1201 may be used to execute the floating-point number processing method of the embodiments of the present disclosure.

[0162] The memory in the electronic device 1200 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage, magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor via a bus. The memory can also be integrated with the processor.

[0163] The computer readable storage medium 1204 may store a computer program including program instructions. When the processor 1201 executes the program instructions stored in the computer readable storage medium 1204 , the processor 1201 may perform any part or all of any steps described in any embodiment of the present disclosure.

[0164] The embodiments of the present disclosure also provide a computer-readable storage medium, wherein the computer-readable storage medium can store a program, and when the program is executed by a processor, the processor can perform part or all of the steps of any one of the above embodiments.

[0165] An embodiment of the present disclosure further provides a computer program, which includes instructions. When the computer program is executed by a multi-core processor, the processor can execute part or all of the steps of any one of the above embodiments.

[0166] Figure 13 A floating-point number processing method according to an embodiment of the present disclosure is shown, which is applied to an electronic device. The floating-point number processing method includes: (1310) obtaining a custom floating-point number, the custom floating-point number including a sign field, an exponent field, and a mantissa field, wherein a value of the custom floating-point number is determined by bits of the sign field, bits of the exponent field, bits of the mantissa field, and a bias value, wherein the bias value is determined by the total number of bits of the exponent field; and (1320) applying the custom floating-point number to numerical calculation.

[0167] Another embodiment of the present disclosure further discloses a floating-point operation method, comprising: receiving a requirement to utilize a neural network to perform floating-point operations, wherein the neural network includes a plurality of weights, and the weights have a customized fault-tolerant floating-point format (FTF); and receiving a neural network input, wherein the neural network utilizes the neural network and the weights to obtain a neural network output.

[0168] In summary, FTT (Fault Tolerant Training) in one embodiment of the present disclosure is a training process designed for training neural network models in the presence of errors in the environment. The core concept is to train in the presence of errors without losing accuracy. This means that FTT in one embodiment of the present disclosure is intended to enable neural network models to maintain robustness and performance in the presence of unavoidable environmental errors.

[0169] Another embodiment of the present disclosure, the FTF (Fault Tolerant Floating Point Format), can be applied to the weights of a neural network model to restrict the numerical range to a limited range of weights. Its core concept is to efficiently utilize each bit of the weight while reducing the impact of exponent bit errors. The goal of the FTF (Fault Tolerant Floating Point Format), another embodiment of the present disclosure, is to reduce numerical uncertainty caused by memory errors while maintaining performance.

[0170] The above-mentioned embodiments of the present disclosure may be applied to memory devices that support special multiply-accumulate (MAC) operations with FTF format, including but not limited to DRAM, NVM, etc.

[0171] While this disclosure may describe many specific details, these should not be construed as limitations on the scope of the claimed invention, but rather as descriptions of features of particular embodiments. Certain features described in the context of a single embodiment in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any appropriate subcombination. Furthermore, while a feature may initially be described as functioning in certain combinations, or even initially described as such a combination, in some cases one or more features may be deleted from the combination, and the described combination may be for a subcombination or variation of a subcombination. Similarly, while operations are depicted in the accompanying drawings as occurring in a particular order, this should not be construed as requiring that the operations must be performed in the particular order or sequence shown, or that all depicted operations must be performed in order to achieve the desired result.

[0172] Although the above embodiments of the present disclosure only disclose some examples and implementations, based on the disclosed content, the examples and implementations and other implementations may be changed, modified, and enhanced.

[0173] In summary, although the present invention has been disclosed above with reference to the embodiments, these are not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A neural network system comprising one or more memory devices storing one or more neural network models, wherein when training the one or more neural network models, the one or more memory devices execute: Start weight training iteration; Determine whether an abnormal weight is found; When the abnormal weight is found, the abnormal weight is set as a reference value; as well as When no abnormal weights are found, the next training iteration is performed.

2. The neural network system according to claim 1, wherein: When determining whether the abnormal weight is found, it is checked whether at least one of a pre-update weight or a post-update weight is abnormal.

3. The neural network system according to claim 1, wherein: When a weight exceeds a numerical range, the weight is determined to be abnormal, and the numerical range depends on a custom floating point format.

4. The neural network system according to claim 1, wherein: When a weight exceeds a numerical range, the weight is determined to be abnormal, and the numerical range depends on a custom floating point format.

5. The neural network system according to claim 1, wherein: The reference value is a median of multiple normal weights.

6. The neural network system according to claim 1, wherein: The reference value is zero.

7. A floating-point number processing method, applied to an electronic device, the floating-point number processing method comprising: Obtaining a custom floating-point number, the custom floating-point number comprising a sign field, an exponent field, and a mantissa field, wherein a value of the custom floating-point number is determined by bits of the sign field, bits of the exponent field, bits of the mantissa field, and a bias value, the bias value being determined by a total number of bits in the exponent field; and Use this custom floating-point number in numerical calculations.

8. The method for processing floating-point numbers according to claim 7, wherein: The exponent field includes X exponent bits, and the mantissa field includes Y mantissa bits, where both X and Y are positive integers; When an exponent field value E of the exponent field and a mantissa field value M of the mantissa field are both 0, the value of the custom floating-point number is positive or negative 0; When the exponent field value E of the exponent field is 0 and the mantissa field value M of the mantissa field is greater than 0, the value of the custom floating-point number is: The value of the custom floating-point number is a denormal number, wherein b is the bias value, and s represents a sign bit of the sign field; When 0 <E<2 X When -1, the value of the custom floating point number is: The value of the custom floating-point number is a normal number; When E=2 X When -1 and M=0, the value of the custom floating-point number is infinitely positive or negative; and When E=2 X When -1 and M>0, the value of the custom floating-point number is not a number.

9. The method for processing floating-point numbers according to claim 8, wherein: The relationship between the deviation value and the total number of bits in the exponent field is: b = 2 X -1.

10. The floating point number processing method according to claim 8, wherein: The relationship between the bias value and the total number of bits in the exponent field is: Round is a rounding function.

11. The method for processing floating-point numbers according to claim 8, wherein: When the deviation value is larger, a numerical range of the custom floating-point number is smaller.

12. A floating point number processing device, comprising a processor, the processor executing: Get a custom floating-point number, which includes a sign field, an exponent field, and a mantissa field, wherein: A value of the custom floating-point number is determined by bits of the sign field, bits of the exponent field, bits of the mantissa field, and a bias value, wherein the bias value is determined by the total number of bits of the exponent field; and Use this custom floating-point number in numerical calculations.

13. The floating point number processing device according to claim 12, wherein: The exponent field includes X exponent bits, and the mantissa field includes Y mantissa bits, where both X and Y are positive integers; When an exponent field value E of the exponent field and a mantissa field value M of the mantissa field are both 0, the value of the custom floating-point number is positive or negative 0; When the exponent field value E of the exponent field is 0 and the mantissa field value M of the mantissa field is greater than 0, the value of the custom floating-point number is: The value of the custom floating-point number is a denormal number, wherein b is the bias value, and s represents a sign bit of the sign field; When 0 <E<2 X When -1, the value of the custom floating point number is: The value of the custom floating-point number is a normal number; When E=2 X When -1 and M=0, the value of the custom floating-point number is infinitely positive or negative; and When E=2 X When -1 and M>0, the value of the custom floating-point number is not a number.

14. The floating point number processing device according to claim 13, wherein: The relationship between the deviation value and the total number of bits in the exponent field is: b = 2 X -1.

15. The floating point number processing device according to claim 13, wherein: The relationship between the bias value and the total number of bits in the exponent field is: Round is a rounding function.

16. The floating point number processing device according to claim 13, wherein: When the deviation value is larger, a numerical range of the custom floating-point number is smaller.