A quantization training method, device, and equipment for a neural network model
By quantizing the input data and parameters of the quantization layer during the training of the neural network model, and freeing memory after obtaining the output data, the problem of excessive memory consumption in deep neural network model training is solved, and a more efficient training process is achieved.
Patent Information
- Application Number
- CN202011645236.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-12-31
AI Technical Summary
As the depth of the neural network model increases, the number of parameters in the model increases, resulting in an increase in the computing and storage capabilities requirements of hardware devices during the training process, especially the memory consumption is too large, making it difficult to complete the training.
During the training process of neural network model, the input data and parameters of the quantization layer are quantized during the forward propagation and backpropagation process, and the amount of calculation is reduced, and the quantized data is released from memory after obtaining the output data, reducing memory consumption.
Through quantitative training methods, the memory usage of data during neural network model training is reduced, memory consumption is reduced, and training efficiency is improved, especially on hardware devices with limited memory can be trained normally.
Smart Images

Figure CN114692824B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a quantization training method, apparatus, and device for a neural network model. Background Art
[0002] A neural network model needs to go through multiple iterative trainings to converge. In each iterative training, it needs to go through three processes: forward propagation, backward propagation, and parameter update. First, each layer in the neural network processes the sample data to obtain the network loss function; then, the gradients of each layer are calculated based on the network loss function; secondly, the parameters of each layer are updated based on the gradients of each layer.
[0003] As the depth of the neural network model becomes deeper and deeper, the number of parameters (weights, biases, etc.) in the neural network model also increases, and the training of the neural network model places higher and higher requirements on the computing and storage capabilities of hardware devices. In a neural network model, there are a large number of numerical operations, and the large amount of generated data consumes a lot of memory. Especially for hardware devices with small memory, it is even difficult to complete the training of the neural network model. Summary of the Invention
[0004] Embodiments of this application disclose a quantization training method, apparatus, and device for a neural network model. The method helps to reduce the memory occupation of the data generated during the training process of the neural network model and reduce memory consumption.
[0005] In a first aspect, embodiments of this application provide a quantization training method for a neural network model. The neural network model includes multiple layers. The method includes: during the forward propagation process, quantizing the first input data of the layer to be quantized and the parameters of the layer to be quantized to obtain quantized first input data and quantized first parameters; performing an operation on the quantized first input data and the quantized first parameters to obtain quantized first output data; performing an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized; and releasing the quantized first input data and the quantized first parameters from the memory.
[0006] It can be seen that during the training process of the neural network model, when forward propagating to the layer to be quantized, the first input data and parameters of the layer to be quantized are quantized, and then the first output data of the layer to be quantized is obtained by performing an operation and inverse quantization on the quantized first input data and quantized first parameters, and then the quantized first input data and quantized first parameters are released. Among them, the parameters include weights and may also include biases, etc.; the first input data, parameters, and first output data may be floating-point data, and the quantized first input data, quantized first parameters, and quantized first output data may be fixed-point integer data.
[0007] In this embodiment, the input data and parameters are first quantized and then the quantized data is calculated, reducing the amount of computation. After the first output data is obtained in the layer to be quantized, the first input data of the intermediate data quantization and the quantized first parameters are released, that is, deleted from the memory, thereby reducing the consumption of the intermediate data for the memory and saving memory space. Based on the first aspect, in a possible implementation, when the layer to be quantized is the first layer of the neural network model, the first input data of the layer to be quantized is the sample data input to the neural network model; otherwise, the first input data of the layer to be quantized is the first output data of the previous layer of the layer to be quantized.
[0008] Based on the first aspect, in a possible implementation, the method further includes: during backpropagation, quantizing the second input data of the layer to be quantized and the parameters of the layer to be quantized respectively to obtain quantized second input data and quantized second parameters; obtaining the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized according to the quantized second input data and the quantized second parameters; performing an inverse quantization operation on the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized to obtain the second output data of the layer to be quantized and the gradient of the second parameters of the layer to be quantized; and releasing the quantized second input data and the quantized second parameters from the memory.
[0009] It can be seen that during backpropagation to the layer to be quantized, the second input data and parameters of the layer to be quantized are quantized respectively, and then through operations and inverse quantization, the second output data and the gradient of the parameters are obtained, and then the quantized second input data and the quantized second parameters are released. Among them, the quantized second input data, the quantized second parameters, the quantized second output data, etc. can be fixed-point integer data, and the second input data, the parameters, and the second output data can be floating-point data.
[0010] In this embodiment, the second input data and parameters are quantized, and the quantized data is operated on to reduce the amount of computation. After the second output data of the layer to be quantized is obtained, the second input data of the intermediate data quantization and the quantized parameters are released from the memory, that is, deleted, reducing the consumption of data for the memory and saving memory.
[0011] Based on the first aspect, in a possible implementation, when the layer to be quantized is the last layer of the neural network model, the second input data includes the first input data and the network loss; where the network loss is obtained by forward propagation; when the layer to be quantized is not the last layer of the neural network model, the second input data includes the first input data and the second output data of the next layer of the layer to be quantized.
[0012] Based on the first aspect, in a possible implementation, after the backpropagation ends, the method further includes: updating the parameters of each layer based on the parameters of each layer in the neural network model and the gradients of the parameters.
[0013] Based on the first aspect, in a possible implementation, the method is applicable to at least one layer in the neural network model.
[0014] It can be understood that the above quantization training method can be applicable to one or more layers of the neural network model. For example, it can be applicable to layers with relatively large computational amounts in the neural network model, such as convolutional layers, deconvolutional layers, fully connected layers, etc. or layers with matrix multiplication operations, etc.
[0015] In a second aspect, an embodiment of the present application provides a method for testing a neural network model, including obtaining test data; using the trained neural network model to test the test data; the neural network model is trained by the method in the above first aspect or any implementation manner of the first aspect.
[0016] In a third aspect, an embodiment of the present application provides a quantization training device for a neural network model. The neural network model includes multiple layers. The device includes: a quantization unit for quantifying the first input data of the layer to be quantized and the parameters of the layer to be quantized respectively during the forward propagation process to obtain the quantized first input data and the quantized first parameters; an operation unit for operating on the quantized input data and the quantized first parameters to obtain the quantized first output data of the layer to be quantized; an inverse quantization unit for performing an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized; a release unit for releasing the quantized first input data and the quantized first parameters from the memory.
[0017] Based on the third aspect, in a possible implementation, when the layer to be quantized is the first layer of the neural network model, the first input data of the layer to be quantized is the sample data input to the neural network model; otherwise, the first input data of the layer to be quantized is the first output data of the previous layer of the layer to be quantized.
[0018] Based on the third aspect, in a possible implementation, the quantization unit is further configured to, during the backpropagation process, quantize the second input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized second input data and quantized second parameters; the operation unit is further configured to, according to the quantized second input data and the quantized second parameters, obtain the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized; the dequantization unit is further configured to perform a dequantization operation on the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized, to obtain the second output data of the layer to be quantized and the gradient of the second parameters of the layer to be quantized; the release unit is further configured to release the quantized second input data and the quantized second parameters from the memory.
[0019] Based on the third aspect, in a possible implementation, when the layer to be quantized is the last layer of the neural network model, the second input data includes the first input data and the network loss; where the network loss is obtained through forward propagation; when the layer to be quantized is not the last layer of the neural network model, the second input data includes the first input data and the second output data of the next layer of the layer to be quantized.
[0020] Based on the third aspect, in a possible implementation, the apparatus further includes: a parameter update unit, configured to update the parameters of each layer based on the parameters of each layer in the neural network model and the gradients of the parameters.
[0021] Each functional unit in the apparatus of the above third aspect is configured to implement the method described in the above first aspect or any implementation manner of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a testing apparatus for a neural network model, including: an acquisition unit, configured to acquire test data; a testing unit, configured to test the test data using a trained neural network model; the neural network model is trained by the method described in the above first aspect or any implementation manner of the first aspect.
[0023] Each functional unit in the apparatus of the fourth aspect is configured to implement the method described in the above second aspect.
[0024] In a fifth aspect, an embodiment of the present application provides a quantization training device for a neural network model, including a memory and a processor, the memory is configured to store instructions, and the processor is configured to call the instructions to execute the method described in the above first aspect or any implementation manner of the first aspect.
[0025] In a sixth aspect, an embodiment of the present application provides a testing device for a neural network model, including a memory and a processor. The memory is used to store instructions, and the processor is used to call the instructions to execute the method described in the second aspect above.
[0026] In a seventh aspect, an embodiment of the present application provides a non-volatile storage medium for storing program instructions. When the program instructions are applied to a training device of a neural network model, they can be used to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0027] In an eighth aspect, an embodiment of the present application provides a non-volatile storage medium for storing program instructions. When the program instructions are applied to a testing device of a neural network model, they can be used to implement the method described in the second aspect or any possible implementation manner of the second aspect.
[0028] In a ninth aspect, an embodiment of the present application provides a computer program product. The computer program product includes program instructions. When the computer program product is executed by a quantization training device of a neural network model, the quantization testing device of the neural network model executes the method described in the first aspect above. The computer program product can be a software installation package. In the case where it is necessary to use the method provided by any possible design of the first aspect, the computer program product can be downloaded and executed on the quantization training device of the neural network model to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0029] In a tenth aspect, an embodiment of the present application provides a computer program product. The computer program product includes program instructions. When the computer program product is executed by a testing device of a neural network model, the testing device of the neural network model executes the method described in the second aspect above. The computer program product can be a software installation package. In the case where it is necessary to use the method provided by any possible design of the second aspect, the computer program product can be downloaded and executed on the testing device of the neural network model to implement the method described in the second aspect or any possible implementation manner of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 Schematic diagram of a quantization training device for a neural network model provided by an embodiment of the present application;
[0032] Figure 2 Schematic diagram of a combined device provided by an embodiment of the present application;
[0033] Figure 3 Schematic diagram of a board structure provided by an embodiment of the present application;
[0034] Figure 4 In (a) and (b) are respectively schematic diagrams of an example of forward propagation and backward propagation;
[0035] Figure 5 Schematic flowchart of a quantization training method for a neural network model provided by an embodiment of the present application;
[0036] Figure 6 Schematic diagram of forward propagation and backward propagation provided by an embodiment of the present application;
[0037] Figure 7 Schematic flowchart of another quantization training method for a neural network model provided by an embodiment of the present application;
[0038] Figure 8 Schematic flowchart of a testing method for a neural network model provided by an embodiment of the present application;
[0039] Figure 9 Schematic diagram of a testing device for a neural network model provided by an embodiment of the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0041] It should be noted that the terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0042] It should be noted that when used in this specification and the appended claims, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a system, product, or device that includes a series of units / devices is not limited to the listed units / devices, but may optionally further include units / devices not listed, or may further optionally include other units / devices inherent to these products or devices.
[0043] It should also be understood that the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" or "in the case of" depending on the context.
[0044] It should be noted that the first and second in this application are only used to distinguish the objects in the forward propagation process and the backward propagation process, rather than for describing a specific order. The first input data and the first output data respectively correspond to the input data and the output data during forward propagation, and the second input data and the second output data respectively correspond to the input data and the output data during backward propagation.
[0045] See Figure 1 , Figure 1 which is a schematic diagram of a quantization training device 100 for a neural network provided by an embodiment of this application. The neural network model includes multiple layers, and the device 100 includes:
[0046] A quantization unit 101, configured to, during the forward propagation process, quantize the first input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized first input data and quantized first parameters;
[0047] An operation unit 102, configured to perform an operation on the quantized input data and the quantized first parameters to obtain quantized first output data of the layer to be quantized;
[0048] An inverse quantization unit 103, configured to perform an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized;
[0049] A release unit 104, configured to release the quantized first input data and the quantized first parameters from the memory.
[0050] In a possible implementation, when the layer to be quantized is the first layer of the neural network model, the first input data of the layer to be quantized is the sample data input to the neural network model; otherwise, the first input data of the layer to be quantized is the first output data of the previous layer of the layer to be quantized.
[0051] In a possible implementation, the quantization unit 101 is further configured to, during the backward propagation process, quantize the second input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized second input data and quantized second parameters;
[0052] The operation unit 102 is further configured to obtain the second quantized output data of the layer to be quantized and the gradient of the second parameter of the layer to be quantized according to the quantized second input data and the quantized second parameter;
[0053] The dequantization unit 103 is further configured to perform a dequantization operation on the second quantized output data of the layer to be quantized and the gradient of the second parameter of the layer to be quantized, so as to obtain the second output data of the layer to be quantized and the gradient of the second parameter of the layer to be quantized;
[0054] The release unit 104 is further configured to release the quantized second input data and the quantized second parameter from the memory.
[0055] In a possible implementation manner, when the layer to be quantized is the last layer of the neural network model, the second input data includes the first input data and the network loss; where the network loss is obtained through forward propagation; when the layer to be quantized is not the last layer of the neural network model, the second input data includes the first input data and the second output data of the subsequent layer of the layer to be quantized.
[0056] In a possible implementation manner, the apparatus 100 further includes: a parameter update unit 105, configured to update the parameters of each layer based on the parameters of each layer in the neural network model and the gradient of the parameters.
[0057] The above functional units of the apparatus 100 can be used to implement the Figure 5 method described in the following Figure 5 embodiment. For specific content, reference can be made to
[0058] Figure 2 FIG. shows a structural diagram of a combined processing apparatus 200 according to an embodiment of the present disclosure. The combined processing apparatus 200 can be used for quantization training of a neural network model or can also be used for testing of a neural network model. As Figure 2 shown in, the combined processing apparatus 200 includes a computing processing apparatus 202, an interface apparatus 204, other processing apparatuses 206, and a storage apparatus 208. According to different application scenarios, the computing processing apparatus may include one or more computing apparatuses 210, and the computing apparatus may be configured to execute the Figure 5 or the Figure 7 operations described herein.
[0059] In various embodiments, the computing processing device of the present disclosure may be configured to perform user-specified operations. In an exemplary application, the computing processing device may be implemented as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing processing device may be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, the computing processing device of the present disclosure may be regarded as having a single-core structure or a homogeneous multi-core structure.
[0060] In an exemplary operation, the computing processing device of the present disclosure may interact with other processing devices through an interface device to jointly complete user-specified operations. Depending on the implementation, the other processing devices of the present disclosure may include one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence processor, which are general-purpose and / or special-purpose processors. These processors may include, but are not limited to, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and the number thereof may be determined according to actual needs. As mentioned above, only with respect to the computing processing device of the present disclosure, it may be regarded as having a single-core structure or a homogeneous multi-core structure. However, when considering the computing processing device and other processing devices together, the two may be regarded as forming a heterogeneous multi-core structure.
[0061] In one or more embodiments, the other processing device may serve as an interface between the computing processing device of the present disclosure (which may be specifically implemented as an operation device related to artificial intelligence, such as neural network operations) and external data and control, and perform basic controls including but not limited to data transfer, turning on and / or off the computing device. In additional embodiments, the other processing device may also cooperate with the computing processing device to jointly complete computing tasks.
[0062] In one or more embodiments, the interface device can be used to transfer data and control instructions between a computing processing device and other processing devices. For example, the computing processing device can obtain input data from other processing devices via the interface device and write it into the storage device (or memory) on the chip of the computing processing device. Further, the computing processing device can obtain control instructions from other processing devices via the interface device and write them into the control cache on the chip of the computing processing device. Alternatively or optionally, the interface device can also read the data in the storage device of the computing processing device and transfer it to other processing devices.
[0063] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is respectively connected to the computing processing device and the other processing device. In one or more embodiments, the storage device can be used to store the data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing devices.
[0064] In some embodiments, the present disclosure also discloses a chip, such as Figure 3 the chip 1302 shown in. In one implementation, the chip is a System on Chip (SoC) and integrates one or more combined processing devices as shown in Figure 2 . The chip can be connected to other related components through an external interface device (such as the external interface device 306 shown in Figure 3 ). The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. In some application scenarios, other processing units (such as a video codec) and / or interface modules (such as a DRAM interface) can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip package structure that includes the above chip. In some embodiments, the present disclosure also discloses a board that includes the above chip package structure. The following will be described in detail with reference to Figure 3 the board.
[0065] Figure 3 is a schematic structural diagram of a board 300 according to an embodiment of the present disclosure. As shown in Figure 3As shown in the figure, the board includes a storage device 304 for storing data, which includes one or more storage units 310. The storage device can be connected to the control device 308 and the chip 302 described above and perform data transmission through means such as a bus. Further, the board also includes an external interface device 306, which is configured to perform data relay or transfer functions between the chip (or the chip in the chip package structure) and an external device 312 (such as a server or a computer, etc.). For example, the data to be processed can be transmitted from the external device to the chip through the external interface device. Another example is that the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms. For example, it can adopt a standard PCIE interface, etc.
[0066] In one or more embodiments, the control device in the disclosed board of the present disclosure can be configured to regulate the state of the chip. For this purpose, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the working state of the chip.
[0067] According to the above combination Figure 2 and Figure 3 of the description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which can include one or more of the above boards, one or more of the above chips, and / or one or more of the above combined processing devices.
[0068] According to different application scenarios, the electronic devices or apparatuses disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things (IoT) terminals, mobile terminals, mobile phones, dash cams, navigators, sensors, cameras, video cameras, projectors, watches, earphones, mobile storage devices, wearable devices, vision terminals, autonomous driving terminals, transportation means, household appliances, and / or medical devices. The transportation means includes airplanes, ships, and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, range hoods; the medical devices include nuclear magnetic resonance (NMR) spectrometers, B-ultrasound devices, and / or electrocardiogram (ECG) devices. The electronic devices or apparatuses disclosed herein may also be applied to fields such as the Internet, the Internet of Things, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Further, the electronic devices or apparatuses disclosed herein may also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as the cloud, the edge, and the terminal. In one or more embodiments, the electronic devices or apparatuses with high computing power according to the disclosed solution may be applied to cloud devices (such as cloud servers), while the electronic devices or apparatuses with low power consumption may be applied to terminal devices and / or edge devices (such as smart phones or cameras). In one or more embodiments, the hardware information of the cloud devices is compatible with the hardware information of the terminal devices and / or edge devices, so that appropriate hardware resources can be matched from the hardware resources of the cloud devices according to the hardware information of the terminal devices and / or edge devices to simulate the hardware resources of the terminal devices and / or edge devices, in order to complete unified management, scheduling, and collaborative work of end-cloud integration or cloud-edge-end integration.
[0069] Before describing the embodiments of the present application, the training process of the neural network model will be introduced first.
[0070] In each iteration of training the neural network model, it is necessary to go through the forward propagation, backward propagation, and parameter update processes. Among them, forward propagation is the process of propagating from the first layer to the last layer of the neural network model, and backward propagation is the process of propagating from the last layer to the first layer of the neural network model. Forward propagation calculates the loss function of the network, backward propagation calculates the gradients of each layer in the network based on the loss function, and parameter update updates the parameters of each layer in the network according to the gradients of each layer. After multiple iterations of training, the neural network model is finally converged to obtain a trained neural network model.
[0071] Refer to Figure 4 as shown in (a) of Figure 4(a) in this document is an example schematic diagram of forward propagation provided by this application. This forward propagation process is applied to the above-mentioned device 100, combined device 200, or board 300. In the figure, when forward propagation reaches the layer to be quantized, data X and W are input into the layer to be quantized. Here, X represents the input data of this layer, and W represents the corresponding parameters (weights, biases, etc.) of this layer. The arithmetic unit 102, computing device 210, or chip 302 calculates X and W in the layer to be quantized to obtain the output data Y. Among them, the input data X of the layer to be quantized is also the output data of the previous layer of the layer to be quantized. Data Y is then used as the input data of the next layer of the layer to be quantized. The arithmetic unit in the next layer of the layer to be quantized calculates based on Y and the parameters of the next layer of the layer to be quantized, and obtains the output data... until forward propagation reaches the last layer of the neural network model, and the loss function is obtained.
[0072] See Figure 4 as shown in (b) in this document. Figure 4 (b) in this document is an example schematic diagram of backward propagation provided by this application. In the figure, when backward propagation reaches the layer to be quantized, the gradients of X and Y are input into the layer to be quantized. Here, the gradients of X and Y represent the input data, and W represents the corresponding parameters (weights, biases, etc.) of this layer. The arithmetic unit in the layer to be quantized calculates the gradients of X, W, and Y to obtain the gradient of the output data X. In addition, the output data of the layer to be quantized also includes the gradient of W (not shown in the figure), and the gradient of W is used in the parameter update process. The gradient of X is input into the previous layer of the layer to be quantized as the input data of the previous layer of the layer to be quantized. The arithmetic unit of the previous layer of the layer to be quantized calculates the gradient of X and the parameters to obtain the output data... until backward propagation reaches the first layer. Through backward propagation, the gradients in each layer of the network are calculated.
[0073] During the forward propagation and backward propagation processes, the output data obtained by each layer of the neural network is stored as intermediate data in the storage device 208 or storage unit 310. When the computing device 202 performs calculations, it obtains the required data from the storage device 208 or storage unit 310.
[0074] Based on the gradients in each layer, the parameters of each layer in the network are updated.
[0075] During the training process of the neural network model, it has the following characteristics. 1) During the backpropagation process, the input data of each layer includes, in addition to the output data of the subsequent layer, the input data of this layer during the normal propagation process. For example, for the layer to be quantized, during backpropagation, the input data of the layer to be quantized includes the gradient of the output data Y of the subsequent layer of the layer to be quantized, and also includes the input data X of the layer to be quantized during the forward propagation process; for the previous layer of the layer to be quantized, during backpropagation, the input data of the previous layer of the layer to be quantized includes the gradient of the output data X of the layer to be quantized, and also includes the input data of the previous layer of the layer to be quantized during the forward propagation process. 2) The output data of a certain layer during the backpropagation process is the gradient of the input data of this layer during the forward propagation process. For example, for the subsequent layer of the layer to be quantized, the input data during forward propagation is Y, and the output data of this layer during backpropagation is the gradient of Y; for the layer to be quantized, the input data during forward propagation is X, and the output data of this layer during backpropagation is the gradient of X. It should be noted that the above Figure 4 (a) and (b) in only exemplarily describe one iteration process during the training process of the neural network model. The neural network model needs to go through many iterations of training to converge. For the sake of simplicity of the specification, it will not be further described.
[0076] During the training process of the neural network model, the data all exists in the form of floating-point numbers, including both the forward propagation and the backpropagation. At the same time, some results of the forward propagation can be used as intermediate results in the backpropagation process. Therefore, a large amount of data will be saved during the operation, occupying a large amount of memory, resulting in a large memory overhead of the device. Especially for hardware devices with relatively small memory, training a neural network model is a more challenging task.
[0077] To solve the above problems, the embodiments of the present application provide a quantization training method for a neural network model. This method can be completed by the computing processing device 202 in the combined device 200 shown in Figure 2 , or can be completed by multiple combined devices 200. The multiple combined devices communicate with each other through the interface device 204; this method can also be completed by Figure 3 one or more chips 302 in the device in. Refer to Figure 5 , Figure 5 which is a schematic flowchart of a quantization training method for a neural network model. The embodiments of this method will be described below in combination with Figure 6 (a) and (b) in, where Figure 6 (a) is a schematic diagram of the forward propagation to the layer to be quantized in the quantization training of the neural network model provided by the embodiments of the present application, Figure 6Among them, (b) is a schematic diagram of the backpropagation to the layer to be quantized in the quantization training of the neural network model provided by the embodiment of the present application. The method embodiment includes but is not limited to the following description.
[0078] S501. During the forward propagation process, quantize the first input data and the parameters of the layer to be quantized respectively to obtain the quantized first input data and the quantized parameters.
[0079] Refer to Figure 6 As shown in (a) among them, the first input data of the layer to be quantized includes X, and the parameters of the layer to be quantized are W, where W includes weights and may also include biases, etc. The quantized first input data includes X1, and the quantized parameters include W1, and W1 includes quantized weights and may also include quantized biases, etc.
[0080] When the forward propagation reaches the layer to be quantized, quantize the first input data X and the parameter W respectively to obtain the quantized first input data X1 and the quantized W1. If the layer to be quantized is the first layer in the neural network model, the first input data is the sample data input to the neural network model; if the layer to be quantized is not the first layer of the neural network model, the first input data is the first output data of the previous layer of the layer to be quantized. For example, Figure 6 In (a) among them, the first input data X of the layer to be quantized can be the first output data of the previous layer of the layer to be quantized.
[0081] Generally speaking, the first input data and the parameters are floating-point data, and the quantized first input data and the quantized parameters are fixed-point integer data. For example, the commonly used floating-point data type is float32, and the commonly used fixed-point data types are int8 and int16. The present application does not limit the specific quantization method.
[0082] It should be noted that the current iteration can be any iteration process in the training of the neural network model. For example, it can be the first iteration, a certain intermediate iteration, or the last iteration.
[0083] S502. Obtain the first output data of the layer to be quantized according to the quantized first input data and the quantized parameters.
[0084] Optionally, obtaining the first output data of the layer to be quantized according to the quantized first input data and the quantized parameters includes: first calculating the quantized first output data according to the quantized first input data and the quantized parameters, and then dequantizing the quantized first output data to obtain the first output data of the layer to be quantized.
[0085] Refer to Figure 6As shown in (a) therein, based on the quantized first input data X1 and the quantized parameter W1, calculation is performed to obtain the quantized first output data Y1, where Y1 is fixed-point integer data; then the fixed-point data Y1 is dequantized to obtain the first output data Y, and Y is floating-point data. This application does not limit the specific dequantization method.
[0086] Through quantization processing, data with a large amount of computation can be transformed into data with a small amount of computation. When performing calculations, the amount of computation can be reduced and the computation speed can be increased. For example, there are many decimal places after the decimal point in floating-point data, while fixed-point integer data is an integer. Therefore, when performing operations on fixed-point integer data, compared with performing operations on floating-point data, the computation speed can be increased and the memory occupancy of the data is relatively small.
[0087] S503: Release the quantized first input data and the quantized parameter from the memory.
[0088] After obtaining the first output data of the layer to be quantized, the quantized first input data and the quantized parameter are released from the memory, that is, deleted. Refer to Figure 6 As shown in (a) therein, after obtaining the first output data Y of the layer to be quantized, the quantized first input data X1 and the quantized parameter W1 (quantized weight, quantized bias, etc.) are released from the memory.
[0089] After the forward propagation of this layer ends, release the quantized first input data and the quantized parameter of this layer, and only store the first data. This is to increase the available memory. Especially when the memory of the hardware device is fixed, it can prevent a lot of intermediate data generated during the training of the neural network model from occupying too much memory and affecting the training of the neural network model. When the backward propagation needs to perform calculations based on the quantized first input data and the quantized parameter, obtain the quantized first input data according to the stored first data.
[0090] It should be noted that what is released here is the quantized first input data and the quantized parameter (fixed-point integer data), rather than the first input data (floating-point data). This is because in some neural network models, the first input data (floating-point data) may be required during the backward propagation process. For example, the first input data X is required during the backward propagation of the previous layer of the layer to be quantized. The first input data X is used as the input data of the previous layer of the layer to be quantized and is input into the previous layer of the layer to be quantized.
[0091] It should be noted that steps S101 to S103 describe the processing process during the forward propagation to the layer to be quantized. After that, the forward propagation continues, that is, the output data Y of the layer to be quantized will be used as the input data of the layer after the layer to be quantized and input to the layer after the layer to be quantized. The layer after the layer to be quantized processes according to the input data and parameters,... until the forward propagation reaches the last layer of the neural network model, calculates the loss function, and the forward propagation ends.
[0092] It can be seen that during the forward propagation process, by quantizing the input data and parameters, the amount of computation during the training of the neural network model can be reduced, and the computing speed can be improved; after the forward propagation of this layer ends, releasing the quantized input data and quantized parameters can reduce the memory occupancy. Under the condition of a certain device memory, the device can normally perform the training of the neural network model and obtain a trained neural network model.
[0093] The embodiment of the present application also provides a quantization training method for a neural network model. Refer to Figure 7 As shown, the method includes but is not limited to the following steps. The content of steps S701 to S703 can refer to the description of S501 to steps S503. For the sake of simplicity of the specification, it will not be repeated here.
[0094] S701. During the forward propagation process, respectively quantize the first input data of the layer to be quantized and the parameters of the layer to be quantized to obtain the quantized first input data and quantized parameters.
[0095] S702. Obtain the first output data of the layer to be quantized according to the quantized first input data and quantized parameters.
[0096] S703. Release the quantized first input data and quantized parameters from the memory.
[0097] S704. During the backward propagation process, respectively quantize the second input data of the layer to be quantized and the parameters of the layer to be quantized to obtain the quantized second input data and the second quantized parameters.
[0098] The backward propagation is from the last layer of the neural network model forward. The last layer calculates according to the first input data and the loss function during the normal propagation to obtain the output data,... and continues to propagate forward until the backward propagation reaches the layer to be quantized.
[0099] During the backpropagation process, although the first quantized input data and the quantized parameters need to be used as the input data for backpropagation, in order to save memory, after the first quantized input data and the quantized parameters generated during the forward propagation are used, these data are released immediately, and only the first input data is saved. When the backpropagation requires the quantized data, the corresponding quantized data (the first quantized input data and the quantized parameters) are generated based on the saved first input data. Therefore, this solution reduces memory by only saving the first input data and releasing the temporarily unnecessary quantized data (the first quantized input data and the quantized parameters).
[0100] When backpropagating to the layer to be quantized, the second input data and the parameters of the layer to be quantized are quantized respectively to obtain the second quantized input data and the second quantized parameters. Refer to Figure 6 As shown in (b) of
[0101] It should be noted that generally, the second quantized parameters (during the backpropagation process) are the same as the first quantized parameters (during the forward propagation process). However, in step S503, the first quantized parameters are released, and there are no first quantized parameters in the memory. In this step (during the backpropagation process), the parameters are quantized again, so they are called the second quantized parameters.
[0102] S705. Obtain the second output data and the gradients of the second parameters of the layer to be quantized according to the second quantized input data and the second quantized parameters.
[0103] Optionally, obtaining the second output data and the gradients of the second parameters of the layer to be quantized according to the second quantized input data and the second quantized parameters includes: calculating according to the second quantized input data and the second quantized parameters to obtain the gradients of the second quantized output data and the second quantized parameters of the layer to be quantized, and then dequantizing the gradients of the second quantized output data and the second quantized parameters to obtain the gradients of the second output data and the second parameters of the layer to be quantized.
[0104] Refer to Figure 6As shown in (b) therein, calculations are performed based on the quantized second input data X1, Y2, and the parameter W1 to be quantized by the layer to be quantized, to obtain the quantized second output data X2 and the gradient W2 of the quantized second parameter. Then, inverse quantization calculations are performed on X2 and W2 to obtain the gradient of X and the gradient of W. Among them, the gradient of X can be used as the input data of the previous layer of the layer to be quantized and continue the backpropagation, and the gradient of W is used to calculate the optimizer in the parameter update process, and the optimizer is the method for parameter update.
[0105] S706. Release the quantized second input data and the quantized second parameter from the memory.
[0106] After obtaining the second output data and the gradient of the second parameter of this layer, release the quantized second input data and the quantized second parameter from the memory to reduce the memory occupancy of the intermediate data and increase the available space of the memory. For example, Figure 6 in (b) therein, after calculating the gradient of the second output data X and the second parameter W of the layer to be quantized, release the quantized second input data X1, Y2, and the quantized second parameter W1 (quantized weight, quantized bias, etc.) from the memory.
[0107] It should be noted that after this step ends, the backpropagation continues. The gradient of the second output data X of the layer to be quantized is used as the second input data of the previous layer of the layer to be quantized,..., until the backpropagation reaches the first layer and the backpropagation ends.
[0108] S707. Update the parameters of each layer of the neural network model.
[0109] Through backpropagation, the gradients of the parameters (weights, biases, etc.) of each layer in the neural network model are obtained. Based on the parameters, the gradients of the parameters, and the optimizer, the parameters of each layer in the neural network model are updated. For example, for a certain layer in the neural network model, calculations are performed according to the weights, the gradients of the weights, and the optimizer (calculated in the backpropagation process) in this layer to obtain new weights, and the new weights are used to replace the original weights to achieve the update of the weights.
[0110] It should be noted that steps S101 to S107 in this embodiment describe the process of a certain iteration in the training process of the neural network model. When actually training, the neural network model needs to be iteratively trained many times to converge.
[0111] It should also be noted that Figure 6 (a) and (b) therein are an exemplary embodiment described in this application. Figure 6The first input data, the first output data, the second input data, the second output data, etc. in [description] are also an exemplary representation. In the actual training of the neural network model, the input and output can be one or multiple, which does not limit this application.
[0112] It should also be noted that the quantization training method of the neural network model described in this application can be applied to one or multiple layers in the neural network model, and this application does not make a limitation. Generally speaking, it can be applied to layers with relatively large computational complexity, such as layers involving convolutional operations and matrix multiplication operations, such as convolutional layers, deconvolutional layers, depth convolutional layers, etc., and also fully connected layers, etc.
[0113] It can be seen that during the forward propagation or backward propagation process, by quantizing the input data and parameters, the computational complexity during the training of the neural network model can be reduced, and the computational speed can be improved; after the forward propagation or backward propagation of this layer is completed, the quantized input data and quantized parameters are released, which can reduce the memory occupancy. Under the condition of a certain device memory, the device can normally perform the training of the neural network model and obtain a trained neural network model.
[0114] The embodiments of this application provide a method for testing a neural network model, referring to Figure 8 the flow schematic diagram of the method for testing the neural network model shown in [figure], and this method includes but is not limited to the following description.
[0115] S801. Obtain test data.
[0116] Obtain test data. For example, the test data can be an image or an image frame, or can also be speech, etc.
[0117] S802. Use the trained neural network model to test the test data.
[0118] Use Figure 5 or Figure 7 the trained neural network model in the method embodiment to test the test data and obtain a test result. Among them, the neural network model is obtained by using a quantization training method, and during the training process, for layers with relatively large computational complexity, the quantization results of the input data and parameters (such as weights and biases) can be released during normal propagation, and during the backward propagation process, the quantization results of the input data and the quantization results of the parameters are released after the propagation of this layer is completed to save memory space and ensure the normal progress of the neural network model training.
[0119] The embodiments of this application also provide a testing device for a neural network model. Refer to Figure 9 the schematic diagram shown in [figure], and the testing device 900 for the neural network model includes:
[0120] An acquisition unit 901 for acquiring test data;
[0121] A test unit 902 for testing the test data using a trained neural network model. The neural network model is obtained by Figure 5 or Figure 7 the method of the embodiments shown.
[0122] It should be noted that, for the purpose of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art can understand that the solutions of this disclosure are not limited by the order of the described actions. Therefore, according to the disclosure or teachings of this disclosure, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in this disclosure can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of some or certain solutions of this disclosure. In addition, according to the differences in the solutions, this disclosure focuses on the descriptions of some embodiments. In view of this, those skilled in the art can understand the parts not detailed in a certain embodiment of this disclosure by referring to the relevant descriptions of other embodiments.
[0123] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented by other means not disclosed herein. For example, for each unit in the foregoing embodiments of the electronic device or apparatus, this disclosure divides them based on the consideration of logical functions, and there may be other division methods in actual implementation. Another example is that multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed in conjunction with the accompanying drawings can be direct or indirect couplings between the units or components. In some scenarios, the foregoing direct or indirect couplings involve communication connections using interfaces, where the communication interfaces can support signal transmissions in electrical, optical, acoustic, magnetic, or other forms.
[0124] In this disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The foregoing components or units may be located at the same position or distributed to multiple network units. In addition, according to actual needs, some or all of the units can be selected to achieve the purpose of the solutions described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure can be integrated into one unit or each unit exists physically separately.
[0125] In some implementation scenarios, the above integrated units can be implemented in the form of software program modules. If implemented in the form of software program modules and sold or used as an independent product, the integrated units can be stored in a computer-readable memory. Based on this, when the solution of this disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of this disclosure. The aforementioned memory may include, but is not limited to, various media that can store program codes, such as USB flash drives, flash memory drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs.
[0126] In some other implementation scenarios, the above integrated units can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include, but is not limited to, physical devices, and the physical devices may include, but are not limited to, devices such as transistors or memristors. In view of this, various devices described herein (such as computing devices or other processing devices) can be implemented by an appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, and an ASIC, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which may be, for example, a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high bandwidth memory (HBM), a hybrid memory cube (HMC), a ROM, and a RAM, etc.
[0127] The foregoing can be better understood according to the following terms:
[0128] Clause A1. A quantization training method for a neural network model, the neural network model including multiple layers, the method comprising: during forward propagation, quantizing respectively the first input data of the layer to be quantized and the parameters of the layer to be quantized to obtain quantized first input data and quantized first parameters; performing an operation on the quantized first input data and the quantized first parameters to obtain quantized first output data; performing an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized; and releasing the quantized first input data and the quantized first parameters from the memory.
[0129] Clause A2. The method according to Clause A1, when the layer to be quantized is the first layer of the neural network model, the first input data of the layer to be quantized is the sample data input to the neural network model; otherwise, the first input data of the layer to be quantized is the first output data of the previous layer of the layer to be quantized.
[0130] Clause A3. The method according to Clause A1 or Clause A2, the method further comprising: during backpropagation, quantizing respectively the second input data of the layer to be quantized and the parameters of the layer to be quantized to obtain quantized second input data and quantized second parameters; obtaining the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized according to the quantized second input data and the quantized second parameters; performing an inverse quantization operation on the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized to obtain the second output data of the layer to be quantized and the gradient of the second parameters of the layer to be quantized; and releasing the quantized second input data and the quantized second parameters from the memory.
[0131] Clause A4. The method according to Clause A3, when the layer to be quantized is the last layer of the neural network model, the second input data includes the first input data and the network loss; wherein the network loss is obtained during forward propagation; when the layer to be quantized is not the last layer of the neural network model, the second input data includes the first input data and the second output data of the next layer of the layer to be quantized.
[0132] Clause A5. The method according to Clause A3 or Clause A4, after the backpropagation ends, the method further comprising: updating the parameters of each layer based on the parameters and the gradients of the parameters of each layer in the neural network model.
[0133] Clause A6. The method according to any one of Clauses A1 - A5, the method is applicable to at least one layer in the neural network model.
[0134] Clause A7. A method for testing a neural network model, comprising: obtaining test data; using the trained neural network model to test the test data; the neural network model being trained by the method according to any one of Clauses A1 - A6.
[0135] Clause A8. A quantization training device for a neural network model, the neural network model including multiple layers, the device comprising: a quantization unit configured to, during forward propagation, respectively quantize the first input data of the layer to be quantized and the parameters of the layer to be quantized to obtain quantized first input data and quantized first parameters; an operation unit configured to perform an operation on the quantized input data and the quantized first parameters to obtain quantized first output data of the layer to be quantized; an inverse quantization unit configured to perform an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized; and a release unit configured to release the quantized first input data and the quantized first parameters from the memory.
[0136] Clause A9. A test device for a neural network model, comprising: an acquisition unit configured to obtain test data; a test unit configured to use the trained neural network model to test the test data; the neural network model being trained by the method according to any one of Clauses A1 - A6.
[0137] Clause A10. A quantization training device for a neural network model, comprising a memory and a processor, the memory being configured to store instructions, and the processor being configured to call the instructions to execute the method according to any one of Claims Clauses A1 - A6.
[0138] Clause A11. A test device for a neural network model, comprising a memory and a processor, the memory being configured to store instructions, and the processor being configured to call the instructions to execute the method according to Clause A7.
[0139] Clause A12. A computer storage medium, comprising program instructions which, when run on a computer, cause the computer to execute the method according to any one of Clauses A1 - A6.
[0140] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and alternative approaches can be envisioned by those skilled in the art without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The appended claims are intended to define the scope of protection of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.
Claims
1. A quantization training method for a neural network model, characterized in that The neural network model includes multiple layers, and the method includes: In the forward propagation process, quantize the first input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized first input data and quantized first parameters; Perform an operation on the quantized first input data and the quantized first parameters to obtain quantized first output data; Perform an inverse quantization operation on the quantized first output data of the layer to be quantized to obtain the first output data of the layer to be quantized; After obtaining the first output data of the layer to be quantized, release the quantized first input data and the quantized first parameters from the memory.
2. The method according to claim 1, wherein: When the layer to be quantized is the first layer of the neural network model, the first input data of the layer to be quantized is the sample data input into the neural network model; Otherwise, the first input data of the layer to be quantized is the first output data of the previous layer of the layer to be quantized.
3. The method according to claim 1 or 2, characterized in that, The method further includes: In the backward propagation process, quantize the second input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized second input data and quantized second parameters; According to the quantized second input data and the quantized second parameters, obtain the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized; Perform an inverse quantization operation on the quantized second output data of the layer to be quantized and the gradient of the quantized second parameters of the layer to be quantized to obtain the second output data of the layer to be quantized and the gradient of the second parameters of the layer to be quantized; Release the quantized second input data and the quantized second parameters from the memory.
4. The method according to claim 3, wherein: When the layer to be quantized is the last layer of the neural network model, the second input data includes the first input data and the network loss; wherein the network loss is obtained from the forward propagation; When the layer to be quantized is not the last layer of the neural network model, the second input data includes the first input data and the second output data of the next layer of the layer to be quantized.
5. The method according to claim 4, wherein After the backward propagation ends, the method further includes: updating the parameters of each layer based on the parameters and the gradients of the parameters of each layer in the neural network model.
6. The method according to claim 1, wherein The method is applicable to at least one layer in the neural network model.
7. A testing method for a neural network model, characterized in that, Including: Obtain test data; Use the trained neural network model to test the test data; The neural network model is trained by the method according to any one of claims 1-6.
8. A quantization training device for a neural network model, characterized in that, The neural network model includes multiple layers, and the device includes: A quantization unit, configured to, in the forward propagation process, quantize the first input data of the layer to be quantized and the parameters of the layer to be quantized respectively, to obtain quantized first input data and quantized first parameters; An operation unit, configured to perform an operation on the quantized first input data and the quantized first parameters to obtain quantized first output data; An inverse quantization unit, configured to perform an inverse quantization operation on the quantized first output data of the layer to be quantized, to obtain the first output data of the layer to be quantized; A release unit, configured to release the quantized first input data and the quantized first parameter from the memory after obtaining the first output data of the layer to be quantized.
9. A testing device for a neural network model, characterized in that, Comprising: An acquisition unit, configured to acquire test data; A test unit, configured to test the test data by using a trained neural network model; The neural network model is trained by the method according to any one of claims 1-6.
10. A quantization training device for a neural network model, characterized in that, Comprising a memory and a processor, the memory is configured to store instructions, and the processor is configured to call the instructions to execute the method according to any one of claims 1-6.
11. A test device for a neural network model, characterized in that, Comprising a memory and a processor, the memory is configured to store instructions, and the processor is configured to call the instructions to execute the method according to claim 7.
12. A computer storage medium, characterized in that, Comprising program instructions, when the program instructions run on a computer, causing the computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method for efficiently clustering massive data
CN102243641A
Image compression architecture and method for reducing memory requirement
CN104219521A