Neural network training method and device, electronic equipment, storage medium and chip

CN116739068BActive Publication Date: 2026-10-09SHANGHAI POWERTENSORS INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310804186.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2026-10-09
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

在利用神经网络进行推理之前需要对神经网络进行训练,一般的,在神经网络的训练过程中采用单精度数据类型,比如float32数据类型,训练得到的神经网络精度较高,但是采用float32数据类型训练神经网络,使得训练过程中计算资源消耗较高、存储空间占用较大

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116739068B_ABST
    Figure CN116739068B_ABST
Patent Text Reader

Abstract

The present disclosure provides a neural network training method and device, electronic equipment, storage medium and chip. The method comprises: in a training process of a neural network to be trained, obtaining feature data corresponding to any target processing layer in the neural network to be trained; performing type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein the data precision of a first data type corresponding to the feature data is higher than the data precision of a second data type corresponding to the converted feature data; performing operation processing on the converted feature data corresponding to the target processing layer to generate output feature data of the target processing layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning technology, and more specifically, to a neural network training method, apparatus, electronic device, storage medium, and chip. Background Technology

[0002] With the development of deep learning research, neural networks are being used more and more widely. Before using a neural network for inference, it needs to be trained. Generally, single-precision data types, such as float32, are used during the training process. This results in a neural network with high precision, but using float32 data types for training leads to high computational resource consumption and large storage space usage during the training process.

[0003] Therefore, there is an urgent need for a neural network training method that can reduce resource consumption, increase computing speed, and at the same time not lose the accuracy of the neural network. Summary of the Invention

[0004] In view of the above, this disclosure provides at least one neural network training method, apparatus, electronic device, storage medium, and chip.

[0005] In a first aspect, this disclosure provides a neural network training method, including:

[0006] During the training process of the neural network to be trained, feature data corresponding to any target processing layer in the neural network to be trained is obtained;

[0007] The feature data of the target processing layer is subjected to type conversion processing to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data;

[0008] The transformed feature data corresponding to the target processing layer is processed to generate the output feature data of the target processing layer.

[0009] Secondly, this disclosure provides a chip, which includes a memory and a computing device;

[0010] The memory is used to store feature data corresponding to the target processing layer in the neural network to be trained;

[0011] The computing device is configured to retrieve the feature data corresponding to the target processing layer from the memory, and perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; and perform computational processing on the converted feature data corresponding to the target processing layer to generate output feature data of the target processing layer.

[0012] Thirdly, this disclosure provides a neural network training apparatus, comprising:

[0013] The acquisition module is used to acquire feature data corresponding to any target processing layer in the neural network to be trained during the training process.

[0014] The first processing module is used to perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data.

[0015] The second processing module is used to perform calculations on the transformed feature data corresponding to the target processing layer to generate the output feature data of the target processing layer.

[0016] Fourthly, this disclosure provides an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, they perform the steps of the neural network training method as described in the first aspect or any embodiment above; or a chip as described in the second aspect or any embodiment above.

[0017] Fifthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the neural network training method as described in the first aspect or any of the embodiments above.

[0018] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a neural network training method provided in an embodiment of this disclosure is shown.

[0021] Figure 2 This illustration shows a schematic diagram of data types in a neural network training method provided by an embodiment of the present disclosure;

[0022] Figure 3 This diagram illustrates the structure of a chip provided in an embodiment of the present disclosure;

[0023] Figure 4 This illustration shows a schematic diagram of another chip structure provided in an embodiment of the present disclosure;

[0024] Figure 5 This diagram illustrates the forward propagation process in a neural network training method provided by an embodiment of the present disclosure.

[0025] Figure 6 This illustration shows a schematic diagram of another chip structure provided in an embodiment of the present disclosure;

[0026] Figure 7a This diagram illustrates the backpropagation process for obtaining input feature gradient values ​​in a neural network training method provided by an embodiment of the present disclosure.

[0027] Figure 7b This diagram illustrates the backpropagation process for obtaining weight feature gradient values ​​in a neural network training method provided by an embodiment of the present disclosure.

[0028] Figure 8 A schematic diagram of the architecture of a neural network training device provided in an embodiment of this disclosure is shown;

[0029] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0031] Artificial Intelligence (AI) computing scenarios are mainly divided into training and inference. For different tasks, AI chips are divided into training chips and inference chips; among them, training chips train neural networks with specific functions using massive amounts of data. Training chips have high requirements for accuracy and computing performance, and the training time is relatively long.

[0032] Generally, single-precision data types (such as Float32) are used in neural network training, resulting in higher precision neural networks. However, training neural networks using the float32 data type consumes a lot of computing resources and storage space, requiring high-performance hardware support. AI inference chips have limited resources in terms of space capacity, bandwidth, and power consumption, which limits their application.

[0033] To alleviate the above problems, this disclosure proposes a neural network training method, apparatus, electronic device, storage medium, and chip. Without sacrificing the accuracy of the neural network, the computational load is reduced by decreasing the data precision, thereby improving computational speed, saving computational resources, and reducing power consumption.

[0034] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0035] To facilitate understanding of the embodiments of this disclosure, a neural network training method disclosed in this disclosure will first be described in detail. The execution entity of the neural network training method provided in this disclosure is generally a computer device with certain computing capabilities, such as a terminal device or a server. In some possible implementations, the neural network training method can be implemented by a processor calling computer-readable instructions stored in memory.

[0036] See Figure 1 The diagram shown is a flowchart of a neural network training method provided in this embodiment of the present disclosure. The method includes: S101-S103, specifically:

[0037] S101, during the training process of the neural network to be trained, acquire the feature data corresponding to any target processing layer in the neural network to be trained.

[0038] S102, perform type conversion on the feature data of the target processing layer to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data.

[0039] S103, perform calculations on the transformed feature data corresponding to the target processing layer to generate the output feature data of the target processing layer.

[0040] In the above method, during the training process of the neural network to be trained, feature data corresponding to any target processing layer in the neural network to be trained is obtained; then, the feature data of the target processing layer can be converted to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; for example, the first data type can be a single-precision data type (Float32), and the second data type can be a half-precision data type (BFloat16); since the data precision of the second data type is lower than that of the first data type, the computational load can be reduced and computational resources can be saved when performing operations on the converted feature data corresponding to the target processing layer, and the output feature data of the target processing layer can be generated more quickly, thereby improving the processing efficiency of the target processing layer and thus improving the training efficiency of the neural network to be trained.

[0041] The following provides a detailed explanation of S101-S103.

[0042] The neural network training process consists of two parts: forward propagation and backward propagation. Forward propagation involves inputting the data required for training into the neural network to be trained, obtaining the output result, and then using the output result and ground truth data to calculate the corresponding loss value for the neural network. This loss value is then used to calculate the gradient of each network parameter through backward propagation, and the network parameters are updated. Forward and backward propagation are iterated repeatedly until the loss function of the neural network converges and the accuracy of the neural network meets expectations.

[0043] The following explains the forward propagation process during neural network training.

[0044] Regarding S101:

[0045] Considering that the operators that consume the most computational resources in neural networks are mainly concentrated in general matrix multiplication (GEMM), such as convolutional layers and fully connected layers, which contain these computational units, acceleration of neural networks is usually aimed at accelerating GEMM.

[0046] Based on this, the target processing layer in this disclosure can be a network processing layer including GEMM computation units, such as a convolutional layer or a fully connected layer. For example, when the target processing layer is a convolutional layer, the acquired feature data can include input feature data and weight feature data. The neural network to be trained can be a neural network with any specific function, such as a semantic segmentation network, a face recognition network, or a classification network. When the target processing layer is an intermediate network processing layer of the neural network to be trained, the input feature data of the target processing layer is the output feature data of the network processing layers preceding it.

[0047] Regarding S102:

[0048] To minimize the loss of precision in neural networks, the feature data corresponding to the target processing layer is typically stored in Float32 data format (the first data type). However, performing dot product operations on data of the first data type consumes significant resources and is slow. Therefore, before performing dot product operations on the feature data corresponding to the target processing layer, a data type conversion can be performed to obtain converted feature data. The precision of the converted feature data (the second data type) is lower than that of the original first data type; for example, the second data type could be BFloat16. Furthermore, the precision loss is significant when performing accumulation operations on data of the second data type. Therefore, before performing accumulation operations on the feature data corresponding to the target processing layer, a data type conversion can be performed on the feature data of the second data type to obtain converted feature data of the first data type. Performing accumulation operations on the converted feature data can minimize the precision loss.

[0049] See Figure 2The diagram illustrates the data types shown. Float32's data format includes 32 binary bits, while BFloat16's data format includes 16 binary bits. Therefore, BFloat16 occupies less space than Float32. In both BFloat16 and Float32 data formats, the exponent is represented by 8 binary bits, so they can represent the same range of data. In Float32's data format, the fraction is represented by 23 binary bits, while in BFloat16's data format, the fraction is represented by 7 binary bits. Therefore, BFloat16 has lower precision than Float32.

[0050] In practice, when converting the data type of feature data from Float32 to BFloat16, the 16 least significant bits of the corresponding binary form of the feature data can be deleted to obtain the converted feature data with data type BFloat16. Conversely, when converting the data type of feature data from BFloat16 to Float32, 16 zeros can be added to the end of the corresponding binary form of the feature data to obtain the converted feature data with data type Float32.

[0051] Regarding S103:

[0052] In implementation, after converting the feature data corresponding to the target processing layer to obtain the converted feature data, operations can be performed on the converted feature data to quickly generate the output feature data of the target processing layer. For example, when the target processing layer is a convolutional layer, these operations can include dot product operations and accumulation operations.

[0053] In one possible implementation, the feature data includes weighted feature data and input feature data, and the transformed feature data includes transformed weighted feature data and transformed input feature data; the transformed feature data corresponding to the target processing layer is processed to generate the output feature data of the target processing layer, including:

[0054] Step A1: Perform a dot product operation between the local feature data within any window of the transformed input feature data and the transformed weighted feature data to obtain the intermediate feature values ​​corresponding to each feature position.

[0055] Step A2 involves summing the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window.

[0056] Step A3: Generate the output feature data of the target processing layer based on the output feature values ​​corresponding to each window.

[0057] If the feature data includes weighted feature data and input feature data, then the transformed feature data includes transformed weighted feature data and transformed input feature data. Local feature data within any window of the transformed input feature data can be processed by performing a dot product operation with the transformed weighted feature data to obtain intermediate feature values ​​corresponding to each feature position. Here, the data types corresponding to the transformed weighted feature data and the transformed input feature data are the second data type, therefore the data type corresponding to the obtained intermediate feature values ​​is also the second data type.

[0058] During implementation, after obtaining the intermediate feature values ​​corresponding to each feature position, the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position can be accumulated to obtain the output feature value corresponding to any window. Based on the output feature values ​​corresponding to each window, the output feature data of the target processing layer can be generated.

[0059] In specific implementation, for steps A2 and A3, the output feature data of the target processing layer can be generated in the following two ways.

[0060] The first method ensures minimal loss of precision in the obtained output feature data. Specifically, it utilizes data of the first data type for accumulation operations. For example, the data type of the intermediate feature values ​​corresponding to each feature position can be converted to the first data type to obtain the converted intermediate feature values. Then, the offset corresponding to the target processing layer (which can be of the first data type) and the converted intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to any window; the data type of this output feature value is the first data type; where the offset is the value corresponding to the offset term of the neural network parameters. Finally, based on the output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated; the data type of this output feature data is the first data type.

[0061] The second approach can improve the computation speed of the target processing layer. Specifically, it can utilize data of a second data type for accumulation operations. For example, the offset corresponding to the target processing layer (which can be of a second data type) and the intermediate feature values ​​corresponding to each feature position can be accumulated to obtain the output feature value corresponding to any window; the data type of this output feature value is a second data type.

[0062] Meanwhile, considering that the feature data corresponding to the target processing layer is usually stored using a high-precision data type, and that the subsequent processing layer has low precision loss when using high-precision data, the data type of the output feature value can be converted into the first data type to obtain the converted output feature value; then, based on the converted output feature value corresponding to each window, the output feature data of the target processing layer is generated; the data type corresponding to this output feature data is the first data type.

[0063] Since the data types of the converted weight feature data and the converted input feature data are both secondary data types, performing a dot product operation between the local feature data within any window of the converted input feature data and the converted weight feature data can quickly obtain the intermediate feature values ​​corresponding to each feature position; this improves computational efficiency and, consequently, the efficiency of obtaining the output feature data.

[0064] The first method will be explained below.

[0065] The intermediate feature values ​​are of the second data type. The intermediate feature values ​​corresponding to the offset of the target processing layer and each feature position are summed to obtain the output feature values ​​corresponding to any window, including:

[0066] Step B1: Convert the intermediate feature values ​​corresponding to each feature position to obtain the converted intermediate feature values; wherein, the data type of the converted intermediate feature values ​​is the first data type.

[0067] Step B2 involves summing the offset corresponding to the target processing layer and the transformed intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window. The data type of the output feature value is the first data type.

[0068] In implementation, following the process described above of converting the data type from BFloat16 to Float32, the intermediate feature values ​​corresponding to each feature position are converted to obtain the converted intermediate feature values. The data type of these converted intermediate feature values ​​is the first data type. Then, the offset corresponding to the target processing layer and the converted intermediate feature values ​​corresponding to each feature position are summed to obtain the output feature value corresponding to any window. The data type of this output feature value is also the first data type. This allows for the direct generation of the target processing layer's output feature data based on the output feature values ​​corresponding to each window.

[0069] Here, in order to ensure that the accuracy loss of the obtained output feature values ​​is low, the intermediate feature values ​​corresponding to each feature position can be converted to obtain the converted intermediate feature values; the data type of the converted intermediate feature values ​​is the first data type; then the offset corresponding to the target processing layer and the converted intermediate feature values ​​corresponding to each feature position are accumulated to ensure that the accuracy of the output feature values ​​corresponding to any window is high.

[0070] In one possible implementation, the intermediate feature values ​​corresponding to each feature position are subjected to type conversion processing to obtain converted intermediate feature values, including:

[0071] Step C1: Based on the network task and / or network structure information corresponding to the neural network to be trained, determine the data processing type corresponding to the accumulation operation in the target processing layer.

[0072] Step C2: In response to the data processing type indication being the first data type, the intermediate feature values ​​corresponding to each feature position are converted to obtain the converted intermediate feature values.

[0073] In practice, the data processing type corresponding to the accumulation operation in the target processing layer can be determined first based on the network task and / or network structure information corresponding to the neural network to be trained. If the network task corresponding to the neural network to be trained is a complex network task that is sensitive to accuracy, such as a semantic segmentation task, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined as the first data type; if the network task corresponding to the neural network to be trained is a simple network task, such as a classification task, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined as the second data type.

[0074] Alternatively, if the network structure information of the neural network to be trained indicates that the number of network processing layers is greater than or equal to the number threshold, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined to be the first data type; if the network structure information of the neural network to be trained indicates that the number of network processing layers is less than the number threshold, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined to be the second data type.

[0075] Alternatively, based on the network structure information corresponding to the neural network to be trained, the accumulation operation of each network processing layer in the neural network to be trained can be determined as a second data type; and after training the neural network according to the above second data type, the first network precision of the neural network can be obtained. When the first network precision meets the precision requirement, the accumulation operation in each network processing layer is determined as a second data type, that is, the data processing type corresponding to the accumulation operation in the target network processing layer is determined as a second data type.

[0076] Considering that when the data processing type corresponding to the accumulation operation in the target network processing layer is the second data type, the obtained accumulation operation result may experience data overflow or rounding errors, resulting in low accuracy of the trained neural network. Therefore, when the first network accuracy does not meet the accuracy requirements, based on the network structure information corresponding to the neural network to be trained, it is determined that the accumulation operations of some network processing layers in the neural network to be trained are of the first data type, and the accumulation operations of some network processing layers are of the second data type. Furthermore, it is determined that the neural network to be trained will be trained according to a combination of some first data type and some second data type data to obtain the second network accuracy. When the second network accuracy meets the accuracy requirements, the data type corresponding to the accumulation operation in each network processing layer is obtained, thus determining the data processing type corresponding to the accumulation operation in the target network processing layer.

[0077] When the accuracy of the second network does not meet the accuracy requirements, the accumulation operation of each network processing layer in the neural network to be trained can be determined as the first data type based on the network structure information corresponding to the neural network to be trained. That is, the data processing type corresponding to the accumulation operation in the target network processing layer is determined as the first data type.

[0078] Alternatively, the data processing type corresponding to the accumulation operation in the target processing layer can be determined based on network task and network structure information. For example, if the network task is a complex network task that is sensitive to precision, and the network structure information indicates that the number of network processing layers is greater than or equal to a threshold, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined to be the first data type. If the network task is a simple network task, and the network structure information indicates that the number of network processing layers is less than the threshold, then the data processing type corresponding to the accumulation operation in the target processing layer can be determined to be the second data type. If the network task is a complex network task, and the network structure information indicates that the number of network processing layers is less than the threshold, then it can be determined that some network processing layers in the neural network to be trained use the first data type for accumulation operations, and some network processing layers use the second data type for accumulation operations, thus determining the data processing type corresponding to the accumulation operation in the target processing layer (which may be the first data type or the second data type).

[0079] In response to the data processing type indication being the first data type, the intermediate feature values ​​corresponding to each feature position can be converted according to the data type conversion process from BFloat16 to Float32 as described above, to obtain the converted intermediate feature values; the data type of the converted intermediate feature values ​​is the first data type.

[0080] Here, the data processing type corresponding to the accumulation operation in the target processing layer can be determined more flexibly based on the network task and / or network structure information corresponding to the neural network to be trained.

[0081] The second method will be explained below.

[0082] Since the intermediate feature values ​​are of type II data, the offset corresponding to the target processing layer (which can be of type II data) and the intermediate feature values ​​corresponding to each feature position can be directly accumulated to obtain the output feature value of type II data for any window. After obtaining the output feature values, the output feature data of the target processing layer is generated based on the output feature values ​​corresponding to each window.

[0083] Specifically, based on the output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated, including:

[0084] Step D1: If the data type of the output feature value is the second data type, perform type conversion on the output feature value to obtain the converted output feature value; wherein, the data type of the converted output feature value is the first data type.

[0085] Step D2: Based on the transformed output feature values ​​corresponding to each window, generate the output feature data of the target processing layer.

[0086] In implementation, when the data type of the output feature value is a second data type, the output feature value can be converted according to the data type conversion process from BFloat16 to Float32 described above, resulting in a converted output feature value. The data type of this converted output feature value is a first data type. Then, based on the converted output feature values ​​corresponding to each window, the output feature data of the target processing layer can be generated. For example, the converted output feature values ​​corresponding to each window can be combined to generate the output feature data of the target processing layer. The data type of this output feature data is a first data type.

[0087] Here, the output feature values ​​with the second data type are converted, which can make the output feature data of the target processing layer with higher accuracy based on the converted output feature values ​​corresponding to each window.

[0088] After obtaining the output feature data of each processing layer of the neural network to be trained, the output result of the neural network to be trained can be obtained based on the output feature data of each processing layer; then, the neural network to be trained can be trained based on the output result to generate the target neural network.

[0089] The backpropagation process during neural network training is explained below.

[0090] In one possible implementation, the method further includes:

[0091] Step E1: Determine the output result of the neural network to be trained based on the output feature data of the target processing layer.

[0092] Step E2: Based on the output results and the corresponding ground truth data, determine the loss value of the neural network to be trained.

[0093] Step E3: Based on the loss value, train the neural network to be trained to generate the target neural network.

[0094] In implementation, after obtaining the output feature data of each target processing layer, the output result of the neural network to be trained can be determined based on the output feature data of the target processing layer; and the loss value of the neural network to be trained can be determined based on the output result and the corresponding ground truth data. The ground truth data corresponding to the output result can be manually labeled data; the loss value can be obtained using a loss function, such as the cross-entropy loss function, mean absolute error loss function, mean squared error loss function, squared loss function, etc. The choice of loss function can be determined according to the network task corresponding to the neural network to be trained; for example, the cross-entropy loss function can be used for classification tasks, and the squared loss function can be used for regression tasks. Furthermore, based on the obtained loss value, the neural network to be trained can be trained to generate the target neural network.

[0095] In this method, based on the output feature data of each target processing layer obtained during the forward propagation process, the output result of the neural network to be trained can be determined relatively quickly; and based on the output result and the corresponding ground truth data, the loss value of the neural network to be trained can be determined relatively quickly; furthermore, based on the obtained loss value, the neural network to be trained can be trained to obtain the target neural network, thus improving the efficiency of the training process.

[0096] In one possible implementation, training the neural network to be trained based on the loss value to generate the target neural network includes: determining the gradient value corresponding to each network processing layer in the neural network to be trained based on the loss value; adjusting the network parameters of each network processing layer based on the gradient value corresponding to each network processing layer until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network.

[0097] In practice, after obtaining the loss value corresponding to the neural network to be trained, the gradient value corresponding to each processing layer in the neural network can be determined based on this loss value. For example, the gradient value of the last processing layer in the neural network can be determined based on the loss value, and the gradient value of the previous processing layer can be determined using the gradient value of the last processing layer, and so on, to obtain the gradient value of each processing layer. The data type of this gradient value is the first data type.

[0098] Based on the gradient values ​​corresponding to each network processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thus generating the target neural network. The training cutoff condition may include the accuracy of the neural network being greater than the accuracy threshold; or the number of training iterations of the neural network to be trained being greater than the training threshold; or the loss function value of the adjusted neural network being less than the loss threshold, etc. The accuracy threshold, training threshold, and loss threshold can be determined according to the complexity of the network task corresponding to the neural network to be trained.

[0099] Here, the gradient value corresponding to each processing layer in the neural network to be trained is determined by the loss value. The data type of the gradient value is the first data type, which makes the data precision of the gradient value corresponding to each processing layer of the network high.

[0100] In one possible implementation, the neural network to be trained includes N network processing layers, where N is a positive integer greater than 1; based on the loss value, the gradient value corresponding to each network processing layer in the neural network to be trained is determined, including:

[0101] Step F1: Based on the loss value and the feature data corresponding to the Nth network processing layer, generate the gradient value corresponding to the Nth network processing layer.

[0102] Step F2: Generate the gradient value corresponding to the i-th network processing layer according to the following steps; based on the gradient value corresponding to the (i+1)-th network processing layer and the feature data corresponding to the i-th network processing layer, generate the gradient value corresponding to the i-th network processing layer; where i is a positive integer less than N and greater than or equal to 1.

[0103] In practice, the neural network to be trained includes N network processing layers, where N is a positive integer greater than 1. The N network processing layers can be regarded as the first to the Nth network processing layers according to the processing order of the forward propagation process of the neural network to be trained. Among them, the N network processing layers include the target processing layer and other processing layers other than the target processing layer.

[0104] The gradient value corresponding to the Nth network processing layer can be generated based on the loss value and the feature data corresponding to the Nth network processing layer; that is, during backpropagation, the data output by the Nth network processing layer is the gradient value corresponding to the Nth network processing layer. The feature data corresponding to the network processing layer is the feature data corresponding to that network processing layer during forward propagation. For example, the input feature data of the i-th network processing layer is the output feature data of the (i-1)-th network processing layer during forward propagation.

[0105] Specifically, when a network processing layer includes weight feature data and input feature data, the gradient value corresponding to the network processing layer can include the input feature gradient value corresponding to the input feature data and the weight feature gradient value corresponding to the weight feature data. The input feature gradient value is used for data transmission between network processing layers, and the weight feature gradient value is used to adjust the weight parameters (i.e., weight feature data) of the network processing layer.

[0106] The gradient value corresponding to the i-th network processing layer can be generated according to the following steps: Based on the input feature gradient value corresponding to the (i+1)-th network processing layer and the input feature data corresponding to the i-th network processing layer, the input feature gradient value corresponding to the i-th network processing layer is generated; and based on the input feature gradient value corresponding to the (i+1)-th network processing layer and the weight feature data corresponding to the i-th network processing layer, the weight feature gradient value corresponding to the i-th network processing layer is generated; the data type of the input feature gradient value and the weight feature gradient value is a first data type; where i is a positive integer less than N and greater than or equal to 1.

[0107] In this embodiment of the disclosure, gradient values ​​corresponding to each network processing layer can be generated based on the loss value and the feature data corresponding to the Nth network processing layer, providing data support for subsequent adjustment of the network parameters of each network processing layer.

[0108] In one possible implementation, when the i-th network processing layer is the target processing layer, the gradient value corresponding to the i-th network processing layer is generated based on the gradient value corresponding to the (i+1)-th network processing layer and the feature data corresponding to the i-th network processing layer. This includes: generating the gradient value corresponding to the i-th network processing layer based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer.

[0109] In implementation, if the i-th network processing layer is the target processing layer, the transformed feature data corresponding to the network processing layer can be stored during the forward propagation process. The data type of the transformed feature data is the second data type, which can save memory space. Then, during the back propagation process, the gradient value corresponding to the i-th network processing layer can be generated based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer.

[0110] When the transformed feature data includes transformed input feature data and transformed weight feature data, the input feature gradient value corresponding to the i-th network processing layer can be generated based on the input feature gradient value corresponding to the (i+1)-th network processing layer and the transformed input feature data corresponding to the i-th network processing layer; and the weight feature gradient value corresponding to the i-th network processing layer can be generated based on the input feature gradient value corresponding to the (i+1)-th network processing layer and the transformed weight feature data corresponding to the i-th network processing layer.

[0111] Alternatively, if the i-th network processing layer is the target processing layer, the data type of the feature data corresponding to the network processing layer can be converted to the second data type according to the data type conversion process from Float32 to BFloat16 described above, to obtain the converted feature data; then, the gradient value corresponding to the i-th network processing layer can be generated based on the gradient value corresponding to the (i+1)-th network processing layer and the converted feature data corresponding to the i-th network processing layer.

[0112] Since the data type of the transformed feature data is the second data type, the data precision is low. When using the second data type for computation, the computation efficiency is high. Therefore, when the i-th network processing layer is the target processing layer, the gradient value corresponding to the i+1-th network processing layer can be generated relatively quickly based on the gradient value corresponding to the i+1-th network processing layer and the transformed feature data corresponding to the i-th network processing layer.

[0113] In one possible implementation, the gradient value corresponding to the i-th network processing layer is generated based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer, including:

[0114] Step G1: Perform type conversion on the gradient value corresponding to the (i+1)th network processing layer to obtain the converted gradient value corresponding to the (i+1)th network processing layer; wherein, the data type of the converted gradient value is the second data type.

[0115] Step G2 involves performing a dot product operation on the transformed gradient value corresponding to the (i+1)th network processing layer and the transformed feature data corresponding to the ith network processing layer to obtain the intermediate data corresponding to the ith network processing layer.

[0116] Step G3 involves summing the offset corresponding to the i-th network processing layer and the intermediate values ​​included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer.

[0117] Here, when the i-th network processing layer is the target processing layer, the gradient value corresponding to the (i+1)-th network processing layer can be first converted to obtain the converted gradient value corresponding to the (i+1)-th network processing layer. The data type of the converted gradient value is the second data type, which has low data precision. Then, the converted gradient value corresponding to the (i+1)-th network processing layer and the converted feature data corresponding to the i-th network processing layer are subjected to a dot product operation. This can quickly obtain the intermediate data corresponding to the i-th network processing layer, improve the efficiency of the dot product operation, and thus quickly obtain the gradient value of the i-th network processing layer, thereby improving the training efficiency of the neural network to be trained.

[0118] During implementation, the data type of the gradient values ​​corresponding to each network processing layer is the first data type. Therefore, before generating the gradient value corresponding to the i-th network processing layer, the data type can be converted from Float32 to BFloat16 as described above. The converted gradient value corresponding to the (i+1)-th network processing layer is then processed to obtain the converted gradient value. The data type of the converted gradient value is the second data type.

[0119] In the case that the i-th network processing layer includes input feature data and weight feature data, the transformed gradient value (i.e., the transformed input feature gradient value) corresponding to the (i+1)-th network processing layer and the transformed input feature data corresponding to the i-th network processing layer are multiplied together to obtain the first intermediate data corresponding to the i-th network processing layer. For example, the transformed gradient value corresponding to the (i+1)-th network processing layer is multiplied by each transformed input feature value in the transformed input feature data corresponding to the i-th network processing layer to obtain each first intermediate value. Each first intermediate value constitutes the first intermediate data, and the data type of the first intermediate data is the second data type.

[0120] Furthermore, the transformed gradient value (i.e., the transformed input feature gradient value) corresponding to the (i+1)th network processing layer and the transformed weight feature data corresponding to the ith network processing layer can be multiplied by a dot product to obtain the second intermediate data corresponding to the ith network processing layer. For example, the transformed gradient value corresponding to the (i+1)th network processing layer can be multiplied by each transformed weight feature value in the transformed weight feature data corresponding to the ith network processing layer to obtain each second intermediate value. Each second intermediate value constitutes the second intermediate data, and the data type of the second intermediate data is the second data type.

[0121] In one approach, the accumulation operation can be performed directly based on the second data type. For example, the offset corresponding to the i-th network processing layer and each of the first intermediate values ​​included in the first intermediate data can be accumulated to generate the input feature gradient value corresponding to the i-th network processing layer; the data type of the input feature gradient value is the second data type. Alternatively, the offset corresponding to the i-th network processing layer and each of the second intermediate values ​​included in the second intermediate data can be accumulated to generate the weight feature gradient value corresponding to the i-th network processing layer; the data type of the weight feature gradient value is the second data type.

[0122] Since the gradient values ​​transmitted during backpropagation are of the first data type, a data type conversion is required when the gradient values ​​are of the second data type.

[0123] After accumulating the offset corresponding to the i-th network processing layer and the intermediate values ​​included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer, the method further includes: if the data type of the gradient value is the second data type, performing a type conversion on the gradient value to obtain the converted gradient value; wherein the data type of the converted gradient value is the first data type.

[0124] Based on the gradient values ​​corresponding to each network processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thus generating the target neural network. This includes: based on the transformed gradient value corresponding to the target processing layer and the gradient values ​​corresponding to the other processing layers in each network processing layer except the target processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thus generating the target neural network.

[0125] In implementation, if the data type of the obtained gradient value is a second data type, the gradient value can be converted from BFloat16 to Float32 according to the conversion process described above. The converted gradient value is then of the first data type. Based on the converted gradient value corresponding to the target processing layer and the gradient values ​​corresponding to other processing layers besides the target processing layer, the network parameters of the network processing layers can be adjusted until the adjusted neural network meets the training cutoff condition, thus generating the target neural network. The data type of the gradient values ​​corresponding to other processing layers can be the first data type; the network parameters can include weight feature data and offsets.

[0126] Here, the data type of the gradient values ​​corresponding to each target processing layer is converted to the first data type, which has higher precision. Based on the converted gradient values ​​corresponding to the target processing layer and the gradient values ​​corresponding to other processing layers in each network processing layer except the target processing layer, the network parameters of the network processing layer can be adjusted more accurately, and the generated target neural network ensures the training accuracy of the neural network training process.

[0127] In another approach, accumulation operations can be performed based on the first data type.

[0128] The intermediate data is of the second data type. The offset corresponding to the i-th network processing layer and all intermediate values ​​included in the intermediate data are summed to generate the gradient value corresponding to the i-th network processing layer. Specifically, this includes:

[0129] Step H1 involves converting the types of each intermediate value in the intermediate data to obtain the converted intermediate value; wherein the data type of the converted intermediate value is the first data type.

[0130] Step H2 involves summing the offset corresponding to the i-th network processing layer and the various transformed intermediate values ​​included in the intermediate data to obtain the gradient value corresponding to the i-th network processing layer, wherein the data type of the gradient value is the first data type.

[0131] During implementation, the intermediate data type obtained is the second data type. To ensure that the accuracy loss of the adjusted neural network is low, the data type can be converted from BFloat16 to Float32 according to the above-described conversion process. First, the intermediate values ​​included in the intermediate data are converted to obtain the converted intermediate values. The data type of the converted intermediate values ​​is the first data type. Then, the offset corresponding to the i-th network processing layer and the converted intermediate values ​​included in the intermediate data can be accumulated to obtain the gradient value corresponding to the i-th network processing layer. The data type of the gradient value is the first data type.

[0132] By using the intermediate values ​​after conversion of the first data type for accumulation, the gradient value of the i-th network processing layer can be obtained more accurately. This gradient value can then be used to adjust the parameters of the neural network to be trained more precisely, thus ensuring the training accuracy of the neural network.

[0133] In one possible implementation, the method further includes: acquiring data to be detected; using a target neural network to detect the data to be detected, and obtaining a detection result, wherein the target neural network is trained based on the neural network training method described in the above implementation.

[0134] In implementation, the target neural network trained using the neural network training method described in the above embodiments can be used to detect the acquired detection data to obtain detection results. The detection data is matched to the application scenario of the target neural network.

[0135] In specific implementation, the target neural network can be applied to road recognition scenarios. The application process of the target neural network can include: acquiring road images collected by the driving device during driving; using the target neural network to perform target detection on target objects in the road images to obtain the object detection results corresponding to the road images; and controlling the driving state of the driving device based on the object detection results corresponding to the road images.

[0136] For example, the driving device can be an autonomous vehicle, a vehicle equipped with an Advanced Driving Assistance System (ADAS), or a robot. The road images can be image data collected in real time by the driving device during its operation.

[0137] By utilizing the generated target neural network to detect objects in road images, object detection results are generated corresponding to the road images. For example, the object detection results can include the location and orientation information of each target object in the road image. The target object can be any object to be detected, such as motor vehicles, non-motor vehicles, pedestrians, animals, road signs, etc. Furthermore, based on the object detection results corresponding to the road images, the driving state of the driving device can be controlled. Specifically, when controlling the driving device, it can be controlled to accelerate, decelerate, stop, steer, brake, and avoid objects. For example, avoiding objects can include bypassing objects or changing the driving route; or voice prompts can be played to guide the driver in controlling the driving state of the driving device.

[0138] Here, considering that the resources consumed by performing dot product operations using feature data of the first data type are greater than the resources consumed by performing two data type conversions, the target processing layer of the target neural network improves the processing efficiency of the target neural network while ensuring the high accuracy of the output results by using two data type conversions. Therefore, using the generated target neural network to detect road images results in high accuracy and fast speed of object detection results corresponding to the road images, thereby improving the safety of the driving device.

[0139] In practice, the target neural network can be applied to face recognition scenarios. The application process of the target neural network can include: acquiring the image to be detected; using the target neural network to detect the image to be detected, and obtaining the face detection result corresponding to the image to be detected.

[0140] The image to be detected can be an image acquired by an image acquisition device. This image is input into a target neural network for detection, yielding the corresponding face detection result. This result can include facial identification information (such as name, employee ID, etc.) or face matching information (such as match or non-match). Using the generated target neural network to detect the image allows for faster and more accurate face detection, thus improving the accuracy of face recognition.

[0141] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0142] Based on the same concept, this disclosure also provides a chip, see [link to relevant documentation]. Figure 3 The diagram shown is a schematic representation of the architecture of a chip provided in this embodiment of the present disclosure, including a memory 301 and a computing device 302. This chip can run the training process of a neural network to be trained. The following is an exemplary description of the training process of the neural network running by this chip:

[0143] The memory 301 is used to store the feature data corresponding to the target processing layer in the neural network to be trained.

[0144] The computing device 302 is used to obtain feature data corresponding to the target processing layer from the memory, and perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; and to perform arithmetic processing on the converted feature data corresponding to the target processing layer to generate output feature data of the target processing layer.

[0145] In specific implementation, the chip includes a memory 301 and a computing device 302. The memory can be used to store feature data corresponding to the target processing layer in the neural network to be trained, and the data type of the feature data is a first data type. The computing device can be used to retrieve the feature data corresponding to the target processing layer from the memory and perform type conversion processing on the feature data of the target processing layer to obtain converted feature data. The data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data. The computing device can also be used to perform computational processing on the converted feature data corresponding to the target processing layer to generate the output feature data of the target processing layer.

[0146] Here, during the training process of the neural network to be trained, the data type of the feature data corresponding to the target processing layer in the stored neural network to be trained is the first data type, which can ensure that the accuracy of the target neural network is high; and the data accuracy of the second data type corresponding to the transformed feature data is lower. By performing calculations on the transformed feature data corresponding to the target processing layer, the output feature data of the target processing layer can be generated more quickly, which can improve the training efficiency of the neural network to be trained.

[0147] In one possible implementation, see Figure 4 As shown, the computing device 302 includes a first precision converter 3021, an accumulator 3022, and a dot product 3023; wherein the dot product 3023 is connected to the accumulator 3022 and the first precision converter 3021 respectively.

[0148] The first precision converter 3021 is used to perform type conversion processing on the feature data of the target processing layer to obtain converted feature data. The converted feature data includes converted weight feature data corresponding to the weight feature data and converted input feature data corresponding to the input feature data. The converted weight feature data and the converted input feature data are then sent to the dot product.

[0149] Dot productor 3023 is used to perform dot product operation on local feature data within any window of the transformed input feature data and the transformed weighted feature data to obtain intermediate feature values ​​corresponding to each feature position; and sends the intermediate feature values ​​corresponding to each feature position to the accumulator.

[0150] Accumulator 3022 is used to accumulate the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window; wherein, the output feature values ​​corresponding to each window constitute the output feature data of the target processing layer.

[0151] In implementation, the computing device 302 includes a first precision converter 3021, an accumulator 3022, and a dot productr 3023. See also... Figure 5 The diagram showing the forward propagation process, combined with... Figure 5 The processing procedure of the computing device is illustrated by way of example. The output feature data of the network processing layer preceding the target processing layer is used as the input feature data of the target processing layer. A first precision converter can perform type conversion processing on the input feature data and weight feature data of the target processing layer to obtain converted input feature data and converted weight feature data. The data types of the converted input feature data and converted weight feature data are a second data type. The converted weight feature data and converted input feature data can then be sent to the dot product.

[0152] The dot product can perform a dot product operation between the local feature data within any window of the transformed input feature data and the transformed weighted feature data to obtain the intermediate feature values ​​corresponding to each feature position; the data type of the intermediate feature values ​​is the second data type; and the intermediate feature values ​​corresponding to each feature position can be sent to the accumulator.

[0153] When the accumulator processes data of the second data type, it can perform an accumulation operation on the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window. Since the data type of the output feature value is the second data type, a second precision converter is needed to convert the data type of the output feature value to obtain the converted output feature value, which is of the first data type (i.e., Float32). The converted output feature value is then used to construct the output feature data, which is of the first data type.

[0154] When the data type processed by the accumulator is the first data type, before performing the accumulation operation, the data type of the intermediate feature values ​​corresponding to each feature position can be converted to the first data type. Then, the offset corresponding to the target processing layer and the converted intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to any window. The data type of the output feature value is the first data type. At this time, there is no need to use the second precision converter for type conversion. The output feature values ​​corresponding to each window are directly used to form the output feature data of the target processing layer. The data type of this output feature data is the first data type.

[0155] Here, before sending the feature data of the target processing layer to the dot product, the feature data of the first data type with higher data precision is first converted into the transformed feature data of the second data type with lower data precision using the first precision converter. Then, the dot product operation is performed on the transformed feature data, which can quickly obtain the intermediate feature values ​​corresponding to each feature position and improve the computational efficiency.

[0156] In one possible implementation, see Figure 6 As shown, when the data type of the output feature value is the second data type, the computing device 302 also includes a second precision converter 3024;

[0157] The second precision converter 2014 is used to perform type conversion processing on the output feature values ​​to obtain the converted output feature values; wherein, each converted output feature value constitutes the output feature data of the target processing layer.

[0158] In implementation, when the data type of the output feature value is a second data type, the computing device 302 may further include a second precision converter 3024, see [link to relevant documentation]. Figure 5 As shown, the second precision converter can first perform type conversion processing on the output feature values ​​to obtain the converted output feature values; wherein, the data type of each converted output feature value is the first data type, so the data type of the output feature data of the target processing layer is the first data type; and then the output feature data with the first data type is sent to the next network processing layer.

[0159] Here, when the data type of the output feature value is the second data type, the data type of the output feature value is first converted to the first data type using the second precision converter, and then the converted feature value is used to form the output feature data. The output feature data has high data precision, which ensures the precision of the trained neural network when the output feature data is used to train the neural network.

[0160] In practice, the loss value of the neural network to be trained can be generated based on the output feature data of each network processing layer. Then, the gradient value corresponding to each network processing layer can be generated based on the loss value and the feature data of each network processing layer. During backpropagation, if the current processing layer is the i-th network processing layer, the data input to the i-th network processing layer is the input feature gradient value output by the (i+1)-th network processing layer; if the current processing layer is the N-th network processing layer, the data input to the N-th network processing layer is the loss value of the neural network to be trained.

[0161] See Figure 7a The diagram shown illustrates the process of obtaining the input feature gradient values ​​during backpropagation. Figure 7a The processing procedure of the computing device is illustrated by example. For the i-th network processing layer, if the i-th network processing layer is the target processing layer, a first precision converter can be used to convert the data type of the input feature gradient value of the (i+1)-th network processing layer to a second data type. Then, the converted input feature gradient value and the converted input feature data of the i-th network processing layer are subjected to a dot product operation to obtain the first intermediate data. When the data type processed by the accumulator is the second data type, each first intermediate value in the first intermediate data can be accumulated with the offset corresponding to the i-th network processing layer to obtain the input feature gradient value of the i-th network processing layer; the data type of this input feature gradient value is the second data type (i.e., BFloat16). At this time, a second precision converter is needed to perform a type conversion process on the input feature gradient value to generate the input feature gradient value of the first data type.

[0162] When the accumulator processes data of the first data type, each first intermediate value in the first intermediate data can be converted to a new data type to generate a converted first intermediate value. Then, the accumulator is used to sum each converted first intermediate value with the offset corresponding to the i-th network processing layer to obtain the input feature gradient value of the i-th network processing layer. Since the input feature gradient value is of the first data type, there is no need to use a second precision converter for data type conversion.

[0163] And, see also Figure 7b The diagram shown illustrates the process of obtaining the gradient values ​​of the weight features during backpropagation. Figure 7b The processing procedure of the computing device is illustrated by way of example. A first precision converter can be used to convert the data type of the input feature gradient value of the (i+1)th network processing layer to a second data type. Then, the converted input feature gradient value and the converted weight feature data of the i-th network processing layer are subjected to a dot product operation to obtain the second intermediate data.

[0164] When the accumulator processes data of the second data type, it can accumulate each second intermediate value in the second intermediate data with the offset corresponding to the i-th network processing layer to obtain the weight feature gradient value of the i-th network processing layer. The data type of this weight feature gradient value is the second data type (i.e., BFloat16). At this time, a second precision converter is needed to perform type conversion processing on the weight feature gradient value to generate the weight feature gradient value of the first data type.

[0165] When the accumulator processes data of the first data type, the types of each second intermediate value in the second intermediate data can be converted to generate the converted second intermediate value of the first data type. Then, the accumulator is used to accumulate each converted second intermediate value in the second intermediate data with the offset corresponding to the i-th network processing layer to obtain the weight feature gradient value of the i-th network processing layer. At this time, the weight feature gradient value is of the first data type, so there is no need to use the second precision converter for data type conversion.

[0166] Finally, the gradient values ​​of the weight features corresponding to each network processing layer can be used to adjust the network parameters of the neural network to be trained until the adjusted neural network meets the training cutoff condition and the target neural network is generated.

[0167] Based on the same concept, this disclosure also provides a neural network training apparatus, see [link to relevant documentation]. Figure 8 The diagram shown is a schematic representation of the neural network training architecture provided in this embodiment of the present disclosure, including an acquisition module 801, a first processing module 802, and a second processing module 803. Specifically:

[0168] The acquisition module 801 is used to acquire feature data corresponding to any target processing layer in the neural network to be trained during the training process.

[0169] The first processing module 802 is used to perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data.

[0170] The second processing module 803 is used to perform calculations on the transformed feature data corresponding to the target processing layer to generate the output feature data of the target processing layer.

[0171] In one possible implementation, the feature data includes weighted feature data and input feature data, and the transformed feature data includes transformed weighted feature data and transformed input feature data;

[0172] The second processing module 803, when performing calculations on the transformed feature data corresponding to the target processing layer to generate the output feature data of the target processing layer, is used to:

[0173] The local feature data within any window of the transformed input feature data is multiplied by the transformed weighted feature data to obtain the intermediate feature values ​​corresponding to each feature position.

[0174] The offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to any window.

[0175] Based on the output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated.

[0176] In one possible implementation, the data type of the intermediate feature value is a second data type, and the second processing module 803, when performing an accumulation operation on the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window, is used to:

[0177] The intermediate feature values ​​corresponding to each feature position are subjected to type conversion processing to obtain converted intermediate feature values; wherein, the data type of the converted intermediate feature values ​​is the first data type;

[0178] The offset corresponding to the target processing layer and the transformed intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to any window, wherein the data type of the output feature value is the first data type.

[0179] In one possible implementation, the second processing module 803, when performing type conversion processing on the intermediate feature values ​​corresponding to each feature position to obtain converted intermediate feature values, is used to:

[0180] Based on the network task and / or network structure information corresponding to the neural network to be trained, determine the data processing type corresponding to the accumulation operation in the target processing layer;

[0181] In response to the data processing type indication being the first data type, the intermediate feature values ​​corresponding to each feature position are subjected to type conversion processing to obtain the converted intermediate feature values.

[0182] In one possible implementation, the second processing module 803, when generating the output feature data of the target processing layer based on the output feature values ​​corresponding to each window, is used to:

[0183] If the data type of the output feature value is the second data type, the output feature value is subjected to type conversion to obtain a converted output feature value; wherein, the data type of the converted output feature value is the first data type;

[0184] Based on the transformed output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated.

[0185] In one possible implementation, the apparatus further includes: a generation module 804, the generation module 804 being configured to:

[0186] Based on the output feature data of the target processing layer, the output result of the neural network to be trained is determined;

[0187] Based on the output results and the corresponding ground truth data, the loss value of the neural network to be trained is determined.

[0188] Based on the loss value, the neural network to be trained is trained to generate the target neural network.

[0189] In one possible implementation, the generation module 804, when training the neural network to be trained based on the loss value to generate the target neural network, is used to:

[0190] Based on the loss value, determine the gradient value corresponding to each processing layer in the neural network to be trained;

[0191] Based on the gradient values ​​corresponding to each network processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network.

[0192] In one possible implementation, the neural network to be trained includes N network processing layers, where N is a positive integer greater than 1; the generation module 804, when determining the gradient value corresponding to each network processing layer in the neural network to be trained based on the loss value, is used to:

[0193] Based on the loss value and the feature data corresponding to the Nth network processing layer, the gradient value corresponding to the Nth network processing layer is generated;

[0194] Generate the gradient value corresponding to the i-th network processing layer according to the following steps:

[0195] Based on the gradient value corresponding to the (i+1)th network processing layer and the feature data corresponding to the ith network processing layer, the gradient value corresponding to the ith network processing layer is generated; where i is a positive integer less than N and greater than or equal to 1.

[0196] In one possible implementation, when the i-th network processing layer is the target processing layer, the generation module 804, when generating the gradient value corresponding to the i-th network processing layer based on the gradient value corresponding to the (i+1)-th network processing layer and the feature data corresponding to the i-th network processing layer, is used to:

[0197] The gradient value corresponding to the (i+1)th network processing layer and the transformed feature data corresponding to the i-th network processing layer are used to generate the gradient value corresponding to the i-th network processing layer.

[0198] In one possible implementation, the generation module 804, when generating the gradient value corresponding to the i-th network processing layer based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer, is used to:

[0199] The gradient value corresponding to the (i+1)th network processing layer is converted to a different type to obtain the converted gradient value corresponding to the (i+1)th network processing layer; wherein the data type of the converted gradient value is the second data type.

[0200] Perform a dot product operation on the transformed gradient value corresponding to the (i+1)th network processing layer and the transformed feature data corresponding to the ith network processing layer to obtain the intermediate data corresponding to the ith network processing layer.

[0201] The offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data are accumulated to generate the gradient value corresponding to the i-th network processing layer.

[0202] In one possible implementation, the data type of the intermediate data is a second data type. The generation module 804, when performing an accumulation operation on the offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer, is used to:

[0203] The intermediate values ​​included in the intermediate data are converted to obtain the converted intermediate values; wherein the data type of the converted intermediate values ​​is the first data type.

[0204] The offset corresponding to the i-th network processing layer and each of the transformed intermediate values ​​included in the intermediate data are accumulated to obtain the gradient value corresponding to the i-th network processing layer, wherein the data type of the gradient value is the first data type.

[0205] In one possible implementation, the apparatus further includes: a third processing module 805, which, after performing an accumulation operation on the offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer, is further configured to:

[0206] If the data type of the gradient value is the second data type, the gradient value is converted to obtain a converted gradient value; wherein, the data type of the converted gradient value is the first data type.

[0207] The generation module 804, when adjusting the network parameters of the network processing layers based on the gradient values ​​corresponding to each network processing layer until the adjusted neural network meets the training cutoff condition and generating the target neural network, is used for:

[0208] Based on the transformed gradient value corresponding to the target processing layer and the gradient values ​​corresponding to the other processing layers in each network processing layer besides the target processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network.

[0209] In one possible implementation, the device further includes a detection module 806, the detection module 806 being configured to:

[0210] Acquire the data to be tested;

[0211] The target neural network is used to detect the data to be detected to obtain the detection result, wherein the target neural network is trained based on the neural network training method described in the first aspect or any embodiment above.

[0212] In some embodiments, the functions or templates of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0213] Based on the same technical concept, this disclosure also provides an electronic device. (See also...) Figure 9 The diagram shows the structure of an electronic device 900 provided in this embodiment of the present disclosure, including a processor 901, a memory 902, and a bus 903. The memory 902 stores execution instructions and includes a main memory 9021 and an external memory 9022. The main memory 9021, also called internal memory, is used to temporarily store computational data in the processor 901 and data exchanged with external memory 9022 such as a hard disk. The processor 901 exchanges data with the external memory 9022 through the main memory 9021. When the electronic device 900 is running, the processor 901 and the memory 902 communicate through the bus 903, causing the processor 901 to execute the following instructions:

[0214] During the training process of the neural network to be trained, feature data corresponding to any target processing layer in the neural network to be trained is obtained;

[0215] The feature data of the target processing layer is subjected to type conversion processing to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data;

[0216] The transformed feature data corresponding to the target processing layer is processed to generate the output feature data of the target processing layer.

[0217] The specific processing flow of the processor 901 can be referred to the description in the above method embodiment, and will not be repeated here.

[0218] Furthermore, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the neural network training method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0219] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the neural network training method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0220] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0221] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0222] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0223] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0224] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0225] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A neural network training method, characterized in that, include: During the training process of the neural network to be trained, feature data corresponding to any target processing layer in the neural network to be trained is obtained; The neural network to be trained is a face recognition network; The feature data of the target processing layer is subjected to type conversion processing to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; the feature data includes weight feature data and input feature data, and the converted feature data includes converted weight feature data and converted input feature data. For local feature data within any window of the transformed input feature data, the local feature data within the window is multiplied by the transformed weighted feature data to obtain intermediate feature values ​​corresponding to each feature position; the data type of the intermediate feature values ​​is a second data type. The offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to the window; the data processing type corresponding to the accumulation operation is determined based on the network task and / or network structure information corresponding to the neural network to be trained. Based on the output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated. Based on the output feature data of the target processing layer, the output result of the neural network to be trained is determined; Based on the output results and the corresponding ground truth data, the loss value of the neural network to be trained is determined. Based on the loss value, determine the gradient value corresponding to each processing layer in the neural network to be trained; Based on the gradient values ​​corresponding to each network processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network; the target neural network is used to detect the data to be detected and obtain the face detection result; the data to be detected is the image acquired by the image acquisition device.

2. The method according to claim 1, characterized in that, The step of accumulating the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to the window includes: The intermediate feature values ​​corresponding to each feature position are converted to obtain the converted intermediate feature values. The offset corresponding to the target processing layer and the transformed intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to the window. Wherein, the data type of the intermediate feature value after conversion and the data type of the output feature value are the first data type.

3. The method according to claim 2, characterized in that, The step of performing type conversion processing on the intermediate feature values ​​corresponding to each feature position to obtain converted intermediate feature values ​​includes: In response to the data processing type indication being the first data type, the intermediate feature values ​​corresponding to each feature position are subjected to type conversion processing to obtain the converted intermediate feature values.

4. The method according to any one of claims 1 to 3, characterized in that, The step of generating the output feature data of the target processing layer based on the output feature values ​​corresponding to each window includes: If the data type of the output feature value is the second data type, the output feature value is subjected to type conversion to obtain a converted output feature value; wherein, the data type of the converted output feature value is the first data type; Based on the transformed output feature values ​​corresponding to each window, the output feature data of the target processing layer is generated.

5. The method according to claim 4, characterized in that, The neural network to be trained includes N network processing layers, where N is an integer greater than 1; determining the gradient value corresponding to each network processing layer in the neural network to be trained based on the loss value includes: Based on the loss value and the feature data corresponding to the Nth network processing layer, the gradient value corresponding to the Nth network processing layer is generated; Generate the gradient value corresponding to the i-th network processing layer according to the following steps: If the i-th network processing layer does not belong to the target processing layer, the gradient value corresponding to the i-th network processing layer is generated based on the gradient value corresponding to the (i+1)-th network processing layer and the feature data corresponding to the i-th network processing layer; where i is an integer less than N and greater than or equal to 1. If the i-th network processing layer belongs to the target processing layer, the gradient value corresponding to the i-th network processing layer is generated based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer.

6. The method according to claim 5, characterized in that, The step of generating the gradient value corresponding to the i-th network processing layer based on the gradient value corresponding to the (i+1)-th network processing layer and the transformed feature data corresponding to the i-th network processing layer includes: The gradient value corresponding to the (i+1)th network processing layer is converted to a different type to obtain the converted gradient value corresponding to the (i+1)th network processing layer; wherein the data type of the converted gradient value is the second data type. Perform a dot product operation on the transformed gradient value corresponding to the (i+1)th network processing layer and the transformed feature data corresponding to the ith network processing layer to obtain the intermediate data corresponding to the ith network processing layer. The offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data are accumulated to generate the gradient value corresponding to the i-th network processing layer.

7. The method according to claim 6, characterized in that, The intermediate data is of a second data type. The step of accumulating the offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer includes: The intermediate values ​​included in the intermediate data are converted to obtain the converted intermediate values; wherein the data type of the converted intermediate values ​​is the first data type. The offset corresponding to the i-th network processing layer and each of the transformed intermediate values ​​included in the intermediate data are accumulated to obtain the gradient value corresponding to the i-th network processing layer, wherein the data type of the gradient value is the first data type.

8. The method according to claim 6, characterized in that, After performing the summation operation on the offset corresponding to the i-th network processing layer and each intermediate value included in the intermediate data to generate the gradient value corresponding to the i-th network processing layer, the method further includes: If the data type of the gradient value is the second data type, the gradient value is converted to obtain a converted gradient value; wherein, the data type of the converted gradient value is the first data type. The step of adjusting the network parameters of each network processing layer based on the gradient values ​​of each processing layer until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network, includes: Based on the transformed gradient value corresponding to the target processing layer and the gradient values ​​corresponding to the other processing layers in each network processing layer besides the target processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network.

9. A chip, characterized in that, The chip includes a memory and a computing device; The memory is used to store feature data corresponding to the target processing layer in the neural network to be trained; the neural network to be trained is a face recognition network. The computing device is configured to retrieve the feature data corresponding to the target processing layer from the memory, and perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; the feature data includes weight feature data and input feature data, and the converted feature data includes converted weight feature data and converted input feature data; and for local feature data within any window of the converted input feature data, perform a dot product operation on the local feature data within the window and the converted weight feature data to obtain intermediate feature values ​​corresponding to each feature position; the data type of the intermediate feature values ​​is the second data type; and perform an accumulation operation on the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position. The output feature value corresponding to the window is obtained; the data processing type corresponding to the accumulation operation is determined based on the network task and / or network structure information corresponding to the neural network to be trained; based on the output feature value corresponding to each window, the output feature data of the target processing layer is generated; based on the output feature data of the target processing layer, the output result of the neural network to be trained is determined; based on the output result and the ground truth data corresponding to the output result, the loss value of the neural network to be trained is determined; based on the loss value, the gradient value corresponding to each network processing layer in the neural network to be trained is determined; based on the gradient value corresponding to each network processing layer, the network parameters of the network processing layer are adjusted until the adjusted neural network meets the training cutoff condition, and the target neural network is generated; the target neural network is used to detect the data to be detected and obtain the face detection result; the data to be detected is the image acquired by the image acquisition device.

10. The chip according to claim 9, characterized in that, The computing device includes a first precision converter, an accumulator, and a dot product; wherein the dot product is connected to the accumulator and the precision converter respectively. The first precision converter is used to perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; the converted feature data includes converted weight feature data corresponding to the weight feature data and converted input feature data corresponding to the input feature data; and sends the converted weight feature data and the converted input feature data to the dot product. The dot productr is used to perform a dot product operation between the local feature data within any window of the transformed input feature data and the transformed weighted feature data to obtain the intermediate feature values ​​corresponding to each feature position; and to send the intermediate feature values ​​corresponding to each feature position to the accumulator. The accumulator is used to perform an accumulation operation on the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position to obtain the output feature value corresponding to any window; wherein, the output feature values ​​corresponding to each window constitute the output feature data of the target processing layer. When the data type of the output feature value is the second data type, the computing device further includes a second precision converter for performing type conversion processing on the output feature value to obtain a converted output feature value; wherein, each of the converted output feature values ​​constitutes the output feature data of the target processing layer.

11. A neural network training device, characterized in that, include: The acquisition module is used to acquire feature data corresponding to any target processing layer in the neural network to be trained during the training process; the neural network to be trained is a face recognition network. The first processing module is used to perform type conversion processing on the feature data of the target processing layer to obtain converted feature data; wherein, the data precision of the first data type corresponding to the feature data is higher than the data precision of the second data type corresponding to the converted feature data; the feature data includes weight feature data and input feature data, and the converted feature data includes converted weight feature data and converted input feature data. The second processing module is used to perform a dot product operation on the local feature data within any window of the transformed input feature data, and the transformed weighted feature data to obtain intermediate feature values ​​corresponding to each feature position; the data type of the intermediate feature values ​​is a second data type; the offset corresponding to the target processing layer and the intermediate feature values ​​corresponding to each feature position are accumulated to obtain the output feature value corresponding to the window; the data processing type corresponding to the accumulation operation is determined based on the network task and / or network structure information corresponding to the neural network to be trained; and the output feature data of the target processing layer is generated based on the output feature values ​​corresponding to each window. The generation module is used to determine the output result of the neural network to be trained based on the output feature data of the target processing layer; determine the loss value of the neural network to be trained based on the output result and the corresponding ground truth data; determine the gradient value corresponding to each network processing layer in the neural network to be trained based on the loss value; adjust the network parameters of the network processing layers based on the gradient values ​​corresponding to each network processing layer until the adjusted neural network meets the training cutoff condition, thereby generating the target neural network; the target neural network is used to detect the data to be detected to obtain the face detection result; the data to be detected is an image acquired by an image acquisition device.

12. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the neural network training method as described in any one of claims 1 to 8; or the chip as described in claim 9 or 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the neural network training method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Neural network training method and device, equipment and computer readable storage medium

    CN113435520A