Object classification method, device, equipment and storage medium
By introducing linear normalized functions as activation functions in object classification networks, the problems of gradient disappearance and internal covariate drift in artificial neural networks are solved, and classification accuracy and training accuracy are improved.
Patent Information
- Application Number
- CN202110821659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-07-20
AI Technical Summary
In the prior art, artificial neural networks are prone to parameter gradient disappearance and internal covariate drift problems during training, resulting in poor classification accuracy and unable to meet the classification needs of users.
A linear normalization function is introduced as an activation function, and the target neurons in the object classification network are activated by determining the linear normalization coefficient based on the target constant value and the number of normalized feature channels.
The problems of gradient vanishing and internal covariate drift are solved, the training accuracy and classification accuracy of the object classification network are improved, and the classification needs of users are met.
Smart Images

Figure CN114925736B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an object classification method, device, equipment and storage medium. Background Art
[0002] With the continuous development of computer technology, deep learning technology has been more and more widely used in image processing, image recognition, object classification and data processing. For example, artificial neural networks (ANN) are usually used to classify images, audio, etc. Artificial neural networks are composed of a large number of interconnected neurons. The output of each neuron is processed by a specific output function, which is an activation function.
[0003] In the prior art, the role of the activation function is to enhance the learning ability of the artificial neural network so that the artificial neural network can adapt to various data processing scenarios. However, when training the artificial neural network with the activation function, the shallow neurons of the artificial neural network are prone to parameter gradient vanishing during the parameter training process, which makes it difficult to fully train the shallow network; and the parameter training of the artificial neural network due to the interference of inter-layer traffic has internal covariate drift problems, which affects the parameter learning efficiency and accuracy of the artificial neural network. Therefore, when classifying objects based on the above artificial neural network, the classification accuracy is poor and cannot meet the classification needs of users. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an object classification method, apparatus, device and storage medium to realize activation based on a linear normalization function in an object classification network, improve the training accuracy of the object classification network, and improve the classification accuracy when using the trained object classification network to perform object classification, thereby meeting the classification needs of users.
[0005] The present disclosure provides an object classification method, the method comprising:
[0006] Get the object to be classified;
[0007] An object to be classified is classified based on an object classification network to obtain a classification result of the object to be classified, at least one target neuron in the object classification network is activated based on a linear normalization function, an output value of the linear normalization function is the product of an input value of the linear normalization function and a linear normalization coefficient, and the linear normalization coefficient is determined according to a target constant value and the number of normalized feature channels.
[0008] The present disclosure provides an object classification device, the device comprising:
[0009] An object acquisition module is used to acquire objects to be classified;
[0010] The object classification module is used to classify the object to be classified based on the object classification network to obtain the classification result of the object to be classified. At least one target neuron in the object classification network is activated based on a linear normalization function. The output value of the linear normalization function is the product of the input value of the linear normalization function and the linear normalization coefficient. The linear normalization coefficient is determined according to the target constant value and the number of normalized feature channels.
[0011] An embodiment of the present invention further provides an object classification device, the device comprising:
[0012] one or more processors;
[0013] A storage device for storing one or more programs;
[0014] When one or more programs are executed by one or more processors, the one or more processors implement the object classification method provided by any embodiment of the present invention.
[0015] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the object classification method provided by any embodiment of the present invention is implemented.
[0016] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0017] The object classification method, device, equipment and storage medium provided by the embodiments of the present disclosure can classify the object to be classified based on the object classification network after obtaining the object to be classified, and obtain the classification result of the object to be classified. Since the object classification network activates the characteristic flow between each neuron based on the linear normalization function of the target neuron, the derivative of the linear normalization function is a constant, and the output feature expression of the weighted input data will not be weakened. The method of linear normalization activation function eliminates the problem of saturation region in the distribution of derivative values inherent in traditional nonlinear activation functions to a certain extent, and can solve the problem of gradient vanishing of shallow neurons in the process of training artificial neural networks, thereby alleviating the problem that the low-level feature parameters of the object classification network cannot be fully trained due to gradient vanishing, and improving the training accuracy of the object classification network. At the same time, since the linear normalization coefficient is determined according to the target constant value and the number of normalized feature channels, the target constant value and the number of normalized feature channels in the linear normalization function in each hidden layer of the neuron make the amplification effect of the hidden layer multiple feature channels of the object classification network on the feature or error transmission normalized, thereby reducing the influence of the small adjustment changes of the parameters of the neurons in each hidden layer of the object classification network during the parameter training and adjustment period on the feature flow changes and error flow changes finally received by the target neuron. Further, the feature extraction gain of the object classification network for multiple different input features can be adjusted to a state of approximately one time by adjusting the target constant value. This activation method further reduces the influence of the small adjustment changes of the parameters of the neurons in each hidden layer of the object classification network during the parameter training and adjustment period on the feature flow changes finally received by the target neuron, thereby alleviating the internal covariate drift problem encountered by the traditional neural network, thereby ensuring the accuracy of the learning gradient of the neuron parameters and improving the training accuracy of the object classification network. In summary, the object classification network that introduces the above linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0020] Figure 1It is a logical schematic diagram of an artificial neural network provided by the prior art;
[0021] Figure 2 is a flow chart of an object classification method provided by an embodiment of the present disclosure;
[0022] Figure 3 is a logical schematic diagram of an object classification network provided by an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of the structure of an object classification network provided by an embodiment of the present disclosure;
[0024] Figure 5 is a structural schematic diagram of an object classification device provided by an embodiment of the present disclosure;
[0025] Figure 6 It is a structural diagram of an object classification device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0028] In related technologies, artificial neural networks are used to classify objects. Figure 1 The figure shows a logic diagram of an artificial neural network provided by the prior art. Figure 1 ,The artificial neural network is composed of three functional layers including input layer, hidden layer and output layer.
[0029] The main function of the input layer is to convert the original object to be classified into one or more data vectors. The data vector is a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values, and can reflect the differentiation of the object. The data vector is the feature data vector to be classified. The input layer passes each feature data vector to be classified to the hidden layer through one or more data channels. Different feature data vectors to be classified of the object to be classified correspond to different data channels. For example, if the object to be classified is a color image, the input layer converts the color image into three feature data vectors to be classified, namely red, green and blue, and passes the above three feature data vectors to be classified to the hidden layer through three data channels. Optionally, the feature data vectors to be classified can be composed of a sequence of feature data vectors to be classified in time sequence and passed to the hidden layer in time sequence.
[0030] The hidden layer of the artificial neural network consists of one hidden layer network, or multiple hidden layer networks connected in stages. A hidden layer network consists of one hidden layer, or multiple hidden layers connected in stages. The hidden layer is connected to the input layer by accessing the feature data vector to be classified of the input layer, and the hidden layer is connected to the previous hidden layer by accessing the feature data vector of the previous hidden layer. Optionally, one hidden layer can be connected to multiple different types of neural network layers at the same time, and different types of neural network layers include input layer, other hidden layers, a certain time series state of this hidden layer, a certain time series state of other hidden layers, etc. Optionally, one hidden layer can be connected to multiple different hidden layer networks at the same time.
[0031] Each hidden layer of the artificial neural network may include multiple feature channels, and each feature channel may correspond to one or more neurons. Neurons refer to entities that extract features from one or more input feature data vectors according to corresponding set weights. The output data of neurons in the hidden layer can be transmitted to the next hidden layer, the output layer, the pooling layer of the current hidden layer, the next time series of the current hidden layer, or the previous time series of the current hidden layer. Different feature channels of the hidden layer refer to neuron feature extraction entities facing all position areas of one or more input feature data vectors with independent neuron weight parameters. Different neurons in the same feature channel are entities that extract features from different position areas of one or more input feature data vectors. Different feature channels of the same hidden layer use the same neuron detection size to extract features from the same one or more input feature data vectors, and obtain independent neuron output data for each feature channel. The feature channel neuron output data forms or forms independent feature channel feature data vectors for each feature channel after pooling through the pooling layer. The feature channel feature data vector is a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values.
[0032] Among them, the types of feature channels include feature channels of single neurons and feature channels of convolutional neural network convolution kernel (or feature map, filter) type of multiple neurons. For feature channels containing single neurons, neurons of the feature channel perform feature extraction on one or more input feature data vectors to obtain a feature channel neuron output data; for feature channels of convolutional neural network convolution kernel type containing multiple neurons, each neuron in the feature channel performs feature extraction on a position area in one or more input feature data vectors to obtain feature channel neuron output data of the corresponding position area of the corresponding feature channel. Feature extraction of neurons refers to that neurons perform weight multiplication and aggregation processing on one or more input feature data vectors according to corresponding set weights to obtain weighted input data, and then the activation function of the neuron performs activation processing on the weighted input data to obtain output data (or output feature value, activation value) of the neuron. Optionally, the feature extraction process of neurons also includes setting bias values in the weighted input data. Optionally, the feature extraction process of neurons also includes numerical normalization processing (such as local response normalization, batch normalization, etc.) on the weighted input data or the activation value obtained by the activation function.
[0033] The output layer of the artificial neural network is connected to the hidden layer, and the output layer includes one or more neurons corresponding to different classification categories, namely, category neurons. The category neurons of the output layer extract category features from the feature data vector output by the hidden layer. Category feature extraction of category neurons refers to the category neurons performing weight multiplication calculations on the data in the feature data vector output by the hidden layer according to the corresponding set weights and summarizing them to obtain weighted input data. The weighted input data of the category neurons of the output layer are the category feature data (or category feature values) corresponding to the category neurons of the output layer. Furthermore, the output layer includes one or more category classifiers, and the category classifiers of the output layer calculate the probability value that the object to be classified belongs to the classification category corresponding to the output layer neurons based on the category feature data. The output layer outputs the classification category and the probability value to obtain the classification result. Optionally, the category feature extraction process of the category neurons also includes setting a bias value in the weighted input data.
[0034] The weight parameters in the artificial neural network are set after the artificial neural network training is completed. When training the artificial neural network, according to the overall error function of the artificial neural network and the current state of the relevant variables, the current partial derivative of the overall error function to the output variable of the neuron is calculated as the output error of the neuron, and then the parameter gradient value of the neuron is calculated according to the output error of the neuron, the derivative of the neuron activation function and the input characteristic value of the neuron, and the parameters of the neuron are adjusted based on the parameter gradient value.
[0035] When training an artificial neural network, in order to prevent the output of each hidden layer neuron from having too great an impact on the feature processing of the next hidden layer, a nonlinear activation function is set for each neuron in the hidden layer of the artificial neural network, and the nonlinear activation function is used to activate the weighted input data so that the output features are maintained within the numerical range allowed by the nonlinear activation function. However, when a nonlinear activation function (such as a Sigmoid function) is used to perform feature activation on weighted input data, after the nonlinear activation function is used to derive the input, the nonlinear activation function causes a saturation problem in the distribution of the weighted input data. For weighted inputs far from the center point, the output feature expression is weak, resulting in the smaller the error of the neuron becomes as it is transmitted to the shallower layer, so that the parameter gradient value calculated based on the error value of the neuron becomes very small as the number of neuron layers decreases, resulting in the problem of gradient vanishing of shallow neurons. In addition, as described above, the parameter learning function is to adjust the parameters of the neuron based on the parameter gradient value after calculating the parameter gradient value according to the output error of the neuron, the derivative of the neuron activation function and the input characteristic value of the neuron. At this time, when adjusting the parameters of the neuron, since the parameters of each neuron in the neural network are adjusted and different neurons are related to each other, the changes in the characteristic values and the changes in the transmission errors caused by the parameter adjustment are superimposed layer by layer on each layer of neurons, resulting in the actual gradient of the current parameter of the neuron changing compared with the calculated value, resulting in the problem of internal covariate shift, which affects the accuracy of parameter training. Obviously, the use of traditional nonlinear activation functions affects the training accuracy of artificial neural networks, and then affects the classification accuracy when classifying objects based on the trained artificial neural networks, and cannot meet the classification needs of users.
[0036] In the related traditional technologies, the nonlinear activation function in the neural network also adopts the modified linear activation function (Relu) activation function method, batch normalization method and gradient clipping method. However, although the use of Relu activation solves the gradient vanishing problem of deep neural networks to a certain extent, the characteristic that the derivative of Relu is zero in the negative interval affects the error backward propagation to a certain extent, and Relu does not solve the problem of internal covariate offset caused by inter-layer flow interference; the disadvantage of using batch normalization method is that when the number of samples is small, the mean and variance of the samples in a batch are greatly different from the overall samples, which will reduce the training accuracy; at the same time, the batch normalization method requires complex normalization calculations and backward error propagation calculations for the training samples, and the amount of calculation in the training process will be greatly increased; at the same time, the batch normalization method is not suitable for recurrent neural network scenarios, that is, the scope of application is small. The disadvantage of using the gradient clipping method is that the idea of the gradient clipping method is to set a clipping threshold. If the gradient exceeds the clipping threshold when updating the gradient, it will be forced to be limited within this range. It has a certain inhibitory effect on the gradient explosion phenomenon, but has no effect on the gradient vanishing phenomenon, and because the gradient is changed, it will also affect the training accuracy to a certain extent.
[0037] In order to solve the above problems, the embodiments of the present disclosure provide an object classification method, apparatus, device and storage medium, which can improve the training accuracy of the object classification network based on the linear normalization function mechanism, and improve the classification accuracy when using the trained object classification network to classify objects, so as to meet the object classification needs of users.
[0038] Below, the object classification method provided by the embodiment of the present disclosure is first described.
[0039] The object classification method provided in this embodiment can be applied to classify objects based on a linear normalization function mechanism. The method can be performed by an object classification device, which can be implemented by software and / or hardware, and the device can be integrated in a device with an object classification function, such as a desktop computer or a server. Figure 2 The method of this embodiment specifically includes the following steps:
[0040] S110: Obtain an object to be classified.
[0041] In the embodiment of the present disclosure, the object classification device can obtain the object to be classified that needs to be classified. The object classification device can obtain the object to be classified by collecting, downloading, uploading, inputting, copying, etc. the object to be classified.
[0042] In some embodiments, the device to be classified may collect data from other devices or equipment that perform data communication with the object to be classified, and use the collected data as the object to be classified.
[0043] In other embodiments, the device to be classified may download data from an open source data set or an open source website, and use the downloaded data as the object to be classified.
[0044] In some other embodiments, the device to be classified can obtain external input data and use the external input data as the object to be classified.
[0045] In some further embodiments, the device to be classified may copy data pre-existing in a data storage device or a data block, and use the copied data as the object to be classified.
[0046] Optionally, the objects to be classified may include images, videos, audios, texts, dynamic images, data vectors, etc., which are not limited here.
[0047] When the object to be classified is an image, a video or a dynamic image, the image may include at least one target object, such as a person, an animal, a scene, an object, etc.
[0048] When the object to be classified is audio, the audio may include at least one audio segment and the like.
[0049] When the object to be classified is text, the text may include text information formed based on at least one font.
[0050] S120 , classifying the object to be classified based on the object classification network to obtain a classification result of the object to be classified.
[0051] In some embodiments of the present disclosure, before executing S120, the object classification method further includes:
[0052] The object to be classified is encoded to obtain a feature data vector to be classified of the object to be classified.
[0053] Accordingly, S120 may include:
[0054] Classify the feature data vector to be classified based on the object classification network.
[0055] In the disclosed embodiment, the object classification network can be connected to an encoder, and the encoder is used to encode the object to be classified to convert the original object to be classified into one or more data vectors, i.e., feature data vectors to be classified, which can be a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values, and can reflect the differentiation of objects. Different feature data vectors to be classified for the object to be classified correspond to different data channels, and the feature data vectors to be classified corresponding to different data channels are classified by the object classification network to obtain the object classification result. For example, the object to be classified is a color image, and the color image is converted into three feature data vectors to be classified based on the encoder, which are red, green and blue, respectively, and the three data channels correspond to the three feature data vectors to be classified, respectively, and the feature data vectors to be classified corresponding to different data channels are classified by the object classification network to obtain the object classification result of the color image.
[0056] In some embodiments of the present disclosure, the object classification network may include an input layer subnetwork and a hidden layer subnetwork.
[0057] Among them, S120 may include:
[0058] Encode the object to be classified based on the input layer sub-network in the object classification network to obtain a feature data vector to be classified of the object to be classified;
[0059] Based on the hidden layer sub-network in the object classification network, the feature data vector to be classified is classified to obtain the classification result of the object to be classified.
[0060] In the disclosed embodiment, after the object classification network obtains the object to be classified, it can encode the object to be classified based on the input layer sub-network, convert the object to be classified into a data vector form, and obtain a feature data vector to be classified of the object to be classified, and further classify the feature data vector to be classified based on the hidden layer sub-network to obtain a classification result of the object to be classified. The input layer sub-network is consistent with the concept of the input layer, and the hidden layer sub-network is consistent with the concept of the hidden layer.
[0061] In the embodiments of the present disclosure, optionally, the feature data vectors to be classified of the objects to be classified can also be combined into a time series sequence based on time series, and the time series sequence can be input into the object classification network or the hidden layer subnetwork of the object classification network in time series, and feature extraction can be performed on the corresponding feature data vectors to be classified based on the object classification network or the hidden layer subnetwork of the object classification network.
[0062] In an embodiment of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork.
[0063] The hidden layer subnetwork may be a hidden layer, that is, the hidden layer subnetwork may be composed of one hidden layer network, or may be composed of multiple hidden layer networks connected step by step. The output layer subnetwork may be an output layer, that is, the output layer subnetwork may include one or more neurons corresponding to different classification categories, that is, category neurons, and the category neurons of the output layer may extract category features from the output data of the hidden layer.
[0064] In some embodiments of the present disclosure, classifying the object to be classified based on the object classification network to obtain the classification result of the object to be classified may include:
[0065] Based on the hidden layer sub-network, feature extraction is performed on the object to be classified to obtain the feature value corresponding to the object to be classified;
[0066] Based on the output layer sub-network, the category features of the eigenvalues are extracted to obtain the category feature data corresponding to the object to be classified. The category feature data is used to characterize the classification result of the object to be classified.
[0067] In an embodiment of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork. After obtaining the object to be classified, the object classification device may input the object to be classified into the object classification network, and perform layer-by-layer feature extraction on the object to be classified based on the neurons of each hidden layer in the hidden layer subnetwork in the object classification network to obtain the feature value corresponding to the object to be classified, and transfer the feature value corresponding to the object to be classified to the output layer subnetwork, and perform category feature extraction on the feature value based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the object to be classified. The category feature data is the numerical value after mapping and summarizing the features reflecting the classification category in the output feature value of the hidden layer.
[0068] For example, the object to be classified is an image data vector, which includes a cat. After the object classification device obtains the image data vector, it inputs the image data vector into the object classification network, and performs layer-by-layer feature extraction on the image data vector based on the neurons of each hidden layer in the object classification network to obtain the eigenvalues corresponding to the image data vector. The eigenvalues corresponding to the image data vector may include the values of the cat head contour points, the values of the cat body contour points, the values of the cat leg contour points, and the like. Further, the object classification device inputs the eigenvalues corresponding to the image data vector into the output layer subnetwork, and performs category feature extraction on the above eigenvalues based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the image data vector, and uses the category feature data as the classification result of the image data vector, that is, the classification result of the image data vector may include cat feature values reflecting the features of the cat head, cat body, cat legs, and the like.
[0069] In an embodiment of the present disclosure, at least one target neuron in the object classification network may be activated based on a linear normalization function.
[0070] It should be noted that the target neuron can be any hidden layer neuron in the object classification network.
[0071] In an embodiment of the present disclosure, a linear normalization function processing unit (LNU) is set for each neuron of the hidden layer of the object classification network, and at least one linear normalization function is set in each linear normalization unit. The output value of the linear normalization function can be the product of the input value of the linear normalization function and the linear normalization coefficient.
[0072] In the disclosed embodiment, after the object classification device obtains the object to be classified, it can extract features of the object to be classified based on the neurons of the feature channel of the target hidden layer of the hidden layer subnetwork in the object classification network, that is, weight multiplication and aggregation are performed according to the corresponding set weights to obtain weighted input data, and the extracted weighted input data is activated by the linear normalization function corresponding to the neuron to obtain an activated eigenvalue result, that is, the activated eigenvalue result is the product of the weighted input data and the linear normalization coefficient, and the feature activation value result is used as the output feature of the corresponding neuron of the corresponding feature channel of the corresponding target hidden layer; then, based on the neurons of the feature channel of the next target hidden layer, feature extraction is performed on the above-mentioned output features, and the extracted weighted input data is activated by the linear normalization function corresponding to the neuron to obtain the feature activation value result as the output feature of the corresponding neuron of the corresponding feature channel of the next target hidden layer, and the output feature of the next hidden layer is further transmitted to a deeper hidden layer until the last hidden layer, and the feature activation value result of the linear normalization function corresponding to the neurons of the feature channel of the last hidden layer is used as the output feature value of the hidden layer.
[0073] In the disclosed embodiment, the linear normalization coefficient of the linear normalization function is determined according to the target constant value and the number of normalized feature channels.
[0074] In the embodiment of the present disclosure, the normalized number of feature channels may include the first number of feature channels of the target hidden layer to which the target neuron belongs or the second number of feature channels of the input neural network layer of the target neuron.
[0075] In the disclosed embodiment, the input neural network layer of the target neuron may be a neural network layer in the form of an input layer of an object classification network connected to the target neuron on the input side, other hidden layers, a certain time sequence state of the hidden layer or a certain time sequence state of other hidden layers, etc. When the input neural network layer of the target neuron is the input layer, the number of normalized feature channels of the input neural network layer of the target neuron may be 1.
[0076] In some other embodiments of the present disclosure, classifying the object to be classified based on the object classification network to obtain the classification result of the object to be classified may include:
[0077] Based on the hidden layer sub-network, feature extraction is performed on the object to be classified to obtain the feature value corresponding to the object to be classified;
[0078] Based on the output layer sub-network, the category feature of the feature value is extracted to obtain the category feature data of the object to be classified;
[0079] According to the category feature data, the classification result of the object to be classified is determined.
[0080] In an embodiment of the present disclosure, the output layer subnetwork may include one or more category neurons and one or more classifiers.
[0081] In the embodiment of the present disclosure, determining the classification result of the object to be classified according to the category feature data can be specifically as follows: based on one or more classifiers in the output layer subnetwork, the category feature data is classified to obtain the classification result of the object to be classified.
[0082] In an embodiment of the present disclosure, after acquiring an object to be classified, the object classification device can input the object to be classified into an object classification network, perform layer-by-layer feature extraction on the object to be classified based on neurons in each hidden layer in the object classification network, obtain feature values corresponding to the object to be classified, and pass the feature values corresponding to the object to be classified to the output layer subnetwork, perform category feature extraction on the feature values based on category neurons in the output layer subnetwork, and obtain category feature data corresponding to the object to be classified. The category feature data is a numerical value that is mapped and summarized by the features reflecting the classification category in the output feature values of the hidden layer. Furthermore, the category feature data is classified based on the classifier in the output layer subnetwork to determine the classification result of the object to be classified.
[0083] For example, the object to be classified is an image data vector, and the image data vector includes a cat. After the object classification device obtains the image data vector, it inputs the image data vector into the object classification network, and performs layer-by-layer feature extraction on the image data vector based on the neurons of each hidden layer in the object classification network to obtain the eigenvalue corresponding to the image data vector. The eigenvalue corresponding to the image data vector may include the numerical value of the cat head contour point, the numerical value of the cat body contour point, the numerical value of the cat leg contour point, and the like; further, the object classification device inputs the eigenvalue corresponding to the image data vector into the output layer subnetwork, and performs category feature extraction on the above eigenvalue based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the image data vector; further, the category feature data is classified based on the classifier in the output layer subnetwork, and the classification result obtained by the classifier is used as the classification result of the image data vector, that is, the classification result of the image data vector may include cat classification category, and the like.
[0084] In the embodiment of the present disclosure, after determining the category feature data, one or more classifiers in the output layer subnetwork can calculate the probability value of the object to be classified belonging to each classification category according to the category feature data; the category with the largest probability value, or the category with the largest probability value that reaches a preset threshold, is used as the classification category of the object to be classified, and the classification result of the object to be classified is obtained. The classification result may include the classification category of the category feature data and the probability value of each classification category.
[0085] In the disclosed embodiment, the category classifier is cross-connected with the category neurons of the output layer. After the category feature data is determined based on the category neurons of the output layer subnetwork, each category classifier can use the Softmax activation function. The category classifier performs exponential processing on the category feature data based on the Softmax function to obtain the corresponding exponential value, and then calculates the proportion of each index value to the total exponential value, and uses the corresponding proportion value as the output value of each classifier. The output value of each category classifier is used as the probability value of each classification category to obtain the classification result of the object to be classified.
[0086] In the embodiment of the present disclosure, the object classification network may also include an input layer subnetwork, a hidden layer subnetwork and an output layer subnetwork.
[0087] In the embodiments of the present disclosure, the input layer subnetwork can encode the object to be classified, convert the object to be classified into a data vector form, and obtain a feature data vector to be classified of the object to be classified; the hidden layer subnetwork can perform feature extraction on the feature data vector to be classified, and obtain feature values corresponding to the object to be classified; the output layer subnetwork can perform category feature extraction on the feature values, and obtain category feature data of the object to be classified, and the category feature data is used to characterize the classification result of the object to be classified, or the output layer subnetwork can classify the category feature data based on one or more classifiers to obtain the classification result of the object to be classified.
[0088] The technical solution provided by the embodiment of the present disclosure can classify the object to be classified based on the object classification network after obtaining the object to be classified, and obtain the classification result of the object to be classified. Since the target neuron in the object classification network is activated based on the linear normalization function, the derivative of the linear normalization function is a constant, and the output feature expression of the weighted input data will not be weakened. The method of linear normalization activation function eliminates the problem of saturation region in the distribution of derivative values inherent in the traditional nonlinear activation function to a certain extent, and can solve the problem of gradient vanishing of shallow neurons in the process of training artificial neural networks, thereby alleviating the problem that the low-level feature parameters of the object classification network cannot be fully trained due to the gradient vanishing, and improving the training accuracy of the object classification network. At the same time, since the linear normalization coefficient is determined according to the target constant value and the number of normalized feature channels, the neurons of each hidden layer are normalized under the mechanism of the target constant value and the number of normalized feature channels in the linear normalization function, so that the amplification gains of all hidden layer multi-feature channels of the object classification network for feature transfer and error reverse transfer are normalized, thereby reducing the influence of small adjustment changes of the parameters of neurons in each hidden layer during the parameter training and adjustment of the object classification network on the feature flow changes and error flow changes finally received by the target neurons. Further, the feature extraction gain of the object classification network for multiple different input features can be adjusted to a state of approximately one time by adjusting the target constant value. This activation method further reduces the influence of small adjustment changes of the parameters of neurons in each hidden layer during the parameter training and adjustment of the object classification network on the feature flow changes finally received by the target neurons, thereby alleviating the internal covariate drift problem encountered by traditional neural networks, thereby ensuring the accuracy of the learning gradient of neuron parameters and improving the training accuracy of the object classification network. In summary, the object classification network that introduces the above-mentioned linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.
[0089] In some other embodiments of the present disclosure, in order to specifically determine the linear normalization coefficient, the linear normalization coefficient in S120 is the product of the target constant value and the multi-feature channel normalization parameter, and the multi-feature channel normalization parameter is the inverse of the number of normalized feature channels.
[0090] In the embodiment of the present disclosure, the normalized number of feature channels includes the first number of feature channels of the target hidden layer to which the target neuron belongs or the second number of feature channels of the input neural network layer of the target neuron.
[0091] In the disclosed embodiment, the input neural network layer of the target neuron may be a neural network layer in the form of an input layer of an object classification network connected to the target neuron on the input side, other hidden layers, a certain time sequence state of the hidden layer or a certain time sequence state of other hidden layers, etc. When the input neural network layer of the target neuron is the input layer, the number of normalized feature channels of the input neural network layer of the target neuron may be 1.
[0092] Optionally, the linear normalization function is expressed as:
[0093] φ(x)=x*(γ / n) (Formula 1)
[0094] Among them, φ(x) is the output value of the linear normalization function, which is the result of the feature activation value and is directly used as the data in the feature data vector of the hidden layer feature channel, or as the data in the feature data vector of the hidden layer feature channel after the pooling operation; x is the weighted input data of the linear normalization function, and the weighted input data can be the data collected and summarized by the neurons to which the linear normalization function belongs to the feature data vector to be classified according to the weights; γ / n is the linear normalization coefficient of the linear normalization function; γ is the unified target constant value used by all target neurons in the target hidden layer of the object classification network; n is the number of normalized feature channels, and the number of normalized feature channels can include the first feature channel number of the target hidden layer to which the target neuron belongs or the second feature channel number of the input neural network layer of the target neuron. When the inverse of the number of normalized feature channels of the input neural network layer of the target neuron is used to normalize multiple feature channels and the input neural network layer of the target neuron is the input layer, the value of n is 1.
[0095] Optionally, the linear normalization function may set a bias value in the weighted input data or in the output value of the linear normalization function.
[0096] It should be noted that for the calculation of the linear normalization coefficient of the object classification network, the reciprocal of the number of the first feature channel can be used as the multi-feature channel normalization parameter, or the reciprocal of the number of the second feature channel can be used as the multi-feature channel normalization parameter. Both methods can achieve the purpose of multi-feature channel normalization of feature and error transmission flow, thereby achieving the effect of alleviating internal covariate drift. However, generally only one of the methods is used in the object classification network to determine the multi-feature channel normalization parameter, and the neuron linear normalization coefficient is set in combination with the target constant value. For the method of using the reciprocal of the number of the first feature channel as the multi-feature channel normalization parameter, the target neuron receives the feature flow of the lower hidden layer, and normalizes the feature flow of multiple feature channels when transmitting to the higher layer. The entire hidden layer only transmits a feature flow of the scale of one feature channel, which will not cause the multi-feature channel amplification of the feature flow. At the same time, the target neuron normalizes the error flow of multiple feature channels when transmitting the reverse error to the lower layer. The entire hidden layer only transmits the error flow of one feature channel to the lower hidden layer, which will not cause the error flow amplification caused by multiple feature channels. For the method of using the inverse of the number of the second feature channels as the normalization parameter of multiple feature channels, the target neuron normalizes the feature flows of multiple feature channels after receiving the feature flows of multiple feature channels of the lower hidden layer. When transmitting the feature flows to the higher hidden layers, it does not cause the layer-by-layer amplification of the feature flows. At the same time, when forwarding the error flows to all feature channels of the lower hidden layers, a copy of the error flows of the feature channels is always transmitted, which does not cause the amplification of the error flows transmitted in the reverse direction.
[0097] In the disclosed embodiment, the target constant value γ used for the linear normalization coefficient in the linear normalization function can be 1, 2, 3, etc., and is not limited here. However, the target constant value used for the linear normalization coefficient of the linear normalization function of all target hidden layer target neurons of the object classification network must be the same. The target constant value is generally adjusted and set according to the influence of the feature value size of the output of the highest hidden layer after the feature data vector to be classified is loaded during the training of the object classification network. When adjusting, it can be adjusted according to the principle that the absolute mean of the feature value output by the highest hidden layer neurons is equivalent to the absolute mean of the feature data vector value to be classified in the input layer. When using the reciprocal normalization method of the first feature channel number, the absolute mean of the feature value output by the highest hidden layer neurons needs to be multiplied by the number of normalized feature channels of the highest hidden layer.
[0098] Therefore, in the disclosed embodiment, the linear normalization coefficient of the linear normalization function of the target neuron adopts the product of the target constant value and the multi-feature channel normalization parameter, and the multi-feature channel normalization parameter is the inverse of the number of normalized feature channels. Under the action of the feature channel parameter mechanism of the linear normalization function, the flow amplification of the feature and error transmission caused by the multiple feature replications of the multiple feature channels of the hidden layer is normalized, thereby reducing the instability of the target neuron receiving the feature flow and error flow caused by the multiple feature channel replications of the hidden layer, thereby alleviating the occurrence of the internal covariate drift problem. Furthermore, under the premise of normalizing the multiple feature channels of the hidden layer, the target constant value can be adjusted according to the principle that the absolute mean of the feature values output by the neurons of the highest hidden layer is roughly equivalent to the absolute mean of the feature data vector values to be classified in the input layer. When the reciprocal normalization method of the first feature channel number is adopted, the absolute mean of the feature values output by the neurons of the highest hidden layer needs to be multiplied by the number of normalized feature channels of the highest hidden layer, so that the feature extraction gain of each hidden layer of the object classification network for multiple different input feature values of the input layer is roughly maintained at a level of approximately one times. This activation method alleviates the object classification network. During the parameter training adjustment period, the small weight parameter changes of the neurons on the feature transmission path affect the changes in the total feature flow received by the target neuron, thereby alleviating the internal covariate drift problem encountered by the traditional neural network, thereby ensuring the accuracy of the learning gradient of the neuron parameters and improving the training accuracy of the object classification network parameters. In summary, the object classification network that introduces the above-mentioned linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.
[0099] In some other embodiments of the present disclosure, in order to adapt to the scenario where the object classification network is a convolutional neural network and the scenario where the target neuron is simultaneously connected to multiple neural network layers, the linear normalization coefficient in S120 may be determined according to, in addition to the target constant value and the number of normalized feature channels, the linear normalization coefficient may also be determined according to the number of first neurons in a single feature channel of the input neural network layer of the target neuron, the number of input data of the single feature channel of the input neural network layer, the number of input neural network layers, and the number of second neurons in a single feature channel of the target hidden layer to which the target neuron belongs.
[0100] In the embodiments of the present disclosure, the input neural network layer of the target neuron may be a neural network layer in the form of an input layer connected to the target neuron on the input side, other hidden layers, a certain timing state of the hidden layer, or a certain timing state of other hidden layers.
[0101] It should be noted that, for the scenario where the target neuron is connected to multiple neural network layers at the same time, the input neural network layer of each target neuron is the input side connected neural network layer of the target neuron. The target neuron needs to set up independent linear normalization functions for each target neuron input neural network layer in the linear normalization unit, and set the corresponding linear normalization coefficient according to the parameter value corresponding to the input neural network layer of the target neuron, which can be specifically the number of normalized feature channels, the number of first neurons, the number of input data of the single feature channel of the input neural network layer, the number of input neural network layers, the number of second neurons and the target constant value. After each linear normalization function performs linear normalization processing on the weighted input data of the corresponding neural network layer, the output values of each linear normalization function are added as the activation value of the target neuron or as the output feature value. Optionally, for the scenario where the target neuron is connected to multiple neural network layers, if the linear normalization coefficients of multiple target neuron input neural network layers are the same, multiple target neurons can use a unified linear normalization function to uniformly perform linear normalization processing on the weighted input data corresponding to different neural network layers. Among them, the scenarios in which the target neuron is simultaneously connected to multiple neural network layers include scenarios in which the target neuron is simultaneously connected to multiple other hidden layer networks, scenarios in which the target neuron is simultaneously connected to other hidden layers and one or more time series states of the current hidden layers, and scenarios in which the target neuron is simultaneously connected to the input layer and one or more time series states of the current hidden layers.
[0102] In the disclosed embodiment, the number of input neural network layers is also the number of input neural network layers of all neural network layers of the object classification network to which the target neurons are connected on the input side.
[0103] In the embodiment of the present disclosure, the number of second neurons in a single feature channel of a target hidden layer to which the target neuron belongs may be the number of neurons in a single feature channel of the target hidden layer to which the target neuron belongs.
[0104] In the embodiment of the present disclosure, for the first number of neurons in a single feature channel of the input neural network layer of the target neuron, when the input neural network layer of the target neuron is a hidden layer, the first number of neurons in the single feature channel of the input neural network layer may be the number of neurons in the single feature channel of the corresponding hidden layer, and when the input neural network layer of the target neuron is an input layer, the first number of neurons in the single feature channel of the input neural network layer may be the product of the matrix area of the feature data vector to be classified of the data channel of the input layer and the number of data channels of the input layer, and when the number of data channels of the input layer is 1, the first number of neurons in the single feature channel of the input neural network layer may be the matrix area of the feature data vector to be classified of the input layer.
[0105] In the disclosed embodiment, the number of input data of a single feature channel of the input neural network layer may be the number of feature data vectors belonging to a single feature channel of the input neural network layer in the input data of the target neuron. Wherein, when the input neural network layer is a hidden layer that does not include a pooling layer, the feature data vector of a single feature channel may be a feature data vector composed of the activation values of the neurons of the feature channel; when the input neural network layer is a hidden layer that includes a pooling layer, the feature data vector of a single feature channel may be a feature data vector composed of the activation values of the neurons of the feature channel after being pooled by the pooling layer.
[0106] In some embodiments, when the input neural network layer of the target neuron is a hidden layer, the amount of input data of a single feature channel of the input neural network layer is the area of the convolution kernel used by the target neuron.
[0107] In other embodiments, when the input neural network layer of the target neuron is the input layer, the amount of input data of a single feature channel of the input neural network layer is the product of the area of the convolution kernel used by the target neuron and the number of data channels of the input layer, where when the number of data channels of the input layer is 1, it is the area of the convolution kernel used by the target neuron.
[0108] Further, in order to specifically calculate the linear normalization coefficient, the linear normalization coefficient in S120 may be the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter, and the multi-input neural network layer normalization parameter;
[0109] in,
[0110] The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels;
[0111] The normalized number of feature channels includes the number of first feature channels of the target hidden layer to which the target neuron belongs or the number of second feature channels of the input neural network layer of the target neuron;
[0112] The convolution pooling normalization parameter is the quotient of the number of the first neurons divided by the product of the number of input data and the number of the second neurons;
[0113] The normalization parameter of a multi-input neural network layer is the inverse of the number of input neural network layers.
[0114] Optionally, the linear normalization function is expressed as:
[0115] φ(x)=x*(γ*i / (n*j*c*m))(Formula 2)
[0116] Among them, φ(x) is the output value of the linear normalization function; x is the weighted input data of the linear normalization function; γ*i / (n*j*c*m) is the linear normalization coefficient of the linear normalization function, among which γ is the unified target constant value used by the linear normalization function of the target neurons of all target hidden layers of the object classification network; i is the number of first neurons of the single feature channel of the input neural network layer of the target neuron; j is the number of second neurons of the single feature channel of the target hidden layer to which the target neuron belongs; c is the number of input data of the single feature channel of the input neural network layer; m is the number of input neural network layers. When the target neuron is connected to only one input layer or hidden layer, m takes the value of 1; n is the number of normalized feature channels, that is, the number of first feature channels or the number of second feature channels. When the reciprocal method of the number of second feature channels is used for normalization and the input neural network layer is the input layer, n takes the value of 1.
[0117] Optionally, the linear normalization function can set a bias value in the weighted input data or in the linearly normalized output values.
[0118] It should be noted that for the calculation of the linear normalization coefficient of the object classification network, the reciprocal of the number of the first feature channel can be used as the multi-feature channel normalization parameter, or the reciprocal of the number of the second feature channel can be used as the multi-feature channel normalization parameter. Both methods can achieve the purpose of multi-feature channel normalization of feature and error transfer flow, thereby achieving the effect of alleviating internal covariate drift. However, generally only one of the methods is used in the object classification network to determine the feature channel parameters, and the neuron linear normalization coefficient is set in combination with the target constant value.
[0119] In the disclosed embodiment, the constant value γ used in the linear normalization coefficient in the linear normalization function is not limited as mentioned above, but the constant value used in the linear normalization coefficient in the linear normalization function of the target neurons of all target hidden layers of the object classification network must be the same. Among them, the target constant value is generally adjusted and set according to the influence of the output feature value size of the highest hidden layer after the feature data vector to be classified is loaded during the training of the object classification network. When adjusting, it can be adjusted according to the principle that the absolute mean of the output feature values of the neurons in the highest hidden layer is roughly equivalent to the absolute mean of the values of the feature data vector to be classified in the input layer. When the reciprocal normalization method of the number of the first feature channels of the target hidden layer is adopted, the absolute mean of the output feature values of the neurons in the highest hidden layer needs to be multiplied by the number of normalized feature channels of the highest hidden layer.
[0120] In the disclosed embodiment, the linear normalization coefficient is the product of the target constant value, the normalization parameter of multiple feature channels, the convolution pooling normalization parameter and the normalization parameter of multiple input neural network layers, wherein the convolution pooling normalization parameter can be divided into a multi-neuron convolution parameter factor and a pooling feature attenuation parameter factor, the multi-neuron convolution parameter factor can be the quotient q / (c*j) of the number of feature data q of a single feature channel of the input neural network layer divided by the number of input data c of a single feature channel of the input neural network layer and the number of second neurons j of a single feature channel of the target hidden layer to which the target neuron belongs, and the pooling feature attenuation parameter factor can be the quotient i / q of the number of first neurons i of a single feature channel of the input neural network layer of the target neuron divided by the number of feature data q of a single feature channel of the input neural network layer. Thus, the convolution pooling normalization parameter can be obtained by multiplying the neuron convolution parameter factor and the pooling feature attenuation parameter factor. Among them, the quotient of the number of feature data q of a single feature channel of the input neural network layer divided by the product of the number of input data c of the single feature channel of the input neural network layer and the number of second neurons j of the single feature channel of the target hidden layer to which the target neuron belongs is used as a multi-neuron convolution parameter factor. Penalty compensation can be performed on the multi-neuron repeated convolution feature extraction of the feature channel to which the target neuron belongs, thereby alleviating the instability of feature flow and internal covariate drift problems caused by multi-neuron repeated convolution. At the same time, it also alleviates the low-level feature gradient explosion phenomenon caused by high-level feature multi-neuron repeated convolution and reverse error propagation amplification, thereby improving the accuracy of the object classification network. The quotient i / q of the number of first neurons i of a single feature channel of the input neural network layer of the target neuron divided by the number of feature data q of a single feature channel of the input neural network layer is used as a pooling feature attenuation parameter factor, which can compensate for the feature pooling attenuation of the pooling layer of the input neural network layer of the target neuron. The pooling feature attenuation normalization processing of different hidden layers not only compensates for the feature pooling attenuation in feature values, but is also beneficial to the linear normalization function feature extraction target constant value of the neurons in each hidden layer, and uniformly adjusts the overall feature extraction gain of the object classification network, thereby alleviating the internal covariate drift problem of the object classification network.
[0121] In the disclosed embodiment, for a target neuron connected to multiple input neural network layers, since the target neuron is connected to multiple input neural network layers at the same time, whether the target neuron is connected to multiple parallel hidden layer network scenarios, or the target neuron is simultaneously connected to the input layer and a hidden layer of a certain time series state scenario, since multiple parallel feature processing neural network layers are copy processing of the same feature of the object to be classified, multiple parallel neural network layers simultaneously input the processed feature data vector to the target neuron, which will inevitably cause the copy and amplification of the same feature of the object to be classified, and this will cause the instability of the feature flow received by the target neuron, causing the occurrence of internal covariate drift. Therefore, the feature flow of multiple neural network layers received by the target neuron is normalized by the number of first neurons of the input neural network layer of the target neuron, the number of input data of the single feature channel of the input neural network layer, and the number of input neural network layers, which can alleviate the internal covariate drift caused by the parallel feature copy processing of multiple neural network layers.
[0122] Therefore, in the embodiment of the present disclosure, the linear normalization coefficient can be the product of the target constant value, the normalization parameter of multiple feature channels, the convolution pooling normalization parameter and the normalization parameter of multiple input neural network layers. In order to further improve the classification accuracy of the neural network, the target constant value can be adjusted and set in a method in which the absolute mean of the output feature values of the neurons of the highest hidden layer is roughly consistent with the absolute mean of the feature data vector data to be classified in the input layer. When the reciprocal normalization method of the first feature channel number is adopted, the absolute mean of the output feature values of the neurons of the highest hidden layer needs to be multiplied by the number of normalized feature channels of the highest hidden layer to maintain the feature extraction gain of each hidden layer of the object classification network for multiple different input features at a level of approximately one times, so as to achieve the overall normalization of the feature extraction gain of multiple input features of the object classification network. This activation method reduces the influence of small adjustment changes of neurons on the feature transfer path during the training adjustment change of the weight parameters of the object classification network on the change of the feature flow received by the target neuron, alleviates the internal covariate drift problem encountered by the traditional neural network, thereby ensuring the accuracy of the learning gradient of the neuron parameters and improving the training accuracy of the object classification network parameters. In summary, the object classification network that introduces the above-mentioned linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.
[0123] In some further embodiments of the present disclosure, in order to adapt to the scenario where each hidden layer of some object classification networks uses different weight initialization weight interval ranges, the linear normalization coefficient in S120 can also be determined according to the maximum weight value of the target hidden layer to which the target neuron belongs.
[0124] In the embodiment of the present disclosure, the normalized number of feature channels may include the first number of feature channels of the target hidden layer to which the target neuron belongs or the second number of feature channels of the input neural network layer of the target neuron.
[0125] In the embodiment of the present disclosure, the input neural network layer of the target neuron can be a neural network layer in the form of an input layer connected to the target neuron on the input side, other hidden layers, a certain time sequence state of the hidden layer, a certain time sequence state of other hidden layers, etc.
[0126] In the disclosed embodiment, the number of input neural network layers is the number of all neural network layers of the object classification network to which the target neuron is connected on the input side.
[0127] Further, in order to specifically calculate the linear normalization coefficient, the linear normalization coefficient in S120 is the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter, the multi-input neural network layer normalization parameter and the weight gain normalization parameter;
[0128] in,
[0129] The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels;
[0130] The convolution pooling normalization parameter is the quotient of the number of the first neurons divided by the product of the number of input data and the number of the second neurons;
[0131] The normalization parameter of the multi-input neural network layer is the inverse of the number of input neural network layers;
[0132] The weight gain normalization parameter is the inverse of the maximum weight value of the target hidden layer.
[0133] Optionally, the linear normalization function is expressed as:
[0134] φ(x)=x*(γ*i / (n*j*c*m*w))(Formula 3)
[0135] Among them, φ(x) is the output value of the linear normalization function; x is the weighted input data of the linear normalization function; γ*i / (n*j*c*m*w) is the linear normalization coefficient of the linear normalization function, among which γ is the unified target constant value used by the linear normalization function of the target neurons of all target hidden layers of the object classification network; i is the number of first neurons of the single feature channel of the input neural network layer of the target neuron; j is the number of second neurons of the single feature channel of the target hidden layer to which the target neuron belongs; c is the number of input data of the single feature channel of the input neural network layer; m is the number of input neural network layers. When the target neuron is only connected to one input layer or hidden layer, m takes the value of 1; n is the number of normalized feature channels, that is, the number of first feature channels or the number of second feature channels. When the inverse method of the number of second feature channels is used for normalization and the input neural network layer is the input layer, n takes the value of 1; w is the maximum weight value of the target hidden layer to which the target neuron belongs.
[0136] Optionally, the linear normalization function can set a bias value in the weighted input data or in the linearly normalized output values.
[0137] It should be noted that for the calculation of the linear normalization coefficient of the object classification network, the inverse of the number of the first feature channels can be used as the multi-feature channel normalization parameter, or the inverse of the number of the second feature channels can be used as the multi-feature channel normalization parameter. Both methods can achieve the purpose of multi-feature channel normalization of feature and error transfer flows, thereby achieving the effect of alleviating internal covariate drift. However, generally only one of the methods is used in the object classification network to determine the multi-feature channel parameters, and the neuron linear normalization coefficient is set in combination with the target constant value.
[0138] In the disclosed embodiment, the constant value γ used in the linear normalization coefficient in the linear normalization function is not limited as mentioned above, but the constant value used in the linear normalization coefficient in the linear normalization function of the target neurons of all target hidden layers of the object classification network must be the same. Among them, the target constant value is generally adjusted and set according to the influence of the output feature value size of the highest hidden layer after the feature data vector to be classified is loaded during the training of the object classification network. When adjusting, it can be adjusted according to the principle that the absolute mean of the output feature values of the neurons in the highest hidden layer is roughly equivalent to the absolute mean of the values of the feature data vector to be classified in the input layer. When the reciprocal normalization method of the number of the first feature channels of the target hidden layer is adopted, the absolute mean of the output feature values of the neurons in the highest hidden layer needs to be multiplied by the number of normalized feature channels of the highest hidden layer.
[0139] In the disclosed embodiments, the maximum weight value of the target hidden layer to which the target neuron belongs is obtained according to the random initialization interval of the weights of the neural network. In some embodiments, when the range of the random initialization interval of the weights of the neural network is [-1, 1], the maximum weight value is 1; in other embodiments, when the range of the random initialization interval of the weights of the neural network is [-2, 2], the maximum weight value is 2; in yet other embodiments, when the range of the random initialization interval of the weights of the neural network is [-3, 3], the maximum weight value is 3. Other weight maximum values are obtained in the same way. Optionally, due to the randomness of the random initialization of the weights of the neural network and the training and adjustment of the weight parameters, the maximum weight value of the target hidden layer to which the target neuron belongs adopts a value close to the maximum value of the hidden layer weight of the actual neural network.
[0140] In the embodiment of the present disclosure, the linear normalization coefficient may be the product of a target constant value, a normalization parameter for multiple feature channels, a normalization parameter for convolution pooling, a normalization parameter for multiple input neural network layers, and a normalization parameter for weight gain. When a neuron extracts features from an input feature data vector according to a weight parameter, the feature value extracted by the neuron may be amplified or reduced due to the weight parameter being too large or too small. Therefore, the weight parameter value based on the target neuron varies between the maximum weight value of the target hidden layer and the negative maximum weight value of the target hidden layer. By using the inverse (1 / w) of the maximum weight value of the target hidden layer to which the target neuron belongs as the weight gain normalization parameter, it is possible to achieve feature value compensation for the differential feature extraction gains of each layer caused by the initialization interval range of different hidden layer weights, reduce the feature extraction value gain variation of each layer of neurons, and also facilitate the unified adjustment of the overall feature extraction gain of the object classification network by the target constant value in the linear normalization function of each hidden layer neuron, thereby alleviating the internal covariate drift problem of the object classification network. Therefore, it is possible to achieve normalization processing for the feature flow attenuation or feature flow amplification problem caused by the initialization weight interval range of different hidden layers of the object classification network being too large or too small. Moreover, in order to further improve the classification accuracy of the neural network, the target constant value can be adjusted and set according to the method that the absolute mean of the output feature value of the neuron of the highest hidden layer is roughly consistent with the absolute mean of the feature data vector data to be classified in the input layer. When the reciprocal normalization method of the number of the first feature channel of the target hidden layer is adopted, the absolute mean of the output feature value of the neuron of the highest hidden layer needs to be multiplied by the number of normalized feature channels of the highest hidden layer, so as to maintain the feature extraction gain of each hidden layer of the object classification network for multiple different input features at a level of about one time, so as to realize the overall normalization of the feature extraction gain of multiple input features of the object classification network. This activation method reduces the influence of the small adjustment changes of neurons on the transmission paths of multiple different input features during the training adjustment of the weight parameters of the object classification network on the changes of the feature flow received by the target neuron, alleviates the internal covariate drift problem encountered by the traditional neural network, and thus ensures the accuracy of the learning gradient of the neuron parameters, and improves the training accuracy of the object classification network parameters. In summary, the object classification network that introduces the above-mentioned linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.
[0141] Figure 3 It is a logical diagram of an object classification network. The object classification network can be a classification network of an artificial neural network, combined with Figure 3 An exemplary explanation of the process of object classification.
[0142] Figure 3The object classification network in may include three functional layers: input layer, hidden layer, and output layer. The input layer is the above-mentioned input layer subnetwork, the hidden layer is the above-mentioned hidden layer subnetwork, and the output layer is the above-mentioned output layer subnetwork. Among them, the hidden layer includes two hidden layers, hidden layer 1 and hidden layer 2, and the neurons in all hidden layers are set with linear normalization functions to perform linear normalization processing on the feature flow and error flow flowing through the neurons. Specifically, neurons B, C, and D of different feature channels in hidden layer 1 are respectively connected to the feature data vectors to be classified in the input layer, and the feature data vectors to be classified are collected and summarized according to the corresponding weight parameters to obtain weighted input data, and the weighted input data are linearly normalized by LNU to obtain the output feature data extracted by neurons B, C, and D. The neuron output data generated by B, C, and D are feature extracted and linearly normalized by neurons E and F in hidden layer 2, and the output feature data extracted by neurons E and F are obtained. Furthermore, the category neurons G and H of the output layer perform category feature extraction on the output feature data of neurons E and F, and obtain their own category feature data, and then the classifier performs classification according to the category feature data to obtain a classification result. Figure 3The solid line in the middle is a schematic diagram of the transmission process of the feature flow, and the dotted line is a schematic diagram of the transmission process of the error flow. In the process of transmitting feature A in the feature data vector to be classified upward, since neurons B, C, and D normalize the received weighted data according to the number of feature channels of hidden layer 1 (3), it is ensured that the feature flow (FeatureE) received by neuron E is still 1 times the feature flow. At the same time, in the process of the error of the category neuron H being transmitted downward, the error is normalized according to the number of feature channels of hidden layer 2 (2), so that the error (CostC) received by neuron C is still 1 times the error flow. This alleviates the internal covariate drift phenomenon caused by the amplification of feature flow or error flow by multiple feature channels in different hidden layers. Furthermore, the target constant value can be adjusted according to the principle that the absolute mean of the output feature values of the highest hidden layer neurons is roughly equivalent to the absolute mean of the feature data vector values to be classified in the input layer, so as to ensure that the feature extraction gain of each hidden layer for multiple input features is basically maintained at a level of approximately one time, thereby alleviating the influence of small changes in the weight parameters of the relevant neurons on the feature transfer path on the changes in the feature flow received by the target neuron during the learning and adjustment of the training parameters of the object classification network. In this way, the internal covariate drift problem encountered by the traditional neural network is alleviated, thereby ensuring the accuracy of the parameter learning gradient calculated based on the feature flow and error flow received by the neuron, and improving the training accuracy of the object classification network and the object classification accuracy. At the same time, since the linear normalization function is used as the activation function, the activation function leads to a constant, which will not cause the disappearance of the reverse transmission error of the neurons in the lower hidden layer, thereby solving the gradient vanishing problem encountered by the nonlinear function mechanism of the traditional neural network and the problem that the low-level feature parameters cannot be fully learned and trained, and improving the training accuracy of the object classification network.
[0143] Figure 4 is a schematic diagram of the structure of an object classification network. The object classification network can be a classification network of an artificial neural network, combined with Figure 4 The classification process of the object to be classified is explained exemplarily. Figure 4 The object classification network in may include an input layer, a hidden layer and an output layer, wherein the input layer is the above-mentioned input layer subnetwork, the hidden layer is the above-mentioned hidden layer subnetwork, and the output layer is the above-mentioned output layer subnetwork. The hidden layer includes 7 convolutional layers and 2 pooling layers, and the output layer includes two category neurons and two classifiers. The input layer may be the input layer subnetwork described above, and the hidden layer and the output layer may constitute the above-mentioned object classification network.
[0144] The input layer encodes the object to be classified and converts it into a feature data vector to be classified, so as to provide it to the hidden layer of the object classification network for access.
[0145] Among them, the hidden layer performs feature extraction on the feature data vector to be classified of the object to be classified based on convolution layers 1 to 7 and pooling layers 2 and 4, and finally outputs the output feature value of the object to be classified through each feature channel of convolution layer 7; the two category neurons of the output layer perform category feature extraction on the output feature value output by each feature channel of convolution layer 7 to obtain the category feature data of the object to be classified.
[0146] The convolution kernel size of convolution layer 1 is 5*5*3, including 10 feature channels of feature mapping type, and the number of neurons in each feature channel is 28*28; the convolution kernel size of convolution layer 2 is 5*5*10, including 10 feature channels of feature mapping type, and the number of neurons in each feature channel is 24*24; the convolution kernel size of convolution layer 3 is 3*3*10, including 10 feature channels of feature mapping type, and the number of neurons in each feature channel is 10*10; the convolution kernel size of convolution layer 4 is 3*3*10, including 10 feature channels of feature mapping type, and the number of neurons in each feature channel is 8*8; the convolution kernel size of convolution layer 5 is 3*3*10, including 10 The convolution layer 6 has a convolution kernel size of 2*2*30, including 30 feature channels of feature mapping type, and the number of neurons in each feature channel is 1*1; the convolution kernel size of convolution layer 7 is 1*1*30, including 30 feature channels of feature mapping type, and the number of neurons in each feature channel is 1*1; the pooling layer 2 has a pooling size of 2*2, including 10 pooling kernels, and the output data vectors of the corresponding feature channels of convolution layer 2 are pooled respectively; the pooling layer 4 has a pooling size of 2*2, including 10 pooling kernels, and the output data vectors of the corresponding feature channels of convolution layer 4 are pooled respectively. Among them, the neurons in the above convolution layer and pooling layer are set with linear normalization functions, and the weighted input data of the connected target neurons are activated based on the linear normalization function, and the linear normalization coefficient of the linear normalization function can be specifically determined based on the above formula 2.
[0147] Furthermore, each category classifier of the output layer is cross-connected with the category neurons of the output layer. Each category classifier can use the Softmax activation function. The category classifier performs exponential processing on the category feature data based on the Softmax activation function to obtain the corresponding exponential value, and then calculates the proportion of each exponential value to the total exponential value as the output value of each classifier. The output value of each classifier is used as the probability value of each classification category to obtain the classification result of the object to be classified.
[0148] Therefore, in the disclosed embodiment, since the object classification network is activated based on the linear normalization function, the expressiveness of its output features will not be weakened, eliminating the saturation problem of the input distribution of the traditional nonlinear activation function, as well as the resulting vanishing gradient phenomenon of the low-level feature parameters of the neural network and the problem that the low-level feature parameters of the neural network cannot be fully trained. The linear normalization coefficient is the product of the target constant value, the normalization parameter of the multi-feature channel, the convolution pooling normalization parameter and the multi-input neural network layer normalization parameter. Based on the linear normalization function, the feature replication and amplification of the multi-feature channels, the feature amplification caused by the multiple convolution replications of the same feature by multiple neurons in the feature channel, and the feature transmission loss caused by the pooling of the hidden layer and the pooling layer are normalized, so as to achieve the stability of the object classification network features and the error transmission flow, thereby ensuring the stability of the learning gradient of each hidden layer parameter, solving the gradient vanishing problem encountered by the traditional neural network, and alleviating the occurrence of the internal covariate problem, thereby improving the training accuracy. Furthermore, in order to further improve the classification accuracy of the neural network, the target constant value can be adjusted and set according to the method that the absolute mean of the output feature value of the neuron of the highest hidden layer is roughly consistent with the absolute mean of the feature data vector data to be classified in the input layer, so as to maintain the feature extraction gain of each hidden layer of the object classification network for the input feature at a level of approximately double. This activation method alleviates the influence of the small weight parameter training adjustment changes of neurons in multiple different input feature transmission paths on the changes of the feature flow received by the target neuron during the training and adjustment of the object classification network parameters, thereby alleviating the internal covariate drift problem encountered by the traditional neural network, thereby ensuring the accuracy of the learning gradient of the neuron parameters and improving the parameter training accuracy of the object classification network. In addition, the object classification network using the linear normalization function does not need to perform complex normalization calculations on the input data. Compared with the object classification network using the batch normalization algorithm, the amount of calculation in the training process is simplified, and the linear normalization activation mechanism is different from the batch normalization mechanism and can be used in the field of recurrent neural networks (recursive neural networks). In addition, compared with the activation function using the gradient clipping method, the linear normalization activation mechanism does not need to force the gradient to be limited to a certain range, therefore, the training accuracy of the object classification network is guaranteed. Obviously, the use of linear normalization activation mechanism improves the processing capability of the object classification network. It can ensure the training accuracy of the object classification network without increasing the training computing power, and can meet the classification needs of users.
[0149] Further, Figure 4 The object classification network test performance shown in Figure 1 is shown in Figure 1. Specifically, 200 test samples are input Figure 4 In the object classification network shown, the parameters such as the computation time, the number of steps used, and the accuracy of the object classification network are calculated, and the performance of the object classification network is tested based on the parameters such as the computation time, the number of steps used, and the accuracy.
[0150] Table 1: Test results of linear normalization function (LNU) and traditional nonlinear activation function
[0151] Activation Function algorithm Use Step Count Usage time (seconds) Accuracy LNU Adam 600 11613 72.52% Relu Adam 646 11080 61.77% Sigmoid Adam 646 11891 46.35%
[0152] It can be seen from Table 1 that the accuracy and convergence efficiency of the object classification network obtained by using the above-mentioned linear normalization function mechanism to activate neurons are better than the results obtained by using the traditional nonlinear activation function for activation.
[0153] The following is an embodiment of an object classification device provided in an embodiment of the present invention. The device and the object classification methods of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the object classification device, please refer to the embodiment of the above-mentioned object classification method.
[0154] This embodiment provides an object classification device. Figure 5 , the device specifically comprises:
[0155] An object acquisition module 510 is used to acquire an object to be classified;
[0156] The object encoding module 520 is used to classify the object to be classified based on the object classification network to obtain the classification result of the object to be classified. At least one target neuron in the object classification network is activated based on a linear normalization function. The output value of the linear normalization function is the product of the input value of the linear normalization function and the linear normalization coefficient. The linear normalization coefficient is determined according to the target constant value and the number of normalized feature channels.
[0157] The technical solution provided by the embodiment of the present disclosure can classify the to-be-classified object based on the object classification network after obtaining the to-be-classified object, and obtain the classification result of the to-be-classified object. Since the target neuron in the object classification network is activated based on the linear normalization function, the derivative of the linear normalization function is a constant, and for the weighted input data, the output feature expression will not be weakened. The method of the linear normalization activation function eliminates the problem of the distribution of the derivative value inherent in the traditional nonlinear activation function to a certain extent. The problem of the gradient vanishing of shallow neurons in the process of training artificial neural networks can be solved, and the problem of the low-level feature parameters not being fully trained caused by the gradient vanishing is alleviated. At the same time, since the linear normalization coefficient is determined according to the target constant value and the number of normalized feature channels, the neurons of each hidden layer are under the mechanism of the target constant value and the number of normalized feature channels in the linear normalization function, so that the amplification effect caused by the feature or error transmission of the hidden layer multiple feature channels of the object classification network is normalized, thereby reducing the influence of the small adjustment changes of the parameters of the neurons of each hidden layer during the parameter training adjustment of the object classification network on the feature flow changes and error flow changes finally received by the target neurons. Furthermore, the feature extraction gain of the object classification network for multiple different input features can be adjusted to a state of approximately one time by adjusting the target constant value. This activation method further reduces the impact of small parameter adjustments of neurons in each hidden layer on the feature flow changes and error flow changes ultimately received by the target neurons during the parameter training adjustment of the object classification network, thereby alleviating the internal covariate drift problem encountered by traditional neural networks, thereby ensuring the accuracy of the learning gradient of neuron parameters and improving the training accuracy of the object classification network. In summary, the object classification network that introduces the above-mentioned linear normalization function can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.
[0158] Optionally, the normalized number of feature channels includes the first number of feature channels of the target hidden layer to which the target neuron belongs or the second number of feature channels of the input neural network layer of the target neuron.
[0159] Optionally, the linear normalization coefficient is the product of a target constant value and a multi-feature channel normalization parameter, and the multi-feature channel normalization parameter is the inverse of the number of normalized feature channels.
[0160] Optionally, the linear normalization coefficient is also determined based on the number of first neurons in the single feature channel of the input neural network layer of the target neuron, the amount of input data of the single feature channel of the input neural network layer, the number of input neural network layers, and the number of second neurons in the single feature channel of the target hidden layer to which the target neuron belongs.
[0161] Optionally, the linear normalization coefficient is the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter, and the multi-input neural network layer normalization parameter;
[0162] in,
[0163] The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels;
[0164] The convolution pooling normalization parameter is the quotient of the number of the first neurons divided by the product of the number of input data and the number of the second neurons;
[0165] The normalization parameter of a multi-input neural network layer is the inverse of the number of input neural network layers.
[0166] Optionally, the linear normalization coefficient is also determined based on the maximum weight value of the target hidden layer;
[0167] Among them, the linear normalization coefficient is the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter, the multi-input neural network layer normalization parameter and the weight gain normalization parameter;
[0168] Among them,
[0169] The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels;
[0170] The convolution pooling normalization parameter is the quotient of the number of the first neurons divided by the product of the number of input data and the number of the second neurons;
[0171] The normalization parameter of the multi-input neural network layer is the inverse of the number of input neural network layers;
[0172] The weight gain normalization parameter is the inverse of the maximum weight value of the target hidden layer.
[0173] Optionally, the object classification network includes a hidden layer subnetwork and an output layer subnetwork;
[0174] The object to be classified is classified based on the object classification network to obtain the classification result of the object to be classified, including:
[0175] Based on the hidden layer sub-network, feature extraction is performed on the object to be classified to obtain the feature value corresponding to the object to be classified;
[0176] Based on the output layer sub-network, the category features of the eigenvalues are extracted to obtain the category feature data corresponding to the object to be classified. The category feature data is used to characterize the classification result of the object to be classified.
[0177] Through an object classification device according to an embodiment of the present invention, a linear normalization function mechanism is implemented to improve the training accuracy of the object classification network. When the trained object classification network is used to classify objects, the object classification accuracy is improved to meet the classification needs of users.
[0178] The object classification device provided in the embodiment of the present invention can execute the object classification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0179] See also Figure 6 This embodiment provides an object classification device 600, which includes: one or more processors 620; a storage device 610, which is used to store one or more programs. When the one or more programs are executed by the one or more processors 620, the one or more processors 620 implement the object classification method provided by the embodiment of the present invention, including:
[0180] Get the object to be classified;
[0181] An object to be classified is classified based on an object classification network to obtain a classification result of the object to be classified, at least one target neuron in the object classification network is activated based on a linear normalization function, an output value of the linear normalization function is the product of an input value of the linear normalization function and a linear normalization coefficient, and the linear normalization coefficient is determined according to a target constant value and the number of normalized feature channels.
[0182] Of course, those skilled in the art can understand that the processor 620 can also implement the technical solution of the object classification method provided by any embodiment of the present invention.
[0183] Figure 6 The object classification device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0184] like Figure 6 As shown, the object classification device 600 includes a processor 620, a storage device 610, an input device 630, and an output device 640; the number of the processor 620 in the device can be one or more. Figure 6 A processor 620 is taken as an example; the processor 620, the storage device 610, the input device 630 and the output device 640 in the device can be connected via a bus or other means. Figure 6 The connection via bus 650 is taken as an example.
[0185] The storage device 610, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the object classification method in the embodiment of the present invention (for example, the object acquisition module, feature extraction module and classification module in the object classification device).
[0186] The storage device 610 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the storage device 610 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device 610 may further include a memory remotely arranged relative to the processor 620, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0187] The input device 630 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the device, for example, it can include at least one of a mouse, a keyboard and a touch screen. The output device 640 can include a display device such as a display screen.
[0188] This embodiment provides a storage medium containing computer executable instructions. When the computer executable instructions are executed by a computer processor, they are used to perform an object classification method. The method includes:
[0189] Get the object to be classified;
[0190] An object to be classified is classified based on an object classification network to obtain a classification result of the object to be classified, at least one target neuron in the object classification network is activated based on a linear normalization function, an output value of the linear normalization function is the product of an input value of the linear normalization function and a linear normalization coefficient, and the linear normalization coefficient is determined according to a target constant value and the number of normalized feature channels.
[0191] Of course, the computer executable instructions of a storage medium including computer executable instructions provided by an embodiment of the present invention are not limited to the above method operations, and can also execute related operations in the object classification method provided by any embodiment of the present invention.
[0192] Through the above description of the implementation method, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute the object classification method provided by each embodiment of the present invention.
[0193] Note that the setting of the target constant value in the above-mentioned embodiments of the present invention is only a recommended setting of the target constant value. Those skilled in the art will understand that for those skilled in the art, adopting other methods of setting the target constant value will not deviate from the protection scope of the present invention.
[0194] Note that the setting of the linear normalization coefficient in the above-mentioned embodiment of the present invention is only the setting of the optimal linear normalization coefficient. Those skilled in the art will understand that for those skilled in the art, adopting a setting value close to the linear normalization coefficient in the above-mentioned embodiment (such as the setting value of the adopted linear normalization coefficient is compared with the setting value of the linear normalization coefficient in the above-mentioned embodiment, and the closeness or degree of conformity is within 90%) will not deviate from the protection scope of the present invention.
[0195] Note that in the above-mentioned embodiment of the present invention, setting a linear normalization function for neurons in the hidden layer is only an optimal deployment embodiment. Those skilled in the art will understand that for those skilled in the art, partially adopting the deployment method in the above-mentioned embodiment (such as adopting a local number of hidden layers or adopting a local number of neurons in the hidden layer to set a linear normalization function), or equivalently modifying the deployment method in the above-mentioned embodiment (such as modifying the original convolution pooling normalization parameter definition of the linear normalization coefficient to: the quotient of the number of feature data of a single feature channel of the input neural network layer divided by the product of the number of input data and the number of feature data of a single feature channel of the target hidden layer), the convolution pooling normalization scheme adjusts the method of normalizing the pooling features of the previous layer from attenuation normalization to attenuation normalization of the pooling features of the current layer, or adjusting on the basis of the deployment method in the above-mentioned embodiment (such as modifying the linear normalization function by adding a bias value to the input data or output data of the original linear normalization function) will not deviate from the protection scope of the present invention.
[0196] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, or substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for object classification, characterized in that: include: Acquire an object to be classified, where the object to be classified is an image, an audio, or a text; Classifying the object to be classified based on an object classification network to obtain a classification result of the object to be classified, wherein at least one target neuron in the object classification network is activated based on a linear normalization function, an output value of the linear normalization function is a product of an input value of the linear normalization function and a linear normalization coefficient, the linear normalization coefficient is determined according to a target constant value and a number of normalized feature channels, and the target neuron is a hidden layer neuron in the object classification network; Among them, the normalized number of characteristic channels includes the first number of characteristic channels of the target hidden layer to which the target neuron belongs or the second number of characteristic channels of the input neural network layer of the target neuron.
2. The method according to claim 1, characterized in that The linear normalization coefficient is the product of the target constant value and a normalization parameter of multiple feature channels, and the normalization parameter of multiple feature channels is the inverse of the number of the normalized feature channels.
3. The method according to claim 1, characterized in that The linear normalization coefficient is also determined based on the number of first neurons in the single feature channel of the input neural network layer of the target neuron, the number of input data of the single feature channel of the input neural network layer, the number of input neural network layers, and the number of second neurons in the single feature channel of the target hidden layer to which the target neuron belongs.
4. The method according to claim 3, characterized in that The linear normalization coefficient is the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter and the multi-input neural network layer normalization parameter; in, The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels; The convolution pooling normalization parameter is the quotient of the first number of neurons divided by the product of the number of input data and the second number of neurons; The multi-input neural network layer normalization parameter is the inverse of the number of input neural network layers.
5. The method according to claim 3, characterized in that: The linear normalization coefficient is also determined according to the maximum weight value of the target hidden layer; Wherein, the linear normalization coefficient is the product of the target constant value, the multi-feature channel normalization parameter, the convolution pooling normalization parameter, the multi-input neural network layer normalization parameter and the weight gain normalization parameter; in, The multi-feature channel normalization parameter is the inverse of the number of normalized feature channels; The convolution pooling normalization parameter is the quotient of the first number of neurons divided by the product of the number of input data and the second number of neurons; The multi-input neural network layer normalization parameter is the inverse of the number of input neural network layers; The weight gain normalization parameter is the inverse of the maximum weight value of the target hidden layer.
6. The method according to claim 1, characterized in that The object classification network includes a hidden layer sub-network and an output layer sub-network; The classifying the object to be classified based on the object classification network to obtain the classification result of the object to be classified includes: Extracting features of the object to be classified based on the hidden layer sub-network to obtain feature values corresponding to the object to be classified; Based on the output layer sub-network, category features are extracted from the feature values to obtain category feature data corresponding to the object to be classified, and the category feature data is used to characterize the classification result of the object to be classified.
7. An object classification device, characterized in that: include: An object acquisition module is used to acquire an object to be classified, wherein the object to be classified is an image, an audio or a text; An object classification module, used for classifying the object to be classified based on an object classification network to obtain a classification result of the object to be classified, wherein at least one target neuron in the object classification network is activated based on a linear normalization function, and an output value of the linear normalization function is a product of an input value of the linear normalization function and a linear normalization coefficient, and the linear normalization coefficient is determined according to a target constant value and a number of normalized feature channels, and the target neuron is a hidden layer neuron in the object classification network; Among them, the normalized number of characteristic channels includes the first number of characteristic channels of the target hidden layer to which the target neuron belongs or the second number of characteristic channels of the input neural network layer of the target neuron.
8. An object classification device, characterized in that: The device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the object classification method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the object classification method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Media classification
CN107851198A