Object classification method, device, equipment and storage medium

By introducing suppression and corrected linear units into artificial neural networks, the problems of parameter gradient vanishing and internal covariate drifting are solved, and classification accuracy is improved.

CN114169405BActive Publication Date: 2025-05-02王树松
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111357998.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-05-02
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

In the prior art, artificial neural networks are prone to the disappearance of parameter gradients and internal covariate drift problems during training, resulting in poor classification accuracy.

Method used

Introduced an inhibition and corrected linear unit (RReLU), which consists of a linear suppression function and a corrected linear class function, is used to activate target hidden layer neurons in the object classification network. The linear suppression function is multiplied by the input by the linear suppression coefficient, and the corrected linear function is corrected in the positive and negative input intervals respectively.

Benefits of technology

By suppressing the use of corrected linear units, the problems of gradient vanishing and internal covariate drift are solved, and the training accuracy and classification accuracy of the object classification network are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114169405B_ABST
    Figure CN114169405B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an object classification method, device, equipment and storage medium. The method includes: obtaining an object to be classified; classifying the object to be classified based on an object classification network to obtain a classification result of the object to be classified, and activating neurons of at least one target hidden layer in the object classification network based on an inhibitory corrected linear unit; the inhibitory corrected linear unit includes a linear inhibition function and a corrected linear class function, and the corrected linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input multiplied by the corresponding linear inhibition coefficient; the corrected linear class function is a positive corrected linear function or a negative corrected linear function; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs. Therefore, the object classification network introducing the inhibitory corrected linear unit improves the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an object classification method, device, equipment and storage medium. Background Art

[0002] With the continuous development of computer technology, deep learning technology has been more and more widely used in image processing, image recognition, object classification and data processing. For example, artificial neural networks (ANN) are usually used to classify images, audio, etc. Artificial neural networks are composed of a large number of interconnected neurons. The output of each neuron is processed by a specific output function, which is an activation function.

[0003] In the prior art, the role of the activation function is to enhance the learning ability of the artificial neural network so that the artificial neural network can adapt to a variety of data processing scenarios. However, when training an artificial neural network that introduces an activation function, the shallow neurons of the artificial neural network are prone to parameter gradient vanishing during the parameter training process, which makes it difficult to fully train the shallow network; and, due to the interference of inter-layer traffic, the artificial neural network has internal covariate drift problems in parameter training, which affects the parameter learning efficiency and accuracy of the artificial neural network. Therefore, when classifying objects based on the above artificial neural network, the classification accuracy is poor and cannot meet the classification needs of users. Summary of the invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an object classification method, apparatus, device and storage medium.

[0005] In a first aspect, the present disclosure provides an object classification method, the method comprising:

[0006] Get the object to be classified;

[0007] Based on the object classification network, the object to be classified is classified to obtain the classification result of the object to be classified, and the neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, and when the input value of the positive correction linear function is greater than or equal to 0, the output value of the positive correction linear function is equal to the input value of the positive correction linear function; when the input value of the negative correction linear function is less than or equal to 0, the output value of the negative correction linear function is equal to the input value of the negative correction linear function, and when the input value of the negative correction linear function is greater than 0, the output value of the negative correction linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0008] In a second aspect, the present disclosure provides an object classification device, the device comprising:

[0009] An object acquisition module is used to acquire objects to be classified;

[0010] The object classification module is used to classify the to-be-classified object based on the object classification network to obtain the classification result of the to-be-classified object, and the neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the positive correction linear function is activated. The output value of the property function is 0. When the input value of the positive corrected linear function is greater than or equal to 0, the output value of the positive corrected linear function is equal to the input value of the positive corrected linear function; when the input value of the negative corrected linear function is less than or equal to 0, the output value of the negative corrected linear function is equal to the input value of the negative corrected linear function, and when the input value of the negative corrected linear function is greater than 0, the output value of the negative corrected linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0011] In a third aspect, an embodiment of the present disclosure further provides an object classification device, the device comprising:

[0012] one or more processors;

[0013] a storage device for storing one or more programs,

[0014] When one or more programs are executed by one or more processors, the one or more processors implement the object classification method provided in the first aspect.

[0015] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the object classification method provided in the first aspect is implemented.

[0016] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0017] An object classification method, device, equipment and storage medium of the disclosed embodiment can classify the object to be classified based on the object classification network after obtaining the object to be classified, and obtain the classification result of the object to be classified. Since the object classification network is activated based on the suppression correction linear unit composed of a linear suppression function and a correction linear class function, the derivative of the suppression correction linear unit in the positive input interval (for the positive correction linear function) or the negative input interval (for the negative correction linear function) is a constant. Since a hidden layer has a large number of neurons that satisfy the positive input interval or the negative input interval at the same time, the output feature expression of the weighted input data will not be weakened. The method of the suppression correction linear unit eliminates the problem of the distribution of the derivative value inherent in the traditional nonlinear activation function to a certain extent. The problem of the gradient vanishing of shallow neurons in the process of training artificial neural networks can be solved, thereby alleviating the problem that the low-level feature parameters of the object classification network cannot be fully trained due to the gradient vanishing, and improving the training accuracy of the object classification network. At the same time, since the linear suppression coefficient is a positive value less than 1, the amplification effect of the feature transfer flow and error transfer flow caused by the feature extraction of the same input data by multiple neurons in the object classification network is suppressed, thereby reducing the influence of the layer-by-layer amplification of the transfer flow change caused by the adjustment of the parameters of the neurons in each hidden layer of the object classification network during the parameter training adjustment on the feature flow and error flow finally received by the target neuron, thereby alleviating the internal covariate drift problem encountered by traditional neural networks, ensuring the accuracy of the learning gradient of neuron parameters, and improving the training accuracy of the object classification network. In summary, the object classification network that introduces the above-mentioned suppression corrected linear unit can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 It is a logical schematic diagram of an artificial neural network provided by the prior art;

[0021] Figure 2 is a flow chart of an object classification method provided by an embodiment of the present disclosure;

[0022] Figure 3 is a logical schematic diagram of an object classification network provided by an embodiment of the present disclosure;

[0023] Figure 4 is a schematic diagram of the structure of an object classification network provided by an embodiment of the present disclosure;

[0024] Figure 5 is a structural schematic diagram of an object classification device provided by an embodiment of the present disclosure;

[0025] Figure 6 It is a structural diagram of an object classification device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0028] In related technologies, artificial neural networks are used to classify objects. Figure 1 The figure shows a logic diagram of an artificial neural network provided by the prior art. Figure 1 ,The artificial neural network is composed of three functional layers including input layer, hidden layer and output layer.

[0029] The main function of the input layer is to convert the original object to be classified into one or more data vectors. The data vector is a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values, and can reflect the differentiation of the object. The data vector is the feature data vector to be classified. The input layer passes each feature data vector to be classified to the hidden layer through one or more data channels. Different feature data vectors to be classified of the object to be classified correspond to different data channels. For example, if the object to be classified is a color image, the input layer converts the color image into three feature data vectors to be classified, namely red, green and blue, and passes the above three feature data vectors to be classified to the hidden layer through three data channels. Optionally, the feature data vectors to be classified can be composed of a sequence of feature data vectors to be classified in time sequence and passed to the hidden layer in time sequence.

[0030] The hidden layer of an artificial neural network consists of one hidden layer network, or multiple hidden layer networks connected in stages. A hidden layer network consists of one hidden layer, or multiple hidden layers connected in stages. The hidden layer is connected to the input layer by accessing the feature data vector to be classified of the input layer, and the hidden layer is connected to the previous hidden layer by accessing the feature data vector of the previous hidden layer. Optionally, one hidden layer can be connected to multiple different types of neural network layers at the same time, and different types of neural network layers include input layer, other hidden layers, a certain time sequence state of this hidden layer, a certain time sequence state of other hidden layers, etc. The neural network layer in the form of input layer, other hidden layers, a certain time sequence state of this hidden layer, a certain time sequence state of other hidden layers, etc. connected to the input side of the target neuron is called the input neural network layer of the target neuron. Optionally, one hidden layer can be connected to multiple different hidden layer networks at the same time.

[0031] Each hidden layer of an artificial neural network can include multiple feature maps, and each feature map can correspond to one or more neurons. A neuron refers to an entity that extracts features from one or more input feature data vectors according to corresponding set weights. The output data of neurons in a hidden layer can be passed to the next hidden layer, the output layer, the pooling layer of the current hidden layer, the next time series of the current hidden layer, or the previous time series of the current hidden layer.

[0032] Different feature maps of the hidden layer refer to neuron feature extraction entities facing all position areas of one or more input feature data vectors with independent neuron weight parameters. Different neurons in the same feature map are entities that perform feature extraction on different position areas of one or more input feature data vectors. The neurons of the feature map can simultaneously use the neuron detection size corresponding to the input neural network layer to perform feature extraction on the feature data of one or more feature maps of one or more different neural network layers. For the same input neural network layer, the neuron detection size of different feature maps of the same hidden layer is the same. Different feature maps of the same hidden layer perform feature extraction on the feature data vectors of the same one or more input neural network layer feature maps to obtain independent neuron output data of each feature map. When different neurons in the feature map perform feature extraction on different position areas of the input feature data vector, the position areas of the corresponding feature data vectors may overlap. All neuron output data of the feature map are directly formed or pooled through the pooling channel of the pooling layer to form independent feature map feature data vectors of each feature map. The feature data vector of the feature map is a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values. The pooling channel of the pooling layer uses a certain pooling kernel size to merge the neuron output data of the feature map corresponding to the number of pooling kernel sizes into a feature map feature data by using algorithms such as the average pooling method and the maximum pooling method. A feature map can use pooling channels corresponding to multiple different pooling kernel sizes (such as when using the spatial pyramid pooling algorithm) to pool the neuron output data of the feature map. At this time, the pooling output data of all the pooling channels obtained separately are spliced ​​to form a feature data vector of the feature map. When a neuron extracts features from feature data vectors of feature maps of multiple input neural network layers at the same time, a separate extraction mode or a splicing extraction mode can be used. When the separate extraction mode is used, the neuron summarizes the feature data of the corresponding areas in the feature data vectors of the feature maps of different input neural network layers according to the corresponding weights to form weighted input data, and then summarizes all the weighted input data from each input neural network layer and activates it by the activation function to form neuron output data. When the splicing extraction mode is adopted, the feature data vectors corresponding to a certain feature mapping number of multiple input neural network layers can be spliced ​​into one feature data vector corresponding to a certain feature mapping number, and unified as the feature data vector corresponding to a certain feature mapping number of multiple input neural network layers, and the neurons then perform feature extraction on the feature data vectors of the multiple input neural network layers.

[0033] The types of feature maps include feature maps of a single neuron and feature maps of a convolutional neural network convolution kernel (or filter) type of multiple neurons. For a feature map containing a single neuron, the neuron of the feature map performs feature extraction on one or more input feature data vectors to obtain a feature map neuron output data; for a feature map of a convolutional neural network convolution kernel type containing multiple neurons, each neuron in the feature map performs feature extraction on a position area in one or more input feature data vectors to obtain feature map neuron output data of the corresponding position area of ​​the corresponding feature map. The feature extraction of neurons refers to the neurons performing weight multiplication and aggregation processing on one or more input feature data vectors according to the corresponding set weights to obtain weighted input data, and then the activation function of the neuron performs activation processing on the weighted input data to obtain the output data of the neuron (or output feature value, activation value). Optionally, the feature extraction process of the neuron also includes setting a bias value in the weighted input data. Optionally, the feature extraction process of the neuron also includes performing numerical normalization processing (such as local response normalization, batch normalization, etc.) on the weighted input data or the activation value obtained by the activation function.

[0034] The output layer of the artificial neural network is connected to the hidden layer, and the output layer includes one or more neurons corresponding to different classification categories, namely, category neurons. The category neurons of the output layer extract category features from the feature data vector output by the hidden layer. Category feature extraction of category neurons refers to the category neurons performing weight multiplication calculations on the data in the feature data vector output by the hidden layer according to the corresponding set weights and summarizing them to obtain weighted input data. The weighted input data of the category neurons of the output layer are the category feature data (or category feature values) corresponding to the category neurons of the output layer. Furthermore, the output layer includes one or more category classifiers, and the category classifiers of the output layer calculate the probability value that the object to be classified belongs to the classification category corresponding to the output layer neurons based on the category feature data. The output layer outputs the classification category and the probability value to obtain the classification result. Optionally, the category feature extraction process of the category neurons also includes setting a bias value in the weighted input data.

[0035] The weight parameters in the artificial neural network are set after the artificial neural network training is completed. When training the artificial neural network, according to the system loss function of the artificial neural network and the current state of the relevant variables, the current partial derivative of the system loss function to the output variable of the neuron is calculated as the output error of the neuron, and then the parameter gradient value of the neuron is calculated according to the output error of the neuron, the derivative of the neuron activation function and the input characteristic value of the neuron, and the parameters of the neuron are adjusted based on the parameter gradient value.

[0036] When training an artificial neural network, in order to prevent the output of each hidden layer neuron from having too much influence on the feature processing of the next hidden layer, a nonlinear activation function is set for each neuron in the hidden layer of the artificial neural network, and the nonlinear activation function is used to activate the weighted input data so that the output features are maintained within the numerical range allowed by the nonlinear activation function. However, when using a nonlinear activation function (such as the Sigmoid function) to activate the features of weighted input data, there are the following two problems:

[0037] 1. After the input is differentiated based on the nonlinear activation function, the nonlinear activation function causes the distribution of weighted input data to be saturated. For weighted input far from the center point, the output feature expression is weak, resulting in the smaller the error of the neuron becomes as it is transmitted to the shallower layer. As a result, the gradient value of the parameter calculated based on the error value of the neuron becomes very small as the number of neuron layers decreases, resulting in the problem of gradient vanishing of shallow neurons.

[0038] 2. The parameter learning function calculates the parameter gradient value based on the output error of the neuron, the derivative of the neuron activation function and the input characteristic value of the neuron, and then adjusts the neuron parameters based on the parameter gradient value. At this time, when adjusting the parameters of the neurons, since the parameters of each neuron in the neural network are adjusted and different neurons are related to each other, the numerical changes of the characteristics and errors transmitted due to the parameter adjustment in each layer of neurons produce a layer-by-layer superposition effect, resulting in a change in the actual gradient of the current parameter of the neuron compared with the calculated value, resulting in the problem of internal covariate shift, which affects the parameter training accuracy.

[0039] In summary, the use of traditional nonlinear activation functions affects the training accuracy of artificial neural networks, which in turn affects the classification accuracy when classifying objects based on the trained artificial neural networks and cannot meet the classification needs of users.

[0040] In order to solve the above problems, in the related traditional technologies, the nonlinear activation function in the neural network also adopts the modified linear activation function (Relu) activation function method, batch normalization method and gradient clipping method. However, although the use of Relu activation solves the gradient vanishing problem of deep neural networks to a certain extent, Relu does not solve the problem of internal covariate shift caused by inter-layer traffic interference; the disadvantage of using the batch normalization method is that when the number of samples is small, the mean and variance of the samples in a batch are greatly different from the overall samples, which will reduce the training accuracy; at the same time, the batch normalization method requires complex normalization calculations and backward error propagation calculations for the training samples, and the amount of calculation in the training process will be greatly increased; at the same time, the batch normalization method is not suitable for recurrent neural network scenarios, that is, the scope of application is small. The disadvantage of using the gradient clipping method is that the idea of ​​the gradient clipping method is to set a clipping threshold. If the gradient exceeds the clipping threshold when updating the gradient, it will be forced to be limited within this range. It has a certain inhibitory effect on the gradient explosion phenomenon, but has no effect on the gradient vanishing phenomenon, and because the gradient is changed, it will also affect the training accuracy to a certain extent.

[0041] In order to solve the above problems, the embodiments of the present disclosure provide an object classification method, apparatus, device and storage medium, which can improve the training accuracy of the object classification network based on the suppressed corrected linear unit mechanism, and improve the classification accuracy when using the trained object classification network to classify objects, so as to meet the object classification needs of users.

[0042] Below, the object classification method provided by the embodiment of the present disclosure is first described.

[0043] The object classification method provided in this embodiment can be applied to classify objects based on the suppressed rectified linear unit mechanism. The method can be performed by an object classification device, which can be implemented by software and / or hardware, and the device can be integrated in a device with object classification function, such as a desktop computer or a server. Figure 2 The method of this embodiment specifically includes the following steps:

[0044] S110: Obtain an object to be classified.

[0045] In the embodiment of the present disclosure, the object classification device can obtain the object to be classified that needs to be classified. The object classification device can obtain the object to be classified by collecting, downloading, uploading, inputting, copying, etc. the object to be classified.

[0046] In some embodiments, the device to be classified may collect data from other devices or equipment that perform data communication with the object to be classified, and use the collected data as the object to be classified.

[0047] In other embodiments, the device to be classified may download data from an open source data set or an open source website, and use the downloaded data as the object to be classified.

[0048] In some other embodiments, the device to be classified can obtain external input data and use the external input data as the object to be classified.

[0049] In some further embodiments, the device to be classified may copy data pre-existing in a data storage device or a data block, and use the copied data as the object to be classified.

[0050] Optionally, the objects to be classified may include images, videos, audios, texts, dynamic images, data vectors, etc., which are not limited here.

[0051] When the object to be classified is an image, a video or a dynamic image, the image may include at least one target object, such as a person, an animal, a scene, an object, etc.

[0052] When the object to be classified is audio, the audio may include at least one audio segment and the like.

[0053] When the object to be classified is text, the text may include text information formed based on at least one font.

[0054] S120, classifying the objects to be classified based on the object classification network to obtain the classification result of the objects to be classified, and activating the neurons of at least one target hidden layer in the object classification network based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, and when the input value of the positive correction linear function is greater than or equal to 0, the output value of the positive correction linear function is equal to the input value of the positive correction linear function; when the input value of the negative correction linear function is less than or equal to 0, the output value of the negative correction linear function is equal to the input value of the negative correction linear function, and when the input value of the negative correction linear function is greater than 0, the output value of the negative correction linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0055] In some embodiments of the present disclosure, before executing S120, the object classification method further includes:

[0056] The object to be classified is encoded to obtain a feature data vector to be classified of the object to be classified.

[0057] Accordingly, S120 may include:

[0058] Classify the feature data vector to be classified based on the object classification network.

[0059] In the disclosed embodiment, the object classification network can be connected to an encoder, and the encoder is used to encode the object to be classified to convert the original object to be classified into one or more data vectors, i.e., feature data vectors to be classified, which can be a zero-dimensional, one-dimensional or two-dimensional numerical matrix composed of one or more numerical values, and can reflect the differentiation of objects. Different feature data vectors to be classified for the object to be classified correspond to different data channels, and the feature data vectors to be classified corresponding to different data channels are classified by the object classification network to obtain the object classification result. For example, the object to be classified is a color image, and the color image is converted into three feature data vectors to be classified based on the encoder, which are red, green and blue, respectively, and the three data channels correspond to the three feature data vectors to be classified, respectively, and the feature data vectors to be classified corresponding to the three different data channels are classified by the object classification network to obtain the object classification result of the color image.

[0060] In some other embodiments of the present disclosure, the object classification network may include an input layer subnetwork and a hidden layer subnetwork.

[0061] Among them, S120 may include:

[0062] Encode the object to be classified based on the input layer sub-network in the object classification network to obtain a feature data vector to be classified of the object to be classified;

[0063] Based on the hidden layer sub-network in the object classification network, the feature data vector to be classified is classified to obtain the classification result of the object to be classified.

[0064] In the disclosed embodiment, after the object classification network obtains the object to be classified, it can encode the object to be classified based on the input layer sub-network, convert the object to be classified into a data vector form, and obtain a feature data vector to be classified of the object to be classified, and further classify the feature data vector to be classified based on the hidden layer sub-network to obtain a classification result of the object to be classified. The input layer sub-network is consistent with the concept of the input layer, and the hidden layer sub-network is consistent with the concept of the hidden layer.

[0065] In the embodiments of the present disclosure, optionally, the feature data vectors to be classified of the objects to be classified can also be combined into a time series sequence based on time series, and the time series sequence can be input into the object classification network or the hidden layer subnetwork of the object classification network in time series, and feature extraction can be performed on the corresponding feature data vectors to be classified based on the object classification network or the hidden layer subnetwork of the object classification network.

[0066] In an embodiment of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork.

[0067] The hidden layer subnetwork may be a hidden layer, that is, the hidden layer subnetwork may be composed of one hidden layer network, or may be composed of multiple hidden layer networks connected step by step. The output layer subnetwork may be an output layer, that is, the output layer subnetwork may include one or more neurons corresponding to different classification categories, that is, category neurons, and the category neurons of the output layer may extract category features from the output data of the hidden layer.

[0068] In some further embodiments of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork.

[0069] Among them, S120 may include:

[0070] Based on the hidden layer sub-network, feature extraction is performed on the object to be classified to obtain the feature value corresponding to the object to be classified;

[0071] Based on the output layer sub-network, the category features of the eigenvalues ​​are extracted to obtain the category feature data corresponding to the object to be classified. The category feature data is used to characterize the classification result of the object to be classified.

[0072] In an embodiment of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork. After obtaining the object to be classified, the object classification device may input the object to be classified into the object classification network, and perform layer-by-layer feature extraction on the object to be classified based on the neurons of each hidden layer in the hidden layer subnetwork in the object classification network to obtain the feature value corresponding to the object to be classified, and transfer the feature value corresponding to the object to be classified to the output layer subnetwork, and perform category feature extraction on the feature value based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the object to be classified. The category feature data is the numerical value after mapping and summarizing the output feature value of the hidden layer to reflect the feature differences of different categories.

[0073] For example, the object to be classified is an image data vector, which includes a cat. After the object classification device obtains the image data vector, it inputs the image data vector into the object classification network, and performs layer-by-layer feature extraction on the image data vector based on the neurons of each hidden layer in the object classification network to obtain the eigenvalues ​​corresponding to the image data vector. The eigenvalues ​​corresponding to the image data vector may include the values ​​of the cat head contour points, the values ​​of the cat body contour points, the values ​​of the cat leg contour points, and the like. Further, the object classification device inputs the eigenvalues ​​corresponding to the image data vector into the output layer subnetwork, and performs category feature extraction on the above eigenvalues ​​based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the image data vector, and uses the category feature data as the classification result of the image data vector, that is, the classification result of the image data vector may include cat feature values ​​reflecting the features of the cat head, cat body, cat legs, and the like.

[0074] In some further embodiments of the present disclosure, the object classification network may include a hidden layer subnetwork and an output layer subnetwork.

[0075] Among them, S120 may include:

[0076] Based on the hidden layer sub-network, feature extraction is performed on the object to be classified to obtain the feature value corresponding to the object to be classified;

[0077] Based on the output layer sub-network, the category feature of the feature value is extracted to obtain the category feature data of the object to be classified;

[0078] According to the category feature data, the classification result of the object to be classified is determined.

[0079] In an embodiment of the present disclosure, the output layer subnetwork may include one or more category neurons and one or more classifiers.

[0080] In the embodiment of the present disclosure, determining the classification result of the object to be classified according to the category feature data can be specifically as follows: based on one or more classifiers in the output layer subnetwork, the category feature data is classified to obtain the classification result of the object to be classified.

[0081] In an embodiment of the present disclosure, after obtaining the object to be classified, the object to be classified can be input into an object classification network, and based on the neurons of each hidden layer in the object classification network, feature extraction is performed layer by layer on the object to be classified to obtain feature values ​​corresponding to the object to be classified, and the feature values ​​corresponding to the object to be classified are passed to the output layer subnetwork, and category feature extraction is performed on the feature values ​​based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the object to be classified. The category feature data is the numerical value after mapping and summarizing the output feature values ​​of the hidden layer to reflect the feature differences of different categories. Furthermore, the category feature data is classified based on the classifier in the output layer subnetwork to determine the classification result of the object to be classified.

[0082] For example, the object to be classified is an image data vector, and the image data vector includes a cat. After the object classification device obtains the image data vector, it inputs the image data vector into the object classification network, and performs layer-by-layer feature extraction on the image data vector based on the neurons of each hidden layer in the object classification network to obtain the eigenvalue corresponding to the image data vector. The eigenvalue corresponding to the image data vector may include the numerical value of the cat head contour point, the numerical value of the cat body contour point, the numerical value of the cat leg contour point, and the like; further, the object classification device inputs the eigenvalue corresponding to the image data vector into the output layer subnetwork, and performs category feature extraction on the above eigenvalue based on the category neurons in the output layer subnetwork to obtain category feature data corresponding to the image data vector; further, the category feature data is classified based on the classifier in the output layer subnetwork, and the classification result obtained by the classifier is used as the classification result of the image data vector, that is, the classification result of the image data vector may include cat classification category, and the like.

[0083] In the embodiment of the present disclosure, after determining the category feature data, one or more classifiers in the output layer subnetwork can calculate the probability value of the object to be classified belonging to each classification category according to the category feature data; the category with the largest probability value, or the category with the largest probability value that reaches a preset threshold, is used as the classification category of the object to be classified, and the classification result of the object to be classified is obtained. The classification result may include the classification category of the category feature data and the probability value of each classification category.

[0084] In the disclosed embodiment, the category classifier is cross-connected with the category neurons of the output layer. After the category feature data is determined based on the category neurons of the output layer subnetwork, each category classifier can use the Softmax activation function. The category classifier performs exponential processing on the category feature data based on the Softmax function to obtain the corresponding exponential value, and then calculates the proportion of each index value to the total exponential value, and uses the corresponding proportion value as the output value of each classifier. The output value of each category classifier is used as the probability value of each classification category to obtain the classification result of the object to be classified.

[0085] In yet another embodiment of the present disclosure, the object classification network may also include an input layer subnetwork, a hidden layer subnetwork, and an output layer subnetwork.

[0086] In the embodiments of the present disclosure, the input layer subnetwork can encode the object to be classified, convert the object to be classified into a data vector form, and obtain a feature data vector to be classified of the object to be classified; the hidden layer subnetwork can perform feature extraction on the feature data vector to be classified, and obtain feature values ​​corresponding to the object to be classified; the output layer subnetwork can perform category feature extraction on the feature values, and obtain category feature data of the object to be classified, and the category feature data is used to characterize the classification result of the object to be classified, or the output layer subnetwork can classify the category feature data based on one or more classifiers to obtain the classification result of the object to be classified.

[0087] It should be noted that in the disclosed embodiment, neurons of at least one target hidden layer in the object classification network are activated based on the suppressed rectified linear unit, that is, all neurons of at least one target hidden layer in the object classification network are activated based on the suppressed rectified linear unit. All neurons of the target hidden layer are all neurons of any hidden layer in the object classification network.

[0088] In the embodiment of the present disclosure, each neuron of the hidden layer of the object classification network is set with a Restrained & Rectified Linear function processing unit (RReLU). Wherein, neurons of at least one target hidden layer in the object classification network are activated based on an inhibition-corrected linear unit; the inhibition-corrected linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, and when the input value of the positive correction linear function is greater than or equal to 0, the output value of the positive correction linear function is equal to the input value of the positive correction linear function; when the input value of the negative correction linear function is less than or equal to 0, the output value of the negative correction linear function is equal to the input value of the negative correction linear function, and when the input value of the negative correction linear function is greater than 0, the output value of the negative correction linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0089] In the disclosed embodiment, the corrected linear class function in the inhibitory corrected linear unit may be a positive corrected linear function or a negative corrected linear function, but the corrected linear class functions of the inhibitory corrected linear units of all neurons in an object classification network generally need to be unified into one function.

[0090] In the disclosed embodiment, after the object classification device obtains the object to be classified, it can extract features of the object to be classified based on the neurons of the feature map of the target hidden layer of the hidden layer subnetwork in the object classification network, that is, perform weight multiplication and aggregation processing according to the corresponding set weights to obtain weighted input data, activate the extracted weighted input data through the inhibition corrected linear unit corresponding to the neuron, obtain the activation feature value result, and use the feature activation value result as the output feature value of the corresponding neuron of the corresponding feature map of the corresponding target hidden layer; then, based on the neurons of the feature map of the next target hidden layer, feature extraction is performed on the above-mentioned output features, and the extracted weighted input data is activated through the inhibition corrected linear unit corresponding to the neuron, and the feature activation value result is used as the output feature value of the corresponding neuron of the corresponding feature map of the next target hidden layer, and the output features of the next hidden layer are further transmitted to a deeper hidden layer until the last hidden layer, and the feature activation value result of the inhibition corrected linear unit corresponding to the neurons of the feature map of the last hidden layer is used as the output feature value of the hidden layer.

[0091] It should be noted that in the disclosed embodiment, neurons of at least one target hidden layer in the object classification network are activated based on the suppressed rectified linear unit, that is, all neurons of at least one target hidden layer in the object classification network are activated based on the suppressed rectified linear unit. All neurons of the target hidden layer are all neurons of any hidden layer in the object classification network.

[0092] It should be noted that the corrected linear class function of the neuron inhibition corrected linear unit of the object classification network can use a positive corrected linear function or a negative corrected linear function, but generally all neurons of the object classification network only select one of the corrected linear class functions.

[0093] Optionally, the suppressed rectified linear unit may set a bias value in the weighted input data or in the output value of the suppressed rectified linear unit.

[0094] In the embodiment of the present disclosure, the suppressed corrected linear unit may set a bias value in the weighted input data or in the output value of the suppressed corrected linear unit.

[0095] In the disclosed embodiment, after obtaining the object to be classified, the object to be classified can be classified based on the object classification network to obtain the classification result of the object to be classified. Since the target neuron in the object classification network is activated based on the suppression correction linear unit composed of a linear suppression function and a correction linear class function, the derivative of the suppression correction linear unit in the positive value interval or the negative value interval is a constant. Since a hidden layer has a large number of neurons that satisfy the positive input interval or the negative input interval at the same time, the output feature expression of the weighted input data will not be weakened. The method of suppressing the correction linear unit eliminates the problem of the distribution of the derivative value inherent in the traditional nonlinear activation function to a certain extent. The problem of the gradient vanishing of shallow neurons in the process of training artificial neural networks can be solved, thereby alleviating the problem that the low-level feature parameters of the object classification network cannot be fully trained due to the gradient vanishing, and improving the training accuracy of the object classification network. At the same time, in the disclosed embodiment, the linear suppression coefficient of the suppression correction linear class function of the target hidden layer neuron is a positive value less than 1. Under the action of the suppression mechanism of the corrected linear unit, the flow amplification of feature and error transmission caused by multiple feature maps of the hidden layer and multiple feature replications of multiple neurons in the feature map is suppressed, thereby reducing the influence of the layer-by-layer amplification of the transmitted features and error changes caused by the parameter adjustment of neurons in each hidden layer during the parameter training adjustment of the object classification network on the feature flow and error flow finally received by the target neuron, thereby alleviating the internal covariate drift problem encountered by traditional neural networks. In summary, the object classification network that introduces the above-mentioned suppression corrected linear unit can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.

[0096] In other embodiments of the present disclosure, in order to specifically determine the linear suppression coefficient, the linear suppression coefficient is determined based on the number of first feature data of a single feature map of the input neural network layer, the number of second feature data of a single feature map of the target hidden layer to which the target neuron belongs, the number of normalized feature maps, the convolution kernel size of the target neuron for the input neural network layer feature map, and a layer expansion parameter, the number of normalized feature maps includes the number of first feature maps of the target hidden layer to which the target neuron belongs or the number of second feature maps of the input neural network layer of the target neuron, the layer expansion parameter is a value not less than 1, and the value of the layer expansion parameter of the target neuron is the same as the value of the layer expansion parameter of other neurons in the target hidden layer to which the target neuron belongs.

[0097] Furthermore, the linear suppression coefficient is the product of the multi-feature map normalization parameter, the convolution normalization parameter and the layer expansion parameter.

[0098] In the disclosed embodiment, the multi-feature map normalization parameter is the inverse of the number of normalized feature maps.

[0099] In the embodiment of the present disclosure, the convolution normalization parameter is the reciprocal of the quotient of the product of the convolution kernel size and the second feature data quantity divided by the first feature data quantity.

[0100] In the embodiments of the present disclosure, the input neural network layer of the target neuron may be a neural network layer in the form of an input layer connected to the target neuron on the input side, other hidden layers, a certain timing state of the hidden layer, or a certain timing state of other hidden layers.

[0101] In the embodiment of the present disclosure, the number of first feature maps of the target hidden layer to which the target neuron belongs may be the number of feature maps of the target hidden layer to which the target neuron belongs.

[0102] In the embodiment of the present disclosure, the number of second feature maps of the input neural network layer of the target neuron may be the number of feature maps of the input neural network layer of the target neuron. When the input neural network layer of the target neuron is the input layer, the number of second feature maps of the input neural network layer of the target neuron is 1.

[0103] In the embodiment of the present disclosure, the first feature data quantity of a single feature map of the input neural network layer may be the feature data quantity of a single feature map of the input neural network layer to which the target neuron is connected.

[0104] In some embodiments, when the feature map of the input neural network layer is not pooled by the pooling layer, the first feature data quantity is the total number of neuron output data of the single feature map of the connected input neural network layer.

[0105] In other embodiments, when the feature map of the input neural network layer is pooled through the pooling channel of the pooling layer, the first feature data quantity is the quantity of feature data of the pooling channel of a single feature map of the input neural network layer.

[0106] In some further embodiments, when the feature map of the input neural network layer contains multiple pooling channels of the pooling layers (such as when using the spatial pyramid pooling algorithm), the first number of feature data is the total number of feature data of all pooling channels of the single feature map of the input neural network layer.

[0107] In some further embodiments, when the input neural network layer is the input layer, the first feature data quantity is the total quantity of data in the feature data vectors to be classified in all data channels of the input layer.

[0108] In the embodiment of the present disclosure, the second feature data quantity of a single feature map of a target hidden layer to which a target neuron belongs may be the feature data quantity of a single feature map of a target hidden layer to which the target neuron belongs.

[0109] In some embodiments, when the feature map is not pooled by a pooling layer, the second feature data quantity is the total quantity of neuron output data of a single feature map.

[0110] In some other embodiments, when a feature map is pooled through a pooling channel of a pooling layer, the second feature data quantity is the quantity of feature data of a pooling channel of a single feature map.

[0111] In some further embodiments, when the feature map includes multiple pooling layers and pooling channels (such as when a spatial pyramid pooling algorithm is used), the second feature data quantity is the total quantity of feature data of all pooling channels of a single feature map.

[0112] In the disclosed embodiment, the convolution kernel size of the target neuron for the input neural network layer feature map is the total number of all feature data in the feature data vector of a single feature map of the input neural network layer to which the target neuron is connected on the input side. It should be noted that, for the input layer, the convolution kernel size of the target neuron for the input neural network layer feature map is the product of the convolution kernel size of the target neuron for a single data channel and the number of data channels of the input layer.

[0113] Optionally, the expression for the suppressed rectified linear unit is:

[0114] φ(z)=max(0,z)(Formula 1)

[0115] η(z)=min(0,z)(Formula 2)

[0116] z=γ*q / (n*k*c)*x(Formula 3)

[0117] Wherein, φ(z) is the output value of the suppression-corrected linear unit when the suppression-corrected linear unit uses the positive correction linear function. When z<0, φ(z) is 0, and when z>=0, φ(z)=z; η(z) is the output value of the suppression-corrected linear unit when the suppression-corrected linear unit uses the negative correction linear function. When z>0, η(z) is 0, and when z<=0, η(z)=z; z is the output value of the linear inhibition function in the suppression-corrected linear unit; x is the weighted input data of the above-mentioned suppression-corrected linear unit; γ*q / (n*k*c) is the linear inhibition coefficient of the linear inhibition function in the suppression-corrected linear unit; γ is the layer expansion parameter of the suppression-corrected linear unit uniformly used in the target hidden layer of the object classification network; q is the number of the first feature data of the single feature map of the input neural network layer of the target neuron. When the input neural network layer is the input layer, the number of the first feature data is the number of the feature data to be classified of all data channels of the input layer The total number of data in the quantity, that is, the product of the number of data of the feature data vector to be classified in a single data channel of the input layer and the number of data channels in the input layer; k is the number of second feature data of a single feature map of the target hidden layer to which the target neuron belongs; c is the convolution kernel size (that is, the convolution kernel area) used by the target neuron when convolving the feature data vector of a single feature map of the input neural network layer, that is, the total number of all feature data in the feature data vector of a single feature map of the input neural network layer to which the target neuron is connected on the input side. When the input neural network layer is the input layer, c is the product of the convolution kernel size of the target neuron for the feature data vector to be classified in a single data channel and the number of data channels in the input layer; n is the number of normalized feature maps, that is, the number of first feature maps or the number of second feature maps. When the inverse method of the second feature map number is used to normalize multiple feature maps and the input neural network layer is the input layer, n takes the value of 1.

[0118] It should be noted that, for the calculation of the normalization parameter of the multi-feature map of the object classification network, the reciprocal of the number of the first feature map can be used as the normalization parameter of the multi-feature map, or the reciprocal of the number of the second feature map can be used as the normalization parameter of the multi-feature map. Both methods can achieve the purpose of normalizing the multi-feature map of the feature and error transmission flow, thereby achieving the effect of alleviating the internal covariate drift, but generally only one of the methods is used to determine the normalization parameter of the multi-feature map in the object classification network. For the method of using the reciprocal of the number of the first feature map as the normalization parameter of the multi-feature map, the target neuron receives the feature flow of the lower hidden layer, and performs the normalization of the feature flow of the multi-feature map when transmitting it to the upper layer. The entire hidden layer only transmits a feature flow of the scale of a feature map, which will not cause the multi-feature map amplification of the feature flow. At the same time, the target neuron performs the multi-feature map error flow normalization when transmitting the reverse error to the lower layer. The entire hidden layer only transmits a copy of the error flow of the feature map to the lower hidden layer, which will not cause the error flow amplification caused by the multi-feature map. For the method of using the inverse of the number of second feature maps as the normalization parameter of multiple feature maps, the target neuron normalizes the feature flow of multiple feature maps after receiving the feature flow of multiple feature maps of the lower hidden layer. When transmitting the feature flow to the higher hidden layer, it does not cause the layer-by-layer amplification of the feature flow. At the same time, when forwarding the error flow to all feature maps of the lower hidden layer, a copy of the error flow of the feature map is always transmitted to all the lower feature maps, which does not cause the amplification of the error flow transmitted in the reverse direction.

[0119] It should be noted that the convolution normalization parameter is the quotient of the number of first feature data of a single feature map of the input neural network layer divided by the product of the convolution kernel size of the target hidden layer and the number of second feature data of the single feature map of the target hidden layer to which the target neuron belongs, wherein the quotient of the number of first feature data of the single feature map of the input neural network layer divided by the convolution kernel size of the target hidden layer is the number of neurons of the single feature map of the target hidden layer required for feature extraction of the feature data of a single input feature map under the premise of non-repetitive feature extraction, so the result obtained by dividing this number by the number of second feature data of the single feature map of the target hidden layer to which the target neuron belongs, that is, the convolution normalization parameter, can be used as a compensating processing factor for feature amplification caused by repeated extraction of the same input feature data by multiple neurons in the single feature map of the target hidden layer, thereby alleviating the problem of feature flow transfer caused by repeated convolution of multiple neurons and amplification of reverse error transfer and internal covariate drift.

[0120] It should be noted that in order to adapt to the scenario where the output value of the output layer neuron is too small due to the input layer data value being too small, the hidden layer weight parameter being too small, the normalized suppression parameters in the hidden layer suppression correction unit, etc., it is necessary to adjust the layer expansion parameters based on the normalized suppression parameters in the suppression correction unit. In the embodiment of the present disclosure, the layer expansion parameter γ used in the linear suppression coefficient in the suppression correction linear unit is an arbitrary value not less than 1, and the specific value is not limited, but the linear suppression coefficient after the introduction of the layer expansion parameter must be less than 1, that is, the product of the multi-feature mapping normalization parameter, the convolution normalization parameter and the layer expansion parameter corresponding to the linear suppression function in the suppression correction linear unit must be less than 1. At the same time, it is required that the layer expansion parameters used by the suppression correction linear unit of the neurons in the same hidden layer in the object classification network must be the same. The same layer expansion parameters ensure the fairness of the processing of different features by each neuron in the hidden layer, thereby not affecting the training accuracy. The specific value of the expansion parameter of this layer is generally adjusted and set according to the influence of the loading of the feature data vector to be classified on the output feature value of the highest hidden layer during the object classification network training. When adjusting, it can be adjusted according to the principle that the absolute mean of the output feature values ​​of the neurons in the highest hidden layer is roughly equivalent to the absolute mean of the values ​​of the feature data vector to be classified in the input layer. When the reciprocal normalization method of the number of the first feature map of the target hidden layer is adopted, the absolute mean of the output feature values ​​of the neurons in the highest hidden layer needs to be multiplied by the number of feature maps of the highest hidden layer. This method basically ensures that the hidden layer neurons will not cause feature amplification when extracting features from the input neural network layer, and also solves the problem of too low extracted feature values.

[0121] In the disclosed embodiment, the linear suppression coefficient is the product of the layer expansion parameter, the multi-feature mapping normalization parameter and the convolution normalization parameter. These normalization parameters can be used as compensation processing factors for feature amplification caused by repeated extraction of the same input feature data by multiple neurons, thereby alleviating the amplification of feature flow transmission and internal covariate drift problems caused by repeated convolution of multiple neurons. At the same time, the introduction of the layer expansion parameter in the linear suppression coefficient compensates for the attenuation of the feature transmission in the neural network, which can avoid the phenomenon that the output layer neuron output feature value is too low due to factors such as too small input value of the input layer, too small value of the hidden layer weight parameter, or the normalization parameter factor in the linear suppression coefficient, thereby further improving the neural network training effect. Therefore, in the disclosed embodiment, the object classification network that introduces the above-mentioned suppression correction linear unit can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.

[0122] In order to further solve the internal covariate drift problem encountered by traditional neural networks, in the disclosed embodiment, the linear suppression coefficient is also determined according to the number of input neural network layers corresponding to the target neuron.

[0123] Specifically, the linear suppression coefficient is the product of the normalization parameter of multiple feature maps, the normalization parameter of convolution, the normalization parameter of multiple input neural network layers, and the layer expansion parameter;

[0124] Among them, the normalization parameter of the multi-input neural network layer is the inverse of the number of input neural network layers;

[0125] The normalization parameter of multiple feature maps is the inverse of the number of normalized feature maps;

[0126] The convolution normalization parameter is the reciprocal of the quotient of the product of the convolution kernel size and the second feature data quantity divided by the first feature data quantity.

[0127] In the disclosed embodiment, the number of input neural network layers is the number of neural network layers of the object classification network to which all target neurons are connected on the input side.

[0128] In the embodiments of the present disclosure, the input neural network layer of the target neuron can be a neural network layer in the form of an input layer of an object classification network to which the target neuron is connected on the input side, other hidden layers, a certain time sequence state of the hidden layer, or a certain time sequence state of other hidden layers.

[0129] In the embodiments of the present disclosure, scenarios in which target neurons are simultaneously connected to multiple neural network layers include scenarios in which target neurons are simultaneously connected to multiple other hidden layer networks, scenarios in which target neurons are simultaneously connected to other hidden layers and one or more time series states of the current hidden layers, and scenarios in which target neurons are simultaneously connected to the input layer and one or more time series states of the current hidden layers.

[0130] Optionally, the expression for the suppressed rectified linear unit is:

[0131] φ(z)=max(0,z)(Formula 4)

[0132] η(z)=min(0,z)(Formula 5)

[0133] z=Σ((γ*q i / (n i *k*c i *m))*x i )(Formula 6)

[0134] Wherein, z is the output value of the linear suppression function used by the suppression correction linear unit. The output value of the linear suppression function is the sum of the output values ​​of one or more linear suppression function subunits. Each subunit uses an independent linear suppression coefficient. When the splicing feature extraction mode is adopted, the number of linear suppression function subunits in the linear suppression function is 1; x i is the weighted input data of the linear inhibition function subunit; γ*q i / (n i *k*c i *m) is the linear inhibition coefficient of each linear inhibition function subunit; γ is the layer expansion parameter of the inhibition correction linear unit uniformly used in the target hidden layer of the object classification network; q i is the number of first feature data of the input neural network layer corresponding to the linear inhibition function subunit. When the input neural network layer is the input layer, the number of first feature data is the total number of data in the feature data vector to be classified of all data channels of the input layer. When the splicing feature extraction mode is adopted, q i is the number of feature data of a single feature map after the multi-input neural network layer is concatenated; k is the number of second feature data of a single feature map of the target hidden layer to which the target neuron belongs; c i is the convolution kernel size (i.e., convolution kernel area) used by the target neuron when convolving the feature data vector of a single feature map of the corresponding input neural network layer, that is, the total number of all feature data in the feature data vector of a single feature map of the corresponding input neural network layer connected to the target neuron on the input side; when the corresponding input neural network layer is the input layer, c i is the product of the convolution kernel size of the target neuron for the feature data vector to be classified in a single data channel of the input layer and the number of data channels in the input layer; when the splicing feature extraction mode is used, c i The convolution kernel size when the target neuron convolves the feature data vector of a single feature map after multiple input neural network layers are concatenated; n i is the number of normalized feature maps, that is, the number of first feature maps or the number of second feature maps. When the inverse method of the first feature map is used, n i is the number of feature maps of the target hidden layer to which the target neuron belongs. When the inverse method of the second feature map number is used, n i is the number of feature maps of the corresponding input neural network layer. When the inverse method of the second feature map number is used and the input neural network layer is the input layer, then n i The value is 1. When using the inverse method of the second feature map number and the multi-input neural network layer splicing feature extraction mode, n i is the number of unified feature maps of the input neural network layer; m is the number of input neural network layers.

[0135] It should be noted that, when the target hidden layer has only one input neural network layer, m in formula (6) is 1, that is, when the target hidden layer has only one input neural network layer, the expression of the suppressed rectified linear unit is the above formula (1), formula (2) and formula (3).

[0136] It should be noted that, for the scenario where the target neuron is connected to multiple neural network layers at the same time, when the target neuron performs a separate feature extraction mode on the feature data vectors of the feature maps of multiple input neural network layers, the target neuron needs to set a corresponding linear inhibition function subunit for each input neural network layer in the linear inhibition function of the suppression correction linear unit, and each linear inhibition function subunit uses the corresponding linear inhibition coefficient to multiply the weighted input of the corresponding input network layer, and then the output values ​​of each linear inhibition function subunit are added and unified as the output value of the linear inhibition function. The linear inhibition coefficient of each linear inhibition function subunit can be different, and the linear inhibition coefficient of the linear inhibition function subunit can be set according to the number of normalized feature maps corresponding to the corresponding input network layer, the number of first feature data, the number of second feature data, the number of input neural network layers, and the convolution kernel size of the target neuron for the feature map of the corresponding input neural network layer. Optionally, for multiple neural network layers connected to the target neuron and in the separate feature extraction mode, if the linear inhibition coefficients of multiple linear inhibition function subunits of the linear inhibition function are the same, the target neuron can use a unified linear inhibition function subunit to simultaneously perform a unified linear inhibition process on multiple weighted input data corresponding to different input neural network layers. For the scenario where the target neuron is connected to multiple neural network layers at the same time, and when the target neuron performs a splicing feature extraction mode on the feature data vectors of the feature maps of multiple input neural network layers, the linear suppression function only needs to set one linear suppression function sub-unit to uniformly perform linear suppression processing on the weighted input data corresponding to the feature data vectors of the feature maps of the spliced ​​multiple-input neural network layers. At this time, the first feature data quantity is the total number of data in the feature data vectors of the single feature maps of the spliced ​​multiple-input neural network layers, the normalized feature map quantity used is the unified feature map quantity of a single input neural network layer (when the inverse of the second feature map quantity method is used) or the feature map quantity of the target hidden layer (when the inverse of the first feature map quantity method is used), and the convolution kernel size used is the convolution kernel size when the target neuron performs convolution processing on the feature data vector of the feature map of the spliced ​​multiple-input neural network layer.

[0137] In the disclosed embodiment, the layer expansion parameter γ used in the linear suppression coefficient in the suppression correction linear unit is an arbitrary value not less than 1, and the specific value is not limited, but the linear suppression coefficient after the introduction of the layer expansion parameter must be less than 1, that is, the product of the multi-feature map normalization parameter, the convolution normalization parameter and the layer expansion parameter corresponding to the linear suppression function in the suppression correction linear unit must be less than 1. At the same time, the layer expansion parameter used by the suppression correction linear unit of the neurons in the same hidden layer of the object classification network must be the same. The same layer expansion parameter ensures the fairness of the processing of different features by each neuron in the hidden layer, thereby not affecting the training accuracy. The specific value of the layer expansion parameter is generally adjusted and set according to the influence of the output feature value size of the highest hidden layer after the feature data vector to be classified is loaded during the training of the object classification network. When adjusting, it can be adjusted according to the principle that the absolute value mean of the output feature value of the neurons in the highest hidden layer is roughly equivalent to the absolute value mean of the value of the feature data vector to be classified in the input layer. When the reciprocal normalization method of the number of the first feature map of the target hidden layer is adopted, the absolute value mean of the output feature value of the neurons in the highest hidden layer needs to be multiplied by the number of feature maps of the highest hidden layer. This method basically ensures that the hidden layer neurons will not cause feature amplification when extracting features from the input neural network layer, based on the replication, amplification and normalization of multi-neuron features in each hidden layer. It also solves the problem of too low extracted feature values.

[0138] Therefore, in the disclosed embodiment, the linear suppression coefficient is the product of the normalization parameter of multiple feature maps, the normalization parameter of convolution and the layer expansion parameter, or the linear suppression coefficient is the product of the normalization parameter of multiple feature maps, the normalization parameter of convolution, the normalization parameter of multi-input neural network layer and the layer expansion parameter. These normalization parameters can be used as compensation processing factors for feature amplification caused by repeated extraction of the same input feature data by multiple neurons, thereby alleviating the amplification of feature flow transmission and internal covariate drift problems caused by repeated convolution of multiple neurons. At the same time, the introduction of the layer expansion parameter in the linear suppression coefficient compensates for the attenuation of the feature transmission in the neural network, which can avoid the phenomenon that the output feature value of the output layer neuron is too low due to factors such as too small input value of the input layer, too small value of the weight parameter of the hidden layer, or the normalization parameter factor in the linear suppression coefficient, thereby further improving the neural network training effect. Therefore, in the disclosed embodiment, the object classification network introducing the above-mentioned suppression correction linear unit can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.

[0139] Figure 3 It is a logical diagram of an object classification network. The object classification network can be a classification network of an artificial neural network, combined with Figure 3An exemplary explanation of the process of object classification.

[0140] Figure 3 The object classification network in may include three functional layers: input layer, hidden layer, and output layer. The input layer is the above-mentioned input layer subnetwork, the hidden layer is the above-mentioned hidden layer subnetwork, and the output layer is the above-mentioned output layer subnetwork. Among them, the hidden layer includes two hidden layers, hidden layer 1 and hidden layer 2, and neurons in all hidden layers are set with suppression correction linear units (linear suppression function and forward correction linear function are used in this example), and linear suppression processing is performed on the feature flow and error flow flowing through this neuron. Specifically, neurons B, C, and D with different feature maps in hidden layer 1 are respectively connected to the feature data vectors to be classified in the input layer, and the feature data vectors to be classified are collected and summarized according to the corresponding weight parameters to obtain weighted input data, and the weighted input data are linearly suppressed by RReLU to obtain the output feature data extracted by neurons B, C, and D. The neuron output data generated by B, C, and D are feature extracted and linearly suppressed by neurons E and F in hidden layer 2, and the output feature data extracted by neurons E and F are obtained. Furthermore, the category neurons G and H of the output layer perform category feature extraction on the output feature data of neurons E and F, and obtain their own category feature data, and then the classifier performs classification according to the category feature data to obtain a classification result. Figure 3The solid line in the middle is a schematic diagram of the transmission process of the feature flow, and the dotted line is a schematic diagram of the transmission process of the error flow. In the process of transmitting feature A in the feature data vector to be classified upward, since neurons B, C, and D normalize the received weighted data according to the number of feature maps of hidden layer 1 (3), it is ensured that the feature flow (Feature E) received by neuron E is still 1 times the feature flow. At the same time, in the process of transmitting the error of the category neuron H downward, the error is normalized according to the number of feature maps of hidden layer 2 (2), so that the error (Error C) received by neuron C is still 1 times the error flow. In this way, the internal covariate drift phenomenon caused by the amplification of the feature flow or error flow by multiple feature maps of different hidden layers is alleviated. Furthermore, the layer expansion parameters of each hidden layer can be adjusted according to the principle that the absolute value mean of the output feature value of the highest hidden layer neuron is roughly equivalent to the absolute value mean of the value of the feature data vector to be classified in the input layer, so as to alleviate the excessive attenuation of the feature flow caused by setting the suppression correction linear class function, so that the output value of the output layer neuron is maintained in a suitable range. This alleviates the internal covariate drift problem encountered by traditional neural networks, thereby ensuring the accuracy of parameter learning gradients calculated based on the characteristic flow and error flow received by neurons, and improving the object classification network training accuracy and object classification accuracy. At the same time, due to the use of suppressed corrected linear units as activation functions, the activation function is a constant in the positive range, which will not cause the disappearance of the reverse propagation error of neurons in the lower hidden layers, thereby solving the gradient vanishing problem encountered by the nonlinear function mechanism of traditional neural networks and the problem that the low-level characteristic parameters cannot be fully learned and trained, and improving the object classification network training accuracy.

[0141] Figure 4 is a schematic diagram of the structure of an object classification network. The object classification network can be a classification network of an artificial neural network, combined with Figure 4 The classification process of the object to be classified is explained exemplarily. Figure 4 The object classification network in may include an input layer, a hidden layer and an output layer, wherein the input layer is the above-mentioned input layer subnetwork, the hidden layer is the above-mentioned hidden layer subnetwork, and the output layer is the above-mentioned output layer subnetwork. The hidden layer includes 7 convolutional layers and 2 pooling layers, and the output layer includes two category neurons and two classifiers. The input layer may be the input layer subnetwork described above, and the hidden layer and the output layer may constitute the above-mentioned object classification network.

[0142] The input layer encodes the object to be classified and converts it into a feature data vector to be classified, so as to provide it to the hidden layer of the object classification network for access.

[0143] Among them, the hidden layer performs feature extraction on the feature data vector to be classified of the object to be classified based on convolution layers 1 to 7 and pooling layers 2 and 4, and finally outputs the output feature value of the object to be classified through each feature map of convolution layer 7; the two category neurons of the output layer perform category feature extraction on the output feature value of each feature map output of convolution layer 7 to obtain the category feature data of the object to be classified.

[0144] The convolution kernel size of convolution layer 1 is 5*5*3, including 10 feature maps, and the number of neurons in each feature map is 28*28; the convolution kernel size of convolution layer 2 is 5*5*10, including 10 feature maps, and the number of neurons in each feature map is 24*24; the convolution kernel size of convolution layer 3 is 3*3*10, including 10 feature maps, and the number of neurons in each feature map is 10*10; the convolution kernel size of convolution layer 4 is 3*3*10, including 10 feature maps, and the number of neurons in each feature map is 8*8; the convolution kernel size of convolution layer 5 is 3*3*10, including 10 Feature map, the number of neurons in each feature map is 2*2; the convolution kernel size of convolution layer 6 is 2*2*30, including 30 feature maps, and the number of neurons in each feature map is 1*1; the convolution kernel size of convolution layer 7 is 1*1*30, including 30 feature maps, and the number of neurons in each feature map is 1*1; pooling layer 2, the pooling size is 2*2, including 10 pooling kernels, respectively pooling the output data of the corresponding feature map of convolution layer 2; pooling layer 4, the pooling size is 2*2, including 10 pooling kernels, respectively pooling the output data of the corresponding feature map of convolution layer 4. Among them, the neurons in the above convolution layer are set with suppression correction linear units, and the weighted input data of the connected target neurons are activated based on the linear suppression function and the correction linear class function. The linear suppression function and the correction linear function of the suppression correction linear unit can be specifically determined based on the above formulas 1, 2, and 3.

[0145] Furthermore, each category classifier of the output layer is cross-connected with the category neurons of the output layer. Each category classifier can use the Softmax activation function. The category classifier performs exponential processing on the category feature data based on the Softmax activation function to obtain the corresponding exponential value, and then calculates the proportion of each exponential value to the total exponential value as the output value of each classifier. The output value of each classifier is used as the probability value of each classification category to obtain the classification result of the object to be classified.

[0146] Therefore, in the disclosed embodiment, since the object classification network is activated based on the suppression correction linear unit, the expressiveness of its output features will not be weakened, eliminating the problem of saturation region in the input distribution of the traditional nonlinear activation function, as well as the resulting vanishing gradient phenomenon of the low-level feature parameters of the neural network and the problem that the low-level feature parameters of the neural network cannot be fully trained. The linear normalization coefficient is the product of the layer expansion parameter, the multi-feature mapping normalization parameter and the convolution normalization parameter. Based on the suppression correction linear unit, the feature amplification caused by the feature replication and amplification of the multi-feature mapping and the feature amplification caused by the multiple convolution replications of the same feature by multiple neurons of the feature mapping is suppressed, and the stability of the object classification network features and error transmission flow is achieved, thereby ensuring the stability of the learning gradient of each hidden layer parameter, solving the gradient vanishing problem encountered by the traditional neural network, and alleviating the occurrence of the internal covariate problem, thereby improving the training accuracy. Further, in order to reduce the output feature value caused by the feature attenuation caused by the suppression, the layer expansion parameter can be adjusted and set according to the method that the absolute value mean of the output feature value of the highest hidden layer neuron is roughly consistent with the absolute value mean of the feature data vector data to be classified in the input layer, further ensuring the parameter training accuracy of the object classification network. In addition, the object classification network using the suppressed corrected linear unit does not need to perform complex normalization calculations on the input data. Compared with the object classification network using the batch normalization algorithm, the amount of calculation in the training process is simplified. The suppressed corrected linear unit activation mechanism is different from the batch normalization mechanism and can be used in the field of recurrent neural networks (recursive neural networks). In addition, compared with the activation function using the gradient clipping method, the suppressed corrected linear unit activation mechanism does not need to force the gradient to be limited to a certain range, thus ensuring the training accuracy of the object classification network. Obviously, the use of the suppressed corrected linear unit activation mechanism improves the processing power of the object classification network, and can ensure the training accuracy of the object classification network without increasing the training computing power, which can meet the classification needs of users.

[0147] Further, Figure 4 The object classification network test performance shown. Specifically, 800 test samples are input Figure 4 In the object classification network shown, the parameters such as the computation time, the number of steps used, and the accuracy of the object classification network are calculated, and the performance of the object classification network is tested based on the parameters such as the computation time, the number of steps used, and the accuracy.

[0148] Table 1: Test results of the Repressed Rectified Linear Unit (RReLU) and traditional nonlinear activation functions

[0149] Activation Function algorithm Iteration cycle (Epoch) Accuracy RReLU Adam 68 83.50% Relu Adam 79 76.65% Sigmoid Adam 89 79.44%

[0150] It can be seen from Table 1 that the accuracy of the object classification network obtained by activating neurons using the above-mentioned suppressed corrected linear unit mechanism is better than the result obtained by activating using the traditional nonlinear activation function.

[0151] The following is an embodiment of an object classification device provided in an embodiment of the present invention. The device and the object classification methods of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the object classification device, please refer to the embodiment of the above-mentioned object classification method.

[0152] This embodiment provides an object classification device. Figure 5 , the device specifically comprises:

[0153] An object acquisition module 510 is used to acquire an object to be classified;

[0154] The object classification module 520 is used to classify the to-be-classified object based on the object classification network to obtain the classification result of the to-be-classified object. The neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the positive correction linear function is activated. The output value of the linear function is 0. When the input value of the positive corrected linear function is greater than or equal to 0, the output value of the positive corrected linear function is equal to the input value of the positive corrected linear function; when the input value of the negative corrected linear function is less than or equal to 0, the output value of the negative corrected linear function is equal to the input value of the negative corrected linear function, and when the input value of the negative corrected linear function is greater than 0, the output value of the negative corrected linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0155] The technical solution provided by the embodiment of the present disclosure can classify the object to be classified based on the object classification network after obtaining the object to be classified, and obtain the classification result of the object to be classified. Since the target neuron in the object classification network is activated based on the suppression correction linear unit, the derivative of the suppression correction linear unit in the positive value interval or the negative value interval is a constant. Since there are a large number of neurons that meet the positive input interval or the negative input interval in a hidden layer at the same time, the output feature expression of the weighted input data will not be weakened. The method of suppressing the correction linear unit eliminates the problem of the distribution of the derivative value inherent in the traditional nonlinear activation function to a certain extent. The problem of the gradient vanishing of shallow neurons in the process of training artificial neural networks can be solved, thereby alleviating the problem that the low-level feature parameters cannot be fully trained due to the gradient vanishing. At the same time, the linear suppression coefficient of the suppression correction linear class function of the target hidden layer neuron is a positive value less than 1. Under the action of the suppression mechanism of the corrected linear unit, the flow amplification of feature and error transmission caused by multiple feature maps of the hidden layer and multiple feature replications of multiple neurons in the feature map is suppressed, thereby reducing the influence of the layer-by-layer amplification of the transmitted features and error changes caused by the parameter adjustment of neurons in each hidden layer during the parameter training adjustment of the object classification network on the feature flow and error flow finally received by the target neuron, thereby alleviating the internal covariate drift problem encountered by traditional neural networks. In summary, the object classification network that introduces the above-mentioned suppression corrected linear unit can improve the training accuracy during the training process. When classifying objects based on the trained object classification network, the classification accuracy can be improved to meet the classification needs of users.

[0156] Optionally, the linear suppression coefficient is determined based on the number of first feature data of a single feature map of the input neural network layer, the number of second feature data of a single feature map of the target hidden layer to which the target neuron belongs, the number of normalized feature maps, the convolution kernel size of the target neuron for the input neural network layer feature map, and a layer expansion parameter, the number of normalized feature maps includes the number of first feature maps of the target hidden layer to which the target neuron belongs or the number of second feature maps of the input neural network layer of the target neuron, the layer expansion parameter is a value not less than 1, and the value of the layer expansion parameter of the target neuron is the same as the value of the layer expansion parameter of other neurons in the target hidden layer to which the target neuron belongs.

[0157] Optionally, the linear suppression coefficient is the product of a multi-feature map normalization parameter, a convolution normalization parameter, and a layer expansion parameter;

[0158] Among them, the multi-feature map normalization parameter is the inverse of the number of normalized feature maps;

[0159] The convolution normalization parameter is the reciprocal of the quotient of the product of the convolution kernel size and the second feature data quantity divided by the first feature data quantity.

[0160] Optionally, the linear inhibition coefficient is also determined based on the number of input neural network layers of the target neuron;

[0161] Among them, the linear suppression coefficient is the product of the multi-feature map normalization parameter, the convolution normalization parameter, the multi-input neural network layer normalization parameter and the layer expansion parameter;

[0162] Among them, the normalization parameter of the multi-input neural network layer is the inverse of the number of input neural network layers;

[0163] The normalization parameter of multiple feature maps is the inverse of the number of normalized feature maps;

[0164] The convolution normalization parameter is the reciprocal of the quotient of the product of the convolution kernel size and the second feature data quantity divided by the first feature data quantity.

[0165] Optionally, the object classification module 520 includes: a first feature extraction unit and a second feature extraction unit;

[0166] A first feature extraction unit is used to extract features of the object to be classified based on the hidden layer sub-network to obtain feature values ​​corresponding to the object to be classified;

[0167] The second feature extraction unit is used to extract category features from the feature values ​​based on the output layer sub-network to obtain category feature data corresponding to the object to be classified, and the category feature data is used to characterize the classification result of the object to be classified.

[0168] Through an object classification device according to an embodiment of the present invention, the training accuracy of the object classification network is improved based on the suppressed corrected linear unit mechanism. When the trained object classification network is used for object classification, the object classification accuracy is improved to meet the classification needs of users.

[0169] The object classification device provided in the embodiment of the present invention can execute the object classification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0170] See also Figure 6 This embodiment provides an object classification device 600, which includes: one or more processors 620; a storage device 610, which is used to store one or more programs. When the one or more programs are executed by the one or more processors 620, the one or more processors 620 implement the object classification method provided by the embodiment of the present invention, including:

[0171] Get the object to be classified;

[0172] Based on the object classification network, the object to be classified is classified to obtain the classification result of the object to be classified, and the neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, and when the input value of the positive correction linear function is greater than or equal to 0, the output value of the positive correction linear function is equal to the input value of the positive correction linear function; when the input value of the negative correction linear function is less than or equal to 0, the output value of the negative correction linear function is equal to the input value of the negative correction linear function, and when the input value of the negative correction linear function is greater than 0, the output value of the negative correction linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0173] Of course, those skilled in the art can understand that the processor 620 can also implement the technical solution of the object classification method provided by any embodiment of the present invention.

[0174] Figure 6 The object classification device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0175] like Figure 6 As shown, the object classification device 600 includes a processor 620, a storage device 610, an input device 630, and an output device 640; the number of the processor 620 in the device can be one or more. Figure 6 A processor 620 is taken as an example; the processor 620, the storage device 610, the input device 630 and the output device 640 in the device can be connected via a bus or other means. Figure 6 The connection via bus 650 is taken as an example.

[0176] The storage device 610, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the object classification method in the embodiment of the present invention (for example, the object acquisition module and object classification module in the object classification device).

[0177] The storage device 610 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the storage device 610 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device 610 may further include a memory remotely arranged relative to the processor 620, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0178] The input device 630 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the device, for example, it can include at least one of a mouse, a keyboard and a touch screen. The output device 640 can include a display device such as a display screen.

[0179] It should be noted that the object classification device in the embodiment of the present invention can also be implemented by a cloud server.

[0180] This embodiment provides a storage medium containing computer executable instructions. When the computer executable instructions are executed by a computer processor, they are used to perform an object classification method. The method includes:

[0181] Get the object to be classified;

[0182] Based on the object classification network, the object to be classified is classified to obtain the classification result of the object to be classified, and the neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, and when the input value of the positive correction linear function is greater than or equal to 0, the output value of the positive correction linear function is equal to the input value of the positive correction linear function; when the input value of the negative correction linear function is less than or equal to 0, the output value of the negative correction linear function is equal to the input value of the negative correction linear function, and when the input value of the negative correction linear function is greater than 0, the output value of the negative correction linear function is 0; the linear inhibition coefficient is a positive value less than 1, and the value of the linear inhibition coefficient of the target neuron is the same as the value of the corresponding linear inhibition coefficient of other neurons in the target hidden layer to which the target neuron belongs.

[0183] Of course, the computer executable instructions of a storage medium including computer executable instructions provided by an embodiment of the present invention are not limited to the above method operations, and can also execute related operations in the object classification method provided by any embodiment of the present invention.

[0184] Through the above description of the implementation method, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, server, cloud server, or network device, etc.) to execute the object classification method provided by each embodiment of the present invention.

[0185] Note that the settings of the layer extension parameters in the above-mentioned embodiments of the present invention are only recommended settings of the layer extension parameters. Those skilled in the art will understand that for those skilled in the art, adopting other methods of setting the layer extension parameters will not deviate from the protection scope of the present invention.

[0186] Note that the setting of the linear suppression coefficient in the above-mentioned embodiment of the present invention is only the setting of the optimal linear suppression coefficient. Those skilled in the art will understand that for those skilled in the art, adopting a setting value close to the linear suppression coefficient in the above-mentioned embodiment (such as the setting value of the adopted linear suppression coefficient is compared with the setting value of the linear suppression coefficient in the above-mentioned embodiment, the closeness or the degree of conformity is within 90%) will not deviate from the protection scope of the present invention.

[0187] Note that the neurons in the hidden layer of the above embodiment of the present invention are set to suppress the corrected linear unit only as the best deployment embodiment. Those skilled in the art will understand that it is convenient for those skilled in the art to partially adopt the deployment method in the above embodiment (such as using a local number of hidden layers or using a local number of neurons in the hidden layer to set the corrected linear unit), or to modify the deployment method in the above embodiment (such as merging the linear suppression function and the corrected linear class function in the original corrected linear unit into And η(x)=min(0,γ*q / (n*k*c*m)*x)) or adjustment based on the deployment method in the above embodiment (such as modifying the corrected linear function by adding a bias value to the input data or output data of the original corrected linear function) will not depart from the scope of protection of the present invention.

[0188] Note that the above are only preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present disclosure. Therefore, although the present disclosure is described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. A method for object classification, characterized in that: include: Acquire an object to be classified, wherein the object to be classified includes a color image; The object to be classified is classified based on the object classification network to obtain the classification result of the object to be classified, and the neurons of at least one target hidden layer in the object classification network are activated based on the inhibition correction linear unit; the inhibition correction linear unit includes a linear inhibition function and a correction linear class function, and the correction linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input and the corresponding linear inhibition coefficient; the correction linear class function is a positive correction linear function or a negative correction linear function; when the input value of the positive correction linear function is less than 0, the output value of the positive correction linear function is 0, when the input value of the positive corrected linear function is greater than or equal to 0, the output value of the positive corrected linear function is equal to the input value of the positive corrected linear function; when the input value of the negative corrected linear function is less than or equal to 0, the output value of the negative corrected linear function is equal to the input value of the negative corrected linear function, and when the input value of the negative corrected linear function is greater than 0, the output value of the negative corrected linear function is 0; the linear suppression coefficient is a positive value less than 1, and the value of the linear suppression coefficient of the target neuron is the same as the value of the corresponding linear suppression coefficient of other neurons in the target hidden layer to which the target neuron belongs; When the object to be classified is the color image, the input layer converts the color image into three feature data vectors to be classified, namely red, green and blue, and transmits the three feature data vectors to be classified to the hidden layer through three data channels.

2. The method according to claim 1, characterized in that The linear suppression coefficient is determined based on the first feature data quantity of a single feature map of the input neural network layer, the second feature data quantity of a single feature map of the target hidden layer to which the target neuron belongs, the normalized feature map quantity, the convolution kernel size of the target neuron for the input neural network layer feature map, and a layer expansion parameter, wherein the normalized feature map quantity includes the first feature map quantity of the target hidden layer to which the target neuron belongs or the second feature map quantity of the input neural network layer of the target neuron, the layer expansion parameter is a value not less than 1, and the value of the layer expansion parameter of the target neuron is the same as the value of the layer expansion parameter of other neurons in the target hidden layer to which the target neuron belongs.

3. The method according to claim 2, characterized in that The linear suppression coefficient is the product of a multi-feature mapping normalization parameter, a convolution normalization parameter, and a layer expansion parameter; Wherein, the multi-feature map normalization parameter is the inverse of the number of normalized feature maps; The convolution normalization parameter is the reciprocal of a quotient of a product of a convolution kernel size and a quantity of the second feature data divided by a quantity of the first feature data.

4. The method according to claim 2, characterized in that: The linear inhibition coefficient is also determined according to the number of input neural network layers of the target neuron; Wherein, the linear suppression coefficient is the product of a multi-feature map normalization parameter, a convolution normalization parameter, a multi-input neural network layer normalization parameter, and a layer expansion parameter; Wherein, the normalization parameter of the multi-input neural network layer is the inverse of the number of the input neural network layers; The multi-feature map normalization parameter is the inverse of the number of normalized feature maps; The convolution normalization parameter is the reciprocal of a quotient of a product of a convolution kernel size and a quantity of the second feature data divided by a quantity of the first feature data.

5. The method according to claim 1, characterized in that The object classification network includes a hidden layer sub-network and an output layer sub-network; The classifying the object to be classified based on the object classification network to obtain the classification result of the object to be classified includes: Extracting features of the object to be classified based on the hidden layer sub-network to obtain feature values ​​corresponding to the object to be classified; Based on the output layer sub-network, category features are extracted from the feature values ​​to obtain category feature data corresponding to the object to be classified, and the category feature data is used to characterize the classification result of the object to be classified.

6. An object classification device, characterized in that: include: An object acquisition module, used to acquire an object to be classified, wherein the object to be classified includes a color image; An object classification module is used to classify the object to be classified based on an object classification network to obtain a classification result of the object to be classified, wherein neurons of at least one target hidden layer in the object classification network are activated based on an inhibition-corrected linear unit; the inhibition-corrected linear unit includes a linear inhibition function and a corrected linear class function, and the corrected linear class function uses the output of the linear inhibition function as input; the linear inhibition function is composed of the product of the input multiplied by the corresponding linear inhibition coefficient; the corrected linear class function is a positive corrected linear function or a negative corrected linear function; when the input value of the positive corrected linear function is less than 0, the positive corrected linear function The output value of the number is 0, when the input value of the positive corrected linear function is greater than or equal to 0, the output value of the positive corrected linear function is equal to the input value of the positive corrected linear function; when the input value of the negative corrected linear function is less than or equal to 0, the output value of the negative corrected linear function is equal to the input value of the negative corrected linear function, and when the input value of the negative corrected linear function is greater than 0, the output value of the negative corrected linear function is 0; the linear suppression coefficient is a positive value less than 1, and the value of the linear suppression coefficient of the target neuron is the same as the value of the corresponding linear suppression coefficient of other neurons in the target hidden layer to which the target neuron belongs; When the object to be classified is the color image, the input layer converts the color image into three feature data vectors to be classified, namely red, green and blue, and transmits the three feature data vectors to be classified to the hidden layer through three data channels.

7. An object classification device, characterized in that: include: processor; A memory for storing executable instructions; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the object classification method described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the object classification method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image classification method and device and related components

    CN112270343A

  • Visual suppression of selective tissue in image data

    WO2013144794A2