Image recognition method and device, equipment and storage medium
By introducing the reverse double-residual module into the image recognition model, the gradient vanishing and overfitting problems are solved, the accuracy and efficiency of image recognition are improved, and more efficient image recognition processing is achieved.
Patent Information
- Application Number
- CN202311867151.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing image recognition model has the problem of gradient disappearing or explosion during the training process, which makes model training difficult and the accuracy of image recognition tasks is low.
The reverse double-residual module is adopted, including the residual computing layer and the reverse residual computing layer. By combining and separating the image information, the overfitting problem during the training process is reduced, and the model computing efficiency is optimized in the application stage.
It improves the training fluency and accuracy of the image recognition model, reduces the time-consuming operation during inference, and improves the generalization ability and execution efficiency of the model.
Smart Images

Figure CN120236324A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly relates to an image recognition method, apparatus, device and storage medium. Background Art
[0002] In image recognition tasks (such as gesture classification models), model training can be used to obtain an image recognition model with relatively high accuracy. However, relying solely on model training to improve image recognition accuracy is too single, resulting in low accuracy of image recognition tasks. Summary of the Invention
[0003] Embodiments of the present invention provide an image recognition method, apparatus, device and storage medium, aiming to effectively improve the accuracy of image recognition tasks.
[0004] In a first aspect, an embodiment of the present invention provides a method, which includes:
[0005] Obtain image information to be recognized;
[0006] Preprocess the image information to be recognized to obtain first processed image information;
[0007] Use an image recognition model to perform recognition processing on the first processed image information to obtain image recognition result information.
[0008] Optionally, the using an image recognition model to perform recognition processing on the first processed image information to obtain image recognition result information includes:
[0009] Use the image recognition model to perform first image processing on the first processed image information to obtain second processed image information;
[0010] Use the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information;
[0011] Based on the target processed image information, determine the image recognition result information.
[0012] Optionally, the using the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information includes:
[0013] Use the residual operation layer of the image recognition model to process the first processed image information and the second processed image information to obtain residual features;
[0014] Use the corresponding inverse residual operation layer of the residual operation layer to process the residual features and the first processed image information, and separately obtain the target processed image information.
[0015] Optionally, the image recognition model is obtained based on the following steps:
[0016] Process the first training image information and the second training image information by using the residual operation layer to obtain first training residual features;
[0017] Delete the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features;
[0018] Based on the reverse residual operation layer, process the second training residual features and the first training image information, separate and obtain the target training image information, and train the image recognition model based on the target training image information.
[0019] Optionally, the step of deleting the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features includes:
[0020] If the data type of the first training residual features is not the target data type, convert the data type of the first training residual features to the target data type;
[0021] Set at least one data feature in the converted first training residual features to the target value through the inactivation layer to obtain the second training residual features.
[0022] Optionally, before using the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information, it further includes:
[0023] Obtain the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, and the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information;
[0024] If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then perform second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information.
[0025] Optionally, before using the image recognition model to perform first image processing on the first processed image information to obtain second processed image information, it further includes:
[0026] Obtain the number of network levels in the image recognition model;
[0027] If the number of levels of the network level is greater than or equal to the first number of levels and less than the second number of levels, perform the first image processing on the first processed image information by using the image recognition model to obtain second processed image information.
[0028] If the number of levels of the network level is less than the first number of levels, output the image recognition result information according to the second processed image information;
[0029] If the number of levels of the network level is greater than or equal to the second number of levels, perform third image processing on the first processed image information and the second processed image information by using a one-way residual module to obtain target processed image information.
[0030] In a second aspect, an embodiment of the present invention provides an image recognition device, and the image recognition device includes:
[0031] An acquisition unit, configured to acquire image information to be recognized;
[0032] A preprocessing unit, configured to perform preprocessing on the image information to be recognized to obtain first processed image information;
[0033] A recognition unit, configured to perform recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information.
[0034] Preferably, the recognition unit performs recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information, including:
[0035] Performing first image processing on the first processed image information by using the image recognition model to obtain second processed image information;
[0036] Performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information;
[0037] Determining the image recognition result information based on the target processed image information;
[0038] Preferably, the recognition unit performs second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information, including:
[0039] Processing the first processed image information and the second processed image information by using a residual operation layer of the image recognition model to obtain residual features;
[0040] Process the residual feature and the first processed image information by using the corresponding inverse residual operation layer of the residual operation layer, and separate to obtain the target processed image information;
[0041] Preferably, the image recognition model is obtained by the recognition unit based on the following steps:
[0042] Process the first training image information and the second training image information by using the residual operation layer to obtain a first training residual feature;
[0043] Delete the connections between at least one data feature in the first training residual feature through the inactivation layer of the image recognition model to obtain a second training residual feature;
[0044] Based on the inverse residual operation layer, process the second training residual feature and the first training image information, separate to obtain the target training image information, and train the image recognition model based on the target training image information;
[0045] Preferably, the recognition unit deletes the connections between at least one data feature in the first training residual feature through the inactivation layer of the image recognition model to obtain a second training residual feature, including:
[0046] If the data type of the first training residual feature is not the target data type, convert the data type of the first training residual feature to the target data type;
[0047] Set at least one data feature in the converted first training residual feature to a target value through the inactivation layer to obtain the second training residual feature;
[0048] Preferably, before the recognition unit uses the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain the target processed image information, it further includes:
[0049] Obtain the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, and the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information;
[0050] If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then perform second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain the target processed image information;
[0051] Preferably, before the recognition unit performs first image processing on the first processed image information by using the image recognition model to obtain second processed image information, the following steps are further included:
[0052] Obtain the number of levels of the network levels in the image recognition model;
[0053] If the number of levels of the network levels is greater than or equal to the first number of levels and less than the second number of levels, perform the first image processing on the first processed image information by using the image recognition model to obtain second processed image information;
[0054] If the number of levels of the network levels is less than the first number of levels, output the image recognition result information according to the second processed image information;
[0055] If the number of levels of the network levels is greater than or equal to the second number of levels, perform third image processing on the first processed image information and the second processed image information by using a one-way residual module to obtain target processed image information.
[0056] In a third aspect, an embodiment of the present invention further provides an image recognition device, including a memory storing multiple computer programs; when the computer programs are executed by the processor, the processor is caused to execute the steps of any one of the image recognition methods provided by the embodiments of the present invention.
[0057] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, the computer-readable storage medium storing multiple computer programs; when the computer programs are run on an electronic device, the processor is caused to execute the steps of any one of the image recognition methods provided by the embodiments of the present invention.
[0058] The present invention obtains image information to be recognized; preprocesses the image information to be recognized to obtain first processed image information; and performs recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information. In this way, before performing model recognition on the image information to be recognized, the image information to be recognized will be preprocessed first to make it more suitable for the processing of the image recognition model, which is beneficial to improving the accuracy of the image recognition model for image recognition. Description of the Drawings
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0060] Figure 1It is a schematic flowchart of an embodiment of the image recognition method provided in the embodiment of the present invention;
[0061] Figure 2 It is a schematic flowchart of another embodiment of the image recognition method provided in the embodiment of the present invention;
[0062] Figure 3 It is a schematic diagram of the application scenario of the image recognition model provided in the embodiment of the present invention;
[0063] Figure 4 It is a schematic structural diagram of the image recognition device provided in the embodiment of the present invention;
[0064] Figure 5 It is a schematic structural diagram of the image recognition device provided in the embodiment of the present invention. Specific embodiments
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention. At the same time, in the description of the embodiments of the present invention, terms such as "first" and "second" are only used for descriptive distinction and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0066] The embodiment of the present invention provides an image recognition method, device, image recognition device and computer-readable storage medium.
[0067] Specifically, this embodiment will be described from the perspective of the image recognition device. This image recognition device can be specifically integrated in the image recognition device, that is, the image recognition method in the embodiment of the present invention can be executed by the image recognition device.
[0068] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution subject is an image recognition device as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from that shown in the drawings.
[0069] According to the description of the background art of the present invention, when the image recognition model containing residual connections runs, the residual connections need to perform element-wise addition operations with the main branch, resulting in additional increase in the operation time of the model and reduction in the fluency of the image recognition model.
[0070] To solve the above problems, the present invention discloses an image recognition method. Please refer to Figure 1 , and the specific process of the image recognition method can be as follows in steps S10 to S40, where:
[0071] Step S10: Obtain the image information to be recognized;
[0072] In this embodiment, the image information to be recognized is an image whose content needs to be recognized. For example, it is a hand image for which a gesture needs to be recognized. The image information to be recognized can be obtained when a preset condition is met. The preset conditions include at least one of the following:
[0073] Receiving an image acquisition instruction triggered by the user;
[0074] Reaching the preset image acquisition time;
[0075] The target to be recognized appears in the image captured by the camera. For example, when recognizing a gesture, if a human hand is recognized in the image captured by the camera, the current image captured by the camera can be used as the image to be recognized.
[0076] Step S20: Preprocess the image information to be recognized to obtain first-processed image information;
[0077] In this embodiment, after obtaining the image information to be recognized, it is necessary to preprocess the image information to be recognized, normalize the image information to between 0 - 1 or between 0 - 255, or adopt other normalization methods to process it into first-processed image information, thereby simplifying the subsequent processing operations on the image information to be recognized.
[0078] Step S30: Use an image recognition model to perform recognition processing on the first-processed image information to obtain image recognition result information.
[0079] In this embodiment, after preprocessing the image information to be recognized to obtain first-processed image information, the first-processed image information is input into the image recognition model. The image recognition mode is a model that has previously learned the correspondence between image information and image recognition result information. After processing the first-processed image information, it can find the corresponding image recognition result information of the first-processed image information as the image recognition result information of the image information to be recognized. For example, obtain a hand image to be recognized, perform normalization processing on the hand image to be recognized to obtain first-processed image information, and use the image recognition model to perform recognition processing on the first-processed image information to obtain gesture information as the image recognition result information of the hand image to be recognized.
[0080] In the technical solution disclosed in this embodiment, the image information to be recognized is obtained; the image information to be recognized is preprocessed to obtain the first processed image information; and the image recognition model is used to perform recognition processing on the first processed image information to obtain the image recognition result information. In this way, before performing model recognition on the image information to be recognized, the image information to be recognized will be preprocessed first, making it more suitable for the processing of the image recognition model, which is beneficial to improving the accuracy of the image recognition model in image recognition.
[0081] Further, step S30 includes:
[0082] The first image processing is performed on the first processed image information by using the image recognition model to obtain the second processed image information;
[0083] The second image processing is performed on the first processed image information and the second processed image information by using the image recognition model to obtain the target processed image information;
[0084] Based on the target processed image information, the image recognition result information is determined.
[0085] In this embodiment, the image recognition model includes a processing module composed of a convolutional layer, a normalization layer, an activation layer, and an attention layer, and a reverse double residual module. The processing module is before the reverse double residual module and preferentially processes the first processed image information. After the first processed image information is input into the processing module, the processing module performs the first image processing on the first processed image information, successively passing through the convolutional layer, the normalization layer, the activation layer, and the attention layer of the processing module. When passing through the convolutional layer, convolutional processing is performed to obtain image information with more prominent features. When passing through the normalization layer, normalization processing is performed to obtain more simplified image information. Activation processing is performed through the activation layer. When passing through the attention layer, weights are set based on the associations between the data features in the image information, making the important information in the image information passing through the attention layer prominent, and finally obtaining the second processed image information after processing. The first processed image information is the input information of the processing module, and the second image information is the output information of the processing module.
[0086] Input the first processed image information as input information and the second processed image information as an output module into the inverse double residual module. Based on the inverse double residual module, perform second image processing on the first processed image and the second processed image information. The second image processing can be inverse double residual processing to obtain the target processed image information. Then, determine the image recognition result information based on the target processed image information. The target processed image information can be input into a fully connected layer for mapping to obtain the image recognition result information corresponding to the target processed image information processed by the inverse double residual module. Alternatively, use the target processed image information processed by the inverse double residual module as the new first processed image information, re-execute the above steps, and input the target processed image information after iterating a preset number of times into the fully connected layer for mapping to obtain the final image recognition result information.
[0087] Optionally, referring to Figure 2 , based on any of the above embodiments, in another embodiment of the image recognition method of the present invention, using the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain the target processed image information may include:
[0088] Step S40: Use the residual operation layer of the image recognition model to process the first processed image information and the second processed image information to obtain residual features;
[0089] In this embodiment, the inverse double residual module includes a residual operation layer and an inverse residual operation layer. The residual operation adopted by the residual operation layer can combine the first processed image information and the second processed image information to obtain residual features. In this way, the residual features contain the features of the input information and the features of the output information.
[0090] Step S50: Use the inverse residual operation layer corresponding to the residual operation layer to process the residual features and the first processed image information, and separately obtain the target processed image information.
[0091] In this embodiment, the image recognition model includes an inverse residual operation layer, which includes a residual operation layer and an inverse residual operation layer. The residual operation layer performs a residual operation, and the inverse residual operation layer performs an inverse residual operation opposite to the residual operation adopted by the residual operation layer. The residual operation can combine the first processed image information and the second processed image information, and the inverse residual operation can separate the first processed image information from the residual features, and the remaining image information is used as the target processed image information. Specifically, input the residual features and the first processed image information into the inverse residual operation layer, and the inverse residual operation layer performs a separation process on the residual features based on the first processed image information, and the remaining image information after the separation process is used as the target processed image information.
[0092] Optionally, the residual operation layer processes the first processed image information and the second processed image information based on the following residual operation formula:
[0093] y = f(x)
[0094] res1 = y + x
[0095] The reverse residual operation layer processes the residual feature and the first processed image information based on the following reverse residual operation formula:
[0096] res2 = res1 - x
[0097] Wherein, X is the first processed image information as the input information of the processing module, f() is the operation function adopted by the processing module, and y and f(X) are the second processed image information as the output information of the processing module, res1 represents the residual feature, and res2 represents the target processed image information.
[0098] In the technical solution disclosed in this embodiment, the residual operation layer of the image recognition model is used to perform the second image processing on the first processed image information and the second processed image information to obtain a residual feature; the reverse residual operation layer corresponding to the residual operation layer is used to process the residual feature and the first processed image information to separately obtain the target processed image information. In this way, by performing combined processing on the first processed image information and the second processed image information to obtain a residual feature, and then using the reverse residual feature to strip the first processed image information from the residual feature to obtain the target processed image information, the target processed image information can be restored to the second processed image information, so that the image recognition result information can be obtained based on the target processed image information.
[0099] Further, the image recognition model is obtained based on the following steps:
[0100] Use the residual operation layer to process the first training image information and the second training image information to obtain the first training residual feature;
[0101] Delete the connection between at least one data feature in the first training residual feature through the inactivation layer of the image recognition model to obtain the second training residual feature;
[0102] Based on the reverse residual operation layer, process the second training residual feature and the first training image information to separately obtain the target training image information, and train the image recognition model based on the target training image information.
[0103] In this embodiment, the image recognition model is divided into an application stage and a training stage. In the application stage, the image recognition model can process the first processed image information corresponding to the image information to be recognized. The image recognition model needs to be trained in the training stage to obtain the image recognition model. First, the training image information with the image recognition result information label is obtained, and preprocessing such as normalization is performed on the training image information to obtain the first training image information. Then, the first training image information is input into the initial image recognition model. The initial image recognition model is the same as the trained image recognition model and also has a processing module and a reverse double residual module. The reverse double residual module is connected after the processing module. After the first training image information, it first passes through the convolutional layer, normalization layer, activation layer, and attention layer of the processing module in sequence to obtain the second training image information. The reverse double residual module will receive the second training image information and obtain the first training image information corresponding to the second training image information. In the training stage, the reverse double residual module includes a residual operation layer, an inactivation layer, and a reverse residual operation layer, that is, the image recognition model also includes an inactivation layer in the training stage.
[0104] In this embodiment, the first training image information is the input information of the processing module, and the second training image information is the output information of the processing module. Based on the residual operation layer, the first training image information and the second training image information can be combined, that is, the features contained in the input information are added to the output information to obtain the first training residual feature, realizing cross-layer connection, solving the problem of difficult training caused by gradient disappearance or explosion, and helping the model better learn features, thereby improving the model performance.
[0105] The residual operation layer can add the first training image information and the second training image information to obtain the residual feature. The specific formula is as follows:
[0106] y = f(x)
[0107] res1 = y + x
[0108] Among them, X is the first training image information that is the input information of the processing module, f() is the operation function adopted by the processing module, and y and f(X) are the second training image information that is the output information of the processing module. res1 represents the first training residual feature.
[0109] Then, at least one connection between data features in the training residual feature is deleted through the inactivation layer to obtain the second training residual feature;
[0110] In this embodiment, the residual features determined based on the input information and output information of the processing module can represent various data features and the connection relationships between these data features. The residual features determined based on the input information and output information of the processing module are input into the inactivation layer, which can be a dropout module. It randomly deletes the connections between at least one of the data features in the training residual features, causing a change in the relationships between the data features included in the first training residual features, and providing richer learnable information. This can reduce the overfitting problem during the training of the image recognition model, improve the generalization ability of the image recognition model, obtain greater performance (accuracy) gains during the training process, and use fewer resources such as content or time consumption during inference. The operation formula of the inactivation layer can be as follows:
[0111] res1_ = dropout(res1, p = a)
[0112] Where dropout() is the operation adopted by the inactivation layer, randomly selecting and deleting the connections between at least one of the data features in res1, a is the parameter of dropout, representing the proportion of the connections between the randomly deleted data features. res1 is input into the inactivation layer to obtain res1_, and res1 is the second training residual feature.
[0113] Furthermore, if the data type of the first training residual feature is not the target data type, then convert the data type of the first training residual feature to the target data type;
[0114] At least one of the data features in the first training residual feature after conversion is set to the target value through the inactivation layer to obtain the second training residual feature.
[0115] In this embodiment, after the first training residual feature is input into the inactivation layer, the inactivation layer randomly selects and deletes the connections between at least one of the data features in the first training residual feature. Due to different representation methods of the features, the ways for the inactivation layer to delete the connections between at least one of the data features are also different. If the data type of the first residual feature is not the target data type, then the first residual feature is converted to the target data type. At least one of the data features in the first training residual feature after conversion is set to the target value through the inactivation layer to obtain the second training residual feature. That is, if the first training residual feature can be correspondingly represented as a feature matrix, then at least one of the data features in the feature matrix is randomly selected and set to the target value, and the target value can be a preset value 0, and then the second training residual feature is obtained, which can more quickly implement the function of the inactivation layer.
[0116] The residual operation layer adds the input information of the processing module to the output information of the processing module. The calculation method of the residual operation layer is opposite to that of its corresponding reverse residual operation layer. After passing through the residual operation layer, without changing the parameters, the reverse residual operation layer processes it, and the first training image information can be separated from the second training residual feature to obtain the target training image information. Restore the parameters to before the residual budget layer processing. For example, the calculation methods of addition and subtraction are opposite, and the calculation methods of multiplication and division are opposite. When the reverse residual operation processes the first training image information, since the second training residual feature is processed by the inactivation layer, compared with the residual feature, some connections between data features are deleted, and the reverse residual operation cannot completely separate the first training image information from the second training residual feature. After the reverse residual operation processes the first training image information, the obtained target training image information still retains some features of the first training image information as input data, which is equivalent to retaining some data features of the previous layer and can still achieve model training, that is, cross-layer connection. It can solve the problem of difficult training caused by gradient disappearance or explosion, help the model better learn features, and thus improve the model performance.
[0117] Specifically, if the residual operation layer adds the first training image information as input data to the second training image information as output information, then the reverse residual operation layer can subtract the second training residual feature and the first training image information. The specific operation formula of the reverse residual operation layer can be as follows:
[0118] res2 = res1_ + x
[0119] Where res2 is the target training image information output by the processing module. Based on the target training image information, the initial image recognition model obtains the training image recognition result information of the image to be trained. The image recognition result information label corresponding to the image to be trained is used to train the initial image recognition model to obtain an applicable image recognition model.
[0120] In this way, based on the residual operation layer, the input information and output information of the processing module are processed to obtain the first training residual feature. The inactivation layer is used to delete the connections between some data features in the first training residual feature, which is equivalent to reducing some features, making them not forward-inferred or backpropagated. Then, based on the reverse residual operation layer, the second training residual feature and the first training image information are processed to basically restore the second training residual feature to the state of the original second training image information, obtaining the target training image information. However, since the second training residual feature has fewer features compared to the first training residual feature, the target training image information still retains some information of the first training image information, enabling it to play the role of residual connection to prevent gradient disappearance. Moreover, the state of the first training image information is the same as that of the second training image information directly output by the processing module. The data iteration does not require element-wise addition to achieve residual connection through the processing module, reducing the model operation time-consuming and improving the training fluency of the image recognition model.
[0121] It should be noted that when the above reverse double-residual module is applied in the image recognition model, its forms in the training stage and the application stage are different, that is, the way of outputting the second training image information of its output processing module is different.
[0122] Reference Figure 3 , if the running stage of the image recognition model is the training stage, then when passing through the reverse double-residual module, the first training image information and the second training image information are processed by the residual operation layer to combine and obtain the first training residual feature; the connections between at least one data feature in the first training residual feature are deleted through the inactivation layer of the reverse double-residual module to obtain the second training residual feature; the second training residual feature and the first training image information are processed by the reverse residual operation layer to separate and obtain the target training image information, and the image recognition model is trained based on the target training image information. This enables the second training image information of the processing module to preserve some information of the previous layer, realizing cross-layer connection. When it needs to be iteratively input into the processing module, the overfitting problem of the model in the training process is reduced, the generalization ability of the model is improved, and thus the training efficiency is enhanced.
[0123] If the image recognition model is in the application stage, then the first processed image information and the second processed image information are processed by the residual operation layer of the reverse double-residual module to combine and obtain the residual feature; the residual feature and the first processed image information are processed by the corresponding reverse residual operation layer of the residual operation layer to separate and obtain the target processed image information.
[0124] During the training phase, the residual connection module applied in the image recognition method includes a dropout layer. When the dropout layer is in effect, the second training image information output by the processing module is different from the target training image information. The second training image information contains the features of some input information, which can achieve the effect of model training, and at the same time can prevent overfitting problems during training and improve the model accuracy. When the dropout layer is not in effect, after the first training image information and the second training image information are operated by the reverse double-residual module, the obtained target training image information is substantially the same as the second training image information. Therefore, in order to make the target processed image information the same as the second processed image information during the application phase, the dropout layer can be ignored during the application phase or removed after the training phase ends. Such a reverse double-residual module can ensure that the output information of the processing module is the same as the input information of the next iteration, reducing the unnecessary noise introduced by model training. At the same time, it makes the results of the image recognition model in the training phase and the application phase as consistent as possible. Thereby improving the accuracy of model application and the accuracy of execution results. It reduces the bypass connections of the model, thereby reducing the memory occupied during inference and the inference time, and avoiding reducing the performance of the terminal device where the model is deployed. It does not require the main and branch to perform element-wise addition operations like a model containing traditional model training, thus not adding extra inference time to the model and improving the fluency when the model performs tasks such as gesture recognition.
[0125] Optionally, based on any of the above embodiments, in another embodiment of the image recognition method of the present invention, before using the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information, it further includes:
[0126] Obtaining the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, as well as the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information;
[0127] If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then perform second image processing on the first processed image information and the second processed image information using the image recognition model to obtain target processed image information.
[0128] In this embodiment, before determining to apply the reverse double residual module provided in this embodiment to the image recognition model that needs to be trained or applied, that is, before performing second image processing on the first processed image information and the second processed image information, or before performing second image processing on the first training image information and the second training image information, it is necessary to determine that the number of channels and the data composition information of the input information and the output information of the processing module of the image recognition model are consistent. Specifically, it is necessary to determine whether the first data composition information of the first processed image information or the first training image information is consistent with the second data composition information of the second processed image information or the second training image information, and to determine whether the first number of channels of the first processed image information or the first training image information is consistent. If both are consistent, then based on the reverse double residual module, the first processed image information or the first training image information of the processing module is processed with its second processed image information or the second training image information. If they are not consistent, then it is not processed based on the reverse double residual module. This is to ensure the consistency between the target training image data and the second training image data, and between the target processed image data and the second processed image data, so as to be used for training or obtaining the image recognition result.
[0129] Optionally, based on any of the above embodiments, in another embodiment of the image recognition method of the present invention, before using the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information, it further includes:
[0130] Obtain the number of levels of the network levels in the image recognition model;
[0131] If the number of levels of the network levels is greater than or equal to the first number of levels and less than the second number of levels, then perform the first image processing on the first processed image information using the image recognition model to obtain the second processed image information;
[0132] If the number of levels of the network levels is less than the first number of levels, then output the image recognition result information according to the second processed image information;
[0133] If the number of levels of the network levels is greater than or equal to the second number of levels, then use the single residual module to perform third image processing on the first processed image information and the second processed image information to obtain the target processed image information.
[0134] In this embodiment, since the second training residual feature reduces some features compared to the residual feature, the second training image information still retains part of the input information, enabling it to play the role of a residual connection to prevent gradient vanishing. The likelihood of gradient vanishing increases as the number of layers in the image recognition model increases. Therefore, obtain the construction information of the image recognition model, and determine the number of layers of the network layer in the image recognition model according to the construction information. When the number of layers of the network layer is greater than or equal to the first number of layers, that is, when the depth of the image recognition model is relatively large, use the image recognition model to perform the first image processing on the first processed image information to obtain the second processed image information. That is, use the residual operation layer of the image recognition model to process the first processed image information and the second processed image information to obtain the residual feature, and use the corresponding reverse residual operation layer of the residual operation layer to process the residual feature and the first processed image information to separate and obtain the target processed image information. In the application stage, it will not increase the difference between the target processed image information and the first processed image information. In the training stage, the number of layers of the network layer in the image recognition model can also be obtained. If the number of layers of the network layer is greater than or equal to the first number of layers and less than the second number of layers, perform the first training process on the first training image information using the image recognition model to obtain the second training image information. That is, if the number of layers of the network layer is greater than or equal to the first number of layers and less than the second number of layers, use the residual operation layer to process the first training image information and the second training image information to obtain the first training residual feature, delete the connection between at least one data feature in the first training residual feature through the inactivation layer of the image recognition model to obtain the second training residual feature, based on the reverse residual operation layer to process the second training residual feature and the first training image information, separate and obtain the target training image information, and train the image recognition model based on the target training image information to delete the connection between at least one data feature in the training residual feature through the inactivation layer to obtain the second training residual feature, based on the reverse residual operation corresponding to the residual operation layer, process the second training residual feature and the first training image information, obtain the target training image information and output it to the next level of the processing module, thereby retaining part of the input information of the processing module during the training process, reducing the possibility of gradient vanishing, and at the same time, it can also reduce the model operation time consumption, shorten the side branch in the inference stage, and reduce the overfitting problem in the training stage.
[0135] In this embodiment, if the number of network levels is less than the first number of layers, that is, when the depth of the image recognition model is relatively shallow, the problem of gradient disappearance is not likely to occur. To improve the operation efficiency of the model, the input information of the processing module in the image recognition model, that is, the second processed image information, can be used as the target training image information and output to the next level of the processing module, and finally the image recognition result information is output. Similarly, in the training stage, if the number of network levels is less than the first number of layers, the input information of the processing module in the image recognition model, that is, the second training image information, can also be used as the target training image information and output to the next level of the processing module, and finally the image recognition result information is output.
[0136] In this embodiment, if the number of network levels is greater than or equal to the first number of layers and less than the second number of layers, that is, when the depth of the image recognition model is relatively high, the problem of gradient disappearance is more likely to occur. Then, the first processed image information and the second processed image information are subjected to third image processing by using a unidirectional residual module to obtain the target processed image information. The unidirectional residual module can be the residual operation layer of the image recognition model. The first processed image information and the second processed image information are subjected to third image processing by using the residual operation layer, and the obtained result is the residual feature, and the residual feature is used as the target processed image information and output to the next level of the processing module, and finally the image recognition result information is output. The target processed image information retains the first processed image information to a greater extent. In the training stage, when the number of network levels is greater than or equal to the first number of layers and less than the second number of layers, that is, when the depth of the image recognition model is relatively high, the training process is more likely to have the problem of gradient disappearance. In this regard, when the number of network levels is greater than or equal to the first number of layers and less than the second number of layers, the first training image information and the second training image information are subjected to third image processing by using a unidirectional residual module to obtain the target training image information. Specifically, the first training image information and the second training image information are subjected to third image processing by using the residual operation layer, and the obtained result is the training residual feature, and the training residual feature is used as the target training image information and output to the next level of the processing module, and finally the image recognition result information is output. The target image information retains the first processed image information to a greater extent. In this way, the problem of gradient disappearance is more effectively prevented.
[0137] In the technical solution disclosed in this embodiment, according to the different intervals corresponding to the number of network levels in the image recognition model, the target training image information retains different degrees of input information, so that image recognition models of various depths can achieve good usage effects.
[0138] This embodiment also provides an image recognition device, which can be specifically integrated in an image recognition device. For example, as Figure 4 shown, the image recognition device may include:
[0139] An acquisition unit 1001, configured to acquire image information to be recognized;
[0140] A preprocessing unit 1002, configured to preprocess the image information to be recognized to obtain first processed image information;
[0141] A recognition unit 1003, configured to perform recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information.
[0142] Preferably, the recognition unit 1003 performs recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information, including:
[0143] Performing first image processing on the first processed image information by using the image recognition model to obtain second processed image information;
[0144] Performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information;
[0145] Determining the image recognition result information based on the target processed image information;
[0146] Preferably, the recognition unit 1003 performs second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information, including:
[0147] Processing the first processed image information and the second processed image information by using a residual operation layer of the image recognition model to obtain residual features;
[0148] Processing the residual features and the first processed image information by using a reverse residual operation layer corresponding to the residual operation layer to separately obtain the target processed image information;
[0149] Preferably, the image recognition model is obtained by the recognition unit 1003 based on the following steps:
[0150] Processing first training image information and second training image information by using the residual operation layer to obtain first training residual features;
[0151] Deleting connections between at least one data feature in the first training residual features through an inactivation layer of the image recognition model to obtain second training residual features;
[0152] Processing the second training residual features and the first training image information by using the reverse residual operation layer to separately obtain the target training image information, and training an image recognition model based on the target training image information;
[0153] Preferably, the recognition unit 1003 deletes the connections between at least one data feature in the first training residual feature through the inactivation layer of the image recognition model to obtain a second training residual feature, including:
[0154] If the data type of the first training residual feature is not the target data type, convert the data type of the first training residual feature to the target data type;
[0155] Set at least one data feature in the converted first training residual feature to a target value through the inactivation layer to obtain the second training residual feature;
[0156] Preferably, before the recognition unit 1003 uses the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information, it further includes:
[0157] Obtain the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, as well as the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information;
[0158] If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then perform second image processing on the first processed image information and the second processed image information using the image recognition model to obtain target processed image information;
[0159] Preferably, before the recognition unit 1003 uses the image recognition model to perform first image processing on the first processed image information to obtain second processed image information, it further includes:
[0160] Obtain the number of network levels in the image recognition model;
[0161] If the number of network levels is greater than or equal to the first number of layers and less than the second number of layers, perform the first image processing on the first processed image information using the image recognition model to obtain second processed image information;
[0162] If the number of network levels is less than the first number of layers, output the image recognition result information according to the second processed image information;
[0163] If the number of network levels is greater than or equal to the second number of layers, use a one-way residual module to perform third image processing on the first processed image information and the second processed image information to obtain target processed image information.
[0164] In this embodiment, image information to be recognized is obtained; the image information to be recognized is preprocessed to obtain first processed image information; and the first processed image information is recognized by an image recognition model to obtain image recognition result information. In this way, before the image information to be recognized is recognized by the model, the image information to be recognized is preprocessed first, making it more suitable for processing by the image recognition model, which is beneficial to improving the accuracy of the image recognition model for image recognition.
[0165] As Figure 5 shown, Figure 5 FIG. is a schematic structural diagram of an image recognition device provided by an embodiment of the present invention. The image recognition device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structural diagram of the image recognition device shown in the figure does not limit the image recognition device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0166] The processor 1101 is the control center of the image recognition device 1100, connecting various parts of the entire image recognition device 1100 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, the processor 1101 executes various functions of the image recognition device 1100 and processes data, thereby monitoring the entire image recognition device 1100. The processor 1101 may be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.
[0167] In the embodiment of the present invention, the processor 1101 in the image recognition device 1100 will load instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions, such as:
[0168] Obtain image information to be recognized;
[0169] Preprocess the image information to be recognized to obtain first processed image information;
[0170] Use the image recognition model to perform recognition processing on the first processed image information to obtain image recognition result information.
[0171] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0172] Optionally, as Figure 5 shown, the image recognition device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art can understand that Figure 5 the structure of the image recognition device shown in
[0173] does not limit the image recognition device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0174] The radio frequency circuit 1104 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other image recognition devices through wireless communication, and to receive and transmit signals between the network device or other image recognition devices.
[0175] The audio circuit 1105 can be used to provide an audio interface between the user and the image recognition device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and then converted into audio data. After the audio data is output to the processor 1101 for processing, it is sent through the radio frequency circuit 1104 to, for example, another image recognition device, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the image recognition device.
[0176] The input unit 1106 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0177] The power supply 1107 is used to supply power to each component of the image recognition device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1107 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0178] Although Figure 5 not shown in the figure, the image recognition device 1100 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0179] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0180] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed through instructions, or through instructions to control relevant hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0181] To this end, an embodiment of the present invention provides a computer-readable storage medium, which stores multiple computer programs that can be loaded by a processor to execute any one of the image recognition methods provided by the embodiments of the present invention. The computer program can execute the steps of the following image recognition method:
[0182] Obtain the image information to be recognized;
[0183] Preprocess the image information to be recognized to obtain the first processed image information;
[0184] Use the image recognition model to perform recognition processing on the first processed image information to obtain the image recognition result information.
[0185] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.
[0186] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0187] Since the computer program stored in the computer-readable storage medium can execute any one of the image recognition methods provided by the embodiments of the present invention, the beneficial effects achievable by any one of the image recognition methods provided by the embodiments of the present invention can be realized. For details, reference may be made to the previous embodiments and will not be elaborated herein.
[0188] In the above embodiments of the image recognition device, computer-readable storage medium, image recognition device, and computer program product, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects brought by the above-described image recognition device, computer-readable storage medium, computer program product, image recognition device, and their corresponding units can refer to the description of the image recognition method in the above embodiments and will not be elaborated herein specifically.
[0189] The above has introduced in detail an image recognition method, an image recognition device, an image recognition device, a computer-readable storage medium, and a computer program product provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image recognition method, characterized in that, The method includes: Obtaining the image information to be recognized; Preprocessing the image information to be recognized to obtain first processed image information; Performing recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information.
2. The image recognition method according to claim 1, wherein The performing recognition processing on the first processed image information by using an image recognition model to obtain image recognition result information includes: Performing first image processing on the first processed image information by using the image recognition model to obtain second processed image information; Performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information; Determining the image recognition result information based on the target processed image information.
3. The image recognition method according to claim 2, characterized in that, The performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information includes: Processing the first processed image information and the second processed image information by using the residual operation layer of the image recognition model to obtain residual features; Processing the residual features and the first processed image information by using the corresponding inverse residual operation layer of the residual operation layer to separately obtain the target processed image information.
4. The image recognition method according to claim 3, wherein The image recognition model is obtained based on the following steps: Processing first training image information and second training image information by using the residual operation layer to obtain first training residual features; Deleting the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features; Processing the second training residual features and the first training image information by using the inverse residual operation layer to separately obtain the target training image information, and training an image recognition model based on the target training image information.
5. The image recognition method according to claim 4, characterized in that, The deleting the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features includes: If the data type of the first training residual features is not the target data type, converting the data type of the first training residual features to the target data type; Setting at least one data feature in the converted first training residual features to a target value through the inactivation layer to obtain the second training residual features.
6. The image recognition method according to claim 2, wherein Before the performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information, it further includes: Obtaining the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, as well as the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information; If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then performing second image processing on the first processed image information and the second processed image information by using the image recognition model to obtain target processed image information.
7. The image recognition method according to claim 2, characterized in that, Before performing the first image processing on the first processed image information using the image recognition model to obtain second processed image information, the following steps are further included: Obtain the number of levels of the network levels in the image recognition model; If the number of levels of the network levels is greater than or equal to the first number of levels and less than the second number of levels, perform the first image processing on the first processed image information using the image recognition model to obtain second processed image information; If the number of levels of the network levels is less than the first number of levels, output the image recognition result information according to the second processed image information; If the number of levels of the network levels is greater than or equal to the second number of levels, perform third image processing on the first processed image information and the second processed image information using a one-way residual module to obtain target processed image information.
8. An image recognition device, characterized in that, The image recognition device includes: An acquisition unit, configured to acquire image information to be recognized; A preprocessing unit, configured to preprocess the image information to be recognized to obtain first processed image information; A recognition unit, configured to perform recognition processing on the first processed image information using an image recognition model to obtain image recognition result information; Preferably, the recognition unit performing recognition processing on the first processed image information using an image recognition model to obtain image recognition result information includes: Performing first image processing on the first processed image information using the image recognition model to obtain second processed image information; Performing second image processing on the first processed image information and the second processed image information using the image recognition model to obtain target processed image information; Determining the image recognition result information based on the target processed image information; Preferably, the recognition unit performing second image processing on the first processed image information and the second processed image information using the image recognition model to obtain target processed image information includes: Processing the first processed image information and the second processed image information using the residual operation layer of the image recognition model to obtain residual features; Processing the residual features and the first processed image information using the corresponding reverse residual operation layer of the residual operation layer to separately obtain the target processed image information; Preferably, the image recognition model is obtained by the recognition unit based on the following steps: Processing first training image information and second training image information using the residual operation layer to obtain first training residual features; Deleting the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features; Processing the second training residual features and the first training image information based on the reverse residual operation layer to separately obtain the target training image information, and training the image recognition model based on the target training image information; Preferably, the recognition unit deleting the connections between at least one data feature in the first training residual features through the inactivation layer of the image recognition model to obtain second training residual features includes: If the data type of the first training residual feature is not the target data type, convert the data type of the first training residual feature to the target data type; Set at least one data feature in the converted first training residual feature to a target value through the inactivation layer to obtain the second training residual feature; Preferably, before the recognition unit uses the image recognition model to perform second image processing on the first processed image information and the second processed image information to obtain target processed image information, it further includes: Obtain the first data composition information of the first processed image information and the first channel number corresponding to the first processed image information, as well as the second data composition information of the second processed image information and the second channel number corresponding to the second processed image information; If the first data composition information and the second data composition information are consistent, and the first channel number and the second channel number are consistent, then perform second image processing on the first processed image information and the second processed image information using the image recognition model to obtain target processed image information; Preferably, before the recognition unit uses the image recognition model to perform first image processing on the first processed image information to obtain second processed image information, it further includes: Obtain the number of network levels in the image recognition model; If the number of network levels is greater than or equal to the first number of layers and less than the second number of layers, perform the first image processing on the first processed image information using the image recognition model to obtain second processed image information; If the number of network levels is less than the first number of layers, output the image recognition result information according to the second processed image information; If the number of network levels is greater than or equal to the second number of layers, perform third image processing on the first processed image information and the second processed image information using a one-way residual module to obtain target processed image information.
9. An image recognition device, characterized in that, It includes a processor and a memory, and the memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the image recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the image recognition method according to any one of claims 1 to 7.