Image recognition method, medium, device and computing equipment

Through a convolutional neural network, the feature map matrix is ​​extracted from the picture to be identified and multiple recognitions are performed, and the problem of insufficient accuracy of picture recognition in the prior art is solved, and high accuracy recognition of preset types of pictures is achieved.

CN113902922BActive Publication Date: 2025-05-13HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111182606.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-05-13
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify whether the picture belongs to the preset type, especially when the picture contains rich elements, it is susceptible to interference, resulting in a low recognition accuracy rate.

Method used

A convolutional neural network is used to extract the feature map matrix from the picture representation matrix of the picture to be identified, and the probability representation value is obtained through feature vector mapping, and a preliminary judgment is made as to whether the picture belongs to the preset type. If the probability characterization value is greater than the preset threshold, the feature map matrix is ​​further cropped to obtain more accurate feature vectors and quadratic recognition is performed to improve accuracy.

Benefits of technology

Through deep learning technology, the deep features of the picture can be mined, and the accurate recognition of preset types of pictures is achieved, which reduces the interference impact during the recognition process and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902922B_ABST
    Figure CN113902922B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method, medium, device and computing device for image recognition. Based on a convolutional neural network, a plurality of first-class feature map matrices are extracted from the image to be recognized, and a first feature vector is determined according to each first-class feature map matrix; the first feature vector is mapped into a first probability representation value; if the first probability representation value is greater than a first preset threshold value, the first feature vector is mapped into a region position coordinate, and the region position coordinate is used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be contained in the first-class feature map matrix; the corresponding prediction region is cut out from each first-class feature map matrix as a second-class feature map matrix, and a second feature vector is determined according to each second-class feature map matrix; the second feature vector is mapped into a second probability representation value and output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of information technology, and more specifically, the embodiments of the present disclosure relate to a method, medium, apparatus and computing device for image recognition. Background Art

[0002] In some businesses, it is necessary to identify whether an image is a preset type image. A preset type image refers to an image that contains preset type elements, and the preset type elements can be set according to actual business needs. For example, the preset type elements can be set to prohibited elements (pornography, violence, terrorism, politics, etc.), and images containing prohibited elements are prohibited images.

[0003] Therefore, a picture recognition method is needed so as to more accurately recognize the preset type of pictures. Summary of the invention

[0004] The present disclosure provides a method, medium, apparatus and computing device for image recognition, so as to more accurately recognize images of preset types.

[0005] In a first aspect of the embodiments of the present disclosure, a method for image recognition is provided, which is applied to a recognition model, and the method comprises:

[0006] Extracting at least one first-category feature map matrix from a picture representation matrix of the picture to be identified based on a convolutional neural network, and determining a first eigenvector according to each first-category feature map matrix;

[0007] Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type;

[0008] If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix;

[0009] Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix;

[0010] The second feature vector is mapped into a second probability representation value and output, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type.

[0011] In one embodiment of the present disclosure, determining the first eigenvector according to each first-category feature map matrix includes:

[0012] A pooling operation is performed on each first-category feature map matrix, and each pooling operation result value is combined into a first feature vector.

[0013] In another embodiment of the present disclosure, determining the second eigenvector according to each second-type feature map matrix includes:

[0014] A pooling operation is performed on each second-category feature map matrix, and each pooling operation result value is combined into a second feature vector.

[0015] In yet another embodiment of the present disclosure, the area location coordinates include a set of coordinate values ​​for defining an area range of the prediction area.

[0016] In yet another embodiment of the present disclosure, according to the region position coordinates, a corresponding prediction region is cropped from each first-category feature map matrix, including:

[0017] If any coordinate value included in the coordinate value set is a non-integer, the coordinate value is adjusted to an integer closest to the non-integer; wherein the area range defined by the adjusted coordinate value set covers the area range defined by the coordinate value set before the adjustment;

[0018] According to the adjusted coordinate value set, a corresponding prediction area is cropped from each first-category feature map matrix.

[0019] In yet another embodiment of the present disclosure, a mapping parameter set used to map to obtain the first probability representation value and a mapping parameter set used to map to obtain the second probability representation value are the same parameter set.

[0020] In yet another embodiment of the present disclosure, it further includes:

[0021] If the first probability representation value is not greater than a first preset threshold, the first probability representation value is output.

[0022] In yet another embodiment of the present disclosure, it further includes:

[0023] If the second probability representation value output by the recognition model is greater than a second preset threshold, it is determined that the image to be recognized belongs to a preset type;

[0024] The second preset threshold is greater than the first preset threshold.

[0025] In a second aspect of the embodiments of the present disclosure, a method for training a recognition model is provided, comprising:

[0026] Obtaining a picture sample and an identification label corresponding to the picture sample, wherein the identification label is used to determine whether the picture sample belongs to a preset type; and performing the following steps to train a recognition model:

[0027] Extracting at least one first-class feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first feature vector and a positioning label according to each first-class feature map matrix; the positioning label is used to determine a corresponding actual area, and the actual area corresponds to a preset type element that may be included in the first-class feature map matrix;

[0028] Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type;

[0029] If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix;

[0030] Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix;

[0031] Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type;

[0032] According to the training objective of the recognition model, the parameter set corresponding to the recognition model is adjusted.

[0033] In one embodiment of the present disclosure, if the image sample belongs to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value, and to reduce the difference between the probability representation value corresponding to the identification label and the second probability representation value, and to reduce the difference between the coordinates corresponding to the positioning label and the regional position coordinates.

[0034] In yet another embodiment of the present disclosure, if the image sample does not belong to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value.

[0035] In yet another embodiment of the present disclosure, the parameter set corresponding to the recognition model includes:

[0036] A mapping parameter set used for mapping to obtain a first probability representation value;

[0037] A mapping parameter set used for mapping to obtain a second probability representation value;

[0038] The mapping parameter set used to map the region location coordinates.

[0039] In yet another embodiment of the present disclosure, the parameter set corresponding to the recognition model further includes:

[0040] The network parameter set of the convolutional neural network.

[0041] In yet another embodiment of the present disclosure, a mapping parameter set used to map to obtain the first probability representation value and a mapping parameter set used to map to obtain the second probability representation value are the same parameter set.

[0042] In yet another embodiment of the present disclosure, the region location coordinates include a set of coordinate values ​​for defining an area range of the prediction region;

[0043] A mapping parameter set for mapping to obtain the coordinates of the regional position, comprising: different weight value sets corresponding to different coordinate values ​​in the coordinate value set, wherein each weight value set comprises respective weight values ​​corresponding one-to-one to respective dimensional values ​​in the first feature vector;

[0044] Mapping the first feature vector into regional position coordinates includes:

[0045] For each weight value set, a weighted sum is calculated according to each dimension value in the first feature vector, the weight value set, and a one-to-one correspondence between the dimension value and the weight value;

[0046] Determine the coordinate value corresponding to the weight value set according to the calculated weighted sum;

[0047] The determined coordinate values ​​are combined into the coordinate value set.

[0048] In yet another embodiment of the present disclosure, a mapping parameter set used for mapping to obtain a first probability representation value includes:

[0049] A first weight value set; the first weight value set includes weight values ​​corresponding one-to-one to each dimension value in the first feature vector;

[0050] Mapping the first feature vector into a first probability representation value includes:

[0051] Calculate a weighted sum according to each dimension value in the first feature vector, the first weight value set, and a one-to-one correspondence between the dimension value and the weight value;

[0052] A first probability characterization value is determined according to the calculated weighted sum.

[0053] In yet another embodiment of the present disclosure, a mapping parameter set used for mapping to obtain a second probability representation value includes:

[0054] A second weight value set; the second weight value set includes weight values ​​corresponding one-to-one to each dimension value in the second feature vector;

[0055] Mapping the second feature vector into a second probability representation value includes:

[0056] Calculate a weighted sum according to each dimension value in the second feature vector, the second weight value set, and a one-to-one correspondence between the dimension value and the weight value;

[0057] A second probability characterization value is determined according to the calculated weighted sum.

[0058] In yet another embodiment of the present disclosure, determining the first eigenvector according to each first-category feature map matrix includes:

[0059] A pooling operation is performed on each first-category feature map matrix, and each pooling operation result value is combined into a first feature vector.

[0060] In yet another embodiment of the present disclosure, each first-type feature map matrix and each weight value of the first weight value set have a one-to-one correspondence;

[0061] The positioning labels are determined according to each first-category feature map matrix, including:

[0062] Get the value of the i-th position in each first-category feature map matrix;

[0063] According to the obtained values ​​of each i-th position, the first weight value set, and the one-to-one correspondence between the first type feature map matrix and the weight value, a weighted sum is calculated, and the calculated weighted sum is used as the replacement value corresponding to the i-th position; i = (1, 2, ..., N), N is the number of positions included in the first type feature map matrix;

[0064] Reassign corresponding replacement values ​​to each position in the first type feature map matrix to obtain a replacement feature map matrix;

[0065] Taking each replacement value in the replacement feature map matrix as a pixel value, a replacement feature map is generated; wherein the pixel value of a pixel point in the replacement feature map is positively correlated with the brightness of the pixel point;

[0066] Determine a high-brightness area in the replacement feature map, and determine the area position coordinates corresponding to the high-brightness area as a positioning label.

[0067] In yet another embodiment of the present disclosure, determining the high brightness area in the replacement feature map includes:

[0068] Converting the replacement feature map into a binary map, and performing connected domain analysis on the binary map;

[0069] According to the analysis results, the high brightness area is determined;

[0070] Determining the area position coordinates corresponding to the high brightness area includes:

[0071] The circumscribed rectangle of the high-brightness area is determined, and the area position coordinates used to define the circumscribed rectangle are obtained.

[0072] In yet another embodiment of the present disclosure, before determining the high brightness area in the replacement feature map, the method further includes:

[0073] The resolution of the replacement feature map is adjusted to the resolution of the image sample.

[0074] In a third aspect of the embodiments of the present disclosure, there is provided a picture recognition device for use in a recognition model, the device comprising:

[0075] A first feature vector determination module extracts at least one first-type feature map matrix from a picture representation matrix of the picture to be identified based on a convolutional neural network, and determines a first feature vector according to each first-type feature map matrix;

[0076] A first classification module maps the first feature vector into a first probability representation value for representing the probability that the first identified image to be identified belongs to a preset type;

[0077] A positioning module, if the first probability representation value is greater than a first preset threshold, maps the first feature vector into a regional position coordinate, the regional position coordinate is used to determine a corresponding prediction area, the prediction area corresponds to a preset type element that may be included in the first type feature map matrix;

[0078] A second eigenvector determination module, which cuts out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determines a second eigenvector according to each second-type feature map matrix;

[0079] The second classification module maps the second feature vector into a second probability representation value and outputs it, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type.

[0080] In a fourth aspect of the embodiments of the present disclosure, a device for training a recognition model is provided, comprising:

[0081] An acquisition module, acquiring a picture sample and an identification tag corresponding to the picture sample, wherein the identification tag is used to determine whether the picture sample belongs to a preset type;

[0082] The training module performs the following steps to train the recognition model:

[0083] Extracting at least one first-class feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first feature vector and a positioning label according to each first-class feature map matrix; the positioning label is used to determine a corresponding actual area, and the actual area corresponds to a preset type element that may be included in the first-class feature map matrix;

[0084] Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type;

[0085] If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix;

[0086] Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix;

[0087] Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type;

[0088] According to the training objective of the recognition model, the parameter set corresponding to the recognition model is adjusted.

[0089] In a fifth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method provided by the present disclosure is implemented.

[0090] In a sixth aspect of an embodiment of the present disclosure, a computing device is provided, comprising a memory and a processor; the memory is used to store computer instructions executable on the processor, and the processor is used to implement the method provided by the present disclosure when executing the computer instructions.

[0091] In the technical solution provided by the present disclosure, a recognition model is used to determine whether a certain picture belongs to a preset type. In the process of recognizing the picture to be recognized by the recognition model, the picture representation matrix of the picture to be recognized is used as input, and the picture representation matrix is ​​first processed into a number of feature map matrices (in order to distinguish in description, called the first type of feature map matrix) based on a convolutional neural network, and then each first type of feature map matrix is ​​further processed into a feature vector (in order to distinguish in description, called the first feature vector), and then the first type of feature vector is mapped into a probability representation value for representing the probability that the picture to be recognized belongs to the preset type (in order to distinguish in description, called the first probability representation value).

[0092] Next, determine whether the first probability representation value is large enough (whether it is greater than the first preset threshold). If it is large enough, it is preliminarily assumed that the image to be identified belongs to the preset type, and the first feature vector is mapped into regional position coordinates. The regional position coordinates are used to determine the corresponding prediction area, and the prediction area corresponds to the preset type elements that may be contained in the first type of feature map matrix.

[0093] Next, a corresponding preset area is cropped from each first-class feature map matrix as several cropped feature map matrices (in order to describe the upper distinction, they are called second-class feature map matrices), and each second-class feature map matrix is ​​further processed into a feature vector (in order to describe the upper distinction, they are called second feature vectors).

[0094] Next, the second feature vector is used again to map the probability representation value, that is, the second feature vector is mapped into a probability representation value (for the purpose of description, referred to as the second probability representation value) used to represent the probability that the image to be identified belongs to the preset type. The second probability representation value is the output of the recognition model.

[0095] It can be seen that through the above technical solution, the convolutional neural network can be used to mine the deep features of the image to be identified, that is, the first feature map matrix can better reflect the features of the image to be identified. The first feature vectors processed by each first-class feature map matrix are used to perform a preliminary classification of the image to be identified (i.e., the first recognition), and the first feature vectors preliminarily classified as a preset type are mapped into regional position coordinates for predicting "areas in the first-class feature map matrix of the preset type that may contain elements of the preset type."

[0096] Cutting out the prediction area from the first type of feature map matrix as the second type of feature map matrix is ​​equivalent to excluding the part of the image to be identified that may not be related to the preset type element. The second feature vectors processed by each second type of feature map matrix are used to perform secondary classification (i.e., re-identification) on the image to be identified, so that the model can focus on the part of the image to be identified that may be related to the preset type element. The second probability representation value obtained in this way can more accurately represent the probability that the image to be identified belongs to the preset type. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0098] Figure 1 An exemplary process of an image recognition method is provided;

[0099] Figure 2An exemplary process of a method for training a recognition model is provided;

[0100] Figure 3 Exemplarily showing a picture sample, replacing a high brightness area in a feature map, and visually replacing a high brightness area in a feature map;

[0101] Figure 4 The process of determining to replace the high brightness area in the feature map is exemplified;

[0102] Figure 5 An exemplary image recognition device is provided;

[0103] Figure 6 An exemplary method of providing a device for training a recognition model is provided;

[0104] Figure 7 is a schematic diagram of a computer-readable storage medium provided by the present disclosure;

[0105] Figure 8 It is a structural schematic diagram of a computing device provided by the present disclosure.

[0106] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. The number of any element in the drawings is for example rather than limitation, and any naming is only for distinction and does not have any limiting meaning. DETAILED DESCRIPTION

[0107] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0108] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0109] According to an embodiment of the present disclosure, a method, medium, apparatus and computing device for image recognition are proposed.

[0110] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.

[0111] The preset type described in the present disclosure may refer to a specific preset type, and the preset type may be set according to actual business needs. For example, the preset type may be set to prohibited, and the preset type element may be a prohibited element (such as pornography, violence, terrorism, politics, etc.), and the picture containing the prohibited element is a prohibited picture.

[0112] Considering that there are a large number of images that need to be identified in actual business, it is not realistic to rely on manual identification. Therefore, we can adopt the idea of ​​artificial intelligence and use models to achieve image recognition tasks.

[0113] The difficulty in using the model to identify whether an image belongs to a preset type is that the elements contained in the image are often relatively rich. Especially for preset type images, they may not only contain preset type elements, but also contain several non-preset type interference elements, which can easily cause greater interference to the recognition algorithm of the recognition model, resulting in low recognition accuracy.

[0114] To this end, an optional technical solution provided by the present disclosure is to first use a target detection model to detect a predicted area that may contain elements of a preset type from an image, and then crop the predicted area from the image and input it into a recognition model (essentially an image classification model) for recognition.

[0115] However, the algorithm structure of the target detection model is often complex and requires multiple feature extraction operations. Therefore, it takes a long time to detect the predicted area that may contain elements of the preset type from the image. In addition, it takes time to input the predicted area into the recognition model for recognition (where the recognition model also needs to perform feature extraction operations on the image of the prediction area, which is also time-consuming). This means that the total time required to implement the image recognition task based on the target detection model and the recognition model is too long.

[0116] Therefore, the present disclosure also aims to provide a technical solution that can not only realize the image recognition task based on the model, but also shorten the total time of the image recognition task as much as possible. In this technical solution, there is no need for a target detection model, and only the recognition model can be used to accurately identify whether the image belongs to a preset type. The recognition model includes not only a feature extraction algorithm and a classification algorithm, but also a positioning algorithm, which is used to predict the area in the image that may contain elements of the preset type.

[0117] In the process of recognizing the picture to be recognized, the recognition model takes the picture representation matrix of the picture to be recognized as input, and first processes the picture representation matrix into several feature map matrices (in order to distinguish in description, called first-type feature map matrices) based on the convolutional neural network, and then further processes each first-type feature map matrix into a feature vector (in order to distinguish in description, called first feature vector), and then maps the first-type feature vector into a probability representation value for representing the probability that the picture to be recognized belongs to a preset type (in order to distinguish in description, called a first probability representation value).

[0118] Next, determine whether the first probability representation value is large enough (whether it is greater than the first preset threshold). If it is large enough, it is preliminarily assumed that the image to be identified belongs to the preset type, and the first feature vector is mapped into regional position coordinates. The regional position coordinates are used to determine the corresponding prediction area, and the prediction area corresponds to the preset type elements that may be contained in the first type of feature map matrix.

[0119] Next, a corresponding preset area is cropped from each first-class feature map matrix as several cropped feature map matrices (in order to describe the upper distinction, they are called second-class feature map matrices), and each second-class feature map matrix is ​​further processed into a feature vector (in order to describe the upper distinction, they are called second feature vectors).

[0120] Next, the second feature vector is used again to map the probability representation value, that is, the second feature vector is mapped into a probability representation value (for the purpose of description, referred to as the second probability representation value) used to represent the probability that the image to be identified belongs to the preset type. The second probability representation value is the output of the recognition model.

[0121] It can be seen that through this technical solution, the convolutional neural network can be used to mine the deep features of the image to be identified, that is, the first feature map matrix can better reflect the features of the image to be identified. The first feature vectors processed by each first-class feature map matrix are used to perform a preliminary classification of the image to be identified (i.e., the first recognition), and the first feature vectors preliminarily classified as a preset type are mapped into regional position coordinates for predicting "areas in the first-class feature map matrix of the preset type that may contain elements of the preset type."

[0122] Cutting out the prediction area from the first type of feature map matrix as the second type of feature map matrix is ​​equivalent to excluding the part of the image to be identified that may not be related to the preset type element. The second feature vectors processed by each second type of feature map matrix are used to perform secondary classification (i.e., re-identification) on the image to be identified, so that the model can focus on the part of the image to be identified that may be related to the preset type element. The second probability representation value obtained in this way can more accurately represent the probability that the image to be identified belongs to the preset type.

[0123] Moreover, during the recognition process, the recognition model only performs one feature extraction operation (i.e., extracting each first-category feature map matrix based on the convolutional neural network), and the total time required to complete the image recognition task is relatively short.

[0124] This technical solution is described in detail below in conjunction with the accompanying drawings.

[0125] Figure 1 An exemplary process of an image recognition method is provided, comprising the following steps:

[0126] S100: extracting at least one first-category feature map matrix from a picture representation matrix of a to-be-recognized picture based on a convolutional neural network, and determining a first eigenvector according to each first-category feature map matrix.

[0127] S102: Mapping the first feature vector into a first probability representation value for representing the probability that the first recognized image to be recognized belongs to a preset type.

[0128] S104: If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates.

[0129] S106: Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix.

[0130] S108: Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type.

[0131] Figure 1 The method flow shown is applied to a recognition model, and the algorithm structure of the recognition model may include a feature extraction algorithm, a classification algorithm, and a positioning algorithm.

[0132] The input of the recognition model can be the mathematical representation corresponding to the image to be recognized. Usually, the matrix composed of the pixel values ​​of each pixel point in the image to be recognized according to the pixel position relationship is used as the mathematical representation corresponding to the image to be recognized, which can be called the image representation matrix.

[0133] Convolutional neural networks can usually be used to extract features from image representation matrices. There are many types of convolutional neural networks, for example, a convolutional neural network such as ResNET-50 can be used. A convolutional neural network has a network parameter set, which generally includes element values ​​in each convolution kernel matrix.

[0134] The input of a convolutional neural network can be a picture representation matrix, and the output can be a number of feature map matrices (for the purpose of description, they are called first-class feature map matrices). Taking ResNET-50 as an example, its output channel number is 2048, so it can output 2048 first-class feature map matrices, and the size of each first-class feature map matrix can be 7*7, that is, it contains 7 rows and 7 columns.

[0135] The classification algorithm in the recognition model can be used to implement binary classification (preset type and non-preset type), and the classification algorithm can be implemented using a vector mapping function. In this way, the parameters in the vector mapping function are called a mapping parameter set, which is used to calculate the vector and map the vector into a probability representation value that can represent the probability that the image to be identified belongs to a preset type (for descriptive distinction, it is called the first probability representation value).

[0136] In some embodiments, the mapping parameter set of the classification algorithm may include two mapping parameter subsets, one mapping parameter subset is used to map the vector into a probability representation value representing the probability that the image to be identified belongs to a preset type, and the other mapping parameter subset is used to map the vector into a probability representation value representing the probability that the image to be identified belongs to a non-preset type.

[0137] Since the vector mapping function of the classification algorithm requires the format of the input data to be a vector format, in step S100, each first-category feature map matrix needs to be further processed into a first feature vector.

[0138] There are many methods for further processing each first-class feature map matrix into a first feature vector. In some embodiments, each first-class feature map matrix can be subjected to a pooling operation (such as an average pooling operation, a maximum pooling operation, etc.), and each pooling operation result value is combined into a first feature vector.

[0139] The first probability representation value is greater than the first preset threshold value, indicating that after preliminary classification judgment, the probability that the image to be identified belongs to the preset type is relatively high. However, since the preliminary classification result is based on the image to be identified as a whole and does not exclude interference elements in the image to be identified, even if the first probability representation value is large, the image to be identified is likely not to belong to the preset type. In other words, the first probability representation value is affected by the interference factors in the image to be identified and is not accurate enough and can only be used as a preliminary classification result.

[0140] Therefore, if the first probability representation value is greater than the first preset threshold, secondary classification may be further performed on the features of possible preset type elements in the picture to be identified, so as to obtain a more accurate classification result.

[0141] In some embodiments, if the first probability representation value is not greater than the first preset threshold, it can be considered that the possibility that the picture to be identified belongs to the preset type is low. Considering that the classification algorithm still outputs a lower first probability representation value even under the influence of rich elements in the picture to be identified, it can be considered that the possibility that the picture to be identified belongs to a non-preset type is very high. The first probability representation value can be directly used as the output result of the recognition model, and the picture to be identified can be determined to belong to a non-preset type based on the first probability representation value.

[0142] In addition, the present disclosure may use a positioning algorithm in the recognition model to determine the features of possible preset type elements in the image to be recognized. The positioning algorithm may be used to locate a prediction area from the first type feature map matrix, and the prediction area corresponds to the preset type elements that may be contained in the first type feature map matrix.

[0143] The positioning algorithm can be implemented based on a vector mapping function. The various parameters of the vector mapping function of the positioning algorithm constitute a mapping parameter set of the positioning algorithm, that is, a mapping parameter set used to map the regional position coordinates. The vector mapping function of the positioning algorithm is different from the vector mapping function of the classification algorithm. The vector mapping function of the positioning algorithm is used to map the first feature vector into the regional position coordinates of the predicted area. Among them, the regional position coordinates may include a set of coordinate values ​​for delimiting the area range of the predicted area. It should be noted here that since the first feature vector can characterize the characteristics of the picture to be identified, and the picture to be identified is preliminarily classified as belonging to a preset type, this may mean that the first feature vector implies the position information of the preset type elements contained in the picture to be identified. Mapping the first feature vector into regional position coordinates can make the position information implied by the first feature vector explicit.

[0144] In some embodiments, the prediction area may be a rectangular area, and the corresponding area position coordinates may include coordinate values ​​corresponding to two vertices on the diagonal line of the rectangle.

[0145] In step S106, the prediction region may be cropped from each first-type feature map matrix according to the region position coordinates output by the positioning algorithm as a second-type feature map matrix.

[0146] In some embodiments, considering that the prediction region belongs to a matrix, the matrix is ​​composed of matrix elements, the position of each matrix element in the matrix corresponds to a pixel point, and the measurement unit of the coordinate value included in the region position coordinate is also a pixel point. If a coordinate value included in the region position coordinate is a non-integer, then the boundary of the prediction region cut out from the first type of feature map matrix using the region position coordinate directly may not cover a complete pixel point, and it is impossible to ensure that each matrix element in the cut prediction region is complete, and some matrix elements may be incomplete.

[0147] To this end, in step S106, if it is determined that any coordinate value included in the coordinate value set is a non-integer, the coordinate value can be adjusted to the integer closest to the non-integer. The area range defined by the adjusted coordinate value set covers the area range defined by the coordinate value set before adjustment. Then, the corresponding prediction area can be cut out from each first-class feature map matrix according to the adjusted coordinate value set. In this way, it can be ensured that each matrix element in the cut prediction area is complete.

[0148] Since the vector mapping function of the classification algorithm requires the format of the input data to be a vector format, in step S106, each second-category feature map matrix needs to be further processed into a second feature vector.

[0149] There are many methods for further processing each second-type feature map matrix into a second feature vector. In some embodiments, each second-type feature map matrix can be subjected to a pooling operation, and each pooling operation result value is combined into a second feature vector.

[0150] In step S108, the second feature vector can be used as an input to the classification algorithm, and the classification algorithm is executed again to map the second feature vector to a second probability representation value. Since the second feature vector corresponding to the prediction area is secondary classified after the interference elements in the image to be identified that are initially classified as the preset type are excluded, the second probability representation value obtained by the secondary classification is more accurate than the first probability representation value.

[0151] In addition, it should be noted that in some embodiments, the recognition model may include only one classification algorithm, which means that the classification algorithm used for the preliminary classification and the secondary classification may be the same classification algorithm. In other words, the mapping parameter set used to map the first probability representation value and the mapping parameter set used to map the second probability representation value may be the same parameter set.

[0152] In other embodiments, the recognition model may include two classification algorithms, and different classification algorithms have their own mapping parameter sets, which means that the classification algorithm used for the preliminary classification and the secondary classification may not be the same classification algorithm. In other words, the mapping parameter set used to map the first probability representation value and the mapping parameter set used to map the second probability representation value may be different parameter sets.

[0153] In addition, in some embodiments, considering that the main purpose of the preliminary classification in the recognition model is to roughly classify the images to be recognized, the images roughly classified as preset types can be further classified more accurately in a secondary classification. Therefore, the recognition standard of the preset type images corresponding to the preliminary classification can be looser than the recognition standard of the preset type images corresponding to the secondary classification. In other words, due to the higher accuracy of the secondary classification, the recognition standard of the preset type images corresponding to the secondary classification can be set more strictly, thereby improving the recognition accuracy of the recognition model.

[0154] A second preset threshold value greater than the first preset threshold value may be set. If the second probability representation value output by the recognition model is greater than the second preset threshold value, it may be determined that the image to be recognized belongs to a preset type.

[0155] Figure 2 An exemplary process of a method for training a recognition model is provided, comprising the following steps:

[0156] S20: Obtain a picture sample and an identification label corresponding to the picture sample.

[0157] S21: Execute steps S211-S216 to train the recognition model.

[0158] S211: extracting at least one first-category feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first eigenvector and a positioning label according to each first-category feature map matrix.

[0159] S212: Mapping the first feature vector into a first probability representation value for representing the probability that the first recognized image to be recognized belongs to a preset type.

[0160] S213: If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates.

[0161] S214: Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix.

[0162] S215: Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type.

[0163] S216: According to the training objective of the recognition model, adjust the parameter set corresponding to the recognition model.

[0164] The process of training the recognition model actually includes several iterations. In one iteration, the image samples with known recognition results (recognition labels) can be input into the recognition model to be trained. The training effect of this iteration is measured according to the output and label of the recognition model. With the goal of improving the training effect, the corresponding parameter set of the recognition model is adjusted, and then the next iteration training is started. After several iterations of training, a recognition model that meets the requirements can be obtained.

[0165] Since the algorithmic process involved in training the recognition model is similar in principle to the algorithmic process involved in applying the recognition model for recognition, the description of the algorithmic process involved in training the recognition model can be understood by referring to the previous text and will not be repeated here.

[0166] The identification labels corresponding to the image samples obtained in step 20 may usually be manually annotated.

[0167] In some embodiments, if the image sample belongs to a preset type, the training goal may be: reducing the difference between the probability representation value corresponding to the identification label and the first probability representation value, reducing the difference between the probability representation value corresponding to the identification label and the second probability representation value, and reducing the difference between the coordinates corresponding to the positioning label and the regional position coordinates.

[0168] Usually, a loss function can be defined, and reducing the loss function value is used as the training goal. The following is an example of a loss function:

[0169] Loss A =w1*Loss1+w2*Loss2+w3*Loss3;

[0170] Among them, Loss A Indicates the loss function of the recognition model used when the image sample belongs to the preset type; Loss1 represents the difference between the probability representation value corresponding to the recognition tag and the first probability representation value, Loss2 represents the difference between the coordinates corresponding to the positioning tag and the coordinates of the regional position, and Loss3 represents the difference between the probability representation value corresponding to the recognition tag and the second probability representation value. w1, w2, and w3 are weight values ​​that can be specified based on experience, for example, they can be 0.5, 0.4, and 0.1 respectively.

[0171] In other embodiments, if the image sample does not belong to the preset type, the training goal may be to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value. The following exemplary loss function is provided:

[0172] Loss B =Loss1;

[0173] Among them, Loss B It represents the loss function of the recognition model used when the image sample does not belong to the preset type, and Loss1 represents the difference between the probability representation value corresponding to the recognition label and the first probability representation value.

[0174] In some embodiments, the set of parameters to be adjusted corresponding to the recognition model may include:

[0175] A mapping parameter set used for mapping to obtain a first probability representation value; a mapping parameter set used for mapping to obtain a second probability representation value; and a mapping parameter set used for mapping to obtain regional position coordinates.

[0176] It should be noted that, in these embodiments, the network parameter set of the convolutional neural network may be pre-set, and the network parameter set of the convolutional neural network may not need to be adjusted during the model training process.

[0177] In other embodiments, the parameter set to be adjusted corresponding to the recognition model may include: a mapping parameter set for mapping to obtain a first probability representation value; a mapping parameter set for mapping to obtain a second probability representation value; a mapping parameter set for mapping to obtain regional location coordinates; and a network parameter set of a convolutional neural network.

[0178] In some embodiments, the mapping parameter set used to map the first probability representation value and the mapping parameter set used to map the second probability representation value may be the same parameter set. In this way, the limited image samples can be fully utilized to train the mapping parameter set shared by the preliminary classification and the secondary classification, thereby improving the training efficiency.

[0179] In some other embodiments, the mapping parameter set used for mapping to obtain the first probability representation value and the mapping parameter set used for mapping to obtain the second probability representation value may be different parameter sets.

[0180] In some embodiments, the region location coordinates may include a set of coordinate values ​​for defining the area range of the prediction region. In this case, the mapping parameter set for mapping the region location coordinates may include: different weight value sets corresponding to different coordinate values ​​in the coordinate value set, wherein each weight value set includes each weight value corresponding to each dimension value in the first feature vector.

[0181] In this way, mapping the first feature vector into regional position coordinates can include: for each weight value set, calculating the weighted sum according to each dimension value in the first feature vector, the weight value set, and according to the one-to-one correspondence between the dimension value and the weight value; determining the coordinate value corresponding to the weight value set according to the calculated weighted sum; and forming the determined coordinate values ​​into the coordinate value set.

[0182] In some embodiments, a mapping parameter set used to map the first probability representation value may include: a first weight value set; the first weight value set includes each weight value corresponding to each dimensional value in the first feature vector.

[0183] In this way, mapping the first feature vector into a first probability representation value may include: calculating a weighted sum based on each dimension value in the first feature vector, the first weight value set, and a one-to-one correspondence between the dimension value and the weight value; and determining the first probability representation value based on the calculated weighted sum.

[0184] In some embodiments, the mapping parameter set used to map to obtain the second probability representation value may include: a second weight value set; the second weight value set includes each weight value corresponding to each dimensional value in the second feature vector.

[0185] In this way, mapping the second feature vector into a second probability representation value may include: calculating a weighted sum based on each dimension value in the second feature vector, the second weight value set, and a one-to-one correspondence between the dimension value and the weight value; and determining the second probability representation value based on the calculated weighted sum.

[0186] In step S211, the manner of determining the positioning label according to each first-category feature map matrix may include at least two of the following:

[0187] In some embodiments, each first-category feature map matrix may be provided to manual labeling of positioning tags.

[0188] In other embodiments, each first-category feature map matrix may be subjected to a pooling operation, and each pooling operation result value may be combined into a first feature vector. Then, according to the classification result given by the preliminary classification algorithm for the first feature vector, the positioning label may be determined. The specific method is as follows:

[0189] Each first-class feature map matrix and each weight value of the first weight value set can be set to have a one-to-one correspondence. The value of the i-th position in each first-class feature map matrix can be obtained; the weighted sum is calculated based on the obtained values ​​of each i-th position, the first weight value set, and the one-to-one correspondence between the first-class feature map matrix and the weight value, and the calculated weighted sum is used as the replacement value corresponding to the i-th position; i = (1, 2, ..., N), N is the number of positions contained in the first-class feature map matrix.

[0190] In this way, each position in the first type feature map matrix can be reassigned with a corresponding replacement value to obtain a replacement feature map matrix. Each replacement value in the replacement feature map matrix can be used as a pixel value to generate a replacement feature map; wherein the pixel value of a pixel point in the replacement feature map is positively correlated with the brightness of the pixel point. Finally, the high brightness area in the replacement feature map can be determined, and the area position coordinates corresponding to the high brightness area can be determined as a positioning label.

[0191] When determining the high brightness area in the replacement feature map, the following method can be selected: converting the replacement feature map into a binary map, and performing a connected domain analysis on the binary map; determining the high brightness area according to the analysis result. Accordingly, the bounding rectangle of the high brightness area can be determined, and the area position coordinates used to delimit the bounding rectangle can be obtained.

[0192] Before determining the high-brightness area in the replacement feature map, the resolution of the replacement feature map can also be adjusted to the resolution of the image sample. The reason for doing this is that if the resolution of the replacement feature map is too small, it means that the resolution of the corresponding high-brightness area is also relatively small, and the edge lines of the high-brightness area are too thick, so the high-brightness area delineated in this way is not accurate enough. Enlarging the resolution of the replacement feature map to the resolution of the image sample can make the edge lines of the corresponding high-brightness area thinner, and the high-brightness area delineated can be more accurate.

[0193] Figure 3 The following are examples of picture samples, replacement of high brightness areas in feature maps, and visualization of replacement of high brightness areas in feature maps. Figure 3 As shown, the handcuffed hands in the picture sample are prohibited elements, and the high-brightness area of ​​the category activation feature map is the area containing the prohibited elements. In the visualized replacement feature map, the high-brightness area can be seen more clearly.

[0194] Figure 4 The process of determining the high brightness area in the replacement feature map is shown as an example. Figure 4 As shown, the replacement feature map is first converted into a binary map, and then the binary map is analyzed by connected domain to determine the circumscribed rectangle of the high-brightness area as the high-brightness area.

[0195] The principle of the above-mentioned method for determining positioning labels without relying on manual annotation is that if the preliminary classification algorithm is required to correctly classify image samples with identification labels as preset types as belonging to the preset types, it means that in the replacement feature map obtained by calculating based on each first-category feature map matrix and the first weight value set, the higher the brightness of the area, the greater the contribution to the preliminary classification result, and the area that contributes more to the preliminary classification result should be the area containing elements of the preset type.

[0196] Therefore, using this principle, during the recognition model training process, the area position coordinates corresponding to the high-brightness area in the replacement feature map are used as positioning labels, and the positioning labels are used as supervisory signals for model training, which can save the need for manual labeling of positioning labels. During the training process, as the first weight value set (the mapping parameter set for preliminary classification) and the network parameter set of the convolutional neural network are iteratively optimized, the determined positioning labels will become more accurate with iterations. Since the positioning labels themselves are also iteratively optimized, the supervisory signals provided by the positioning labels can be considered to be weak supervisory signals, which are different from the strong supervisory signals manually labeled.

[0197] In addition, through Figure 2 The recognition model training method shown integrates the classification algorithm and the positioning algorithm into the same recognition model. The classification algorithm and the positioning algorithm share a unified feature encoding. In addition, reducing the loss of the classification algorithm and reducing the loss of the positioning algorithm are simultaneously used as training goals. This can achieve implicit collaboration between the classification algorithm and the positioning algorithm, and improve the classification algorithm's perception of the predicted area in the picture that may contain preset type elements, thereby improving the training effect of the recognition model.

[0198] Figure 5 An image recognition device is exemplarily provided, which is applied to a recognition model, and the device comprises:

[0199] A first feature vector determination module 501 extracts at least one first-type feature map matrix from a picture representation matrix of a picture to be identified based on a convolutional neural network, and determines a first feature vector according to each first-type feature map matrix;

[0200] A first classification module 502 maps the first feature vector into a first probability representation value for representing the probability that the first identified image to be identified belongs to a preset type;

[0201] A positioning module 503, if the first probability representation value is greater than a first preset threshold, maps the first feature vector into a region position coordinate, the region position coordinate is used to determine a corresponding prediction region, the prediction region corresponds to a preset type element that may be included in the first type feature map matrix;

[0202] A second eigenvector determining module 504 cuts out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determines a second eigenvector according to each second-type feature map matrix;

[0203] The second classification module 505 maps the second feature vector into a second probability representation value and outputs the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type.

[0204] In some embodiments, the first feature vector module 501 performs a pooling operation on each first-category feature map matrix, and combines the pooling operation result values ​​into a first feature vector.

[0205] In some embodiments, the second feature vector determination module 504 performs a pooling operation on each second-type feature map matrix, and combines the pooling operation result values ​​into a second feature vector.

[0206] In some embodiments, the region location coordinates include a set of coordinate values ​​for defining an area range of the prediction region.

[0207] In some embodiments, the second feature vector determination module 504 adjusts any coordinate value included in the coordinate value set to an integer closest to the non-integer if the coordinate value included in the coordinate value set is a non-integer; wherein the area range defined by the adjusted coordinate value set covers the area range defined by the coordinate value set before the adjustment; and according to the adjusted coordinate value set, a corresponding prediction area is cropped out from each first-category feature map matrix.

[0208] In some embodiments, the mapping parameter set used to map to obtain the first probability representation value and the mapping parameter set used to map to obtain the second probability representation value are the same parameter set.

[0209] In some embodiments, the first classification module 502 outputs the first probability representation value if the first probability representation value is not greater than a first preset threshold.

[0210] In some embodiments, it also includes:

[0211] The identification module 506 determines that the image to be identified belongs to a preset type if the second probability representation value output by the recognition model is greater than a second preset threshold; wherein the second preset threshold is greater than the first preset threshold.

[0212] Figure 6 An exemplary method for training a recognition model is provided, comprising:

[0213] An acquisition module 601 acquires a picture sample and an identification tag corresponding to the picture sample, wherein the identification tag is used to determine whether the picture sample belongs to a preset type;

[0214] The training module 602 performs the following steps to train the recognition model:

[0215] Extracting at least one first-class feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first feature vector and a positioning label according to each first-class feature map matrix; the positioning label is used to determine a corresponding actual area, and the actual area corresponds to a preset type element that may be included in the first-class feature map matrix;

[0216] Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type;

[0217] If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix;

[0218] Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix;

[0219] Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type;

[0220] According to the training objective of the recognition model, the parameter set corresponding to the recognition model is adjusted.

[0221] In some embodiments, if the image sample belongs to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value, and to reduce the difference between the probability representation value corresponding to the identification label and the second probability representation value, and to reduce the difference between the coordinates corresponding to the positioning label and the coordinates of the area position.

[0222] In some embodiments, if the image sample does not belong to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value.

[0223] In some embodiments, the parameter set corresponding to the recognition model includes:

[0224] A mapping parameter set used for mapping to obtain a first probability representation value;

[0225] A mapping parameter set used for mapping to obtain a second probability representation value;

[0226] The mapping parameter set used to map the region location coordinates.

[0227] In some embodiments, the parameter set corresponding to the recognition model also includes:

[0228] The network parameter set of the convolutional neural network.

[0229] In some embodiments, the mapping parameter set used to map to obtain the first probability representation value and the mapping parameter set used to map to obtain the second probability representation value are the same parameter set.

[0230] In some embodiments, the area location coordinates include a set of coordinate values ​​for defining an area range of the prediction area;

[0231] A mapping parameter set for mapping to obtain the coordinates of the regional position, comprising: different weight value sets corresponding to different coordinate values ​​in the coordinate value set, wherein each weight value set comprises respective weight values ​​corresponding one-to-one to respective dimensional values ​​in the first feature vector;

[0232] The training module 602 calculates the weighted sum for each weight value set according to the dimension values ​​in the first feature vector, the weight value set, and the one-to-one correspondence between the dimension values ​​and the weight values; determines the coordinate value corresponding to the weight value set according to the calculated weighted sum; and organizes the determined coordinate values ​​into the coordinate value set.

[0233] In some embodiments, a mapping parameter set used to map to obtain a first probability representation value includes:

[0234] A first weight value set; the first weight value set includes weight values ​​corresponding one-to-one to each dimension value in the first feature vector;

[0235] The training module 602 calculates a weighted sum according to each dimension value in the first feature vector, the first weight value set, and a one-to-one correspondence between the dimension value and the weight value; and determines a first probability representation value according to the calculated weighted sum.

[0236] In some embodiments, a mapping parameter set used to map to obtain a second probability representation value includes:

[0237] A second weight value set; the second weight value set includes weight values ​​corresponding one-to-one to each dimension value in the second feature vector;

[0238] The training module 602 calculates a weighted sum according to each dimension value in the second feature vector, the second weight value set, and a one-to-one correspondence between the dimension value and the weight value; and determines a second probability representation value according to the calculated weighted sum.

[0239] In some embodiments, the training module performs a pooling operation on each first-category feature map matrix, and combines the values ​​of each pooling operation result into a first feature vector.

[0240] In some embodiments, each first-type feature map matrix and each weight value of the first weight value set have a one-to-one correspondence;

[0241] The training module 602 obtains the value of the i-th position in each first-class feature map matrix; calculates the weighted sum according to the obtained value of each i-th position, the first weight value set, and the one-to-one correspondence between the first-class feature map matrix and the weight value, and uses the calculated weighted sum as the replacement value corresponding to the i-th position; i = (1, 2, ..., N), N is the number of positions contained in the first-class feature map matrix; reassigns corresponding replacement values ​​to each position in the first-class feature map matrix to obtain a replacement feature map matrix; uses each replacement value in the replacement feature map matrix as a pixel value to generate a replacement feature map; wherein the pixel value of the pixel point in the replacement feature map is positively correlated with the brightness of the pixel point; determines the high-brightness area in the replacement feature map, and determines the area position coordinates corresponding to the high-brightness area as a positioning label.

[0242] In some embodiments, the training module 602 converts the replacement feature map into a binary map and performs a connected domain analysis on the binary map; determines a high-brightness area based on the analysis results; determines a bounding rectangle of the high-brightness area, and obtains the area position coordinates used to delineate the bounding rectangle.

[0243] In some embodiments, the training module 602 adjusts the resolution of the replacement feature map to the resolution of the image sample before determining the high-brightness area in the replacement feature map.

[0244] It should be noted that, although several units / modules or subunits / modules of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.

[0245] Figure 7 1 is a schematic diagram of a computer-readable storage medium provided by the present disclosure. A computer program is stored on the medium 140. When the program is executed by a processor, the method of any embodiment of the present disclosure is implemented.

[0246] The present disclosure also provides a computing device, including a memory and a processor; the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement the method of any embodiment of the present disclosure when executing the computer instructions.

[0247] Figure 8It is a structural diagram of a computing device provided by the present disclosure. The computing device 15 may include but is not limited to: a processor 151, a memory 152, and a bus 153 connecting different system components (including the memory 152 and the processor 151).

[0248] The memory 152 stores computer instructions, which can be executed by the processor 131, so that the processor 151 can execute the method of any embodiment of the present disclosure. The memory 152 may include a random access memory unit RAM 1521, a cache memory unit 1522 and / or a read-only memory unit ROM 1523. The memory 152 may also include: a program tool 1525 having a set of program modules 1524, the program modules 1524 include but are not limited to: an operating system, one or more application programs, other program modules and program data, and one or more combinations of these program modules may include the implementation of a network environment.

[0249] The bus 153 may include, for example, a data bus, an address bus, and a control bus. The computing device 15 may also communicate with an external device 155 through an I / O interface 154, and the external device 155 may be, for example, a keyboard, a Bluetooth device, etc. The computing device 15 may also communicate with one or more networks through a network adapter 156, and for example, the network may be a local area network, a wide area network, a public network, etc. The network adapter 156 may also communicate with other modules of the computing device 15 through the bus 153.

[0250] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0251] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. A picture recognition method, applied to a recognition model, the method comprising: Extracting at least one first-category feature map matrix from a picture representation matrix of the picture to be identified based on a convolutional neural network, and determining a first eigenvector according to each first-category feature map matrix; Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type; If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix; Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix; Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type; The mapping parameter set used for mapping to obtain the first probability representation value and the mapping parameter set used for mapping to obtain the second probability representation value are the same parameter set.

2. The method of claim 1, wherein determining the first eigenvector according to each first-type feature map matrix comprises: A pooling operation is performed on each first-category feature map matrix, and each pooling operation result value is combined into a first feature vector.

3. The method of claim 1, wherein determining the second eigenvector according to each second-type feature map matrix comprises: A pooling operation is performed on each second-category feature map matrix, and each pooling operation result value is combined into a second feature vector.

4. The method according to claim 1, wherein the region location coordinates include a set of coordinate values ​​used to define an area range of the prediction region.

5. The method of claim 4, wherein the corresponding prediction region is cut out from each first-category feature map matrix according to the region position coordinates, comprising: If any coordinate value included in the coordinate value set is a non-integer, the coordinate value is adjusted to an integer closest to the non-integer; wherein the area range defined by the adjusted coordinate value set covers the area range defined by the coordinate value set before the adjustment; According to the adjusted coordinate value set, a corresponding prediction area is cropped from each first-category feature map matrix.

6. The method of claim 1, further comprising: If the first probability representation value is not greater than a first preset threshold, the first probability representation value is output.

7. The method of claim 1, further comprising: If the second probability representation value output by the recognition model is greater than a second preset threshold, it is determined that the image to be recognized belongs to a preset type; The second preset threshold is greater than the first preset threshold.

8. A method for training a recognition model, comprising: Obtaining an image sample and an identification tag corresponding to the image sample, wherein the identification tag is used to determine whether the image sample belongs to a preset type; and Perform the following steps to train the recognition model: Extracting at least one first-class feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first feature vector and a positioning label according to each first-class feature map matrix; the positioning label is used to determine a corresponding actual area, and the actual area corresponds to a preset type element that may be included in the first-class feature map matrix; Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type; If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix; Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix; Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type; According to the training objective of the recognition model, adjusting the parameter set corresponding to the recognition model; The parameter set corresponding to the recognition model includes: A mapping parameter set used for mapping to obtain a first probability representation value; A mapping parameter set used for mapping to obtain a second probability representation value; A mapping parameter set used to map the coordinates of the region location; The mapping parameter set used for mapping to obtain the first probability representation value and the mapping parameter set used for mapping to obtain the second probability representation value are the same parameter set.

9. The method of claim 8, wherein: If the image sample belongs to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value, and to reduce the difference between the probability representation value corresponding to the identification label and the second probability representation value, and to reduce the difference between the coordinates corresponding to the positioning label and the regional position coordinates.

10. The method of claim 8 or 9, wherein if the image sample does not belong to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value.

11. The method according to claim 8, wherein the parameter set corresponding to the recognition model further comprises: The network parameter set of the convolutional neural network.

12. The method of claim 8, wherein the region position coordinates include a set of coordinate values ​​for defining an area range of the prediction region; The mapping parameter set used to map the region location coordinates includes: Different weight value sets corresponding to different coordinate values ​​in the coordinate value set, wherein each weight value set includes respective weight values ​​corresponding one-to-one to respective dimensional values ​​in the first feature vector; Mapping the first feature vector into regional position coordinates includes: For each weight value set, a weighted sum is calculated according to each dimension value in the first feature vector, the weight value set, and a one-to-one correspondence between the dimension value and the weight value; Determine the coordinate value corresponding to the weight value set according to the calculated weighted sum; The determined coordinate values ​​are combined into the coordinate value set.

13. The method of claim 8, wherein the mapping parameter set used to map the first probability representation value comprises: A first weight value set; the first weight value set includes weight values ​​corresponding one-to-one to each dimension value in the first feature vector; Mapping the first feature vector into a first probability representation value includes: Calculate a weighted sum according to each dimension value in the first feature vector, the first weight value set, and a one-to-one correspondence between the dimension value and the weight value; A first probability characterization value is determined according to the calculated weighted sum.

14. The method of claim 8, wherein the mapping parameter set used to map the second probability representation value comprises: A second weight value set; the second weight value set includes weight values ​​corresponding one-to-one to each dimension value in the second feature vector; Mapping the second feature vector into a second probability representation value includes: Calculate a weighted sum according to each dimension value in the second feature vector, the second weight value set, and a one-to-one correspondence between the dimension value and the weight value; A second probability characterization value is determined according to the calculated weighted sum.

15. The method of claim 13, wherein determining the first eigenvector according to each first-type feature map matrix comprises: A pooling operation is performed on each first-category feature map matrix, and each pooling operation result value is combined into a first feature vector.

16. The method of claim 15, wherein each first-type feature map matrix and each weight value of the first weight value set have a one-to-one correspondence; The positioning labels are determined according to each first-category feature map matrix, including: Get the value of the i-th position in each first-category feature map matrix; According to the obtained values ​​of each i-th position, the first weight value set, and the one-to-one correspondence between the first type feature map matrix and the weight value, a weighted sum is calculated, and the calculated weighted sum is used as the replacement value corresponding to the i-th position; i = (1, 2, ..., N), N is the number of positions included in the first type feature map matrix; Reassign corresponding replacement values ​​to each position in the first type feature map matrix to obtain a replacement feature map matrix; Taking each replacement value in the replacement feature map matrix as a pixel value, a replacement feature map is generated; wherein the pixel value of a pixel point in the replacement feature map is positively correlated with the brightness of the pixel point; Determine a high-brightness area in the replacement feature map, and determine the area position coordinates corresponding to the high-brightness area as a positioning label.

17. The method of claim 16, wherein determining the high brightness area in the replacement feature map comprises: Converting the replacement feature map into a binary map, and performing connected domain analysis on the binary map; According to the analysis results, the high brightness area is determined; Determining the area position coordinates corresponding to the high brightness area includes: The circumscribed rectangle of the high-brightness area is determined, and the area position coordinates used to define the circumscribed rectangle are obtained.

18. The method of claim 16, before determining the high brightness area in the replacement feature map, the method further comprises: The resolution of the replacement feature map is adjusted to the resolution of the image sample.

19. A picture recognition device, applied to a recognition model, comprising: A first feature vector determination module extracts at least one first-type feature map matrix from a picture representation matrix of the picture to be identified based on a convolutional neural network, and determines a first feature vector according to each first-type feature map matrix; A first classification module maps the first feature vector into a first probability representation value for representing the probability that the first identified image to be identified belongs to a preset type; A positioning module, if the first probability representation value is greater than a first preset threshold, maps the first feature vector into a regional position coordinate, the regional position coordinate is used to determine a corresponding prediction area, the prediction area corresponds to a preset type element that may be included in the first type feature map matrix; A second eigenvector determination module, which cuts out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determines a second eigenvector according to each second-type feature map matrix; A second classification module maps the second feature vector into a second probability representation value and outputs the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type; The mapping parameter set used for mapping to obtain the first probability representation value and the mapping parameter set used for mapping to obtain the second probability representation value are the same parameter set.

20. The device as claimed in claim 19, wherein the first feature vector module performs a pooling operation on each first-type feature map matrix respectively, and combines the values ​​of each pooling operation result into a first feature vector.

21. In the device as described in claim 19, the second feature vector determination module performs pooling operations on each second-type feature map matrix respectively, and composes each pooling operation result value into a second feature vector.

22. The apparatus as claimed in claim 19, wherein the region location coordinates include a set of coordinate values ​​for defining an area range of the prediction region.

23. The device of claim 22, wherein the second feature vector determination module is configured to adjust any coordinate value included in the coordinate value set to an integer closest to the non-integer if the coordinate value is a non-integer; wherein: The area range defined by the adjusted coordinate value set covers the area range defined by the coordinate value set before adjustment; and according to the adjusted coordinate value set, a corresponding prediction area is cropped out from each first-category feature map matrix.

24. The device as claimed in claim 19, wherein the first classification module outputs the first probability characterization value if the first probability characterization value is not greater than a first preset threshold.

25. The apparatus of claim 19, further comprising: The identification module determines that the image to be identified belongs to a preset type if the second probability representation value output by the identification model is greater than a second preset threshold; wherein the second preset threshold is greater than the first preset threshold.

26. A device for training a recognition model, comprising: An acquisition module, acquiring a picture sample and an identification tag corresponding to the picture sample, wherein the identification tag is used to determine whether the picture sample belongs to a preset type; The training module performs the following steps to train the recognition model: Extracting at least one first-class feature map matrix from the picture representation matrix of the picture sample based on a convolutional neural network, and determining a first feature vector and a positioning label according to each first-class feature map matrix; the positioning label is used to determine a corresponding actual area, and the actual area corresponds to a preset type element that may be included in the first-class feature map matrix; Mapping the first feature vector into a first probability representation value, which is used to represent the probability that the first identified image to be identified belongs to a preset type; If the first probability representation value is greater than a first preset threshold, mapping the first feature vector into regional position coordinates, the regional position coordinates are used to determine a corresponding prediction region, and the prediction region corresponds to a preset type element that may be included in the first type feature map matrix; Cut out a corresponding prediction region from each first-type feature map matrix as a second-type feature map matrix, and determine a second eigenvector according to each second-type feature map matrix; Mapping the second feature vector into a second probability representation value and outputting the second probability representation value, which is used to represent the probability that the re-recognized image to be recognized belongs to a preset type; According to the training objective of the recognition model, adjusting the parameter set corresponding to the recognition model; The parameter set corresponding to the recognition model includes: A mapping parameter set used for mapping to obtain a first probability representation value; A mapping parameter set used for mapping to obtain a second probability representation value; A mapping parameter set used to map the coordinates of the region location; The mapping parameter set used for mapping to obtain the first probability representation value and the mapping parameter set used for mapping to obtain the second probability representation value are the same parameter set.

27. The device of claim 26, wherein: If the image sample belongs to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value, and to reduce the difference between the probability representation value corresponding to the identification label and the second probability representation value, and to reduce the difference between the coordinates corresponding to the positioning label and the regional position coordinates.

28. The device of claim 26 or 27, wherein if the image sample does not belong to a preset type, the training goal is to reduce the difference between the probability representation value corresponding to the identification label and the first probability representation value.

29. The apparatus of claim 26, wherein the parameter set corresponding to the recognition model further comprises: The network parameter set of the convolutional neural network.

30. The apparatus of claim 26, wherein the region location coordinates include a set of coordinate values ​​for defining an area range of the prediction region; A mapping parameter set for mapping to obtain the coordinates of the regional position, comprising: different weight value sets corresponding to different coordinate values ​​in the coordinate value set, wherein each weight value set comprises respective weight values ​​corresponding one-to-one to respective dimensional values ​​in the first feature vector; The training module calculates the weighted sum for each weight value set based on the dimensional values ​​in the first feature vector, the weight value set, and the one-to-one correspondence between the dimensional values ​​and the weight values; determines the coordinate value corresponding to the weight value set based on the calculated weighted sum; and organizes the determined coordinate values ​​into the coordinate value set.

31. The apparatus of claim 26, wherein the mapping parameter set for mapping to obtain the first probability representation value comprises: A first weight value set; the first weight value set includes weight values ​​corresponding one-to-one to each dimension value in the first feature vector; The training module calculates a weighted sum based on each dimension value in the first feature vector, the first weight value set, and a one-to-one correspondence between the dimension value and the weight value; and determines a first probability representation value based on the calculated weighted sum.

32. The apparatus of claim 26, wherein the mapping parameter set used for mapping to obtain the second probability representation value comprises: A second weight value set; the second weight value set includes weight values ​​corresponding one-to-one to each dimension value in the second feature vector; The training module calculates the weighted sum according to each dimension value in the second feature vector, the second weight value set, and the one-to-one correspondence between the dimension value and the weight value; and determines the second probability representation value according to the calculated weighted sum.

33. In the device as described in claim 31, the training module performs pooling operations on each first-category feature map matrix respectively, and combines the values ​​of each pooling operation result into a first feature vector.

34. The device of claim 33, wherein each first-type feature map matrix and each weight value of the first weight value set have a one-to-one correspondence; The training module obtains the value of the i-th position in each first-class feature map matrix; calculates the weighted sum according to the obtained value of each i-th position, the first weight value set, and the one-to-one correspondence between the first-class feature map matrix and the weight value, and uses the calculated weighted sum as the replacement value corresponding to the i-th position; i = (1, 2, ..., N), N is the number of positions included in the first-class feature map matrix; reassigns the corresponding replacement value to each position in the first-class feature map matrix to obtain a replacement feature map matrix; uses each replacement value in the replacement feature map matrix as a pixel value to generate a replacement feature map; wherein, The pixel value of the pixel point in the replacement feature map is positively correlated with the brightness of the pixel point; the high-brightness area in the replacement feature map is determined, and the area position coordinates corresponding to the high-brightness area are determined as the positioning label.

35. As claimed in the device of claim 34, the training module converts the replacement feature map into a binary map and performs a connected domain analysis on the binary map; determines the high-brightness area based on the analysis result; determines the circumscribed rectangle of the high-brightness area, and obtains the area position coordinates for delimiting the circumscribed rectangle.

36. In the apparatus as claimed in claim 34, the training module adjusts the resolution of the replacement feature map to the resolution of the image sample before determining the high-brightness area in the replacement feature map.

37. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 18 is implemented.

38. A computing device, comprising a memory and a processor; the memory is used to store computer instructions executable on the processor, and the processor is used to implement any one of the methods described in claims 1 to 18 when executing the computer instructions.

Citation Information

Patent Citations

  • Method and device for identification of object node in image, terminal and readable medium

    CN108520247A