Image recognition method, device and electronic equipment

By performing dual classification recognition of target objects and picture styles in the picture recognition model and integrating the results, the problems of low recognition accuracy and high call-missing rate in the prior art are solved, and higher recognition accuracy and lower call-rate are achieved.

CN115035347BActive Publication Date: 2025-05-23MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210725302.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-05-23
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

When identifying sensitive pictures in the Internet, the detection algorithm adopts a detection box method, which can only recognize local information of the picture, resulting in low recognition accuracy and high error call rate.

Method used

Classification and recognition processing is performed by obtaining the target image to be identified and inputting it into the pre-trained image recognition model. The model simultaneously performs the first classification recognition of the target object and the second classification recognition of the picture style, and combines the results of the two to improve the recognition accuracy.

Benefits of technology

By taking into account the global features of the picture and the characteristics of the local target objects, the limitations of picture recognition are reduced, the accuracy of picture recognition is improved, and the error call rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035347B_ABST
    Figure CN115035347B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device and electronic device for image recognition, including: obtaining a target image to be recognized; inputting the target image into a pre-trained image recognition model for classification and recognition processing, and outputting a classification and recognition result of the target image, wherein the image recognition model is used to perform a first classification and recognition on the category to which a target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition are fused to obtain the classification and recognition result; and determining the target category to which the target image and its target object belong in common according to the classification and recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a picture recognition method, device and electronic equipment. Background Art

[0002] With the rapid development of the Internet industry, there are often malicious users on the Internet who deliberately post sensitive pictures to incite public opinion and seek illegal profits. In order to protect the rights and interests of users, it is necessary to identify sensitive pictures on the Internet and have reviewers review the sensitive pictures to determine whether to post the sensitive pictures on the Internet.

[0003] In some scenarios, the detection algorithm for identifying sensitive images uses detection frames to identify the images to be identified, and each detection frame only contains local information of the image to be identified. Since it can only identify local information of the image to be identified, the image recognition has great limitations and the image recognition accuracy is low, resulting in a large number of false calls in image recognition. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method, device and electronic device for image recognition, which improve the recognition accuracy of images.

[0005] In order to solve the above technical problems, the embodiments of the present application are implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides an image recognition method, including: obtaining a target image to be identified; inputting the target image into a pre-trained image recognition model for classification and recognition processing, and outputting a classification and recognition result of the target image, wherein the image recognition model is used to perform a first classification and recognition on the category to which a target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition are fused to obtain the classification and recognition result; and determining the target category to which the target image and its target object belong in common according to the classification and recognition result.

[0007] In a second aspect, an embodiment of the present application provides a method for training an image recognition model, comprising obtaining multiple sample images; generating a training sample set based on the multiple sample images, wherein each training sample in the training sample set is annotated with a label, the label comprising a first category label of a category to which a target object in the training sample belongs and a second category label of a style to which the training sample belongs, and the first category label and the second category label have a corresponding relationship; inputting the training sample set into the image recognition model to be trained for iterative training until the loss function corresponding to the image recognition model converges, thereby obtaining a trained image recognition model, the loss function representing the error between the predicted value of the target category to which the sample image and its target object belong output by the image recognition model and the true value, and the true value is determined based on the first category label and the second category label.

[0008] In a third aspect, an embodiment of the present application provides an image recognition device, comprising: an acquisition module, used to acquire a target image to be recognized; a recognition module, used to input the target image into a pre-trained image recognition model for classification and recognition processing, and output a classification and recognition result of the target image, wherein the image recognition model is used to perform a first classification and recognition on the category to which a target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and to perform a fusion process on a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition to obtain the classification and recognition result;

[0009] A determination module is used to determine the target category to which the target image and its target object belong together according to the classification and recognition result.

[0010] In a fourth aspect, an embodiment of the present application provides a training device for an image recognition model, comprising: an acquisition module for acquiring multiple sample images; a generation module for generating a training sample set based on the multiple sample images, wherein each training sample in the training sample set is annotated with a label, and the label includes a first category label of a category to which a target object in the training sample belongs and a second category label of a style to which the training sample belongs, and the first category label and the second category label have a corresponding relationship; a training module for inputting the training sample set into the image recognition model to be trained for iterative training until the loss function corresponding to the image recognition model converges, thereby obtaining a trained image recognition model, wherein the loss function represents the error between the predicted value of the target category to which the sample image and its target object belong output by the image recognition model and the true value, and the true value is determined based on the first category label and the second category label.

[0011] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory communicate with each other through a bus; the memory is used to store computer programs; the processor is used to execute the programs stored in the memory to implement the image recognition method described in the first aspect or the training method steps of the image recognition model described in the second aspect.

[0012] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the training method steps of the image recognition method described in the first aspect or the image recognition model described in the second aspect are implemented.

[0013] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the image recognition method described in the first aspect or the training method steps of the image recognition model described in the second aspect.

[0014] It can be seen from the technical solutions provided by the above embodiments of the present application that by obtaining a target image to be identified;

[0015] The target image is input into an image recognition model for classification and recognition processing, and a classification and recognition result of the target image output by the image recognition model is obtained. The image recognition model is used to perform a first classification and recognition on the category to which the target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition are fused; the target category to which the target image and its target object belong together is determined based on the classification and recognition result. In this way, the image recognition model takes into account both the global features of the target image and the local features of the target object, reduces the limitations of image recognition, improves the recognition accuracy of the image, and reduces the false call rate of image recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0017] Figure 1 A schematic diagram of a first process flow of the image recognition method provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of a method for training an image recognition model provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of a clustering result is provided for an embodiment of the present application;

[0020] Figure 4 A schematic diagram of a model training process is provided for an embodiment of the present application;

[0021] Figure 5 A schematic diagram of a probability graph of a model output is provided for an embodiment of the present application;

[0022] Figure 6 A schematic diagram of the functional modules of a picture recognition device is provided for an embodiment of the present application;

[0023] Figure 7 A schematic diagram of functional modules of a training device for an image recognition model is provided for an embodiment of the present application;

[0024] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The purpose of the embodiments of the present application is to provide a method, device and electronic device for image recognition, which improve the recognition accuracy of images.

[0026] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0027] As mentioned above, the detection algorithm for identifying sensitive images uses a detection frame to identify the image to be identified, and each detection frame only contains local information of the image to be identified. Since it can only recognize local information of the image to be identified, it lacks consideration of the global features of the entire image. Although some local features of the image meet the target, the local features do not meet the target from the perspective of the entire image. The image recognition has great limitations and the image recognition accuracy is low, which leads to a large number of false calls in image recognition.

[0028] In order to solve the above technical problems, the embodiments of the present application provide a method, device and electronic device for image recognition. The method, device and electronic device for image recognition provided by the embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0029] like Figure 1 As shown, the embodiment of the present application provides an image recognition method, and the execution subject of the method can be a server, wherein the server can be an independent server or a server cluster composed of multiple servers. The image recognition method can specifically include the following steps S101-S105:

[0030] In step S101, a target image to be identified is obtained.

[0031] Specifically, the target image to be identified can be an image with a target object to be detected in the image, and identifying the target image specifically involves identifying the category of the target object in the image and the category of the style to which the target image belongs. Among them, the target image can be a sensitive image with sensitive information. For such a sensitive image, the embodiment of the present application needs to identify the sensitive information in the image and the category of the sensitive image. For example, the target image is a comic-style image with a spoof character. When identifying the image, it is necessary to identify the image style to which the image belongs and the category to which the spoof character in the image belongs.

[0032] In step S103, the target image is input into the image recognition model for classification and recognition processing, and the classification and recognition result of the target image is output.

[0033] Among them, the image recognition model is used to perform a first classification recognition on the category to which the target object in the target image belongs and a second classification recognition on the style to which the target image belongs, and to fuse the first sub-classification recognition result of the first classification recognition and the second sub-classification result of the second classification recognition to obtain a classification recognition result.

[0034] Specifically, the target object in the target image can be an object that needs to be detected in the image, such as a person, a landscape, or a sensitive symbol (such as funny text, funny illustrations, etc.). The target image's style category includes, but is not limited to, a cartoon category, a landscape category, a person category, a text category, etc. For the image recognition model, it can classify and recognize the category to which the target object in the target image belongs, and it can also classify and recognize the style of the overall target image, so that the image recognition model takes into account both the global features of the image and the local features of the target object, effectively solving the problem of reducing the limitations of image recognition, improving the recognition accuracy of the image, and thus reducing the false call rate of image recognition.

[0035] In a possible implementation, the image recognition model includes a backbone feature extraction layer, a target object detection layer, a whole-image feature classification layer and a feature fusion layer; in the classification and recognition processing, the backbone feature extraction layer is used to extract the backbone image features of the target image to obtain the backbone image features of the target image; the target object detection layer is used to perform a first classification recognition on the category to which the target object in the backbone image features belongs, and obtain a first sub-classification recognition result of the target object in the target image, and the first sub-classification recognition result indicates the probability of each category to which the target object in the target image belongs; the whole-image feature classification layer is used to perform a second classification recognition on the style to which the target image belongs according to the backbone image features, and obtain a second sub-classification recognition result of the target image, and the second sub-classification recognition result indicates the probability of each style to which the target image belongs; the feature fusion layer is used to fuse the first sub-classification recognition result and the second sub-classification recognition result to obtain a classification recognition result of the target image, and the classification recognition result indicates the target category to which the target image and its target object belong together.

[0036] Specifically, the backbone feature extraction layer is a backbone feature extraction network constructed by stacking structures such as convolution, batch normalization, activation function, and residual structure. First, the target image is preprocessed and sent to the backbone feature extraction layer for forward propagation to obtain the backbone feature map of the target image, and then the backbone feature map is sent to the target object detection layer and the whole image feature classification layer respectively. Among them, the whole image feature classification layer is a whole image feature classifier constructed by a layer of convolution, batch normalization, pooling, linear classification, and Softmax function. In the whole image feature classifier, an additional category for images that are not classified or recognized needs to be set. The backbone feature map is sent to the whole image feature classification layer for forward propagation, and a probability map containing the classification information of the style of the whole target image can be obtained (the second sub-classification result); the target object detection layer is composed of convolution, batch normalization, activation function, upsampling, concatenation, etc. Among them, three target object feature detection maps of different scales can be formed according to different upsampling ratios. After the backbone feature map is sent to the target object detection layer for forward propagation, a probability map containing the location information of the local target object of the target image and the classification information of the style of the target image is obtained (the first sub-classification result). The feature fusion layer is used to fuse the first sub-classification result and the second sub-classification result, and perform fusion operations according to the corresponding rules to obtain a probability map containing the local information of the target object and the overall characteristics of the image (classification and recognition result).

[0037] In a possible implementation, the feature fusion layer is also used to filter out target categories to which the target object in the target image belongs whose probability is greater than a set threshold from the first sub-classification recognition result, determine the target correspondence between the target categories to which the target object in the target image belongs and the styles to which the target image belongs according to the preset correspondence between the object category and the style of the image, and fuse the first target sub-classification recognition result and the second sub-classification recognition result according to the target correspondence.

[0038] Specifically, the feature fusion layer is used to preliminarily filter the probability map of the classification information of the local target object of the target image to obtain a probability map with a higher probability of containing the target object, and fuse the filtered probability map with the above-mentioned probability map containing the classification information of the entire image. Specifically, according to the correspondence between the preset object category and the second category of the style to which the image belongs, the target correspondence between the target category to which the target object belongs and the various styles to which the target image belongs is determined, and the probability map of the category to which the target image belongs and the probability map of the target object in the target image are multiplied according to the target object relationship to obtain a fused probability map.

[0039] In step S105, the target category to which the target image and its target object belong is determined according to the classification and recognition result.

[0040] Specifically, according to the probability map corresponding to the classification recognition result, after performing NMS operation on the probability map, the style category of the target image and the category of the target object in the target image are determined.

[0041] It can be seen from the technical solution provided by the above embodiments of the present application that by obtaining a target image to be identified; inputting the target image into an image recognition model for classification and identification processing, obtaining a classification and identification result of the target image output by the image recognition model, the image recognition model is used to perform a first classification and identification on the category to which the target object in the target image belongs and a second classification and identification on the style to which the target image belongs, and to fuse the first sub-classification and identification results of the first classification and identification and the second sub-classification and identification results of the second classification and identification; determining the target category to which the target image and its target object belong according to the classification and identification results, in this way, the image recognition model takes into account both the global features of the target image and the local features of the target object, reduces the limitations of image recognition, improves the recognition accuracy of the image, and reduces the false call rate of image recognition.

[0042] like Figure 2 As shown, the embodiment of the present application provides a training method for an image recognition model, and the execution subject of the method can be a server, wherein the server can be an independent server or a server cluster composed of multiple servers. The image recognition method can specifically include the following steps S201-S205:

[0043] In step S201, multiple sample images are obtained.

[0044] Specifically, the image data that needs to be identified are collected as sample images, and these sample images are constructed into a data set D. For the sample images in the data set, there are a total of N categories of target objects (target objects) that need to be identified, each object that needs to be identified is recorded as Ni, and the category of the entire sample image can be divided into C categories, each category is recorded as Ci, where i=1,2,3…n. In the embodiment of the present application, for the convenience of explanation, any sample image in the data set D is recorded as d. Assuming that there are 7 objects that need to be detected in the data set D, the categories of the sample images in the data set can be divided into 5 categories. In addition, according to actual needs, an additional category can be added to be identified, and the categories of the sample images in the data set can be divided into 6 categories in total. Among them, there are three ways to determine the categories of sample images in the data set D and the categories of objects in the pictures, namely, clustering, picture detection category voting, and expert prior classification. In order to improve the training efficiency and effect of the picture recognition model, clustering and picture detection category voting can be used for classification. These two machine automatic category selection methods are used.

[0045] In step S203, a training sample set is generated according to a plurality of sample images.

[0046] Each training sample in the training sample set is annotated with a label, which includes a first category label of the category to which the target object in the training sample belongs and a second category label of the style to which the training sample belongs, and the first category label and the second category label have a corresponding relationship.

[0047] Specifically, generating a training sample set based on multiple sample images includes: labeling the multiple sample images, clustering the labeled sample images through a clustering algorithm, obtaining a second category label of the style to which the sample images belong and a first category label of the category to which the target object in the training sample belongs; establishing a corresponding relationship between the first category label and the second category label to generate a training sample set.

[0048] Specifically, referring to the records in the above embodiment, the sample images in the data set D are first labeled, specifically, the category of the style to which the sample image belongs, the category of the target object in the sample image, the center point, the width information and the height information are used as labels. Then, a clustering algorithm is used to cluster the sample images processed by the above labeling, specifically, a k-means algorithm is used to cluster the sample images processed by the above labeling, for example, Figure 3As shown, after the sample images are clustered using the k-means algorithm, the styles of the sample images are divided into 5 categories, namely, category C1 to category C5, and the target objects in the sample images are divided into 7 categories, namely, target N1 (category N1) to target N7 (category N7). Each sample image in the data set is classified and labeled according to category C1 to category C5 and category N1 to category N7. Each sample image has its second category label, and each target object in the sample image has its first category label. A correspondence between the first category label and the second category label is established, and the second category label, the first category label and the above correspondence of each sample image are stored. Exemplary, as Figure 3 The figure shows the correspondence between the first category label and the second category label. In addition, in order to meet the subsequent requirements of additional sample image categories that need to be identified, if the image in the dataset D does not belong to the classified or identified image, an additional classification task category can be manually marked in the category of the sample image, that is, Figure 3 Category C6 shown in , thereby making the trained image recognition model more scalable and further reducing the limitations of the image recognition model.

[0049] In step S205, the training sample set is input into the image recognition model to be trained for iterative training until the loss function corresponding to the image recognition model converges, thereby obtaining a trained image recognition model.

[0050] The loss function represents the error between the predicted value and the true value of the target category to which the sample image and its target object output by the image recognition model belong, and the true value is determined based on the first category label and the second category label.

[0051] Combine the following Figure 4 The training process of the image recognition model of the embodiment of the present application is described in detail. Figure 4 As shown, after the above steps of constructing a training sample set (constructing image labels), the training sample set is input into the image recognition model to be trained. The image recognition model to be trained includes: a backbone feature extraction layer, a target object detection layer, a whole image feature classification layer, and a feature fusion layer.

[0052] Among them, the backbone feature extraction layer is a feature extraction network constructed by stacking structures such as convolution, batch normalization, activation function and residual structure. First, each training sample in the training sample set is preprocessed and sent to the backbone feature extraction layer for forward propagation to obtain the backbone feature map of the image, and then the backbone feature map is sent to the target object detection layer and the whole image feature classification layer respectively. Among them, the whole image feature classification layer is a whole image feature classifier constructed by a layer of convolution, batch normalization, pooling, linear classification (LinearLayer) and Softmax function. In the whole image feature classifier, it is necessary to set an additional category for pictures that do not belong to the classified or identified pictures. The backbone feature map is sent to the whole image feature classification layer for forward propagation, and a probability map containing the classification information of the style of the whole image can be obtained; the target object detection layer is composed of convolution, batch normalization, activation function, upsampling (Upsampling), concatenation (Concat), etc. Among them, three target object feature detection maps of different scales can be constructed according to different upsampling ratios. After the backbone feature map is sent to the target object detection layer for forward propagation, a probability map containing the location information of the local target object of the picture and the classification information of the style of the whole picture can be obtained. The feature fusion layer is used to preliminarily filter the probability map of the classification information of the local target object in the image to obtain a probability map with a higher probability of containing the target object, fuse the filtered probability map with the above probability map containing the classification information of the entire image, and perform fusion operations according to the corresponding rules to obtain a probability map containing the local information of the target object and the overall features of the image. Figure 5 As shown, after the image d is forward propagated through the whole image feature classification layer, the probability maps containing the classification information of the whole image are obtained as follows Figure 5 As shown in C1 to C6, the backbone feature map is sent to the target object detection layer for forward propagation to obtain a probability map containing the location information of the local target object of the image and the classification information of the style of the image. Figure 5 As shown in N1 to N7, the feature fusion layer Figure 5 C1 to C6 and Figure 5 After N1 to N7 are fused, the fusion probability map is as follows Figure 5 As shown in n1 to n7.

[0053] After the training of the image recognition model is completed, the image recognition model is inferred based on the image data. Specifically, the classification result of the image is used to determine whether the manually marked category C6 is hit. If it is hit, the category is recalled. In this way, since local detection cannot capture the global information of the image, such as the need to recall comic-style or landscape images, the labels of local detection are insufficient. Through this application, such requirements can be placed in the manually marked classification module C6 for recognition, further improving the recognition reliability, scalability and recognition accuracy of the image recognition model.

[0054] Through the technical solution disclosed in the embodiments of the present application, there is a corresponding relationship between the training samples in the training sample set and the target objects of the training samples. The image recognition model obtained based on the training samples takes into account both the global features of the target image and the local features of the target object, thereby reducing the limitations of image recognition, improving the recognition accuracy of the image, and reducing the false call rate of image recognition.

[0055] In one possible implementation, Figure 5 As shown, the specific steps of each iterative training of the image recognition model include: extracting the backbone image features of the sample image to obtain the backbone image features of the sample image, and determining a first probability map of the category to which the target object of the sample image belongs; determining a second probability map of the style to which the sample image belongs based on the backbone image features; fusing the first probability map and the second probability map to obtain a third probability map of the target category to which the sample image and its target object belong together; and adjusting the model parameters of the image recognition model based on the target category corresponding to the third probability map, the first category label, the second category label, and the loss function corresponding to the image recognition model.

[0056] Specifically, according to the description in the above embodiment, after the first probability map of the category to which the target object of the sample image belongs and the second probability map of the style to which the sample image belongs are fused according to the backbone image features through the feature fusion layer, a third probability map is obtained, and the third probability map obtained by the feature fusion layer is combined with the position information of the target object in the image for fusion feature detection, specifically, the category of the style to which the sample image belongs obtained through the backbone feature extraction layer, the target object detection layer, the whole image feature classification layer and the feature fusion layer and the category to which the target object in the sample image belongs are detected, and the detection result is compared with the first category label and the second category label annotated by each training sample in the training sample set pre-constructed in the above embodiment, and a loss function is constructed according to the similarity between the detection result and the first category label and the second category label of the pre-established sample image. In this way, the loss function is constructed according to the similarity between the category corresponding to the third probability map and the first category label and the second category label, which can further improve the training accuracy of the image recognition model, and when the trained image recognition model is applied to the scene of image recognition, the recognition accuracy of the image recognition model in recognizing images can be improved.

[0057] Among them, in a possible implementation, the loss function corresponding to the image recognition model includes a target object detection loss function and a whole image feature classification loss function, wherein the target object detection loss function represents the error between the predicted value of the category to which the target object in the training sample belongs and the first category label, and the whole image feature classification loss function represents the error between the predicted value of the style to which the training sample belongs and the second category label. Before inputting the training sample set into the image recognition model to be trained for iterative training, the method also includes: determining a first weight corresponding to the target object detection loss function and a second weight corresponding to the whole image feature classification loss function; and constructing a loss function according to the first weight, the target object detection loss function, the second weight, and the whole image feature classification loss function.

[0058] Specifically, the target object detection loss function and the whole image feature classification loss function can be added according to the corresponding weight coefficients to construct a loss function. Among them, the first weight and the second weight can be set according to the focus of image recognition. If the focus is on identifying the category of the style of the overall image, the second weight is set higher than the first weight. If the focus is on identifying the category of the target object in the image, the second weight is set lower than the first weight. The specific setting can be based on the actual situation, and the embodiment of the present application is not limited here.

[0059] Furthermore, the target object detection loss function includes the target detection classification loss function, the target detection foreground classification loss function and the target detection box regression loss function.

[0060] Corresponding weight parameters can be set for the whole image feature classification loss function, target detection classification loss function, target detection foreground classification loss function, and target detection frame regression loss function, so as to make full use of multiple loss functions to improve the training accuracy of the image recognition model. When the trained image recognition model is applied to the scene of image recognition, the recognition accuracy of the image recognition model can be improved. For example, the weight values ​​of the whole image feature classification loss function, the target detection classification loss function, the target detection foreground classification loss function, and the target detection frame regression loss function are set to 0.3, 0.7, 0.3, and 0.05 respectively, and the whole image feature classification loss function, the target detection classification loss function, the target detection foreground classification loss function, and the target detection frame regression loss function are multiplied by their respective weight values ​​and then added together, and the image recognition model to be trained is trained using the gradient descent algorithm.

[0061] Corresponding to the image recognition method provided in the above embodiment, based on the same technical concept, the embodiment of the present application also provides an image recognition device, Figure 6 Schematic diagram of the module composition of the image recognition device provided in the embodiment of the present application, the image recognition device is used to execute the image recognition method described in the above embodiment, such as Figure 6 As shown, the image recognition device 600 includes: an acquisition module 601, which is used to acquire a target image to be recognized; an identification module 602, which is used to input the target image into a pre-trained image recognition model for classification and recognition processing, and output a classification and recognition result of the target image, wherein the image recognition model is used to perform a first classification and recognition on the category to which the target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and to fuse the first sub-classification and recognition results of the first classification and recognition and the second sub-classification and recognition results of the second classification and recognition to obtain a classification and recognition result; a determination module 603, which is used to determine the target category to which the target image and its target object belong together according to the classification and recognition result.

[0062] Through the technical solution disclosed in the embodiments of the present application, the image recognition model takes into account both the global features of the target image and the local features of the target object, reduces the limitations of image recognition, improves the recognition accuracy of the image, and reduces the false call rate of image recognition.

[0063] In a possible implementation, the image recognition model includes: a backbone feature extraction layer, a target object detection layer, a whole-image feature classification layer and a feature fusion layer; in the classification and recognition processing, the backbone feature extraction layer is used to extract the backbone image features of the target image to obtain the backbone image features of the target image; the target object detection layer is used to perform a first classification recognition on the category to which the target object in the backbone image features belongs, and obtain a first sub-classification recognition result of the target object in the target image, and the first sub-classification recognition result indicates the probability of each category to which the target object in the target image belongs; the whole-image feature classification layer is used to perform a second classification recognition on the style to which the target image belongs according to the backbone image features, and obtain a second sub-classification recognition result of the target image, and the second sub-classification recognition result indicates the probability of each style to which the target image belongs; the feature fusion layer is used to fuse the first sub-classification recognition result and the second sub-classification recognition result to obtain a classification recognition result of the target image, and the classification recognition result indicates the target category to which the target image and its target object belong together.

[0064] In a possible implementation, the feature fusion layer is also used to filter out target categories to which the target object in the target image belongs whose probability is greater than a set threshold from the first sub-classification recognition result, determine the target correspondence between the target categories to which the target object in the target image belongs and the styles to which the target image belongs according to the preset correspondence between the object category and the style of the image, and fuse the first target sub-classification recognition result and the second sub-classification recognition result according to the target correspondence.

[0065] The image recognition device provided in the embodiment of the present application can implement each process in the embodiment corresponding to the above-mentioned image recognition method. To avoid repetition, it will not be repeated here.

[0066] It should be noted that the image recognition device provided in the embodiment of the present application and the image recognition method provided in the embodiment of the present application are based on the same inventive concept and have the same technical effect. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned image recognition method, and the repeated parts will not be repeated.

[0067] Corresponding to the training method of the image recognition model provided in the above embodiment, based on the same technical concept, the embodiment of the present application also provides a training device for the image recognition model. Figure 7 A schematic diagram of the module composition of a training device for an image recognition model provided in an embodiment of the present application, wherein the training device for an image recognition model is used to execute the training method for an image recognition model described in the above embodiment, such as Figure 7As shown, the training device 700 of the image recognition model includes: an acquisition module 701, used to acquire multiple sample images; a generation module 702, used to generate a training sample set based on the multiple sample images, wherein each training sample in the training sample set is marked with a label, and the label includes a first category label of the category to which the target object in the training sample belongs and a second category label of the style to which the training sample belongs, and the first category label and the second category label have a corresponding relationship; a training module 703, used to input the training sample set into the image recognition model to be trained for iterative training until the loss function corresponding to the image recognition model converges, and the trained image recognition model is obtained, and the loss function represents the error between the predicted value of the target category to which the sample image and its target object output by the image recognition model belong and the true value, and the true value is determined according to the first category label and the second category label.

[0068] In one possible implementation, the generation module 702 is further used to label multiple sample images, cluster the labeled sample images through a clustering algorithm, obtain a second category label of the style to which the sample image belongs and a first category label of the category to which the target object in the training sample belongs; establish a correspondence between the first category label and the second category label, and generate a training sample set.

[0069] In a possible implementation, it also includes: an extraction module, which is used to extract the backbone image features of the sample image, obtain the backbone image features of the sample image, and determine a first probability map of the category to which the target object of the sample image belongs; a determination module, which is used to determine a second probability map of the style to which the sample image belongs based on the backbone image features; a fusion module, which is used to fuse the first probability map and the second probability map to obtain a third probability map of the target category to which the sample image and its target object belong together; and a construction module, which is used to adjust the model parameters of the image recognition model based on the target category, the first category label, the second category label, and the loss function corresponding to the third probability map.

[0070] In one possible implementation, the loss function includes a target object detection loss function and a whole image feature classification loss function determination module, and is also used for the target object detection loss function to represent the error between the predicted value of the category to which the target object in the training sample belongs and the first category label, and the whole image feature classification loss function to represent the error between the predicted value of the style to which the training sample belongs and the second category label.

[0071] The training device for the image recognition model provided in the embodiment of the present application can implement each process in the embodiment corresponding to the training method for the above-mentioned image recognition model. To avoid repetition, it will not be repeated here.

[0072] It should be noted that the training device for the image recognition model provided in the embodiment of the present application and the training method for the image recognition model provided in the embodiment of the present application are based on the same inventive concept and have the same technical effect. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned image recognition model training method, and the repeated parts will not be repeated.

[0073] Corresponding to the method embodiment provided in the above embodiment, based on the same technical concept, the embodiment of the present application further provides an electronic device, which is used to execute the method embodiment. Figure 8 A schematic diagram of the structure of an electronic device for implementing various embodiments of the present invention is shown in FIG. Figure 8 As shown. The electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors 801 and memory 802, and the memory 802 may store one or more storage applications or data. Among them, the memory 802 can be a temporary storage or a permanent storage. The application stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the electronic device. Furthermore, the processor 801 can be configured to communicate with the memory 802 to execute a series of computer executable instructions in the memory 802 on the electronic device. The electronic device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input and output interfaces 805, and one or more keyboards 806.

[0074] In this embodiment, the electronic device includes a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory communicate with each other through the bus; the memory is used to store computer programs; the processor is used to execute the programs stored in the memory to implement the steps described in the above method embodiment.

[0075] It should be noted that the electronic device provided in the embodiment of the present application and the method embodiment provided in the embodiment of the present application are based on the same inventive concept and have the same technical effect. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned method embodiment, and the repeated parts will not be repeated.

[0076] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps described in the above method embodiment are implemented.

[0077] It should be noted that the computer-readable storage medium provided in the embodiment of the present application and the method provided in the above-mentioned method embodiment are based on the same inventive concept and have the same technical effect. Therefore, the specific implementation of this embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0078] In a specific embodiment, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps described in the above method embodiment.

[0079] It should be noted that the chip provided in the embodiment of the present application and the method embodiment provided in the embodiment of the present application are based on the same inventive concept and have the same technical effect. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned method embodiment, and the repeated parts will not be repeated.

[0080] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0081] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0082] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0084] In a typical configuration, an electronic device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0085] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0086] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0087] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0088] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0089] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for image recognition, It is characterized in that The image recognition method comprises: Get the target image to be identified; Input the target image into a pre-trained image recognition model for classification and recognition processing, and output the classification and recognition result of the target image, wherein the image recognition model is used to perform a first classification and recognition on the category to which the target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition are fused to obtain the classification and recognition result; Determine the target category to which the target image and its target object belong together according to the classification and recognition result; The step of fusing the first sub-classification recognition result of the first classification recognition and the second sub-classification recognition result of the second classification recognition includes: According to the preset correspondence between the category to which the object belongs and the style to which the picture belongs, determining the target correspondence between each target category to which the target object belongs and each style to which the target picture belongs; each target category is a category whose probability in the first sub-classification recognition result is greater than a set threshold; The first sub-classification recognition result and the second sub-classification recognition result are fused according to the target corresponding relationship.

2. The image recognition method according to claim 1, It is characterized in that The image recognition model includes: a backbone feature extraction layer, a target object detection layer, a whole image feature classification layer and a feature fusion layer; In the classification and recognition process, the backbone feature extraction layer is used to extract the backbone image features of the target image to obtain the backbone image features of the target image; The target object detection layer is used to perform a first classification recognition on the category to which the target object in the trunk image features belongs, and obtain a first sub-classification recognition result of the target object in the target image, wherein the first sub-classification recognition result indicates the probability of each category to which the target object in the target image belongs; The whole-image feature classification layer is used to perform a second classification recognition on the style to which the target image belongs according to the trunk image feature, and obtain a second sub-classification recognition result of the target image, wherein the second sub-classification recognition result indicates the probability of each style to which the target image belongs; The feature fusion layer is used to fuse the first sub-classification recognition result and the second sub-classification recognition result to obtain a classification recognition result of the target image, and the classification recognition result indicates a target category to which the target image and its target object belong.

3. A training method for an image recognition model, It is characterized in that include: Get multiple sample images; Generate a training sample set according to the multiple sample images, wherein each training sample in the training sample set is annotated with a label, the label including a first category label of a category to which a target object in the training sample belongs and a second category label of a style to which the training sample belongs, and the first category label and the second category label have a corresponding relationship; Inputting the training sample set into the image recognition model to be trained for iterative training until the loss function corresponding to the image recognition model converges, thereby obtaining a trained image recognition model, wherein the loss function represents the error between the predicted value of the target category to which the sample image and its target object belong together output by the image recognition model and the true value, wherein the true value is determined according to the first category label and the second category label; Among them, the predicted value is determined by the image recognition model according to the preset correspondence between the object category and the style of the image, determining the target correspondence between the target categories to which the target object belongs indicated by the first sub-classification recognition result of the sample image and the styles to which the sample image belongs indicated by the second sub-classification recognition result, and fusing the first sub-classification recognition result and the second sub-classification recognition result of the sample image according to the target correspondence; the target categories are categories in the first sub-classification recognition result whose probability is greater than a set threshold.

4. The method for training the image recognition model according to claim 3, It is characterized in that Generating a training sample set according to the plurality of sample images comprises: The plurality of sample images are labeled, and the labeled sample images are clustered by a clustering algorithm to obtain a second category label of the style to which the sample images belong and a first category label of the category to which the target object in the training sample belongs; A corresponding relationship between the first category label and the second category label is established to generate the training sample set.

5. The method for training the image recognition model according to claim 3, It is characterized in that The specific steps of each iterative training of the image recognition model include: Extracting the main image features of the sample image to obtain the main image features of the sample image, and determining a first probability map of the category to which the target object of the sample image belongs; Determine a second probability map of the style to which the sample image belongs according to the trunk image feature; The first probability map and the second probability map are fused to obtain a third probability map of the target category to which the sample image and its target object belong; Adjust the model parameters of the image recognition model according to the target category corresponding to the third probability map, the first category label, the second category label, and the loss function.

6. The method for training an image recognition model according to claim 3, It is characterized in that The loss function includes a target object detection loss function and a whole image feature classification loss function; wherein the target object detection loss function represents the error between the predicted value of the category to which the target object in the training sample belongs and the first category label, and the whole image feature classification loss function represents the error between the predicted value of the style to which the training sample belongs and the second category label.

7. A picture recognition device, It is characterized in that The picture recognition includes: An acquisition module is used to acquire the target image to be identified; A recognition module, used for inputting the target image into a pre-trained image recognition model for classification and recognition processing, and outputting the classification and recognition result of the target image, wherein the image recognition model is used for performing a first classification and recognition on the category to which the target object in the target image belongs and a second classification and recognition on the style to which the target image belongs, and for fusing a first sub-classification and recognition result of the first classification and recognition and a second sub-classification and recognition result of the second classification and recognition to obtain the classification and recognition result; A determination module, used to determine the target category to which the target image and its target object belong according to the classification and recognition result; The recognition module is specifically used to determine the target correspondence between each target category to which the target object belongs and each style to which the target picture belongs according to the preset correspondence between the category to which the object belongs and the style to which the picture belongs; each target category is a category whose probability in the first sub-classification recognition result is greater than a set threshold; The first sub-classification recognition result and the second sub-classification recognition result are fused according to the target corresponding relationship.

8. An electronic device comprising a processor, a communication interface, a memory and a communication bus; in, The processor, the communication interface and the memory communicate with each other via a bus; the memory is used to store computer programs; the processor is used to execute the programs stored in the memory to implement the image recognition method as described in any one of claims 1-2 or the training method steps of the image recognition model as described in any one of claims 3-6.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the training method of the image recognition method described in any one of claims 1 to 2 or the image recognition model described in any one of claims 3 to 6 are implemented.

Citation Information

Patent Citations

  • Article recognition method and device, vending system and storage medium

    CN109754009A

  • Object detection method and device, model training method and device and electronic equipment

    CN113326796A