A method for image recognition based on artificial intelligence and related devices

The attention-guided image enhancement and multi-level tagging method enhances image recognition accuracy by focusing on key regions and refining labels, addressing the inefficiencies and inaccuracies of existing systems.

CN113569889BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110083832.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-21
Publication Date
2025-07-15
Estimated Expiration
2041-01-21

AI Technical Summary

Technical Problem

In the prior art, vulgar image recognition relies on massive labeled data, which is time-consuming and labor-intensive, and is prone to missed details, affecting the accuracy of image recognition.

Method used

The attention-guided method is adopted to obtain the attention map of the input image, image adjustment and enhancement are performed, and the target recognition network is trained, combined with the hierarchical recognition network, and the first and second types of tags are obtained to improve the model's data learning and detailed identification of key parts.

Benefits of technology

The accuracy of image recognition is improved, especially in vulgar image recognition scenes, and the details in the image can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569889B_ABST
    Figure CN113569889B_ABST
Patent Text Reader

Abstract

The present application discloses a method for image recognition based on artificial intelligence and related devices, which is applied to the computer vision technology. By obtaining an input image; then inputting the input image into a target recognition network obtained by attention guidance in a target model to obtain an attention map and an image feature map; and further inputting the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label. Thus, an image hierarchical recognition process guided by an attention area is realized. Since the attention area in the attention map is used to enhance the acquisition of the image, the model focuses on the data learning of the key parts, and the hierarchical recognition method is used for display, so that the details in the image can be recognized, and the accuracy of image recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for image recognition based on artificial intelligence and related devices. Background Art

[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence.

[0003] Vulgar picture recognition is an important application direction of artificial intelligence. In the solutions provided by related technologies, an end-to-end model is usually constructed relying on a large amount of labeled data, that is, all processing of the picture to be recognized is completed by an end-to-end model, and the output is a result of vulgar or non-vulgar.

[0004] However, the process of obtaining a large amount of labeled data is time-consuming and laborious, and details of vulgar pictures may be omitted, affecting the accuracy of image recognition. Summary of the Invention

[0005] In view of this, this application provides a method for image recognition based on artificial intelligence, which can effectively improve the accuracy of image recognition.

[0006] The first aspect of this application provides a method for image recognition based on artificial intelligence, which can be applied to a system or program with an image recognition function in a terminal device, and specifically includes:

[0007] Obtain an input image;

[0008] Input the input image into a preset recognition network in a target model to obtain an attention map, where the attention map contains attention regions,

[0009] Perform image adjustment on the attention map based on the attention regions to obtain an enhanced image, and train the preset recognition network according to the enhanced image to obtain a target recognition network;

[0010] Input the input image into the target recognition network to obtain an image feature map;

[0011] Input the image feature map into the hierarchical recognition network in the target model to obtain a first type of label and a second type of label. The hierarchical recognition network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than that of the first type of label for the target object.

[0012] Optionally, in some possible implementation manners of the present application, the method of performing image adjustment on the attention map based on the attention region to obtain an enhanced image, and training the preset recognition network according to the enhanced image to obtain a target recognition network includes:

[0013] Mask the attention region to update the attention map to obtain a first adjusted image, and adjust the label corresponding to the first adjusted image;

[0014] Strengthen the weight parameter corresponding to the attention region to update the attention map to obtain a second adjusted image, and keep the label corresponding to the second adjusted image unchanged;

[0015] Train the preset recognition network according to the first adjusted image and the second adjusted image to obtain the target recognition network.

[0016] Optionally, in some possible implementation manners of the present application, the method of training the preset recognition network according to the first adjusted image and the second adjusted image to obtain the target recognition network includes:

[0017] Perform region perturbation based on the first adjusted image to generate a negative sample sequence;

[0018] Perform weight parameter perturbation based on the second adjusted image to generate a positive sample sequence;

[0019] Train the preset recognition network according to the negative sample sequence and the positive sample sequence to obtain the target recognition network.

[0020] Optionally, in some possible implementation manners of the present application, the method further includes:

[0021] Determine the attention first-level label and the attention second-level label corresponding to the attention region;

[0022] Perform constraint based on the region corresponding to the attention first-level label and the region corresponding to the attention second-level label to obtain attention loss information;

[0023] Adjust the parameters of the target recognition network according to the attention loss information.

[0024] Optionally, in some possible implementation manners of the present application, the method further includes:

[0025] Obtain first-level label training data;

[0026] Determine the classification loss in the first-level label training data to train the first-level label branch;

[0027] Obtain second-level label training data;

[0028] Input the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples;

[0029] Train the second-level label branch based on the second-level label positive samples and second-level label negative samples.

[0030] Optionally, in some possible implementation manners of the present application, the inputting the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples includes:

[0031] Determine the target samples in the second-level label training data;

[0032] Calculate the moving average based on the batch data corresponding to the target samples to obtain dynamic threshold information, where the dynamic threshold information includes a positive sample threshold and a negative sample threshold;

[0033] Input the target samples into the binary classifier to obtain prediction values;

[0034] Compare the prediction values with the dynamic threshold information to determine the second-level label positive samples and second-level label negative samples in the second-level label training data.

[0035] Optionally, in some possible implementation manners of the present application, the comparing the prediction values with the dynamic threshold information to determine the second-level label positive samples and second-level label negative samples in the second-level label training data includes:

[0036] Compare the prediction values with the positive sample threshold in the dynamic threshold information;

[0037] If the prediction value is greater than the positive sample threshold, determine that the target sample is the second-level label positive sample;

[0038] Compare the prediction values with the negative sample threshold in the dynamic threshold information;

[0039] If the predicted value is less than the negative sample threshold, determine that the target sample is a negative sample of the secondary label.

[0040] Optionally, in some possible implementation manners of the present application, the method further includes:

[0041] If the predicted value is greater than the negative sample threshold and less than the positive sample threshold, determine that the target sample is a noise sample;

[0042] Set the noise sample not to participate in the training of the secondary label branch.

[0043] Optionally, in some possible implementation manners of the present application, the obtaining the input image includes:

[0044] Obtain an instant media data stream;

[0045] Extract the images in the media data stream according to a target time sequence to obtain the input image, and the input image is published according to the target time sequence after being recognized.

[0046] Optionally, in some possible implementation manners of the present application, the method further includes:

[0047] Extract first key information in the first type of label;

[0048] Extract second key information in the second type of label;

[0049] Associate the first key information and the second key information to obtain description information of the input image;

[0050] Mark the input image based on the description information.

[0051] Optionally, in some possible implementation manners of the present application, the method further includes:

[0052] Respond to a target operation to trigger a call process of the input image;

[0053] Cache the input image based on the call process and recognize the label of the input image;

[0054] If the label of the input image meets a preset condition, display the input image.

[0055] Optionally, in some possible implementation manners of the present application, the target model is used for identifying vulgar images, the first type of label is used to indicate the individual type of the target object, and the second type of label is used to indicate the part type of the target object.

[0056] The second aspect of the present application provides an image recognition device, including:

[0057] An acquisition unit for acquiring an input image;

[0058] An input unit for inputting the input image into a preset recognition network in a target model to obtain an attention map, where the attention map contains attention regions,

[0059] An adjustment unit for performing image adjustment on the attention map based on the attention regions to obtain an enhanced image, and training the preset recognition network according to the enhanced image to obtain a target recognition network;

[0060] The input unit is further configured to input the input image into the target recognition network to obtain an image feature map;

[0061] A recognition unit for inputting the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch, the first-level label branch is used to determine the first type of label of the input image, the second-level label branch is used to recognize the second type of label of the input image, the first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object.

[0062] Optionally, in some possible implementation manners of the present application, the adjustment unit is specifically configured to mask the attention regions to update the attention map to obtain a first adjusted image, and adjust the label corresponding to the first adjusted image;

[0063] The adjustment unit is specifically configured to strengthen the weight parameters corresponding to the attention regions to update the attention map to obtain a second adjusted image, and keep the label corresponding to the second adjusted image unchanged;

[0064] The adjustment unit is specifically configured to train the preset recognition network according to the first adjusted image and the second adjusted image to obtain the target recognition network.

[0065] Optionally, in some possible implementation manners of the present application, the adjustment unit is specifically configured to perform regional perturbation on the first adjusted image to generate a negative sample sequence;

[0066] The adjustment unit is specifically configured to perform weight parameter perturbation on the second adjusted image to generate a positive sample sequence;

[0067] The adjustment unit is specifically configured to train the preset recognition network according to the negative sample sequence and the positive sample sequence to obtain the target recognition network.

[0068] Optionally, in some possible implementation manners of the present application, the adjustment unit is specifically configured to determine an attention first-level label and an attention second-level label corresponding to the attention area;

[0069] The adjustment unit is specifically configured to perform constraints based on the area corresponding to the attention first-level label and the area corresponding to the attention second-level label to obtain attention loss information;

[0070] The adjustment unit is specifically configured to adjust the parameters of the target recognition network according to the attention loss information.

[0071] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to obtain first-level label training data;

[0072] The recognition unit is specifically configured to determine the classification loss in the first-level label training data to train the first-level label branch;

[0073] The recognition unit is specifically configured to obtain second-level label training data;

[0074] The recognition unit is specifically configured to input the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples;

[0075] The recognition unit is specifically configured to train the second-level label branch based on the second-level label positive samples and the second-level label negative samples.

[0076] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to determine target samples in the second-level label training data;

[0077] The recognition unit is specifically configured to perform a sliding mean calculation based on the batch data corresponding to the target samples to obtain dynamic threshold information, and the dynamic threshold information includes a positive sample threshold and a negative sample threshold;

[0078] The recognition unit is specifically configured to input the target samples into the binary classifier to obtain prediction values;

[0079] The recognition unit is specifically configured to compare the prediction values with the dynamic threshold information to determine the second-level label positive samples and the second-level label negative samples in the second-level label training data.

[0080] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to compare the predicted value with the positive sample threshold in the dynamic threshold information;

[0081] If the predicted value is greater than the positive sample threshold, it is determined that the target sample is a positive sample of the secondary label;

[0082] The recognition unit is specifically configured to compare the predicted value with the negative sample threshold in the dynamic threshold information;

[0083] The recognition unit is specifically configured to, if the predicted value is less than the negative sample threshold, determine that the target sample is a negative sample of the secondary label.

[0084] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to, if the predicted value is greater than the negative sample threshold and less than the positive sample threshold, determine that the target sample is a noise sample;

[0085] The recognition unit is specifically configured to set the noise sample not to participate in the training of the secondary label branch.

[0086] Optionally, in some possible implementation manners of the present application, the acquisition unit is specifically configured to acquire an instant media data stream;

[0087] The acquisition unit is specifically configured to extract images in the media data stream according to a target time sequence to obtain the input image, and the input image is published according to the target time sequence after being recognized.

[0088] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to extract first key information in the first type of label;

[0089] The recognition unit is specifically configured to extract second key information in the second type of label;

[0090] The recognition unit is specifically configured to associate the first key information and the second key information to obtain description information of the input image;

[0091] The recognition unit is specifically configured to mark the input image based on the description information.

[0092] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to trigger a calling process of the input image in response to a target operation;

[0093] The recognition unit is specifically configured to cache the input image based on the calling process and recognize the mark of the input image;

[0094] The recognition unit is specifically configured to display the input image if the label of the input image meets a preset condition.

[0095] A third aspect of the present application provides a computer device, including: a memory, a processor, and a bus system; the memory is used to store program codes; the processor is used to execute the image recognition method described in the first aspect or any one of the first aspects according to the instructions in the program codes.

[0096] A fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is caused to execute the image recognition method described in the first aspect or any one of the first aspects.

[0097] According to an aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image recognition method provided in the first aspect or various optional implementation manners of the first aspect.

[0098] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0099] By obtaining an input image; then inputting the input image into a preset recognition network in a target model to obtain an attention map, the attention map includes an attention area, and based on the attention area, the attention map is adjusted to obtain an enhanced image, and the preset recognition network is trained according to the enhanced image to obtain a target recognition network; further, the input image is input into the target recognition network to obtain an image feature map; and then the image feature map is input into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch, the first-level label branch is used to determine the first type of label of the input image, the second-level label branch is used to identify the second type of label of the input image, the first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object. Thus, an image hierarchical recognition process guided by the attention area is realized. Since the attention area in the attention map is used to obtain the enhanced image, the model focuses on the data learning of the key parts, and the hierarchical recognition method is used for display, so that the details in the image can be recognized, and the accuracy of image recognition is improved. Description of the Drawings

[0100] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0101] Figure 1 It is a network architecture diagram for the operation of an image recognition system;

[0102] Figure 2 It is a process architecture diagram for image recognition provided by an embodiment of the present application;

[0103] Figure 3 It is a flowchart of a method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0104] Figure 4 It is a scenario schematic diagram of a method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0105] Figure 5 It is a scenario schematic diagram of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0106] Figure 6 It is a scenario schematic diagram of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0107] Figure 7 It is a scenario schematic diagram of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0108] Figure 8 It is a scenario schematic diagram of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0109] Figure 9 It is a flowchart of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0110] Figure 10 It is a flowchart of another method for image recognition based on artificial intelligence provided by an embodiment of the present application;

[0111] Figure 11 It is a structural schematic diagram of an image recognition device provided by an embodiment of the present application;

[0112] Figure 12 It is a structural schematic diagram of a terminal device provided by an embodiment of the present application;

[0113] Figure 13A structural schematic diagram of a server provided by an embodiment of the present application. Specific embodiments

[0114] An embodiment of the present application provides a method for image recognition based on artificial intelligence and related devices, which can be applied to a system or program with an image recognition function in a terminal device. By obtaining an input image; then inputting the input image into a preset recognition network in a target model to obtain an attention map, the attention map includes attention regions, and based on the attention regions, the attention map is adjusted to obtain an enhanced image, and the preset recognition network is trained according to the enhanced image to obtain a target recognition network; further, the input image is input into the target recognition network to obtain an image feature map; and then the image feature map is input into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch, the first-level label branch is used to determine the first type of label of the input image, the second-level label branch is used to identify the second type of label of the input image, the first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object. Thus, an image hierarchical recognition process guided by attention regions is realized. Since the attention regions in the attention map are used to obtain the enhanced image, the model focuses on the data learning of key parts, and a hierarchical recognition method is used for display, so that the detailed parts in the image can be recognized, and the accuracy of image recognition is improved.

[0115] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0116] First, some nouns that may appear in the embodiments of the present application are explained.

[0117] Attention mechanism: It means that in the process of image classification by a classification model, different regions in the image will have different weight preferences, which can be represented by a heat map. This attention mechanism is learned by the classification model through a large amount of data.

[0118] Fine-grained recognition: It refers to a subfield in image classification. The image classification problem is to divide images into different large categories as required, such as lions, dogs, airplanes, etc. Fine-grained recognition, however, is to make further distinctions within a category. For example, face recognition is a special fine-grained recognition problem that requires finding the face of the person you want from a vast number of faces. Currently, the mainstream fine-grained recognition dataset is CUB-200, which is used to identify different categories from birds.

[0119] Data noise: It refers to the situation where, during the labeling process by annotators, the wrong label is given because it is impossible to determine which category it belongs to; or some label information in the image is missed, resulting in a deterioration of the model training effect.

[0120] It should be understood that the image recognition method provided in this application can be applied to systems or programs with image recognition functions in terminal devices. For example, a vulgar image detection tool. Specifically, the image recognition system can run in a network architecture as Figure 1 shown. As Figure 1 shown, it is the network architecture diagram for the operation of the image recognition system. As can be seen from the figure, the image recognition system can provide the process of image recognition with multiple information sources, that is, obtain multimedia data through a triggering operation on the terminal side, and then perform the recognition of the multimedia data on the terminal side or the server side to obtain the vulgar images therein and perform processing; it can be understood that Figure 1 shows various terminal devices. The terminal device can be a computer device. In actual scenarios, more or fewer types of terminal devices may participate in the image recognition process. The specific quantity and types depend on the actual scenario and are not limited here. Additionally, Figure 1 shows a server, but in actual scenarios, multiple servers may also participate. The specific number of servers depends on the actual scenario.

[0121] In this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods. The terminal and the server can be connected to form a blockchain network, which is not limited in this application.

[0122] It can be understood that the above image recognition system can run on a personal mobile terminal. For example, as an application such as a vulgar image detection tool, it can also run on a server, or can run on a third-party device to provide image recognition to obtain the processing results of image recognition of the information source. The specific image recognition system can run in the above device in the form of a program, or can run as a system component in the above device, or can also be a type of cloud service program. The specific operation mode depends on the actual scenario and is not limited here.

[0123] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence.

[0124] Computer Vision (CV) is included in artificial intelligence technology. Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, trace tracing, and measurement on targets, and further perform graphic processing to make the computer process into an image that is more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0125] Among them, vulgar picture recognition is an important application direction of computer vision technology. In the solutions provided by related technologies, it usually relies on a large amount of labeled data to build an end-to-end model, that is, all processing of the picture to be recognized is completed by an end-to-end model, and the output is a result of vulgar or not vulgar.

[0126] However, the process of obtaining a large amount of labeled data is time-consuming and laborious, and details of vulgar pictures may be omitted, affecting the accuracy of image recognition.

[0127] To solve the above problems, this application proposes a method for image recognition based on artificial intelligence. This method is applied to Figure 2 the shown image recognition process framework, asFigure 2 As shown in the figure, it is a process architecture diagram for image recognition provided by an embodiment of the present application. The user obtains multimedia data through interactive operations on the interface layer, and converts this multimedia data into image input on the server side for image recognition, so as to obtain hierarchical labels corresponding to the multimedia data, and perform vulgarity determination to judge whether to publish or display on the terminal.

[0128] The image recognition process of the present application adopts an attention-guided method to help the model improve the attention to fine discriminative regions in the image, ensuring the accuracy of the model in the image vulgarity recognition scenario. In the process of determining the hierarchical labels, deep learning trains through a large amount of data, enabling the model to have a rough attention area and weak localization information for the image. However, the learning of attention in the classification model during training is essentially passive model parameter learning. The present invention introduces attention guidance to actively assist the classification model in learning the attention area, helping the model achieve better results on data in small violation areas. Moreover, for the noisy labels in the data, the present invention uses an adaptive threshold method with different categories to judge the data reliability. Through the model judgment results, selective learning of the data is performed.

[0129] It can be understood that the method provided by the present application can be a program written as a processing logic in a hardware system, or can be an image recognition device that implements the above processing logic in an integrated or external connection manner. As an implementation method, the image recognition device obtains an input image; then inputs the input image into a preset recognition network in the target model to obtain an attention map, which contains an attention area, and performs image adjustment on the attention map based on the attention area to obtain an enhanced image, and trains the preset recognition network according to the enhanced image to obtain a target recognition network; further inputs the input image into the target recognition network to obtain an image feature map; and then inputs the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than that of the first type of label for the target object. Thus, the image hierarchical recognition process guided by the attention area is realized. Since the attention area in the attention map is used to obtain the enhanced image, the model focuses on the data learning of the key parts, and uses the hierarchical recognition method for display, enabling the details in the image to be recognized and improving the accuracy of image recognition.

[0130] The solution provided in the embodiments of this application relates to the computer vision technology of artificial intelligence, and is specifically illustrated through the following embodiments:

[0131] Combined with the above process architecture, the method for image recognition in this application will be introduced below. Please refer to Figure 3 , Figure 3 is a flowchart of a method for image recognition based on artificial intelligence provided in the embodiments of this application. This management method can be executed by a terminal device, or by a server, or jointly executed by a terminal device and a server. The embodiments of this application at least include the following steps:

[0132] 301. Obtain the input image.

[0133] In this embodiment, the input image can be an image for target model training, and this image has corresponding feature labels, so as to facilitate the learning of the target model's image feature extraction ability.

[0134] In a possible scenario, when the target model is a trained model, the input image can be an image in the instant data stream, such as the new data corresponding to the refresh of the Moments, so as to identify vulgar information for the new data to guide the process of online release.

[0135] Specifically, this application takes the target model for identifying vulgar images as an example for illustration, and the first type of label obtained in the subsequent identification is used to indicate the individual type of the target object, and the second type of label is used to indicate the part type of the target object, that is, the result of hierarchical identification. The specific identification scenario depends on the actual situation.

[0136] It can be understood that in the scenario of identifying vulgar images, vulgar identification is different from pornographic identification. Pornographic data is data that triggers the red line such as genital exposure. Vulgar data is data with revealing and sexy clothing. Pornographic data is more distinguishable in terms of problems, while vulgar data is often easily confused with normal data and is more difficult. Therefore, this application adopts a combination of attention guidance and hierarchical identification.

[0137] Specifically, further label division is carried out on vulgar data, including: sexual suggestion, child exposure, animal exposure, art exposure, female sexuality (first-level label) - chest (second-level label), female sexuality - legs, female sexuality - hips, female sexuality - figure, male sexuality, acg sexuality, etc. The specific label types depend on the actual scenario and are not limited here.

[0138] 302. Input the input image into the preset recognition network in the target model to obtain an attention map.

[0139] In this embodiment, the attention map contains attention regions. Herein, the attention map is also referred to as the attention heat map, and different-weight image features are shown by the depth of color in the map. For example, the weight set for the vulgar image region is large, and the corresponding color is deep. The corresponding attention region, that is, a part of the input image, can be obtained through the color distribution (weight distribution) in the attention map.

[0140] Specifically, for the architecture of the target model in this application, it can be as Figure 4 shown. Figure 4 It is a schematic diagram of the scenario of another method for image recognition based on artificial intelligence provided by the embodiment of this application; that is, after the input image is input into the preset recognition network, an attention map is obtained, and then the image is adjusted based on the attention map to obtain an enhanced image, so as to further train the preset recognition network; in addition, for the part of label recognition, the feature map of the image is extracted based on the trained target recognition network, and then feature fusion is performed and input into different task branches for recognition.

[0141] It can be understood that the process of multi-task hierarchical recognition is the recognition process of different granularities for the same object. As Figure 5 shown. Figure 5 It is a schematic diagram of the scenario of another method for image recognition based on artificial intelligence provided by the embodiment of this application; the recognized object A1 corresponding to the first-level label and the recognized object A2 corresponding to the second-level label are shown in the figure. It can be seen that the recognized object A2 corresponding to the second-level label is a part of the recognized object A1 corresponding to the first-level label. For example, if the recognized object A1 corresponding to the first-level label is the female body, then the recognized object A2 corresponding to the second-level label is the buttocks, so as to realize the process of hierarchical recognition and judgment of vulgar scenes.

[0142] 303. Perform image adjustment on the attention map based on the attention region to obtain an enhanced image, and train the preset recognition network according to the enhanced image to obtain a target recognition network.

[0143] In this embodiment, the process of recognizing the preset recognition network based on the enhanced image is the process of guiding attention; specifically, attention guidance is a data augmentation technique, that is, the process as Figure 6 shown. Figure 6 It is a schematic diagram of the scenario of another method for image recognition based on artificial intelligence provided by the embodiment of this application; that is, through the attention region learned by the current model, further data augmentation is performed on the original image (such as masking the attention region and strengthening the attention region), and the enhanced image is further learned (when masking the attention region, the label will become normal; when strengthening the attention region, the label remains unchanged). In this way, it can actively help the model learn the distinguishable regions it needs to focus on.

[0144] Specifically, in the process of image adjustment, the attention area can be masked first to update the attention map to obtain a first adjusted image, and the label corresponding to the first adjusted image is adjusted; then the weight parameter corresponding to the attention area is strengthened to update the attention map to obtain a second adjusted image, and the label corresponding to the second adjusted image remains unchanged; furthermore, the preset recognition network is trained according to the first adjusted image and the second adjusted image to obtain a target recognition network. For example, in the scenario shown in Figure 7, Figure 7 This is a schematic diagram of a scenario of another method for image recognition based on artificial intelligence provided by an embodiment of this application; the attention area of the original image is shown in the figure, and then the attention area is respectively covered and enhanced, and its label is adjusted accordingly, so as to improve the recognition ability of the target model for the attention area.

[0145] Optionally, in order to further improve the recognition ability of the target model for the attention area, perturbations can also be performed on the basis of the enhanced image to expand the data volume; specifically, first, regional perturbations are performed on the first adjusted image to generate a negative sample sequence, that is, an image that does not contain vulgar areas; then, weight parameter perturbations are performed on the second adjusted image to generate a positive sample sequence, that is, an image that contains vulgar areas; furthermore, the preset recognition network is trained according to the negative sample sequence and the positive sample sequence to obtain a target recognition network.

[0146] Optionally, during the training process based on the enhanced image, the attention area needs to be constrained. That is, first determine the attention primary label and attention secondary label corresponding to the attention area; then, based on the area corresponding to the attention primary label and the area corresponding to the attention secondary label, perform constraints to obtain attention loss information; furthermore, adjust the parameters of the target recognition network according to the attention loss information, thereby improving the training effect of the target recognition network.

[0147] Specifically, for the process of constraining the attention area, that is, the attention areas of the primary and secondary labels should be as consistent as possible. Through this constraint, the learning effect of the attention area is strengthened. Specifically, it can be carried out with reference to the following formula:

[0148]

[0149] where (x, y) is any point on the attention map, represents the value at (x, y) in the i-th channel of the attention map. 1(condition) means that when the condition is true, the output is 1, and when the condition is false, the output is 0.

[0150] 304. Input the input image into the target recognition network to obtain an image feature map.

[0151] In this embodiment, after completing the attention-guided training for the preset recognition network to obtain the target recognition network, the image features can be extracted based on the target recognition network.

[0152] Specifically, the target recognition network can be a network of the Resnet series, such as the Resnet18 recognition network. The specific network type depends on the actual scenario and is not limited here.

[0153] 305. Input the image feature map into the hierarchical recognition network in the target model to obtain the first type of label and the second type of label.

[0154] In this embodiment, as Figure 4 shown in the architecture, the hierarchical recognition network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than that of the first type of label for the target object, so as to implement the hierarchical detailed recognition process.

[0155] Optionally, for the training processes of the first-level label branch and the second-level label branch, they can be carried out separately, that is, obtain the first-level label training data; then determine the classification loss in the first-level label training data to train the first-level label branch; and obtain the second-level label training data; then input the second-level label training data into a binary classifier to obtain the second-level label positive samples and the second-level label negative samples; and then train the second-level label branch based on the second-level label positive samples and the second-level label negative samples.

[0156] It should be noted that during the training process of the second-level label branch, since the second-level labels of a piece of vulgar data are often not fully labeled. The annotators often only focus on the second-level labels that they are most concerned about in the image and ignore some other second-level labels that also exist at the same time. For example, if there is both chest sensuality and leg sensuality in the picture. And if the annotator only labels one chest sensuality, according to normal deep learning, when the model learns this data, the chest label is positive, while the leg label is negative, which will confuse the model's learning of the leg label. The samples that confuse the model's learning can be called noise samples.

[0157] To avoid noise samples from participating in the training process of the secondary label branch, a dynamic threshold judgment process can be adopted. That is, first determine the target samples in the secondary label training data; then calculate the moving average based on the batch data corresponding to the target samples to obtain dynamic threshold information, which includes the positive sample threshold and the negative sample threshold; further input the target samples into the binary classifier of logistic regression to obtain prediction values; and compare the prediction values with the dynamic threshold information to determine the secondary label positive samples and secondary label negative samples in the secondary label training data, thus ensuring the accuracy of sample labeling.

[0158] Specifically, the determination process of the secondary label positive samples and secondary label negative samples is obtained by comparing the prediction values with the dynamic threshold. For example, compare the prediction value with the positive sample threshold in the dynamic threshold information; if the prediction value is greater than the positive sample threshold, determine the target sample as a secondary label positive sample; or compare the prediction value with the negative sample threshold in the dynamic threshold information; if the prediction value is less than the negative sample threshold, determine the target sample as a secondary label negative sample.

[0159] It can be understood that for the determination of noise samples, that is, when the prediction value is greater than the negative sample threshold and less than the positive sample threshold, then determine the target sample as a noise sample, which can also be called an ignored sample; then set the noise sample not to participate in the training of the secondary label branch. Specifically, as Figure 8 shown Figure 8 is a schematic diagram of the scenario of another method for image recognition based on artificial intelligence provided by an embodiment of the present application; for the training data input into the secondary label branch, the dynamic threshold is updated based on each sample divided into batch data, so that secondary label positive samples, secondary label negative samples, and secondary label ignored samples (noise samples) can be obtained, and the noise samples are ignored, that is, they do not participate in the calculation of the loss function, thereby improving the accuracy of the secondary label branch recognition.

[0160] In a possible scenario, category thresholds are set for the positive and negative parts in the loss function. The positive threshold of each secondary label is initialized to 1, and the negative threshold is initialized to 0. By the scores of the secondary label predictions corresponding to different samples in each batch training, the positive and negative thresholds of the corresponding secondary labels are adjusted by the method of moving average. In addition, according to the comparison between the model output result and the positive and negative thresholds of the corresponding secondary labels, mislabeled samples and wrongly labeled samples are distinguished. For example, the positive threshold of the female chest label is 0.7, and the female chest label predicted by the sample model is 0.9, then this sample model considers it as a correct sample and participates in the model training; if the chest label predicted by another sample model is 0.3, and the true annotation of this sample is female chest, then the model considers this sample as a wrongly labeled sample (noise sample) and does not participate in the model training.

[0161] Specifically, the loss function can refer to the following formula:

[0162] L(x, y) = 1(p(x) > θ p ) y log(p(x)) - 1(p(x) ≤ θ n ) * (1 - y) log(1 - p(x)),

[0163] θ p = min(μ P + α * σ P , 1),

[0164] θ n = max(μ n - α * σ n , 0),

[0165] where, for the secondary labels of a certain category, p(x) represents the corresponding model prediction result, and θ p represents the positive threshold of this category, and θ n represents the negative threshold of this category, and y represents the true label (0 or 1) of the sample for this category. α is the threshold iteration rate, which is 0.1 in this formula, and σ P is the model prediction score for the corresponding positive samples, and σ n is the model prediction score for the negative samples. The update of the positive / negative thresholds is performed through the sample scores after model screening for each batch. The purpose of this formula is to only train the samples screened by the model, and other samples do not participate in the training.

[0166] In another possible scenario, after obtaining each hierarchical label, image marking can be performed, that is, first extract the first key information in the first type of label; and extract the second key information in the second type of label; then associate the first key information and the second key information to obtain the description information of the input image; and then mark the input image based on the description information.

[0167] Specifically, for the marked image, the call process of the input image can be triggered in response to the target operation; then the input image is cached based on the call process, and the marking of the input image is recognized; when the marking of the input image meets the preset conditions (for example, does not contain exposed parts of the legs), the input image is displayed.

[0168] Optionally, in the above embodiment, the target recognition network can be a lightweight network (such as the chostnet network) deployed in the terminal, and the hierarchical recognition network is an image recognition network (such as resnet18) deployed on the server side, so as to improve the algorithm performance deployed on the service through the cascade framework. The specific performance includes the accuracy and speed of image recognition.

[0169] As can be seen from the above embodiments, by obtaining an input image; then inputting the input image into a preset recognition network in a target model to obtain an attention map, where the attention map contains attention regions, and adjusting the attention map based on the attention regions to obtain an enhanced image, and training the preset recognition network according to the enhanced image to obtain a target recognition network; further inputting the input image into the target recognition network to obtain an image feature map; and then inputting the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch, the first-level label branch is used to determine the first type of label of the input image, the second-level label branch is used to identify the second type of label of the input image, the first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object. Thus, an image hierarchical recognition process guided by attention regions is realized. Since the attention regions in the attention map are used to obtain the enhanced image, the model focuses on the data learning of key parts, and uses the hierarchical recognition method for display, so that the details in the image can be recognized, and the accuracy of image recognition is improved.

[0170] Next, the above image recognition process will be described in combination with the scenario of online release of multimedia data, as Figure 9 shown. Figure 9 FIG. is a flowchart of another method for image recognition based on artificial intelligence provided by an embodiment of the present application; the embodiment of the present application at least includes the following steps:

[0171] 901. Obtain an instant data stream.

[0172] In this embodiment, the instant data stream may be a data stream obtained by an instant messaging software, such as a WeChat Moments refresh data, a short video application refresh data, or a live video stream data.

[0173] Specifically, for the recognition process of video data, one or more frames of images in the video stream may be selected for recognition, so as to realize the recognition of video data.

[0174] 902. Input the image data in the instant data stream into a target model for recognition to obtain a first type of label and a second type of label.

[0175] In this embodiment, for the process of inputting into the target model for recognition, refer to the description of the embodiment shown in Figure 3 and details are not described herein.

[0176] It can be understood that for the application scenarios of short videos or live broadcasts, the video data stream can be converted into image data for recognition; specifically, each frame can be used as the image data, or the start or end frame of the video can be used as the image data, or a fixed image acquisition interval can be used to extract video frames to obtain image data; due to the certain correlation between the contents of video frames, the intermittent acquisition method will not miss possible vulgar images, improving the recognition efficiency while ensuring the recognition accuracy.

[0177] 903. Determine vulgar images based on the first type of tag.

[0178] In this embodiment, the first type of tag can be directly used to determine whether vulgar images are included, such as whether images of the "female sensuality" type are included, so as to make markings and determine whether to push.

[0179] 904. Process and publish vulgar parts based on the second type of tag.

[0180] In this embodiment, the second type of tag can obtain specific recognition parts. If the first type of tag or the second type of tag indicates that the image contains vulgar content, the vulgar parts can be processed based on the second type of tag, such as mosaic processing, and then the processed image can be published to ensure the accuracy of the online information.

[0181] It can be understood that for the order of information going online, it can be based on the corresponding time sequence when obtaining, that is, obtaining the real-time media data stream; then extracting images from the media data stream according to the target time sequence to obtain input images, and the input images are published according to the target time sequence after recognition.

[0182] In this embodiment, through the introduction of the guided attention mechanism, category-based adaptive threshold learning, and the joint training of double-branch and double-task, the recognition accuracy of vulgar information is high. Among hundreds of millions of data per day in instant messaging software, millions of vulgar data can be accurately resisted, and new scenarios such as video accounts and live broadcasts are also ensured to be unaffected by vulgar image information.

[0183] The following will explain the above image recognition process in combination with data maintenance on the server side. Please refer to Figure 10 , Figure 10 which is the flowchart of another method for artificial intelligence-based image recognition provided by the embodiments of the present application. The embodiments of the present application at least include the following steps:

[0184] 1001. Obtain the real-time data stream.

[0185] In this embodiment, the obtaining of the real-time data stream is similar to step 901 of the embodiment shown in Figure 9 and will not be elaborated here.

[0186] 1002. Input the image data in the real-time data stream into the target model for recognition to obtain the first type of label and the second type of label.

[0187] In this embodiment, for the process of inputting into the target model for recognition, refer to Figure 3 the description of the illustrated embodiment, which will not be elaborated here.

[0188] 1003. Determine the normal images in the real-time data stream and publish them.

[0189] In this embodiment, for the data indicated by the first type of label as normal data, immediate upper limit publishing can be performed.

[0190] 1004. Mark the abnormal images based on the first type of label and the second type of label, and upload them to the server.

[0191] In this embodiment, since the second type of label records the specific position of the vulgar part of the image, it can be marked and corresponding description information can be generated. For example: Since this image contains exposed legs, so on.

[0192] Optionally, for the pictures that record the vulgar part of the image, during the process of recognizing similar pictures, check the attention area that records the vulgar part of the image, thereby improving the efficiency of vulgar image recognition and realizing the dynamic recognition process, that is, the continuous collection of the attention area of vulgar images can guide the subsequent image recognition.

[0193] Specifically, there may be images that are incorrectly uploaded, so the model parameters can be adjusted based on the incorrectly uploaded images to improve the accuracy of the target model.

[0194] As can be seen from the above embodiments, the image recognition process of the present application has strong interpretability. According to the attention area of the visualization model, the reason analysis of the model's prediction result for the image can be given, laying a foundation for the subsequent iteration of the model with higher indicators.

[0195] To better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Please refer to Figure 11 , Figure 11 is a schematic structural diagram of an image recognition device provided by an embodiment of the present application. The recognition device 1100 includes:

[0196] An acquisition unit 1101, configured to acquire an input image;

[0197] An input unit 1102, configured to input the input image into a preset recognition network in the target model to obtain an attention map, and the attention map includes an attention area.

[0198] An adjustment unit 1103, configured to perform image adjustment on the attention map based on the attention region to obtain an enhanced image, and train the preset recognition network according to the enhanced image to obtain a target recognition network;

[0199] The input unit 1102 is further configured to input the input image into the target recognition network to obtain an image feature map;

[0200] A recognition unit 1104, configured to input the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label. The hierarchical recognition network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object.

[0201] Optionally, in some possible implementation manners of the present application, the adjustment unit 1103 is specifically configured to mask the attention region to update the attention map to obtain a first adjusted image, and adjust the label corresponding to the first adjusted image;

[0202] The adjustment unit 1103 is specifically configured to strengthen the weight parameter corresponding to the attention region to update the attention map to obtain a second adjusted image, and keep the label corresponding to the second adjusted image unchanged;

[0203] The adjustment unit 1103 is specifically configured to train the preset recognition network according to the first adjusted image and the second adjusted image to obtain the target recognition network.

[0204] Optionally, in some possible implementation manners of the present application, the adjustment unit 1103 is specifically configured to perform regional perturbation based on the first adjusted image to generate a negative sample sequence;

[0205] The adjustment unit 1103 is specifically configured to perform weight parameter perturbation based on the second adjusted image to generate a positive sample sequence;

[0206] The adjustment unit 1103 is specifically configured to train the preset recognition network according to the negative sample sequence and the positive sample sequence to obtain the target recognition network.

[0207] Optionally, in some possible implementation manners of the present application, the adjustment unit 1103 is specifically configured to determine an attention first-level label and an attention second-level label corresponding to the attention area;

[0208] The adjustment unit 1103 is specifically configured to perform constraints based on the area corresponding to the attention first-level label and the area corresponding to the attention second-level label to obtain attention loss information;

[0209] The adjustment unit 1103 is specifically configured to adjust the parameters of the target recognition network according to the attention loss information.

[0210] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to obtain first-level label training data;

[0211] The recognition unit 1104 is specifically configured to determine the classification loss in the first-level label training data to train the first-level label branch;

[0212] The recognition unit 1104 is specifically configured to obtain second-level label training data;

[0213] The recognition unit 1104 is specifically configured to input the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples;

[0214] The recognition unit 1104 is specifically configured to train the second-level label branch based on the second-level label positive samples and the second-level label negative samples.

[0215] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to determine target samples in the second-level label training data;

[0216] The recognition unit 1104 is specifically configured to perform a sliding mean calculation based on the batch data corresponding to the target samples to obtain dynamic threshold information, and the dynamic threshold information includes a positive sample threshold and a negative sample threshold;

[0217] The recognition unit 1104 is specifically configured to input the target samples into the binary classifier to obtain a predicted value;

[0218] The recognition unit 1104 is specifically configured to compare the predicted value with the dynamic threshold information to determine the second-level label positive samples and the second-level label negative samples in the second-level label training data.

[0219] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to compare the predicted value with the positive sample threshold in the dynamic threshold information;

[0220] If the predicted value is greater than the positive sample threshold, determine that the target sample is a positive sample of the secondary label;

[0221] The recognition unit 1104 is specifically configured to compare the predicted value with the negative sample threshold in the dynamic threshold information;

[0222] The recognition unit 1104 is specifically configured to, if the predicted value is less than the negative sample threshold, determine that the target sample is a negative sample of the secondary label.

[0223] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to, if the predicted value is greater than the negative sample threshold and less than the positive sample threshold, determine that the target sample is a noise sample;

[0224] The recognition unit 1104 is specifically configured to set the noise sample not to participate in the training of the secondary label branch.

[0225] Optionally, in some possible implementation manners of the present application, the acquisition unit 1101 is specifically configured to acquire an instant media data stream;

[0226] The acquisition unit 1101 is specifically configured to extract images in the media data stream according to a target time sequence to obtain the input image, and the input image is published according to the target time sequence after being recognized.

[0227] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to extract first key information in the first type of label;

[0228] The recognition unit 1104 is specifically configured to extract second key information in the second type of label;

[0229] The recognition unit 1104 is specifically configured to associate the first key information and the second key information to obtain description information of the input image;

[0230] The recognition unit 1104 is specifically configured to mark the input image based on the description information.

[0231] Optionally, in some possible implementation manners of the present application, the recognition unit 1104 is specifically configured to trigger a call process of the input image in response to a target operation;

[0232] The recognition unit 1104 is specifically configured to cache the input image based on the call process and recognize the mark of the input image;

[0233] The recognition unit 1104 is specifically configured to display the input image if the label of the input image meets a preset condition.

[0234] By obtaining an input image; then inputting the input image into a preset recognition network in a target model to obtain an attention map, where the attention map contains an attention area, and based on the attention area, performing image adjustment on the attention map to obtain an enhanced image, and training the preset recognition network according to the enhanced image to obtain a target recognition network; further inputting the input image into the target recognition network to obtain an image feature map; and then inputting the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label, where the hierarchical recognition network includes a first-level label branch and a second-level label branch, the first-level label branch is used to determine the first type of label of the input image, the second-level label branch is used to identify the second type of label of the input image, the first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is smaller than the description granularity of the first type of label for the target object. Thus, an image hierarchical recognition process guided by an attention area is realized. Since the attention area in the attention map is used to obtain the enhanced image, the model focuses on data learning of key parts, and a hierarchical recognition method is adopted for display, so that the detailed parts in the image can be recognized, and the accuracy of image recognition is improved.

[0235] An embodiment of the present application further provides a terminal device, such as Figure 12 shown, which is a schematic structural diagram of another terminal device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc. Taking the terminal as a mobile phone as an example:

[0236] Figure 12 shown is a block diagram of a part of the structure of the mobile phone related to the terminal provided by an embodiment of the present application. Referring to Figure 12 the mobile phone includes components such as a radio frequency (RF) circuit 1210, a memory 1220, an input unit 1230, a display unit 1240, a sensor 1250, an audio circuit 1260, a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290. Those skilled in the art can understand, Figure 12The mobile phone structure shown does not constitute a limitation on the mobile phone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0237] The following will specifically introduce each component of the mobile phone in conjunction with Figure 12 :

[0238] The RF circuit 1210 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 1280 for processing; in addition, the uplink data designed is sent to the base station. Generally, the RF circuit 1210 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1210 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0239] The memory 1220 can be used to store software programs and modules. The processor 1280 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1220. The memory 1220 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 1220 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0240] The input unit 1230 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1230 may include a touch panel 1231 and other input devices 1232. The touch panel 1231, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 1231, and air-touch operations within a certain range on the touch panel 1231), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1231 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1280, and can receive and execute commands sent by the processor 1280. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1231. In addition to the touch panel 1231, the input unit 1230 may further include other input devices 1232. Specifically, the other input devices 1232 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0241] The display unit 1240 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1240 may include a display panel 1241. Optionally, the display panel 1241 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1231 can cover the display panel 1241. After the touch panel 1231 detects a touch operation thereon or nearby, it is transmitted to the processor 1280 to determine the type of touch event. Subsequently, the processor 1280 provides a corresponding visual output on the display panel 1241 according to the type of touch event. Although in Figure 12 the touch panel 1231 and the display panel 1241 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1231 and the display panel 1241 can be integrated to realize the input and output functions of the mobile phone.

[0242] The mobile phone may further include at least one sensor 1250, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1241 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1241 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0243] The audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface between the user and the mobile phone. The audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, and the speaker 1261 converts it into a sound signal for output; on the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1260 and then converted into audio data. After the audio data is output to the processor 1280 for processing, it is sent through the RF circuit 1210 to, for example, another mobile phone, or the audio data is output to the memory 1220 for further processing.

[0244] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 1270, which provides users with wireless broadband Internet access. Although Figure 12 the WiFi module 1270 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0245] The processor 1280 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1220, and by calling the data stored in the memory 1220, it executes various functions of the mobile phone and processes data. Optionally, the processor 1280 may include one or more processing units; optionally, the processor 1280 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1280 either.

[0246] The mobile phone further includes a power supply 1290 (such as a battery) for supplying power to each component. Optionally, the power supply can be logically connected to the processor 1280 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0247] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.

[0248] In the embodiment of the present application, the processor 1280 included in the terminal further has the function of executing each step of the page processing method as described above.

[0249] The embodiment of the present application further provides a server. Please refer to Figure 13 , Figure 13 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1322 (for example, one or more processors) and a memory 1332, and one or more storage media 1330 for storing application programs 1342 or data 1344 (for example, one or more mass storage devices). Among them, the memory 1332 and the storage media 1330 may be transient storage or persistent storage. The program stored in the storage media 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1322 may be configured to communicate with the storage media 1330 and execute a series of instruction operations in the storage media 1330 on the server 1300.

[0250] The server 1300 may further include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1358, and / or one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0251] The steps executed by the management device in the above embodiments may be based on the Figure 13 shown server structure.

[0252] The embodiment of the present application further provides a computer-readable storage medium, in which instructions for image recognition are stored. When it runs on a computer, it causes the computer to execute the steps executed by the image recognition device in the method described in the foregoing Figures 3 to 10 shown embodiments.

[0253] In an embodiment of the present application, a computer program product including an instruction for image recognition is further provided. When it runs on a computer, it causes the computer to execute the steps performed by the image recognition device in the method described in the foregoing Figures 3 to 10 embodiment shown.

[0254] In an embodiment of the present application, an image recognition system is further provided. The image recognition system may include Figure 11 the image recognition device in the described embodiment, or Figure 12 the terminal device in the described embodiment, or Figure 13 the described server.

[0255] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0256] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0257] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0258] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0259] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an image recognition device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0260] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A method for image recognition based on artificial intelligence, characterized in that, Including: Obtain an input image; Input the input image into a preset recognition network in a target model to obtain an attention map, where the attention map contains attention regions; Mask the attention regions to update the attention map to obtain a first adjusted image, and adjust the label corresponding to the first adjusted image; Reinforce the weight parameters corresponding to the attention regions to update the attention map to obtain a second adjusted image, and keep the label corresponding to the second adjusted image unchanged; Train the preset recognition network according to the first adjusted image and the second adjusted image to obtain a target recognition network; Input the input image into the target recognition network to obtain an image feature map; Input the image feature map into a hierarchical recognition network in the target model to obtain a first type of label and a second type of label. The hierarchical recognition network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is less than the description granularity of the first type of label for the target object.

2. The method according to claim 1, wherein The training of the preset recognition network according to the first adjusted image and the second adjusted image to obtain the target recognition network includes: Perform regional perturbation based on the first adjusted image to generate a negative sample sequence; Perform weight parameter perturbation based on the second adjusted image to generate a positive sample sequence; Train the preset recognition network according to the negative sample sequence and the positive sample sequence to obtain the target recognition network.

3. The method according to claim 1, characterized in that The method further includes: Determine an attention first-level label and an attention second-level label corresponding to the attention regions; Perform constraint based on the regions corresponding to the attention first-level label and the regions corresponding to the attention second-level label to obtain attention loss information; Adjust the parameters of the target recognition network according to the attention loss information.

4. The method according to claim 1, characterized in that The method further includes: Obtain first-level label training data; Determine the classification loss in the first-level label training data to train the first-level label branch; Obtain second-level label training data; Input the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples; Train the second-level label branch based on the second-level label positive samples and the second-level label negative samples.

5. The method according to claim 4, wherein The inputting of the second-level label training data into the binary classifier to obtain second-level label positive samples and second-level label negative samples includes: Determine the target samples in the second-level label training data; Perform sliding mean calculation based on the batch data corresponding to the target samples to obtain dynamic threshold information, where the dynamic threshold information includes a positive sample threshold and a negative sample threshold; Input the target samples into the binary classifier to obtain prediction values; Compare the predicted value with the dynamic threshold information to determine the positive and negative samples of the secondary label in the secondary label training data.

6. The method according to claim 5, wherein The comparing the predicted value with the dynamic threshold information to determine the positive and negative samples of the secondary label in the secondary label training data includes: Compare the predicted value with the positive sample threshold in the dynamic threshold information; If the predicted value is greater than the positive sample threshold, determine that the target sample is the positive sample of the secondary label; Compare the predicted value with the negative sample threshold in the dynamic threshold information; If the predicted value is less than the negative sample threshold, determine that the target sample is the negative sample of the secondary label.

7. The method according to claim 6, characterized in that, The method further includes: If the predicted value is greater than the negative sample threshold and less than the positive sample threshold, determine that the target sample is a noise sample; Set the noise sample not to participate in the training of the secondary label branch.

8. The method according to claim 1, characterized in that The obtaining the input image includes: Obtain an instant media data stream; Extract images from the media data stream according to the target time sequence to obtain the input image, and the input image is published according to the target time sequence after being recognized.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Extract the first key information from the first type of label; Extract the second key information from the second type of label; Associate the first key information and the second key information to obtain the description information of the input image; Mark the input image based on the description information.

10. The method according to claim 9, wherein The method further includes: Respond to a target operation to trigger the call process of the input image; Cache the input image based on the call process and recognize the label of the input image; If the label of the input image meets the preset conditions, display the input image.

11. The method according to claim 1, characterized in that The target model is used for the recognition of vulgar images. The first type of label is used to indicate the individual type of the target object, and the second type of label is used to indicate the part type of the target object.

12. An image recognition device, characterized in that, It includes: An acquisition unit for acquiring an input image; An input unit for inputting the input image into a preset recognition network in the target model to obtain an attention map, and the attention map contains an attention area, An adjustment unit for masking the attention area to update the attention map to obtain a first adjusted image and adjust the label corresponding to the first adjusted image; strengthening the weight parameter corresponding to the attention area to update the attention map to obtain a second adjusted image and keep the label corresponding to the second adjusted image unchanged; training the preset recognition network according to the first adjusted image and the second adjusted image to obtain a target recognition network; The input unit is further used for inputting the input image into the target recognition network to obtain an image feature map; An identification unit, configured to input the image feature map into a hierarchical identification network in the target model to obtain a first type of label and a second type of label. The hierarchical identification network includes a first-level label branch and a second-level label branch. The first-level label branch is used to determine the first type of label of the input image, and the second-level label branch is used to identify the second type of label of the input image. The first type of label and the second type of label are used to indicate the same target object, and the description granularity of the second type of label for the target object is less than the description granularity of the first type of label for the target object.

13. The device according to claim 12, characterized in that, The adjustment unit is specifically configured to: Perform regional perturbation based on the first adjusted image to generate a negative sample sequence; Perform weight parameter perturbation based on the second adjusted image to generate a positive sample sequence; Train the preset identification network according to the negative sample sequence and the positive sample sequence to obtain the target identification network.

14. The device according to claim 12, characterized in that, The adjustment unit is specifically configured to: Determine an attention first-level label and an attention second-level label corresponding to the attention region; Perform constraint based on the region corresponding to the attention first-level label and the region corresponding to the attention second-level label to obtain attention loss information; Adjust the parameters of the target identification network according to the attention loss information.

15. The device according to claim 12, characterized in that, The identification unit is specifically configured to: Obtain first-level label training data; Determine the classification loss in the first-level label training data to train the first-level label branch; Obtain second-level label training data; Input the second-level label training data into a binary classifier to obtain second-level label positive samples and second-level label negative samples; Train the second-level label branch based on the second-level label positive samples and the second-level label negative samples.

16. The device according to claim 15, characterized in that, The identification unit is specifically configured to: Determine the target samples in the second-level label training data; Perform sliding mean calculation based on the batch data corresponding to the target samples to obtain dynamic threshold information, where the dynamic threshold information includes a positive sample threshold and a negative sample threshold; Input the target samples into the binary classifier to obtain prediction values; Compare the prediction values with the dynamic threshold information to determine the second-level label positive samples and the second-level label negative samples in the second-level label training data.

17. The device according to claim 16, characterized in that, The identification unit is specifically configured to: Compare the prediction values with the positive sample threshold in the dynamic threshold information; If the prediction value is greater than the positive sample threshold, determine that the target sample is the second-level label positive sample; Compare the prediction values with the negative sample threshold in the dynamic threshold information; If the prediction value is less than the negative sample threshold, determine that the target sample is the second-level label negative sample.

18. The device according to claim 17, wherein The identification unit is specifically configured to: If the prediction value is greater than the negative sample threshold and less than the positive sample threshold, determine that the target sample is a noise sample; Set the noise sample not to participate in the training of the second-level label branch.

19. The device according to claim 12, characterized in that, The acquisition unit is specifically configured to: Acquire an instant media data stream; Extract the images in the media data stream according to the target time sequence to obtain the input images, and publish the input images according to the target time sequence after recognition.

20. The device according to any one of claims 12-19, characterized in that, The recognition unit is specifically configured to: Extract the first key information in the first type of tag; Extract the second key information in the second type of tag; Associate the first key information and the second key information to obtain the description information of the input image; Mark the input image based on the description information.

21. The device according to claim 20, characterized in that, The recognition unit is specifically configured to: Respond to a target operation to trigger the call process of the input image; Cache the input image based on the call process and recognize the mark of the input image; If the mark of the input image meets the preset conditions, display the input image.

22. The device according to claim 12, wherein The target model is used for the recognition of vulgar images. The first type of tag is used to indicate the individual type of the target object, and the second type of tag is used to indicate the part type of the target object.

23. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store program code; the processor is used to execute the image recognition method according to any one of claims 1 to 11 based on the instructions in the program code.

24. A computer-readable storage medium storing instructions that, when run on a computer, cause the computer to execute the image recognition method according to any one of claims 1 to 11 above.

25. A computer program product comprising computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to execute the image recognition method according to any one of claims 1 to 11 above.

Citation Information

Patent Citations

  • Training image recognition network, image recognition searching method and related device

    CN111553372A