An image classification method and device, electronic equipment and storage medium

CN117351252BActive Publication Date: 2026-09-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210750528.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2026-09-15
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

现有技术的不足之处至少包括:不仅容易引入与分类无关的背景信息,而且也并未充分利用细节信息,导致分类误检率较高

Benefits of technology

[0015] The technical solution of this disclosure uses a prediction model to predict the confidence level of each key point in the original image that belongs to the target object; based on the confidence level of each key point, it determines whether the category of the original image contains the target object. By using the confidence level of key points in the image for classification, background information can be avoided, image detail information can be effectively utilized, and the false positive rate can be greatly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351252B_ABST
    Figure CN117351252B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an image classification method and device, electronic equipment and storage medium, wherein the method comprises: predicting, by a prediction model, a confidence of each key point belonging to a target object in an original image; determining whether a category of the original image is a category containing the target object based on the confidence of each key point. By using the confidence of the key points in the image for classification, background information can be avoided, image detail information can be effectively utilized, and the classification false detection rate can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image classification method, apparatus, electronic device, and storage medium. Background Technology

[0002] In existing technologies, classifiers often classify images based on global information. The shortcomings of existing technologies include at least the tendency to introduce background information irrelevant to the classification process, and the failure to fully utilize detailed information, leading to a high false positive rate. Summary of the Invention

[0003] This disclosure provides an image classification method, apparatus, electronic device, and storage medium that can classify images based on the confidence level of key points, effectively reducing the false detection rate.

[0004] In a first aspect, embodiments of this disclosure provide an image classification method, including:

[0005] The prediction model is used to predict the confidence level of each key point in the original image that belongs to the target object.

[0006] Based on the confidence level of each key point, determine whether the category of the original image is a category containing the target object.

[0007] Secondly, embodiments of this disclosure also provide an image classification apparatus, comprising:

[0008] The prediction module is used to predict the confidence level of each key point belonging to the target object in the original image through a prediction model;

[0009] A classification module is used to determine whether the category of the original image is a category containing the target object based on the confidence level of each key point.

[0010] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0011] One or more processors;

[0012] Storage device for storing one or more programs.

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the image classification method as described in any of the embodiments of this disclosure.

[0014] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image classification method as described in any of the embodiments of this disclosure.

[0015] The technical solution of this disclosure uses a prediction model to predict the confidence level of each key point in the original image that belongs to the target object; based on the confidence level of each key point, it determines whether the category of the original image contains the target object. By using the confidence level of key points in the image for classification, background information can be avoided, image detail information can be effectively utilized, and the false positive rate can be greatly reduced. Attached Figure Description

[0016] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0017] Figure 1 This is a schematic flowchart of an image classification method provided in an embodiment of the present disclosure;

[0018] Figure 2 This is a schematic diagram of the structure of a prediction model in an image classification method provided in an embodiment of the present disclosure;

[0019] Figure 3 This is a schematic diagram of key points in the original image in an image classification method provided in an embodiment of this disclosure.

[0020] Figure 4 This is a schematic diagram of the structure of an image classification device provided in an embodiment of the present disclosure;

[0021] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0022] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0024] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0025] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0026] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0027] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0028] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0029] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0030] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0031] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0032] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0033] Figure 1 This is a schematic flowchart illustrating an image classification method provided in an embodiment of this disclosure. This disclosure is applicable to image classification scenarios, such as classifying whether an image is a face image before outputting facial key points. The method can be executed by an image classification device, which can be implemented in software and / or hardware and can be configured in an electronic device.

[0034] like Figure 1 As shown, the image classification method provided in this embodiment may include:

[0035] S110. Using a prediction model, predict the confidence level of each key point in the original image that belongs to the target object.

[0036] In this embodiment of the disclosure, the original image can be preliminarily determined by a target object detection model before being input into the prediction model. Then, the original image can be cropped based on the detection boxes, and the cropped image can be input into the prediction model. Compared to inputting the entire original image into the prediction model, cropping the original image before inputting it can avoid introducing background information to a certain extent and improve the model's prediction performance.

[0037] The prediction model can be a deep learning model, such as a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a transformer model. The prediction model can be trained under supervised supervision using sample data. The trained prediction model can perform the following operations on the input image: extract image features; perform average pooling on the extracted features; and output a confidence vector based on the average pooling result through a fully connected layer.

[0038] The length of the confidence vector is equal to the number of keypoints in the target object, and each element in the vector represents the confidence level of the corresponding keypoint. The number of keypoints in the target object can be preset based on different application scenarios. For example, the common number of keypoints for facial objects can be 21, 68, 72, 128, or 150. The confidence level of a keypoint represents the accuracy of its predicted location. If the target object in the original image is partially occluded, the confidence level of keypoints in the occluded portion is lower, while the confidence level of keypoints in the unoccluded portion is higher.

[0039] S120. Based on the confidence level of each key point, determine whether the category of the original image is a category containing the target object.

[0040] The original image can be categorized into categories that contain the target object and categories that do not contain the target object.

[0041] The more accurately the keypoint locations are predicted, the more clearly the keypoints can represent the detailed locations of the target object, and the higher the confidence level of the keypoints. Based on this, the confidence level can be used to determine the category of the original image. For example, the average or median confidence level of each keypoint can be used to determine the category of the original image. When the average or median is greater than a certain preset value, it can be considered that most keypoint locations can represent the detailed locations of the target object, and the corresponding original image belongs to the category containing the target object.

[0042] In some optional implementations, determining whether the category of the original image contains the target object based on the confidence level of each keypoint may include: counting the number of keypoints with a confidence level greater than a third value; if the number is greater than a fourth value, then the category of the original image is determined to contain the target object.

[0043] The third value can be the median or other quantile of the confidence score range. For example, when the confidence score range is [0,1], the third value can be 0.5. The fourth value can be the median or other quantile of the number of keypoints. For example, when the number of keypoints is 128, the fourth value can be 64. If the number of keypoints with a confidence score greater than the third value is greater than the fourth value, then most keypoints in the original image can be considered keypoints of the target object, and the original image can be classified as containing the target object. Conversely, if the third value is less than or equal to the fourth value, the original image can be classified as not containing the target object. In these optional implementations, the original image category is determined through a confidence score voting method.

[0044] In some alternative implementations, the image classification method may also include: predicting the positions of key points belonging to the target object in the original image using a prediction model; and outputting the positions of the key points if the category of the original image is a category that includes the target object.

[0045] For example, Figure 2 This is a schematic diagram of the prediction model in an image classification method provided in this embodiment of the disclosure. See also... Figure 2The prediction model can include a feature extraction layer (denoted by FE in the figure), a pooling layer (denoted by AP in the figure), a fully connected layer 1 (denoted by FC1 in the figure), and a fully connected layer 2 (denoted by FC2 in the figure); among which the fully connected layer 1 and the fully connected layer 2 can be connected in parallel and independently after the pooling layer. Through the feature extraction layer, features of the input image can be extracted; through the pooling layer, the extracted features can be averaged; through the fully connected layer 1, the location of key points can be predicted; through the fully connected layer 2, the confidence of key points can be predicted.

[0046] Because the detection bounding boxes determined by the detection model may contain false detections, the image cropped based on the detection bounding boxes (i.e., the input image) may not contain any "real" target objects. In this case, the keypoint positions will inevitably not meet expectations. In these optional implementations, before outputting the positions of each keypoint, the confidence level of each keypoint can be used to determine whether the original image contains the target object. Furthermore, the keypoint positions can be output only if the original image contains the target object, and not output if the original image does not contain the target object, thus ensuring that the output keypoint positions meet expectations.

[0047] In some optional implementations, the target object includes a facial object, and the key points include facial key points; after outputting the position of each key point, it also includes: performing at least one of the following operations based on the position of each key point: identity recognition, expression recognition, and adding special effects.

[0048] In existing technologies, image classification can be performed based on global information before outputting keypoints. However, this classification method is prone to misclassifying the image as not containing facial objects when faces are occluded. In such cases, even if the keypoint predictions match expectations, the keypoint positions cannot be output correctly.

[0049] In this embodiment, even when faces are occluded in the image, the confidence level of each key point is considered to determine whether the original image contains a facial object. This avoids introducing background information and fully utilizes detailed information for judgment, significantly reducing the false detection rate. When the original image is determined to contain a facial object, the positions of each facial key point are output correctly. Actual experiments have shown that image classification using this embodiment can reduce the false detection rate by five times compared to existing image classification methods, thus increasing image recall.

[0050] For example, Figure 3 This is a schematic diagram of key points in the original image in an image classification method provided in an embodiment of this disclosure. Figure 3 The mid-face object is occluded by an object, and there are 21 facial landmarks. Of these, 14 round dots represent the unoccluded facial landmarks with high confidence; the 7 triangular dots represent the occluded facial landmarks with lower confidence. Figure 3In the scenario shown, when the number of facial landmarks with a confidence level greater than the third value is greater than the fourth value, the original image can be successfully determined to belong to a category containing facial objects, even if the facial object is occluded. Furthermore, different markers such as dots and triangles can represent the levels of confidence, enabling the output of the facial landmark positions while reflecting the quality of each landmark.

[0051] After outputting the positions of each facial landmark, different visual tasks can be performed based on these positions. These tasks may include, but are not limited to, identity recognition, expression recognition, and adding special effects (such as stickers, beautification, etc.). In addition, other visual tasks (such as face swapping) can be performed based on facial landmarks, which will not be listed here.

[0052] The technical solution of this disclosure uses a prediction model to predict the confidence level of each key point in the original image that belongs to the target object; based on the confidence level of each key point, it determines whether the category of the original image contains the target object. By using the confidence level of key points in the image for classification, background information can be avoided, image detail information can be effectively utilized, and the false positive rate can be greatly reduced.

[0053] This embodiment can be combined with various optional schemes in the image classification method provided in the above embodiments. The image classification method provided in this embodiment details the training steps of the prediction model. By segmenting the sample image to include the target object, confidence labels for each keypoint can be set based on the segmented region, thus preparing training data. Furthermore, the prediction model can be trained based on the confidence labels, so that the trained prediction model can predict the confidence of each keypoint in the image.

[0054] In this embodiment of the disclosure, the prediction model is trained based on the following steps: obtaining the location labels of each key point belonging to the target object in the sample image; segmenting the target object in the sample image using a segmentation model to obtain segmented regions; setting confidence labels for each key point based on the location labels and segmented regions; and training the prediction model based on each confidence label.

[0055] The system can obtain sample images and pre-annotated location labels for keypoints belonging to the target object in each sample image from an open-source database. An open-source target object segmentation model can be used to segment the target object in the sample images, obtaining regions containing the target object (i.e., segmented regions). The location labels can be classified based on the segmented regions to set confidence labels for corresponding keypoints, thus preparing the training data. Furthermore, the prediction model can be trained based on the confidence labels, enabling the trained model to predict the confidence level of each keypoint in the image.

[0056] In some optional implementations, setting a confidence label for each key point based on the location label and the segmentation region may include: if the location represented by the location label is within the segmentation region, then setting the confidence label of the key point corresponding to the location label to a first value; if the location represented by the location label is outside the segmentation region, then setting the confidence label of the key point corresponding to the location label to a second value; wherein the first value is greater than the second value.

[0057] In these optional implementations, the confidence level of a keypoint can be set based on whether its location, as represented by the location label, is within the segmentation region. If it is within the segmentation region, the corresponding keypoint is considered visible in the sample image, and its confidence level is set to a first value; if it is not within the segmentation region, the corresponding keypoint is considered invisible in the sample image, and its confidence level is set to a second value. The first value can be the maximum value in the confidence level range, and the second value can be the minimum value in the confidence level range; for example, the first value can be 1, and the second value can be 0. This can be understood as the confidence level of the keypoint being set according to the occlusion status of the keypoint in the image, thus facilitating image classification prediction based on keypoints in the unoccluded portion when the target object is occluded.

[0058] In some optional implementations, the prediction model is trained based on each confidence label, including: predicting the confidence of each key point in the sample image belonging to the target object using the prediction model; and training the prediction model based on the loss value between the predicted confidence and each confidence label.

[0059] In these optional implementations, before inputting the sample image into the prediction model, a detection model can first determine the bounding boxes containing the target object in the sample image. Then, the sample image can be cropped based on the detection boxes, and the cropped image can be input into the prediction model. The prediction model can also perform operations such as feature extraction, average pooling, and outputting a confidence vector on the input image. Each element in the confidence vector should conform to the range of confidence values; for example, when the confidence value range is [0-1], the elements of the confidence vector can be normalized to between 0 and 1 using a normalization function (such as the sigmoid function).

[0060] The loss function between vectors can be used to determine the predicted confidence level and the loss value between each confidence label. For example, when the confidence level ranges from [0-1], the loss function between vectors can be binary cross-entropy (BCE) or logarithmic loss, etc. The parameters of the prediction model can be iteratively adjusted to make the loss value less than a certain preset threshold, thereby obtaining the trained prediction model.

[0061] The technical solution of this disclosure provides a detailed description of the training steps of the prediction model. By segmenting the sample image to identify the region containing the target object, confidence labels for each keypoint can be set based on the segmented region, thus preparing the training data. Furthermore, the prediction model can be trained based on the confidence labels, enabling the trained model to predict the confidence level of each keypoint in the image. In addition, the image classification method provided in this disclosure belongs to the same concept as the image classification method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.

[0062] Figure 4 This is a schematic diagram of the structure of an image classification device provided in an embodiment of this disclosure. The image classification device provided in this embodiment is applicable to image classification scenarios, such as classifying whether an image is a face image before outputting facial key points.

[0063] like Figure 4 As shown, the image classification apparatus provided in this embodiment may include:

[0064] The prediction module 410 is used to predict the confidence level of each key point belonging to the target object in the original image through the prediction model;

[0065] The classification module 420 is used to determine whether the category of the original image is a category containing the target object based on the confidence level of each key point.

[0066] In some alternative implementations, the image classification device may further include:

[0067] The training module is used to train the prediction model based on the following steps:

[0068] Obtain the location labels of each key point belonging to the target object in the sample image;

[0069] The target object in the sample image is segmented using a segmentation model to obtain the segmented region;

[0070] Set confidence labels for each key point based on location labels and segmentation regions;

[0071] The prediction model is trained based on each confidence label.

[0072] In some alternative implementations, the training module can be used for:

[0073] If the location represented by the location label is within the segmented region, then the confidence label of the key point corresponding to the location label is set to the first value;

[0074] If the location represented by the location label is outside the segmentation region, then the confidence label of the key point corresponding to the location label is set to the second value;

[0075] The first value is greater than the second value.

[0076] In some alternative implementations, the training module can be used for:

[0077] The prediction model is used to predict the confidence level of each key point belonging to the target object in the sample image.

[0078] The prediction model is trained based on the loss values ​​between the predicted confidence levels and the confidence level labels.

[0079] In some alternative implementations, the classification module can be used for:

[0080] The number of key points with a confidence level greater than the third value;

[0081] If the number is greater than the fourth value, then the category of the original image is determined to be the category containing the target object.

[0082] In some alternative implementations, the prediction module can also be used to predict the positions of key points belonging to the target object in the original image using a prediction model;

[0083] The image classification device may further include an output module for outputting the positions of key points if the category of the original image is a category containing the target object.

[0084] In some alternative implementations, the target object includes a face object, and the key points include facial key points;

[0085] The image classification device may further include: an operation module for performing at least one of the following operations based on the position of each key point after outputting the position of each key point: identity recognition, expression recognition, and adding special effects.

[0086] The image classification apparatus provided in this disclosure can execute the image classification method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0087] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0088] The following is for reference. Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0089] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0090] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0091] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined above in the image classification method of embodiments of this disclosure.

[0092] The electronic device provided in this embodiment and the image classification method provided in the above embodiments belong to the same disclosed concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0093] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image classification method provided in the above embodiments.

[0094] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory (FLASH), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0095] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0096] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0097] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: predict the confidence level of each key point in the original image belonging to the target object using a prediction model; and determine whether the category of the original image is a category containing the target object based on the confidence level of each key point.

[0098] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0100] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units and modules do not, in certain circumstances, constitute a limitation on the unit or module itself.

[0101] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0102] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0103] According to one or more embodiments of this disclosure, [Example 1] provides an image classification method, the method comprising:

[0104] The prediction model is used to predict the confidence level of each key point in the original image that belongs to the target object.

[0105] Based on the confidence level of each key point, determine whether the category of the original image is a category containing the target object.

[0106] According to one or more embodiments of this disclosure, [Example 2] provides an image classification method, which further includes:

[0107] In some alternative implementations, the prediction model is trained based on the following steps:

[0108] Obtain the location labels of each key point belonging to the target object in the sample image;

[0109] The target object in the sample image is segmented using a segmentation model to obtain segmented regions;

[0110] Confidence labels are set for each key point based on the location labels and the segmented regions;

[0111] The prediction model is trained based on the confidence labels.

[0112] According to one or more embodiments of this disclosure, [Example 3] provides an image classification method, which further includes:

[0113] In some optional implementations, setting confidence labels for each key point based on the location labels and the segmented regions includes:

[0114] If the location represented by the location label is within the segmented region, then the confidence label of the key point corresponding to the location label is set to the first value;

[0115] If the location represented by the location label is outside the segmentation region, then the confidence label of the key point corresponding to the location label is set to the second value;

[0116] The first value is greater than the second value.

[0117] According to one or more embodiments of this disclosure, [Example 4] provides an image classification method, further comprising:

[0118] In some optional implementations, training the prediction model based on each of the confidence labels includes:

[0119] The prediction model is used to predict the confidence level of each key point in the sample image that belongs to the target object.

[0120] The prediction model is trained based on the loss value between the predicted confidence level and each of the confidence level labels.

[0121] According to one or more embodiments of this disclosure, [Example 5] provides an image classification method, which further includes:

[0122] In some optional implementations, determining whether the category of the original image is a category containing the target object based on the confidence level of each key point includes:

[0123] Count the number of key points whose confidence level is greater than the third value;

[0124] If the number is greater than the fourth value, then the category of the original image is determined to be the category containing the target object.

[0125] According to one or more embodiments of this disclosure, [Example Six] provides an image classification method, further comprising:

[0126] In some optional implementations, the prediction model is used to predict the positions of key points belonging to the target object in the original image;

[0127] If the category of the original image contains the target object, then the positions of each key point are output.

[0128] According to one or more embodiments of this disclosure, [Example Seven] provides an image classification method, further comprising:

[0129] In some optional implementations, the target object includes a face object, and the key points include facial key points; after outputting the positions of the key points, the method further includes:

[0130] Perform at least one of the following operations based on the location of each key point: identity recognition, facial expression recognition, and adding special effects.

[0131] According to one or more embodiments of this disclosure, [Example Eight] provides an image classification apparatus, the apparatus comprising:

[0132] The prediction module is used to predict the confidence level of each key point belonging to the target object in the original image through a prediction model;

[0133] A classification module is used to determine whether the category of the original image is a category containing the target object based on the confidence level of each key point.

[0134] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0135] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0136] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image classification method, characterized in that, include: The prediction model is used to predict the confidence level of each key point in the original image that belongs to the target object. Based on the confidence level of each key point, determine whether the category of the original image is a category containing the target object; The prediction model is trained based on the following steps: Obtain the location labels of each key point belonging to the target object in the sample image; The target object in the sample image is segmented using a segmentation model to obtain segmented regions; Confidence labels are set for each key point based on the location labels and the segmented regions; The prediction model is trained based on each of the confidence labels. The step of setting confidence labels for each key point based on the location labels and the segmented regions includes: If the location represented by the location label is within the segmented region, then the confidence label of the key point corresponding to the location label is set to the first value; If the location represented by the location label is outside the segmentation region, then the confidence label of the key point corresponding to the location label is set to the second value; The first value is greater than the second value.

2. The method according to claim 1, characterized in that, The step of training the prediction model based on each of the confidence labels includes: The prediction model is used to predict the confidence level of each key point in the sample image that belongs to the target object. The prediction model is trained based on the loss value between the predicted confidence level and each of the confidence level labels.

3. The method according to claim 1, characterized in that, Determining whether the category of the original image contains the target object based on the confidence level of each key point includes: Count the number of key points whose confidence level is greater than the third value; If the number is greater than the fourth value, then the category of the original image is determined to be the category containing the target object.

4. The method according to claim 1, characterized in that, Also includes: The prediction model is used to predict the positions of key points belonging to the target object in the original image. If the category of the original image contains the target object, then the positions of each key point are output.

5. The method according to claim 4, characterized in that, The target object includes a facial object, and the key points include facial key points; after outputting the positions of the key points, the method further includes: Perform at least one of the following operations based on the location of each key point: identity recognition, facial expression recognition, and adding special effects.

6. An image classification device, characterized in that, include: The prediction module is used to predict the confidence level of each key point belonging to the target object in the original image through a prediction model; A classification module is used to determine whether the category of the original image is a category containing the target object based on the confidence level of each key point; The image classification device further includes: The training module is used to train the prediction model based on the following steps: Obtain the location labels of each key point belonging to the target object in the sample image; The target object in the sample image is segmented using a segmentation model to obtain segmented regions; Confidence labels are set for each key point based on the location labels and the segmented regions; The prediction model is trained based on each of the confidence labels. The step of setting confidence labels for each key point based on the location labels and the segmented regions includes: If the location represented by the location label is within the segmented region, then the confidence label of the key point corresponding to the location label is set to the first value; If the location represented by the location label is outside the segmentation region, then the confidence label of the key point corresponding to the location label is set to the second value; The first value is greater than the second value.

7. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image classification method as described in any one of claims 1-5.

8. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image classification method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Mask wearing detection method, device and equipment and storage medium

    CN113947795A

  • Human body identification method, electronic device and storage medium

    US20210312172A1