Method, device and equipment for classifying human body images

By selecting the optimal image from multiple frames of human body images and performing feature fusion, the problem of the inability to classify the human body images collected by the camera is solved, effectively classifying and identifying human body images, and improving the accuracy of management and analysis.

CN114359957BActive Publication Date: 2025-08-29HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111511302.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-08-29
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The prior art cannot effectively classify human images collected by cameras, resulting in human images being unable to be included in the archive for management and analysis.

Method used

By selecting the optimal human image from multiple frames of human body images, the fused features are determined based on the feature similarity between the optimal human image and the candidate human body image, and the similarity is matched with the cover image of the created human body file, thereby classifying the human body image into the target human body file.

Benefits of technology

It realizes effective classification of human body images, improves personnel classification effect and recognition accuracy, improves the accuracy and universality of image classification, and can quickly obtain useful information from massive images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359957B_ABST
    Figure CN114359957B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, and apparatus for classifying human images. The method includes: selecting the optimal human image from multiple human images of a target object, selecting candidate human images from the multiple human images based on the initial similarity between the optimal human image and each human image frame; determining the fusion features corresponding to the multiple human images of the target object based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images; determining the target similarity between the fusion features and the cover features corresponding to the human cover image based on the human cover image corresponding to each created human profile; selecting a target human profile from all human profiles based on the target similarity, and classifying the optimal human image of the target object as the human image corresponding to the target human profile. Through the technical solution of the present application, human images can be classified to improve human recognition effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a method, device and apparatus for classifying human body images. Background Art

[0002] Image classification refers to identifying multiple images of an object (such as a user) and assigning the same label to multiple images of the same object. This label serves as a unique identifier for the object, and a profile is created for the object. The profile can include the label, the object's structured information (such as ID number, mobile phone number, gender, home address, etc.), multiple images of the object, the capture time and longitude and latitude coordinates corresponding to each image (indicating that the object in the image was at that latitude and longitude coordinate at that time of capture), etc. Obviously, the label can be used to obtain the object's profile, and then the content of the profile can be queried. When it is necessary to manage an object, the object's real-time location and walking route can be analyzed by querying the profile content.

[0003] In order to achieve image classification, for the image captured by the camera, the similarity between the image and the reference image of any file can be calculated. If the similarity is greater than the similarity threshold, the image is classified as an image of the file. If the similarity is not greater than the similarity threshold, the image is not classified as an image of the file, and the similarity between the image and the reference image of another file is continued to be calculated, and so on.

[0004] In the above process, the images captured by the camera are facial images (i.e., images that include human faces), and the reference images in the archive are also facial images, meaning that image classification needs to be based on facial images. However, in actual applications, cameras may capture not only facial images but also human images. For human images captured by cameras, human images cannot be classified, meaning that human images are not included in the archive. Summary of the Invention

[0005] The present application provides a method for classifying human body images, the method comprising:

[0006] Selecting an optimal human body image from multiple frames of human body images of the target object, and selecting a candidate human body image from the multiple frames of human body images based on an initial similarity between the optimal human body image and each frame of human body image;

[0007] Determining fusion features corresponding to multiple frames of human body images of the target object based on the optimal human body features corresponding to the optimal human body image and the candidate human body features corresponding to the candidate human body images;

[0008] Based on the human body cover image corresponding to each created human body file, determining the target similarity between the fusion feature and the cover feature corresponding to the human body cover image;

[0009] A target human body file is selected from all human body files based on the target similarity, and the optimal human body image of the target object is classified as the human body image corresponding to the target human body file.

[0010] The present application provides a human body image classification device, the device comprising: an acquisition module, configured to select an optimal human body image from multiple human body images of a target object, and select a candidate human body image from the multiple human body images based on an initial similarity between the optimal human body image and each human body image;

[0011] a determination module configured to determine, based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images, a fusion feature corresponding to the multiple human images of the target object; and, based on the human cover image corresponding to each created human profile, determine a target similarity between the fusion feature and the cover feature corresponding to the human cover image;

[0012] A classification module is used to select a target human body file from all human body files based on the target similarity, and classify the optimal human body image of the target object as the human body image corresponding to the target human body file.

[0013] The present application provides a back-end device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the human body image classification method disclosed in the above example of the present application.

[0014] It can be seen from the above technical solution that in the embodiment of the present application, the target human body file can be determined based on multiple frames of human body images of the target object, and the optimal human body image of the target object can be classified as the human body image corresponding to the target human body file, so as to classify the human body image, that is, the human body file includes human body images, improves the personnel classification effect, improves the human body recognition effect, improves the accuracy and universality of image classification, and collaboratively realizes intelligent analysis of personnel archiving, which can quickly obtain useful information from massive images. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the embodiments of the present application.

[0016] Figure 1 1 is a flow chart of a method for classifying human body images in one embodiment of the present application;

[0017] Figure 2 This is a schematic diagram of the system structure in one embodiment of the present application;

[0018] Figure 3 This is a flow chart of a method for classifying facial images in one embodiment of the present application;

[0019] Figure 4 is a schematic diagram of a graph convolutional network model in one embodiment of the present application;

[0020] Figure 5 1 is a flow chart of a method for classifying human body images in one embodiment of the present application;

[0021] Figure 6 1 is a schematic structural diagram of a human body image classification device in one embodiment of the present application;

[0022] Figure 7 This is a hardware structure diagram of the back-end device in one embodiment of the present application. DETAILED DESCRIPTION

[0023] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.

[0024] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used may also be interpreted as "at the time of" or "when" or "in response to determining".

[0025] In the embodiment of the present application, a method for classifying human images is proposed, which can also be called a method for archiving human images, and can be applied to back-end devices such as management devices, analysis devices, etc., without limitation. Figure 1 FIG. 4 is a flow chart of the method, which may include the following steps:

[0026] Step 101: Select an optimal human body image from multiple human body images of a target object, and select a candidate human body image from the multiple human body images based on an initial similarity between the optimal human body image and each human body image.

[0027] Exemplarily, selecting the optimal human image from multiple human images of a target object may include, but is not limited to: for each human image of the target object, determining a target frame within the human image, and determining a score for the human image based on at least one feature information corresponding to the target frame and a weight value corresponding to each feature information; wherein the feature information may include, but is not limited to, at least one of the following: clarity, illumination, brightness, completeness, pitch angle, posture, area, and proximity to an edge. Selecting the optimal human image from the multiple human images based on the score corresponding to each human image frame.

[0028] Exemplarily, selecting a candidate human image from multiple frames of human images based on the initial similarity between the optimal human image and each frame of human image may include but is not limited to: determining the initial similarity between the optimal human image and each frame of human image in the multiple frames of human images except the optimal human image; if the initial similarity is greater than a first similarity threshold (which can be configured based on experience), the human image may be selected as a candidate human image, or, if the initial similarity is not greater than the first similarity threshold, the human image may be prohibited from being selected as a candidate human image.

[0029] Step 102: Determine fusion features corresponding to multiple frames of human body images of the target object based on the optimal human body features corresponding to the optimal human body image and the candidate human body features corresponding to the candidate human body images.

[0030] Exemplarily, based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images, a first neighbor matrix centered on the optimal human features can be established, and a second neighbor matrix centered on the candidate human features can be established. The first neighbor matrix is ​​input into the trained target network model to obtain the first sub-feature, and the second neighbor matrix is ​​input into the target network model to obtain the second sub-feature. Based on the first sub-feature and the first weighting coefficient (i.e., the weighting coefficient corresponding to the first sub-feature), the second sub-feature and the second weighting coefficient (i.e., the weighting coefficient corresponding to the second sub-feature), the fusion features corresponding to the multiple frames of human images of the target object are determined; wherein the first weighting coefficient is greater than the second weighting coefficient.

[0031] Exemplarily, the total number of candidate human images may be M, where M is a positive integer greater than 1, and the M candidate human images may correspond to M candidate human features. Based on this, a first nearest neighbor matrix centered on the optimal human feature and a second nearest neighbor matrix centered on the candidate human feature are established based on the optimal human feature and the candidate human features. This may include, but is not limited to: sorting the M candidate human features based on the similarity between the optimal human feature and each candidate human feature; selecting K candidate human features with high similarity based on the sorting results, where K can be a positive integer and is less than or equal to M; and establishing the first nearest neighbor matrix based on the optimal human feature and the K candidate human features. Furthermore, for each candidate human feature, sorting all reference human features (i.e., M reference human features) based on the similarity between the candidate human feature and each reference human feature, where all reference human features may include the optimal human feature and the remaining candidate human features of the M candidate human features excluding the candidate human feature; selecting K reference human features with high similarity based on the sorting results; and establishing a second nearest neighbor matrix corresponding to the candidate human feature based on the candidate human feature and the K reference human features.

[0032] Step 103: Based on the human body cover image corresponding to each created human body file, determine the target similarity between the fusion feature and the cover feature corresponding to the human body cover image.

[0033] Step 104 : selecting a target human body profile from all human body profiles based on the target similarity, and classifying the optimal human body image of the target object as the human body image corresponding to the target human body profile.

[0034] For example, if a human profile corresponds to N human cover images, and N is a positive integer, then the human profile corresponds to N target similarities. Based on this, for each human profile, based on the N target similarities corresponding to the human profile, a first number of target similarities greater than a second similarity threshold and a second number of target similarities greater than a third similarity threshold can be counted. Based on the first number and the weight value of the first number, the second number and the weight value of the second number, the number of targets corresponding to the human profile is determined. Wherein, if the third similarity threshold is greater than the second similarity threshold, the weight value of the second number is greater than the weight value of the first number. A target human profile is selected from all human profiles based on the number of targets corresponding to each human profile.

[0035] Exemplarily, a target human body image can also be selected from multiple frames of human body images of the target object, where the target human body image includes the facial area of ​​the target object; a target face file that matches the facial features corresponding to the facial area is selected from all face files, and a target human body file corresponding to the target object is generated based on the target face file, and the optimal human body image of the target object is classified as the human body image corresponding to the target human body file.

[0036] In one possible implementation, after the optimal human body image of the target object is classified as the human body image corresponding to the target human body file, the optimal human body image of the target object can also be used as the human body cover image corresponding to the target human body file, that is, participate in subsequent comparison and analysis, and there is no restriction on this process.

[0037] It can be seen from the above technical solution that in the embodiment of the present application, the target human body file can be determined based on multiple frames of human body images of the target object, and the optimal human body image of the target object can be classified as the human body image corresponding to the target human body file, so as to classify the human body image, that is, the human body file includes human body images, improves the personnel classification effect, improves the human body recognition effect, improves the accuracy and universality of image classification, and collaboratively realizes intelligent analysis of personnel archiving, which can quickly obtain useful information from massive images.

[0038] The technical solutions of the embodiments of the present application are described below in conjunction with specific embodiments.

[0039] Before introducing the technical solutions of the embodiments of the present application, the following technical terms related to the present application are introduced:

[0040] Face images and body images: For images captured by a camera, facial regions can be identified. If correlation analysis is performed based on facial regions, the image is called a face image. For images captured by a camera, human body regions can be identified. If correlation analysis is performed based on human body regions, the image is called a human body image. Images captured by a camera can be either a face image, a human body image, or both.

[0041] Face profile: A face profile can be created for an object (such as a user). The face profile may include a label for the object, which serves as the unique identifier of the object, structured information of the object (such as ID number, mobile phone number, gender, home address, etc.), multiple face images of the object, and the collection time and longitude and latitude coordinates corresponding to each face image, indicating that the object was at the longitude and latitude coordinates at the collection time.

[0042] Obviously, when the object needs to be managed, the content of the object's face file can be queried through the object's label, and then the object's real-time location and walking route can be analyzed.

[0043] Image classification: Image classification can also be called image archiving or image clustering. That is, for a facial image captured by a camera, the facial profile to which the facial image belongs is determined, and the facial image is classified as a facial image of that facial profile, thereby identifying multiple facial images of the same subject. To achieve image classification, for a facial image captured by a camera, the similarity between the facial image and a reference image in any facial profile is calculated. If the similarity is greater than a similarity threshold, the facial image is classified as belonging to that facial profile. If the similarity is not greater than the similarity threshold, the facial image is not classified as belonging to that facial profile, and the similarity between the facial image and a reference image in another facial profile is calculated again, and so on.

[0044] Reference images for a face profile: A face profile includes multiple facial images of a subject. These images may include a base image and a cover image. These base and cover images are referred to as reference images. During image classification, the similarity between the face image and the reference image is calculated. Based on this similarity, the facial image is classified as belonging to the face profile or not.

[0045] Among them, the base image is an image obtained through information collection, such as an ID card image, etc. The base image is relatively clear. When the base image is obtained, the base image already has the identity information of the face file.

[0046] The facial profile may include multiple facial images of an object. For each facial image other than the base image, a determination is made as to whether the facial image meets the conditions for adding a cover image (such as quantity restrictions, quality restrictions, and similarity restrictions, which are not limited). If the conditions for adding a cover image are met, the facial image is used as the facial cover image of the object, i.e., as a reference image. If the conditions for adding a cover image are not met, the facial image is not used as the facial cover image of the object, i.e., it is not used as a reference image.

[0047] To sum up, the base image and the face cover image can be selected from multiple face images in the face archive, and the base image and the face cover image can be used as reference images. The reference image is an image used for image classification of face images. That is to say, in the image classification process, it is necessary to compare the similarity between the face image and the reference image, and then classify the face image.

[0048] During image classification, facial images captured by cameras are classified, meaning the facial profile to which the facial image belongs must be determined. However, in practice, cameras can capture a large number of human images, making it difficult to classify these images. In response to this finding, embodiments of the present application propose a method for classifying human images that can be used to classify human images.

[0049] See also Figure 2 , which is a schematic diagram of the system structure of an embodiment of the present application. The system structure may include an image acquisition unit, a network transmission unit, an image classification and analysis unit, and a data storage and management unit.

[0050] Among them, the image acquisition unit may include several cameras (such as analog cameras or network cameras, etc.), which can capture images of the target object (such as the target user), such as facial images or body images, etc. For each camera, when the target object is within the field of view of the camera, the camera can capture multiple frames of images of the target object, and identify these multiple frames of images as multiple frames of images of the same target object based on a certain algorithm. There is no restriction on this process. For example, the camera can capture multiple frames of human body images of the target object, and combine the multiple frames of human body images into a human body image sequence, that is, the human body image sequence includes multiple frames of human body images. Alternatively, the camera can capture multiple frames of facial images of the target object, and combine the multiple frames of facial images into a facial image sequence, that is, the facial image sequence includes multiple frames of facial images.

[0051] The network transmission unit may include industrial switches and fiber optic transceivers, etc. The network transmission unit is responsible for building a local area network at the intersection to implement functions such as data transmission and data exchange. For example, the network transmission unit can obtain a human image sequence and / or a facial image sequence from the image acquisition unit. For example, each camera can send its own captured human image sequence and / or facial image sequence to the network transmission unit. The network transmission unit can send the human image sequence and / or facial image sequence to the image classification and analysis unit. The network transmission unit can send the human image sequence and / or facial image sequence to the data storage and management unit.

[0052] The data storage and management unit may include a data server, a management client, and a fiber optic transceiver. The data storage and management unit is responsible for storing human image sequences and / or facial image sequences. The data storage and management unit is responsible for configuring and managing each camera in the image acquisition unit.

[0053] The image classification and analysis unit is a backend device, such as a management device or an analysis device. The image classification and analysis unit is used to implement image classification, such as the classification of facial images or the classification of human body images. For example, upon receiving a facial image sequence, the image classification and analysis unit can classify multiple facial images within the facial image sequence. Upon receiving a human body image sequence, the image classification and analysis unit can classify multiple human body images within the human body image sequence. The following describes the classification process for facial images and the classification process for human body images, in conjunction with specific application scenarios.

[0054] Case 1: Based on the facial image sequence, the image classification analysis unit classifies the facial images. Figure 3 FIG. 1 is a flow chart of a method for classifying facial images, which may include:

[0055] Step 301: Acquire a facial image sequence of a target object. The facial image sequence includes multiple frames of facial images of the target object. For example, the facial image sequence includes n frames of facial images such as a1, a2, ..., an.

[0056] Step 302: Select the optimal face image from all face images in the face image sequence.

[0057] For example, for each frame of facial image, the target frame where the target object is located can be determined from the facial image, and the score value corresponding to the facial image can be determined based on at least one feature information corresponding to the target frame and the weight value corresponding to each feature information; based on the score value corresponding to each frame of facial image, the optimal facial image is selected from multiple frames of facial images, for example, the facial image with the largest score value is selected as the optimal facial image.

[0058] In a possible implementation, the feature information may include but is not limited to at least one of the following: clarity, illumination, brightness, completeness, pitch angle, posture, area, and proximity to an edge.

[0059] For each frame of facial image, taking facial image a1 as an example, the target frame (i.e., target rectangular frame) where the target object is located in facial image a1 can be determined, and the feature information corresponding to the target frame can be determined, such as the clarity s1 corresponding to the target frame, the illumination s2 corresponding to the target frame, the brightness s3 corresponding to the target frame, the completeness s4 corresponding to the target frame, the pitch angle s5 corresponding to the target frame (i.e., the pitch angle of the target object in the target frame), the posture s6 corresponding to the target frame (i.e., the posture of the target object in the target frame), the area s7 corresponding to the target frame, and the proximity to the edge s8 corresponding to the target frame. For the convenience of description, the determination of the above 8 types of feature information will be taken as an example. Regarding the method of determining clarity s1, illumination s2, brightness s3, completeness s4, pitch angle s5, posture s6, area s7, and proximity to the edge s8, there is no limitation in this embodiment. It can be obtained by analyzing the facial image a1, or it can be obtained through a deep learning algorithm, that is, the facial image a1 is input into the network model of the deep learning algorithm, and the network model of the deep learning algorithm outputs the above feature information.

[0060] After obtaining the above eight features, the score corresponding to the facial image a1 can be determined using the following formula: Score = w1*s1+w2*s2+w3*s3+w4*s4+w5*s5+w6*s6+w7*s7+w8*s8. In the above formula, w1 represents the weight corresponding to clarity s1, w2 represents the weight corresponding to illumination s2, w3 represents the weight corresponding to brightness s3, w4 represents the weight corresponding to completeness s4, w5 represents the weight corresponding to pitch angle s5, w6 represents the weight corresponding to posture s6, w7 represents the weight corresponding to area s7, and w8 represents the weight corresponding to edge proximity s8.

[0061] The values ​​of w1, w2, w3, w4, w5, w6, w7, and w8 can be determined empirically or algorithmically, without limitation. For example, multiple sample sequences can be obtained, each of which can include multiple frames of facial images. The optimal facial image in each sample sequence can be manually calibrated, and machine learning methods can be used to train the values ​​of w1, w2, w3, w4, w5, w6, w7, and w8. Of course, the above method is only an example, and there is no limitation on the values ​​of w1, w2, w3, w4, w5, w6, w7, and w8.

[0062] In summary, the score value corresponding to the face image a1 can be determined. Similarly, the score value corresponding to the face image a2 can be determined, ..., and the score value corresponding to the face image an can be determined.

[0063] Then, the face image with the largest score value can be used as the optimal face image. For example, assuming that the score value corresponding to face image a1 is the largest, the optimal face image can be face image a1.

[0064] Step 303: Based on the initial similarity between the optimal facial image and each facial image in the facial image sequence, select candidate facial images from multiple facial image frames. Assume that the total number of candidate facial images is M, where M is a positive integer greater than 1, that is, select M candidate facial images from multiple facial image frames.

[0065] For example, for each frame of face image in the face image sequence except the optimal face image (such as face image a1), taking face image a2 as an example, based on the facial features corresponding to face image a1 and the facial features corresponding to face image a2, the initial similarity between face image a1 and face image a2 can be determined.

[0066] If the initial similarity is greater than a first similarity threshold τ (which can be configured based on experience, such as τ can be 0.5, which is not limited to this), facial image a2 is selected as a candidate facial image. If the initial similarity is not greater than the first similarity threshold, facial image a2 is considered to be an error frame or an inappropriate frame. An error frame is a facial image with a jump, and an inappropriate frame is a facial image in which the target object is occluded or truncated, or a facial image with poor imaging quality. Therefore, facial image a2 is not selected as a candidate facial image.

[0067] In summary, after performing the above processing on each face image frame in the face image sequence, a candidate face image can be selected from the face image sequence. It is assumed that M candidate face images are obtained.

[0068] Step 304: Based on the optimal facial features corresponding to the optimal facial image and the candidate facial features corresponding to the candidate facial images, a first nearest neighbor matrix centered on the optimal facial features is established.

[0069] For example, the facial features corresponding to the optimal facial image are recorded as the optimal facial features, and the facial features corresponding to the candidate facial images are recorded as candidate facial features. Thus, M candidate facial images correspond to M candidate facial features. Based on this, the similarity between the optimal facial features and each candidate facial feature can be determined, resulting in M ​​similarities corresponding to the M candidate facial features. Based on the similarities between the optimal facial features and each candidate facial feature, the M candidate facial features can be ranked. Based on the ranking results, K candidate facial features with the highest similarity are selected, where K can be a positive integer and less than or equal to M.

[0070] For example, if M candidate facial features are sorted in descending order of similarity, the top K candidate facial features are selected based on the sorting results. Alternatively, if M candidate facial features are sorted in ascending order of similarity, the bottom K candidate facial features are selected based on the sorting results. Of course, the above sorting methods are merely examples and are not intended to be limiting.

[0071] After selecting K candidate facial features, the K candidate facial features and the optimal facial feature can be combined into a nearest neighbor matrix (i.e., a K nearest neighbor matrix). This nearest neighbor matrix is ​​called the first nearest neighbor matrix. The first nearest neighbor matrix includes the optimal facial feature and the K candidate facial features. For example, the value of K can be configured based on experience. For example, if K is 50, then the first nearest neighbor matrix includes 50 candidate facial features.

[0072] Since K candidate facial features are selected based on the similarity between the optimal facial feature and each candidate facial feature, the first nearest neighbor matrix is ​​a nearest neighbor matrix centered on the optimal facial feature.

[0073] Step 305: Based on the optimal facial features corresponding to the optimal facial image and the candidate facial features corresponding to the candidate facial images, for each candidate facial feature (i.e., each candidate facial feature in the M candidate facial features), a second nearest neighbor matrix centered on the candidate facial feature is established.

[0074] For example, assuming the optimal facial feature is f0, the M candidate facial features are f1,…,f m On this basis, for the candidate face feature f1, f0 and f2,…,f m As the reference facial feature corresponding to f1. The similarity between the candidate facial feature f1 and each reference facial feature can be determined. Based on the similarity between the candidate facial feature f1 and each reference facial feature, all reference facial features are ranked; and based on the ranking results, K reference facial features with the highest similarity are selected. For example, if all reference facial features are ranked from highest to lowest similarity, the top K reference facial features are selected. Alternatively, if all reference facial features are ranked from lowest to highest similarity, the bottom K reference facial features are selected. After selecting the K reference facial features, the K reference facial features and the candidate facial feature f1 can be combined to form a nearest neighbor matrix. This nearest neighbor matrix serves as the second nearest neighbor matrix corresponding to the candidate facial feature f1. That is, the second nearest neighbor matrix includes the candidate facial feature f1 and the K reference facial features. Since the K reference facial features are selected based on the similarity between the candidate facial feature f1 and each reference facial feature, this second nearest neighbor matrix is ​​a nearest neighbor matrix centered on the candidate facial feature f1.

[0075] For the candidate facial feature f2, f0, f1 and f3,…,f m As the reference facial features corresponding to f2, K reference facial features can be selected from all reference facial features, and the second nearest neighbor matrix corresponding to the candidate facial feature f2 can be formed based on the candidate facial feature f2 and the K reference facial features. Similarly, each candidate facial feature corresponds to a second nearest neighbor matrix, that is, M candidate facial features correspond to M second nearest neighbor matrices.

[0076] Step 306: Input the first nearest neighbor matrix into the trained target network model to obtain the first sub-feature, and input the second nearest neighbor matrix into the target network model to obtain the second sub-feature. For example, since there are M second nearest neighbor matrices, each second nearest neighbor matrix needs to be input into the target network model to obtain the second sub-feature corresponding to the second nearest neighbor matrix, i.e., M second sub-features are obtained.

[0077] For example, the target network model can be an N-layer graph convolutional network model, such as an 8-layer graph convolutional network model. The first nearest neighbor matrix corresponding to the optimal facial feature f0 can be input into the graph convolutional network model, and the graph convolutional network model processes the first nearest neighbor matrix without any restrictions on the processing process to obtain the first sub-feature corresponding to the optimal facial feature f0. The second nearest neighbor matrix corresponding to the candidate facial feature f1 can be input into the graph convolutional network model, and the graph convolutional network model processes the second nearest neighbor matrix to obtain the second sub-feature corresponding to the candidate facial feature f1 By analogy, the candidate facial features f m The corresponding second nearest neighbor matrix is ​​input to the graph convolutional network model, which processes the second nearest neighbor matrix to obtain the candidate facial features f m The corresponding second sub-feature In summary, based on the first nearest neighbor matrix and the second nearest neighbor matrix, the following sub-features can be obtained:

[0078] For example, see Figure 4 As shown, it is a schematic diagram of the N-layer graph convolutional network model, F l Represents the feature matrix of the lth layer of the graph convolutional network model, W l represents the weight parameter matrix of the lth layer of the graph convolutional network model, σ(·) represents the nonlinear activation function, represents the Laplace matrix, and the Laplace matrix is ​​determined as follows: A is the adjacency matrix, i.e., the first neighbor matrix or the second neighbor matrix in the above embodiment, and D is the degree matrix of the adjacency matrix A. Of course, in practical applications, Figure 4 The N-layer graph convolutional network model shown can also use the adjacency matrix A to replace the Laplacian matrix There is no restriction on this.

[0079] based on Figure 4 The N-layer graph convolutional network model shown can be trained first, and there are no restrictions on the training process of the N-layer graph convolutional network model. Obviously, based on the trained N-layer graph convolutional network model, the adjacency matrix A (such as the first nearest neighbor matrix or the second nearest neighbor matrix) can be input into the N-layer graph convolutional network model to obtain the sub-features corresponding to the adjacency matrix A.

[0080] Step 307: Determine a fused feature based on the first sub-feature and the first weighting coefficient corresponding to the first sub-feature, and the second sub-feature and the second weighting coefficient corresponding to the second sub-feature. For example, the fused feature can be obtained by performing a weighted operation on the first sub-feature and the second sub-feature based on the first weighting coefficient and the second weighting coefficient. For example, the first weighting coefficient can be greater than the second weighting coefficient.

[0081] For example, the first sub-feature is The second sub-characteristics include The first sub-feature The corresponding first weighting coefficient is α0, which can be configured based on experience, and the first weighting coefficient α0 needs to be greater than the second weighting coefficient corresponding to each second sub-feature, such as the first weighting coefficient α0 is 0.5. The corresponding second weighting coefficient is α1, which can be configured based on experience. The corresponding second weighting coefficient is α2, which can be configured based on experience, and so on. The second weighting coefficients corresponding to different second sub-features can be the same or different. Taking the second weighting coefficients corresponding to different second sub-features as an example, the second weighting coefficient corresponding to each second sub-feature is (1-α0) / m, that is, the sum of the second weighting coefficients corresponding to m second sub-features is (1-α0). Obviously, if the first weighting coefficient α0 is 0.5, then the second weighting coefficient corresponding to each second sub-feature is 0.5 / m.

[0082] For example, based on the first sub-feature and the first weighting coefficient, the second sub-feature and the second weighting coefficient, the fusion feature f can be determined using the following formula: * : Of course, the above formula is just an example and is not limited to this. In the above formula, when i is 0, represents the first sub-feature, α0 represents the first weighting coefficient, when i is 1, represents the second sub-feature, α1 represents The corresponding second weighting coefficient, when i is 2, represents the second sub-feature, α2 represents The corresponding second weighting coefficient, and so on.

[0083] In summary, the weighted fusion method can be used to obtain fusion features with stronger expressive ability. When performing face recognition based on fusion features, the face recognition effect can be significantly improved and the classification accuracy can be higher.

[0084] Step 308: Classify the target object's facial image sequence based on the fused feature, that is, determine the facial profile to which the multiple facial images in the facial image sequence belong. For example, determine the similarity between the fused feature and the facial features of a reference image of any facial profile. If the similarity is greater than a similarity threshold, then the multiple facial images in the facial image sequence are classified as the facial profile or the optimal facial image in the facial image sequence is classified as the facial profile. If the similarity is not greater than the similarity threshold, then the multiple facial images in the facial image sequence are not classified as the facial profile, nor is the optimal facial image in the facial image sequence classified as the facial profile. It is necessary to continue calculating the similarity between the fused feature and the facial features of a reference image of another facial profile, and so on, until the facial profile to which the multiple facial images in the facial image sequence belong is found, or all facial profiles are not the facial profiles to which the multiple facial images in the facial image sequence belong.

[0085] Case 2: Based on the human body image sequence, the image classification and analysis unit classifies the human body images. For example, for the image captured by the camera, the image may be a face image, a human body image, or a face image and a human body image at the same time, that is, the same image includes a face area and a human body area. On this basis, a human body image sequence of the target object can be obtained, and the human body image sequence includes multiple frames of human body images of the target object. A target human body image can be selected from the multiple frames of human body images. The target human body image includes the face area of ​​the target object, that is, the target human body image is a human body image including the face area of ​​the target object. Then, a target face file that matches the facial features corresponding to the facial area is selected from all face files. For example, if the similarity between the facial features corresponding to the facial area and the facial features of the reference image of a certain face file is greater than the similarity threshold, the face file is used as the target face file.

[0086] For example, after obtaining the target human image, a facial image sequence having the target human image (the target human image is also regarded as a facial image) can be determined. On this basis, the fusion features corresponding to the facial image sequence can be determined (i.e., the facial features corresponding to the facial region in the target human image), and the facial profile to which the facial image sequence belongs can be determined based on the fusion features. This facial profile is regarded as the target facial profile. The process of determining the target facial profile is described in detail in the following. Figure 3 The process shown.

[0087] For example, after obtaining the target face file, a target body file corresponding to the target object can be generated based on the target face file, and multiple frames of body images of the target object (i.e., multiple frames of body images in a body image sequence) can be classified as body images corresponding to this target body file, or the optimal body image among the multiple frames of body images of the target object can be classified as the body image corresponding to this target body file.

[0088] The target body file includes the label of the target object (obtained from the target face file), the structured information of the target object (obtained from the target face file, such as ID number, mobile phone number, gender, home address, etc.), the body image of the target object, the collection time and longitude and latitude coordinates corresponding to the body image, indicating that the target object is at the longitude and latitude coordinates at the collection time, and there is no restriction on this.

[0089] For example, after classifying multiple human body images of the target object (i.e., multiple human body images in a human body image sequence) as human body images corresponding to the target human body profile, or after classifying the optimal human body image of the target object as human body images corresponding to the target human body profile, the optimal human body image can also be used as the human body cover image corresponding to the target human body profile. For example, if the image quality of the optimal human body image is relatively good, the optimal human body image can be used as the human body cover image.

[0090] Case 3: Based on the human body image sequence, the image classification analysis unit classifies the human body images. Figure 5 FIG. 1 is a flow chart of a method for classifying human body images, which may include:

[0091] Step 501: Acquire a human body image sequence of a target object. The human body image sequence includes multiple frames of human body images of the target object. For example, the human body image sequence includes n frames of human body images such as b1, b2, ..., bn.

[0092] Step 502: Select the optimal human body image from all human body images in the human body image sequence.

[0093] For example, for each frame of a human body image, the target frame where the target object is located can be determined from the human body image, and the score value corresponding to the human body image can be determined based on at least one feature information corresponding to the target frame and the weight value corresponding to each feature information; based on the score value corresponding to each frame of the human body image, the optimal human body image is selected from multiple frames of human body images, for example, the human body image with the largest score value is selected as the optimal human body image.

[0094] In a possible implementation, the feature information may include but is not limited to at least one of the following: clarity, illumination, brightness, completeness, pitch angle, posture, area, and proximity to an edge.

[0095] For each frame of human body image, taking human body image b1 as an example, the target frame (i.e., target rectangular frame) where the target object is located in the human body image b1 can be determined, and the feature information corresponding to the target frame can be determined, such as the clarity s1 corresponding to the target frame, the illumination s2 corresponding to the target frame, the brightness s3 corresponding to the target frame, the completeness s4 corresponding to the target frame, the pitch angle s5 corresponding to the target frame (i.e., the pitch angle of the target object in the target frame), the posture s6 corresponding to the target frame (i.e., the posture of the target object in the target frame), the area s7 corresponding to the target frame, and the proximity to the edge s8 corresponding to the target frame. For the convenience of description, the determination of the above 8 types of feature information is taken as an example. Regarding the method of determining clarity s1, illumination s2, brightness s3, completeness s4, pitch angle s5, posture s6, area s7, and proximity to the edge s8, there is no restriction in this embodiment. It can be obtained by analyzing the human body image b1, or it can be obtained through a deep learning algorithm, that is, the human body image b1 is input into the network model of the deep learning algorithm, and the network model of the deep learning algorithm outputs the above feature information.

[0096] After obtaining the above eight features, the score corresponding to the human body image b1 can be determined using the following formula: Score = w1*s1+w2*s2+w3*s3+w4*s4+w5*s5+w6*s6+w7*s7+w8*s8. In the above formula, w1 is used to represent the weight value corresponding to clarity s1, w2 is used to represent the weight value corresponding to illumination s2, w3 is used to represent the weight value corresponding to brightness s3, w4 is used to represent the weight value corresponding to completeness s4, w5 is used to represent the weight value corresponding to pitch angle s5, w6 is used to represent the weight value corresponding to posture s6, w7 is used to represent the weight value corresponding to area s7, and w8 is used to represent the weight value corresponding to proximity to edge s8.

[0097] The values ​​of w1, w2, w3, w4, w5, w6, w7, and w8 can be empirically configured or algorithmically derived, with no limitation imposed. For example, multiple sample sequences can be obtained, each of which can include multiple frames of human images. The optimal human image in each sample sequence can be manually calibrated, and machine learning methods can be used to train the values ​​of w1, w2, w3, w4, w5, w6, w7, and w8. Of course, the above methods are merely examples, and there is no limitation imposed on the values ​​of w1, w2, w3, w4, w5, w6, w7, and w8.

[0098] In summary, the score value corresponding to the human body image b1 can be determined. Similarly, the score value corresponding to the human body image b2 can be determined, ..., and the score value corresponding to the human body image bn can be determined.

[0099] Then, the human body image with the largest score value can be used as the optimal human body image. For example, assuming that the score value corresponding to human body image b1 is the largest, the optimal human body image can be human body image b1.

[0100] Step 503: Based on the initial similarity between the optimal human image and each frame of the human image in the human image sequence, select candidate human images from the multiple frames of human images. Assume that the total number of candidate human images is M, where M is a positive integer greater than 1, that is, M candidate human images are selected from the multiple frames of human images.

[0101] For example, for each frame of human body image in the human body image sequence except the optimal human body image (such as human body image b1), taking human body image b2 as an example, based on the human body features corresponding to human body image b1 and the human body features corresponding to human body image b2, the initial similarity between human body image b1 and human body image b2 can be determined.

[0102] If the initial similarity is greater than a first similarity threshold τ (which can be configured based on experience, such as τ can be 0.5, which is not limited to this), human image b2 is selected as a candidate human image. If the initial similarity is not greater than the first similarity threshold, human image b2 is considered to be an error frame or an inappropriate frame. An error frame is a human image with jumps, and an inappropriate frame is a human image in which the target object is occluded or truncated, or a human image with poor imaging quality. Therefore, human image b2 is not selected as a candidate human image.

[0103] In summary, after performing the above processing on each frame of the human body image in the human body image sequence, a candidate human body image can be selected from the human body image sequence. It is assumed that M candidate human body images are obtained.

[0104] Step 504: Based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images, a first nearest neighbor matrix centered on the optimal human features is established.

[0105] For example, the human feature corresponding to the optimal human image is recorded as the optimal human feature, and the human feature corresponding to the candidate human images is recorded as the candidate human feature. Thus, M candidate human images correspond to M candidate human features. Based on this, the similarity between the optimal human feature and each candidate human feature can be determined, resulting in M ​​similarities corresponding to the M candidate human features. Based on the similarities between the optimal human feature and each candidate human feature, the M candidate human features can be ranked. Based on the ranking results, K candidate human features with the highest similarity are selected, where K can be a positive integer and less than or equal to M.

[0106] For example, if M candidate human features are sorted in descending order of similarity, the top K candidate human features are selected based on the sorting results. Alternatively, if M candidate human features are sorted in ascending order of similarity, the bottom K candidate human features are selected based on the sorting results. Of course, the above sorting methods are merely examples and are not intended to be limiting.

[0107] After selecting K candidate human features, the K candidate human features and the optimal human features can be combined into a nearest neighbor matrix (i.e., a K nearest neighbor matrix). This nearest neighbor matrix is ​​referred to as the first nearest neighbor matrix. The first nearest neighbor matrix includes the optimal human features and the K candidate human features. For example, the value of K can be configured based on experience. For example, when K is 50, the first nearest neighbor matrix includes 50 candidate human features.

[0108] Since K candidate human features are selected based on the similarity between the optimal human feature and each candidate human feature, the first nearest neighbor matrix is ​​a nearest neighbor matrix centered on the optimal human feature.

[0109] Step 505: Based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images, for each candidate human feature (ie, each candidate human feature in the M candidate human features), a second nearest neighbor matrix centered on the candidate human feature is established.

[0110] For example, assuming the optimal human feature is f0, the M candidate human features are f1,…,f m On this basis, for the candidate human feature f1, f0 and f2,…,f m As the reference human feature corresponding to f1. The similarity between the candidate human feature f1 and each reference human feature can be determined. Based on the similarity between the candidate human feature f1 and each reference human feature, all reference human features are sorted; based on the sorting result, K reference human features with the largest similarity are selected. For example, if all reference human features are sorted in order from large to small in similarity, the K reference human features with the highest sorting are selected. Alternatively, if all reference human features are sorted in order from small to large in similarity, the K reference human features with the lowest sorting are selected. After selecting the K reference human features, the K reference human features and the candidate human feature f1 can be combined into a neighbor matrix, which is the second neighbor matrix corresponding to the candidate human feature f1, that is, the second neighbor matrix includes the candidate human feature f1 and the K reference human features. Since the K reference human features are selected based on the similarity between the candidate human feature f1 and each reference human feature, the second neighbor matrix is ​​a neighbor matrix centered on the candidate human feature f1.

[0111] For the candidate human feature f2, f0, f1 and f3,…,f m As the reference human features corresponding to f2, K reference human features can be selected from all reference human features, and the second nearest neighbor matrix corresponding to the candidate human feature f2 can be formed based on the candidate human feature f2 and the K reference human features. Similarly, each candidate human feature corresponds to a second nearest neighbor matrix, that is, M candidate human features correspond to M second nearest neighbor matrices.

[0112] Step 506: Input the first nearest neighbor matrix into the trained target network model to obtain the first sub-feature, and input the second nearest neighbor matrix into the target network model to obtain the second sub-feature. For example, since there are M second nearest neighbor matrices, each second nearest neighbor matrix needs to be input into the target network model to obtain the second sub-feature corresponding to the second nearest neighbor matrix, i.e., M second sub-features are obtained.

[0113] For example, the target network model can be an N-layer graph convolutional network model, such as an 8-layer graph convolutional network model. The first nearest neighbor matrix corresponding to the optimal human feature f0 can be input into the graph convolutional network model, and the graph convolutional network model processes the first nearest neighbor matrix without any restrictions on the processing process to obtain the first sub-feature corresponding to the optimal human feature f0. The second nearest neighbor matrix corresponding to the candidate human feature f1 can be input into the graph convolutional network model, and the graph convolutional network model processes the second nearest neighbor matrix to obtain the second sub-feature corresponding to the candidate human feature f1 By analogy, the candidate human feature f m The corresponding second nearest neighbor matrix is ​​input to the graph convolutional network model, which processes the second nearest neighbor matrix to obtain the candidate human feature f m The corresponding second sub-feature In summary, based on the first nearest neighbor matrix and the second nearest neighbor matrix, the following sub-features can be obtained:

[0114] Step 507: Determine a fused feature based on the first sub-feature and the first weighting coefficient corresponding to the first sub-feature, and the second sub-feature and the second weighting coefficient corresponding to the second sub-feature. For example, the fused feature can be obtained by performing a weighted operation on the first sub-feature and the second sub-feature based on the first weighting coefficient and the second weighting coefficient. For example, the first weighting coefficient can be greater than the second weighting coefficient.

[0115] For example, the first sub-feature is The second sub-characteristics include The first sub-feature The corresponding first weighting coefficient is α0, which can be configured based on experience, and the first weighting coefficient α0 needs to be greater than the second weighting coefficient corresponding to each second sub-feature, such as the first weighting coefficient α0 is 0.5. The corresponding second weighting coefficient is α1, which can be configured based on experience. The corresponding second weighting coefficient is α2, which can be configured based on experience, and so on. The second weighting coefficients corresponding to different second sub-features can be the same or different. Taking the second weighting coefficients corresponding to different second sub-features as an example, the second weighting coefficient corresponding to each second sub-feature is (1-α0) / m, that is, the sum of the second weighting coefficients corresponding to m second sub-features is (1-α0). Obviously, if the first weighting coefficient α0 is 0.5, then the second weighting coefficient corresponding to each second sub-feature is 0.5 / m.

[0116] For example, based on the first sub-feature and the first weighting coefficient, the second sub-feature and the second weighting coefficient, the fusion feature f can be determined using the following formula: * : Of course, the above formula is just an example and is not limited to this. In the above formula, when i is 0, represents the first sub-feature, α0 represents the first weighting coefficient, when i is 1, represents the second sub-feature, α1 represents The corresponding second weighting coefficient, when i is 2, represents the second sub-feature, α2 represents The corresponding second weighting coefficient, and so on.

[0117] In summary, the weighted fusion method can be used to obtain fusion features with stronger expressive ability. When human body recognition is performed based on the fusion features, the human body recognition effect can be significantly improved and the classification accuracy can be higher.

[0118] Step 508: Based on the human body cover image corresponding to each created human body file, determine the target similarity between the fusion feature and the cover feature corresponding to the human body cover image.

[0119] For example, multiple human body files can be created in advance, such as using the second scenario to create a human body file, or other methods can be used to create a human body file. There is no restriction on the creation method of the human body file.

[0120] For each human body profile that has been created, the human body profile can correspond to at least one human body cover image. Assuming that the human body profile corresponds to N human body cover images, and N is a positive integer, the target similarity between the fusion feature and the cover feature corresponding to each human body cover image can be determined, that is, N target similarities are obtained. In other words, the human body profile can correspond to N target similarities.

[0121] For example, if human body file 1 corresponds to human body cover image 11, human body cover image 12, and human body cover image 13, that is, the value of N is 3, then the target similarity between the fused feature and the cover feature corresponding to human body cover image 11 (the human body feature corresponding to the human body cover image is called the cover feature) is determined, the target similarity between the fused feature and the cover feature corresponding to human body cover image 12 is determined, and the target similarity between the fused feature and the cover feature corresponding to human body cover image 13 is determined. That is, human body file 1 can correspond to 3 target similarities. For another example, if human body file 2 corresponds to human body cover image 21, human body cover image 22, human body cover image 23, and human body cover image 24, that is, the value of N is 4, then the target similarity between the fused feature and the cover feature corresponding to human body cover image 21 is determined, the target similarity between the fused feature and the cover feature corresponding to human body cover image 22 is determined, the target similarity between the fused feature and the cover feature corresponding to human body cover image 23 is determined, and the target similarity between the fused feature and the cover feature corresponding to human body cover image 24 is determined. That is, human body file 2 can correspond to 4 target similarities, and so on.

[0122] Step 509: Select a target human body file from all human body files based on the target similarity, and classify the optimal human body image of the target object as the human body image corresponding to the target human body file, or classify multiple frames of human body images of the target object as human body images corresponding to the target human body file, that is, classify the optimal human body image sequence of the target object and determine the target human body file to which the optimal human body image belongs.

[0123] Exemplarily, for each human profile, N target similarities corresponding to the human profile can be obtained (see step 508). Based on the N target similarities corresponding to the human profile, a first number of target similarities greater than a second similarity threshold and a second number of target similarities greater than a third similarity threshold can be counted. Based on the first number and the weight value of the first number, the second number and the weight value of the second number, the number of targets corresponding to the human profile is determined; wherein, if the third similarity threshold is greater than the second similarity threshold, the weight value of the second number can be greater than the weight value of the first number. On this basis, a target human profile can be selected from all human profiles based on the number of targets corresponding to each human profile.

[0124] For example, a second similarity threshold T1 and a third similarity threshold T2 can be pre-configured. Both the second similarity threshold T1 and the third similarity threshold T2 can be configured based on experience and are not limited thereto. The third similarity threshold T2 can be greater than the second similarity threshold T1, and the second similarity threshold T1 can be greater than the first similarity threshold. For example, the third similarity threshold T2 is 0.95, and the second similarity threshold T1 is 0.8.

[0125] For example, a first number of weight values ​​(denoted as weight value α1) and a second number of weight values ​​(denoted as weight value α2) can be pre-configured. Both weight value α1 and weight value α2 can be configured based on experience and are not limited thereto. Weight value α2 is greater than weight value α1, and the sum of weight value α2 and weight value α1 is 1. For example, weight value α2 can be 0.7 or 0.8, and weight value α1 can be 0.3 or 0.2.

[0126] Human profile 1 corresponds to three target similarities. A first number of target similarities greater than the second similarity threshold T1 can be counted. Assuming all three target similarities are greater than the second similarity threshold T1, the first number is 3. A second number of target similarities greater than the third similarity threshold T2 can also be counted. Assuming only one target similarity is greater than the third similarity threshold T2, the second number is 1. Then, based on the first number and weight value α1, the second number and weight value α2, and the total number N of human cover images corresponding to human profile 1 (i.e., the total number of target similarities corresponding to human profile 1 is 3), the number of targets corresponding to human profile 1 can be determined. For example, the following formula can be used to determine the number of targets: S = (α1 × S1 + α2 × S2) ÷ N.

[0127] In the above formula, S represents the target number corresponding to body profile 1, S1 represents the first number, S2 represents the second number, and N represents the total number of body cover images corresponding to body profile 1. By dividing by the total number of body cover images N, we avoid the problem of a larger target number S as the value of N increases.

[0128] Human profile 2 corresponds to four target similarities. A first number of target similarities greater than the second similarity threshold T1 can be counted, and a second number of target similarities greater than the third similarity threshold T2 can be counted. The number of targets corresponding to human profile 2 is determined based on the first number and weight value α1, the second number and weight value α2, and the total number N of human cover images corresponding to human profile 2 (i.e., 4). The determination method is described in the above formula.

[0129] To sum up, the number of targets corresponding to each human body file can be obtained. Then, the human body file with the largest number of targets can be used as the target human body file, and the optimal human body image in the human body image sequence can be classified as the human body image corresponding to the target human body file, thus completing the classification process of the human body image sequence.

[0130] For example, after the optimal human image in the human image sequence is classified as the human image corresponding to the target human profile, the optimal human image can also be used as the human cover image corresponding to the target human profile. For example, if the quality of the optimal human image is relatively good, then this optimal human image can be used as the human cover image; if the quality of the optimal human image is relatively poor, then this optimal human image is not used as the human cover image. There is no restriction on this process.

[0131] It can be seen from the above technical solutions that in the embodiment of the present application, the target human file can be determined based on multiple frames of human images of the target object, and the optimal human image of the target object can be classified as the human image corresponding to the target human file, so as to classify the human images, improve the personnel classification effect, improve the human recognition effect, improve the accuracy and universality of image classification, and collaboratively realize the intelligent analysis of personnel archiving, so as to quickly obtain useful information from massive images. The personnel archiving effect is improved by means of optimal frame images, graph convolution multi-frame feature fusion, face and human body association information, etc. The optimal frame image uses attributes such as clarity, lighting, brightness, completeness, pitch angle, posture, area, and proximity to the edge to fuse into a score representing quality, and the one with the best score is selected as the optimal frame image. By selecting the optimal frame image and K frame candidate images that are relatively similar to the optimal frame image and fusing them in a graph convolution manner, the fused features are obtained to improve the human or face recognition effect. The archived human body is classified into the cover file with the highest similarity ratio.

[0132] Based on the same application concept as the above method, a human body image classification device is proposed in the embodiment of the present application, see Figure 6 FIG. 1 is a schematic structural diagram of the device, which may include:

[0133] An acquisition module 61 is used to select an optimal human body image from multiple frames of human body images of a target object, and select a candidate human body image from the multiple frames of human body images based on the initial similarity between the optimal human body image and each frame of human body image; a determination module 62 is used to determine the fusion features corresponding to the multiple frames of human body images of the target object based on the optimal human features corresponding to the optimal human body image and the candidate human features corresponding to the candidate human images; based on the human cover image corresponding to each created human body file, the target similarity between the fusion features and the cover features corresponding to the human cover image is determined; a classification module 63 is used to select a target human body file from all human body files based on the target similarity, and classify the optimal human body image of the target object as the human body image corresponding to the target human body file.

[0134] Exemplarily, when the acquisition module 61 selects the optimal human body image from multiple frames of human body images of the target object, it is specifically used to: for each frame of the human body image of the target object, determine the target frame where the target object is located from the human body image, and determine the score value corresponding to the human body image based on at least one feature information corresponding to the target frame and the weight value corresponding to each feature information; the feature information includes at least one of the following: clarity, lighting, brightness, completeness, pitch angle, posture, area, and proximity to the edge; and select the optimal human body image from the multiple frames of human body images based on the score value corresponding to each frame of the human body image.

[0135] Exemplarily, the acquisition module 61 is specifically used to select a candidate human image from multiple frames of human images based on the initial similarity between the optimal human image and each frame of human image: for each frame of human image in the multiple frames of human images except the optimal human image, determine the initial similarity between the optimal human image and the human image; if the initial similarity is greater than a first similarity threshold, select the human image as the candidate human image; or, if the initial similarity is not greater than the first similarity threshold, prohibit the human image from being selected as the candidate human image.

[0136] Exemplarily, when the determination module 62 determines the fusion features corresponding to multiple frames of human body images of the target object based on the optimal human features corresponding to the optimal human body image and the candidate human body features corresponding to the candidate human body image, it is specifically used to: establish a first neighbor matrix centered on the optimal human body features based on the optimal human body features and the candidate human body features, and establish a second neighbor matrix centered on the candidate human body features; input the first neighbor matrix to the trained target network model to obtain a first sub-feature, and input the second neighbor matrix to the target network model to obtain a second sub-feature; determine the fusion features based on the first sub-feature and the first weighting coefficient, the second sub-feature and the second weighting coefficient; wherein the first weighting coefficient is greater than the second weighting coefficient.

[0137] Exemplarily, the total number of candidate human images is M, where M is a positive integer greater than 1, and the M candidate human images correspond to M candidate human features; the determination module 62 establishes a first nearest neighbor matrix centered on the optimal human feature based on the optimal human feature and the candidate human feature, and establishes a second nearest neighbor matrix centered on the candidate human feature, which is specifically used to: sort the M candidate human features based on the similarity between the optimal human feature and each candidate human feature; select K candidate human features with large similarity based on the sorting result; establish the first nearest neighbor matrix based on the optimal human feature and the K candidate human features; for each candidate human feature, sort all reference human features based on the similarity between the candidate human feature and each reference human feature, all reference human features include the optimal human feature and the remaining candidate human features in the M candidate human features except the candidate human feature; select K reference human features with large similarity based on the sorting result; and establish a second nearest neighbor matrix corresponding to the candidate human feature based on the candidate human feature and the K reference human features.

[0138] Exemplarily, if a human body profile corresponds to N human cover images, and N is a positive integer, then the human body profile corresponds to N target similarities; the classification module 63 is specifically used to select a target human body profile from all human body profiles based on the target similarities: for each human body profile, based on the N target similarities corresponding to the human body profile, count a first number whose target similarity is greater than a second similarity threshold and a second number whose target similarity is greater than a third similarity threshold; determine the number of targets corresponding to the human body profile based on the first number and the weight value of the first number, the second number and the weight value of the second number; the third similarity threshold is greater than the second similarity threshold, and the weight value of the second number is greater than the weight value of the first number; select a target human body profile from all human body profiles based on the number of targets corresponding to each human body profile.

[0139] Exemplarily, the classification module 63 is further used to select a target human body image from multiple frames of human body images of the target object, where the target human body image includes the facial area of ​​the target object; select a target face file that matches the facial features corresponding to the facial area from all face files, generate a target human body file corresponding to the target object based on the target face file, and classify the optimal human body image of the target object as the human body image corresponding to the target human body file.

[0140] Exemplarily, after the classification module 63 classifies the optimal human body image of the target object as the human body image corresponding to the target human body profile, it is further configured to: use the optimal human body image as the human body cover image corresponding to the target human body profile.

[0141] Based on the same application concept as the above method, a back-end device is proposed in the embodiment of the present application, see Figure 7 As shown, the back-end device includes a processor 71 and a machine-readable storage medium 72, wherein the machine-readable storage medium 72 stores machine-executable instructions that can be executed by the processor 71; the processor 71 is used to execute the machine-executable instructions to implement the human body image classification method disclosed in the above example of the application.

[0142] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the human body image classification method disclosed in the above example of the present application can be implemented.

[0143] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0144] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0145] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0146] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0148] Furthermore, these computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0150] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for classifying human body images, characterized in that: The method comprises: Selecting an optimal human body image from multiple frames of human body images of the target object, and selecting a candidate human body image from the multiple frames of human body images based on an initial similarity between the optimal human body image and each frame of human body image; Based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human image, establishing a first nearest neighbor matrix centered on the optimal human features and establishing a second nearest neighbor matrix centered on the candidate human features; Inputting the first neighbor matrix into the trained target network model to obtain a first sub-feature, and inputting the second neighbor matrix into the target network model to obtain a second sub-feature; Determining a fusion feature corresponding to multiple frames of human body images of the target object based on the first sub-feature and the first weighting coefficient, the second sub-feature and the second weighting coefficient; wherein the first weighting coefficient is greater than the second weighting coefficient; Based on the human body cover image corresponding to each created human body file, determining the target similarity between the fusion feature and the cover feature corresponding to the human body cover image; A target human body file is selected from all human body files based on the target similarity, and the optimal human body image of the target object is classified as the human body image corresponding to the target human body file.

2. The method according to claim 1, characterized in that The selecting the optimal human body image from multiple frames of human body images of the target object includes: For each human body image of the target object, a target frame in which the target object is located is determined from the human body image, and a score value corresponding to the human body image is determined based on at least one feature information corresponding to the target frame and a weight value corresponding to each feature information; wherein the feature information includes at least one of the following: clarity, lighting, brightness, completeness, pitch angle, posture, area, and proximity to the edge; An optimal human body image is selected from the multiple frames of human body images based on the score value corresponding to each frame of human body image.

3. The method according to claim 1, characterized in that The selecting a candidate human image from multiple frames of human images based on the initial similarity between the optimal human image and each frame of human image includes: For each human body image in the multiple human body images except the optimal human body image, an initial similarity between the optimal human body image and the human body image is determined; if the initial similarity is greater than a first similarity threshold, the human body image is selected as a candidate human body image; if the initial similarity is not greater than the first similarity threshold, the human body image is prohibited from being selected as a candidate human body image.

4. The method according to claim 1, wherein The total number of candidate human images is M, where M is a positive integer greater than 1, and the M candidate human images correspond to M candidate human features. The step of establishing, based on the optimal human feature and the candidate human feature, a first nearest neighbor matrix centered on the optimal human feature and a second nearest neighbor matrix centered on the candidate human feature includes: Sorting the M candidate human features based on the similarity between the optimal human feature and each candidate human feature; selecting K candidate human features with large similarity based on the sorting results; and establishing the first nearest neighbor matrix based on the optimal human feature and the K candidate human features; For each candidate human feature, all reference human features are sorted based on the similarity between the candidate human feature and each reference human feature, wherein all reference human features include the optimal human feature and the remaining candidate human features of the M candidate human features except the candidate human feature; K reference human features with large similarity are selected based on the sorting result; and a second nearest neighbor matrix corresponding to the candidate human feature is established based on the candidate human feature and the K reference human features.

5. The method according to claim 1, wherein If a human profile corresponds to N human cover images, and N is a positive integer, then the human profile corresponds to N target similarities; The selecting a target human profile from all human profiles based on the target similarity comprises: For each human profile, based on the N target similarities corresponding to the human profile, a first number of target similarities greater than a second similarity threshold and a second number of target similarities greater than a third similarity threshold are counted; based on the first number and a weight value of the first number, the second number and the weight value of the second number, the number of targets corresponding to the human profile is determined; wherein, if the third similarity threshold is greater than the second similarity threshold, the weight value of the second number is greater than the weight value of the first number; A target human profile is selected from all human profiles based on the number of targets corresponding to each human profile.

6. The method according to claim 1, wherein The method further includes: selecting a target human body image from multiple frames of human body images of the target object, the target human body image including a facial region of the target object; selecting a target face file that matches facial features corresponding to the facial region from all face files, generating a target human body file corresponding to the target object based on the target face file, and classifying the optimal human body image of the target object as the human body image corresponding to the target human body file; After classifying the optimal human body image of the target object as the human body image corresponding to the target human body file, the method further includes: using the optimal human body image as the human body cover image corresponding to the target human body file.

7. A human body image classification device, characterized in that: The device includes: an acquisition module for selecting an optimal human body image from multiple frames of human body images of a target object, and selecting a candidate human body image from the multiple frames of human body images based on an initial similarity between the optimal human body image and each frame of human body image; A determination module is configured to establish, based on the optimal human features corresponding to the optimal human image and the candidate human features corresponding to the candidate human images, a first neighbor matrix centered on the optimal human features and a second neighbor matrix centered on the candidate human features; input the first neighbor matrix into a trained target network model to obtain a first sub-feature, and input the second neighbor matrix into the target network model to obtain a second sub-feature; determine, based on the first sub-feature and a first weighting coefficient, the second sub-feature and a second weighting coefficient, a fusion feature corresponding to multiple frames of human images of the target object; wherein the first weighting coefficient is greater than the second weighting coefficient; and determine, based on the human cover image corresponding to each created human profile, a target similarity between the fusion feature and the cover feature corresponding to the human cover image; A classification module is used to select a target human body file from all human body files based on the target similarity, and classify the optimal human body image of the target object as the human body image corresponding to the target human body file.

8. The device according to claim 7, It is characterized by: in, The acquisition module is specifically configured to select the optimal human body image from multiple frames of human body images of the target object by: for each frame of the human body image of the target object, determining a target frame in which the target object is located from the human body image, and determining a score value corresponding to the human body image based on at least one feature information corresponding to the target frame and a weight value corresponding to each feature information; wherein the feature information includes at least one of the following: clarity, illumination, brightness, completeness, pitch angle, posture, area, and proximity to an edge; and selecting the optimal human body image from the multiple frames of human body images based on the score value corresponding to each frame of the human body image; Wherein, the acquisition module is specifically configured to select a candidate human image from multiple frames of human images based on the initial similarity between the optimal human image and each frame of human image: for each frame of human image other than the optimal human image in the multiple frames of human images, determine the initial similarity between the optimal human image and the human image; if the initial similarity is greater than a first similarity threshold, select the human image as the candidate human image; or if the initial similarity is not greater than the first similarity threshold, prohibit the human image from being selected as the candidate human image; Wherein, the total number of candidate human images is M, M is a positive integer greater than 1, and the M candidate human images correspond to M candidate human features; the determination module establishes a first nearest neighbor matrix centered on the optimal human feature based on the optimal human feature and the candidate human feature, and establishes a second nearest neighbor matrix centered on the candidate human feature, specifically for: sorting the M candidate human features based on the similarity between the optimal human feature and each candidate human feature; selecting K candidate human features with large similarity based on the sorting result; establishing the first nearest neighbor matrix based on the optimal human feature and the K candidate human features; for each candidate human feature, sorting all reference human features based on the similarity between the candidate human feature and each reference human feature, all reference human features including the optimal human feature and the remaining candidate human features in the M candidate human features except the candidate human feature; selecting K reference human features with large similarity based on the sorting result; and establishing a second nearest neighbor matrix corresponding to the candidate human feature based on the candidate human feature and the K reference human features; Among them, if the human body file corresponds to N human cover images, and N is a positive integer, then the human body file corresponds to N target similarities; when the classification module selects the target human body file from all human body files based on the target similarities, it is specifically used to: for each human body file, based on the N target similarities corresponding to the human body file, count a first number whose target similarity is greater than a second similarity threshold and a second number whose target similarity is greater than a third similarity threshold; based on the first number and the weight value of the first number, the second number and the weight value of the second number, determine the number of targets corresponding to the human body file; the third similarity threshold is greater than the second similarity threshold, and the weight value of the second number is greater than the weight value of the first number; select the target human body file from all human body files based on the number of targets corresponding to each human body file; The classification module is further configured to select a target human image from the multiple human images of the target object, the target human image including a facial region of the target object; select a target facial profile that matches facial features corresponding to the facial region from all facial profiles, generate a target human profile corresponding to the target object based on the target facial profile, and classify the optimal human image of the target object as the human image corresponding to the target human profile; Wherein, after the classification module classifies the optimal human body image of the target object as the human body image corresponding to the target human body file, it is further used to: use the optimal human body image as the human body cover image corresponding to the target human body file.

9. A backend device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method steps described in any one of claims 1-6.

Citation Information

Patent Citations

  • Information generation method and device

    CN108921138A

  • Image archiving method, device and equipment

    CN112528078A