Face and human body association method, electronic device, and storage medium

By determining the tracking trajectory and initial image of the target object in video images, and combining face and human body similarity scores, the association between the target face and human body is established, which solves the problem of low face and human body association accuracy in dense scenes and achieves higher association accuracy.

CN115457595BActive Publication Date: 2026-02-17ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210983599.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2026-02-17
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of face-body association for the same target object is low, especially in scenarios with dense target occlusion, where misassociations are prone to occur.

Method used

By acquiring the tracking trajectory of the target object in the video image and multiple initial images, the target face image and target body image of the target object are determined based on face similarity and body similarity, and the correlation is established, including quality score and similarity comparison.

Benefits of technology

It improves the accuracy of associating the target object's face with the human body, reduces false associations, and effectively enhances the accuracy of association, especially in dense scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457595B_ABST
    Figure CN115457595B_ABST
Patent Text Reader

Abstract

The application discloses a face and human body association method, an electronic device and a storage medium, wherein the face and human body association method comprises the following steps: acquiring a video image, determining a tracking trajectory of a target object in the video image and a plurality of initial images of the target object, and the initial images comprising an initial face image, an initial human body image and an initial face and human body image; in response to the tracking stability of the tracking trajectory being not greater than a stability threshold, determining a target face image and a target human body image of the target object based on face similarity and human body similarity between the plurality of initial images of the target object, and establishing an association relationship between the target face image and the target human body image. Through the above manner, the application can improve the association accuracy of the face and the human body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to methods for associating human faces with human bodies, electronic devices, and storage media. Background Technology

[0002] With the rapid development of science and technology and the arrival of the big data era, information security has become increasingly important. Image recognition, as a secure, contactless, convenient, user-friendly, and efficient method of identity authentication, has been widely applied to all aspects of social life.

[0003] The application scenarios of face and body association are becoming increasingly widespread. For example, in intelligent monitoring systems, due to issues such as the number of cameras, their placement, and image information, it is difficult to capture all faces; at a certain time, only human bodies may be captured. Even if no clear face is captured, the captured human body can still be searched in the face-body association database. After finding a matching human body, the associated face information can be obtained, thereby determining the identity of that human body.

[0004] The commonly used method at present is to determine whether there is a correlation between the face and the body of the target object through object tracking. However, in scenarios with dense object occlusion, object tracking methods can easily associate the face of the same target object with the body of other pedestrians, resulting in low accuracy in associating the face and body of the same target object. Summary of the Invention

[0005] This invention provides a method, electronic device, and storage medium for associating a face with a human body, in order to solve the problem of low accuracy in associating the face with the human body of the same target object.

[0006] To address the aforementioned technical problems, this invention provides a method for associating a face with a human body, comprising: acquiring a video image and determining the tracking trajectory of a target object in the video image and multiple initial images of the target object, the initial images including an initial face image, an initial human body image, and an initial face-human body image; in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, determining the target face image and the target human body image of the target object based on the face similarity and human body similarity among the multiple initial images of the target object, and establishing an association relationship between the target face image and the target human body image.

[0007] Specifically, in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, the target face image and target body image of the target object are determined based on the face similarity and body similarity among multiple initial images of the target object, and the association between the target face image and the target body image is established. This includes: performing quality scoring on each initial image of the target object to determine the quality score of each initial image; in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, determining multiple quality images from multiple initial images based on the quality scores of each initial image; and determining the target face image and target body image of the target object based on the face similarity and body similarity among multiple quality images of the target object, and establishing the association between the target face image and the target body image.

[0008] The process includes: scoring the quality of each initial image of the target object to determine its quality score; determining the quality scores of each initial face image, each initial human body image, and each initial face / human body image within each initial image; and determining multiple quality images from the initial images based on their quality scores, provided that the tracking stability of the tracking trajectory does not exceed a stability threshold. This includes determining the optimal face quality image and the optimal face quality image based on the quality scores of each initial face image, each initial human body image, and each initial face / human body image. The optimal image corresponds to the human body image, and the face and human body images are found in the best overall quality image of the face and human body. Based on the face similarity and human body similarity among multiple quality images of the target object, the target face image and target human body image of the target object are determined, and the association between the target face image and the target human body image is established. This includes: based on the face similarity and human body similarity among the best quality face image of the target object, the human body image corresponding to the best quality face image, and the face and human body images in the best overall quality image of the face and human body, the target face image and target human body image of the target object are determined, and the association between the target face image and the target human body image is established.

[0009] The method of determining the target face image and target body image of the target object based on the optimal face image of the target object, the body image corresponding to the optimal face image, and the face and body similarity between the face image and body image in the optimal combined face and body image includes: in response to the face similarity between the optimal face image and the face image in the optimal combined face and body image not being less than a preset face similarity, determining the face image in the optimal face image and the face image in the optimal combined face and body image as the target face image of the target object, and determining whether the body similarity between the body image corresponding to the optimal face image and the body image in the optimal combined face and body image is less than a preset body similarity; in response to the body similarity between the body image corresponding to the optimal face image and the body image in the optimal combined face and body image being less than a preset body similarity, determining the body image in the optimal combined face and body image as the target body image of the target object, and determining the body image corresponding to the optimal face image as the body image of other target objects.

[0010] The method of determining the target face image and target body image of the target object based on the optimal face image of the target object, the body image corresponding to the optimal face image, and the face and body similarity between the face image and body image in the optimal combined face and body image includes: in response to the face similarity between the optimal face image and the face image in the optimal combined face and body image being less than a preset face similarity, determining the face image in the optimal combined face and body image as the target face image of the target object, and determining the body image corresponding to the optimal face image and the body image in the optimal combined face and body image. If the similarity between the human bodies in the images is less than the preset similarity, and the similarity between the human body image corresponding to the best-quality face image and the human body image in the best-quality face-body composite image is less than the preset similarity, then the human body image in the best-quality face-body composite image is determined as the target human body image of the target object; if the similarity between the human body image corresponding to the best-quality face image and the human body image in the best-quality face-body composite image is not less than the preset similarity, then the human body image in the best-quality face-body composite image and the human body image corresponding to the best-quality face image are determined as the target human body images of the target object.

[0011] The process includes determining the optimal face image, the corresponding human body image, and the face and human body images of the target object based on the quality scores of each initial face image, each initial human body image, and each initial face and human body image. It also includes determining the human body quality image of the target object from multiple initial images based on the quality scores of each initial human body image in each initial image; determining the target face image and target human body image of the target object based on the face similarity and human body similarity between the optimal face image, the corresponding human body image, and the face and human body images in the optimal face and human body image. Furthermore, it includes determining the human body quality image as the target human body image of the target object in response to a human body similarity greater than a preset human body similarity.

[0012] Specifically, in response to the tracking stability of the tracking trajectory not being greater than a stability threshold, multiple quality images are determined from multiple initial images based on the quality scores of each initial image, including: determining the tracking stability of the tracking trajectory based on the tracking trajectory; and in response to the tracking stability not being greater than a stability threshold, determining the optimal face quality image of the target object, the body image corresponding to the optimal face quality image, and the face image and body image in the optimal combined face and body quality image based on the quality scores of each initial face image, each initial body image, and each initial face and body image.

[0013] The process involves determining multiple quality images from multiple initial images based on their quality scores. This includes: sorting the initial face images, initial human body images, and initial face-human body images in descending order of quality score to obtain a face image sequence, a human body image sequence, and a face-human body image sequence; extracting face features from the first preset number of initial face images in the face image sequence to obtain a set of face image features including multiple face images; clustering the face image feature set to obtain at least one face image cluster; determining the face image with the highest quality score in the largest face image cluster as the optimal face quality image and obtaining the corresponding human body image; determining the initial human body images in the human body image sequence that meet the quality conditions as human body quality images; and determining the initial face-human body image with the highest quality score in the face-human body image sequence as the optimal face-human body overall quality image, thus obtaining the face image and human body image in the optimal face-human body overall quality image.

[0014] The human body quality images include: frontal human body quality images, side human body quality images, and back human body quality images; the initial human body images in the human body image sequence that meet the quality conditions are determined as human body quality images, including: determining the frontal human body images in the human body image sequence that meet the quality conditions as frontal human body quality images; determining the side human body images in the human body image sequence that meet the quality conditions as side human body quality images; and determining the back human body images in the human body image sequence that meet the quality conditions as back human body quality images; wherein, the quality conditions include the highest quality score or the quality score exceeding a preset quality score.

[0015] The process involves acquiring video images and determining the tracking trajectory of the target object in the video images, as well as multiple initial images of the target object. The initial images include face image frames, initial human body images, and initial face and human body images. This includes: performing target tracking on the video images to determine the tracking trajectory of the target object in the video images; and performing target detection on each image frame of the video images to determine multiple initial images of the target object.

[0016] To address the aforementioned technical problems, the present invention also provides an electronic device comprising: a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the face-body association method described above.

[0017] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium storing program data that can be executed to implement the face-body association method as described above.

[0018] The beneficial effects of this invention are as follows: Unlike the prior art, this invention acquires video images, determines the tracking trajectory of the target object in the video images and multiple initial images of the target object, and then, in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, determines the target face image and target body image of the target object based on the face similarity and body similarity between the multiple initial images of the target object, and establishes the association relationship between the target face image and the target body image. Based on the initial images, it can effectively reduce the misassociation of faces and bodies through similarity comparison, thereby improving the accuracy of the association between the face and body of the target object. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating an embodiment of the method for associating a face with a human body provided by the present invention;

[0020] Figure 2 This is a flowchart illustrating another embodiment of the method for associating a face with a human body provided by the present invention;

[0021] Figure 3 This is a schematic diagram of the framework of an embodiment of the face and body association device of the present invention;

[0022] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device provided by the present invention;

[0023] Figure 5 This is a schematic diagram of an embodiment of the computer-readable storage medium provided by the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the method for associating a face with a human body provided by the present invention.

[0026] Step S11: Obtain the video image and determine the tracking trajectory of the target object in the video image and multiple initial images of the target object.

[0027] In one specific application scenario, a fixed surveillance camera can capture video images over a period of time. In another specific application scenario, a mobile camera, such as a smart mobile robot or a handheld camera, can also capture video images over a period of time. The specific method of acquiring video images is not limited here.

[0028] After acquiring the video image, the tracking trajectory of the target object in the video image and multiple initial images of the target object are obtained. In this embodiment, the target object refers to a specific person. The face and body association method of this embodiment determines the face and body images of this person in the video image, thereby facilitating subsequent steps such as feature analysis or identification of this person.

[0029] In one specific application scenario, the DeepSort object tracking algorithm, based on deep learning, can be used to track target objects in video images and determine their tracking trajectory. Specifically, this algorithm uses an appearance feature extraction network to perform ReID (Re-identification) on the detection boxes, and then uses Kalman filtering and the Hungarian algorithm to associate the current target detection box with the tracking trajectory. In another specific application scenario, object tracking algorithms can also be used to track target objects in video images and determine their tracking trajectory. Target tracking methods can include region matching, feature point tracking, active contour-based tracking algorithms, optical flow methods, etc. The most commonly used method is feature matching, which first extracts target features and then finds the most similar features in subsequent frames for target localization. Commonly used features include SIFT features, SURF features, and Harris corner points.

[0030] In another specific application scenario, it can also receive manual confirmation and marking of target objects in video images to obtain the tracking trajectory of the target objects. The specific method for obtaining the tracking trajectory of the target objects is not limited here.

[0031] The initial image of the target object includes the initial face image, the initial body image, and the initial face-body image. Specifically, the initial face image refers to a partial image of the target object's face within a video frame; the initial body image refers to a partial image of the target object's body within a video frame; and the initial face-body image refers to a partial image of both the target object's face and body within a video frame. For example, if a video frame contains five faces, one of which is the target object's face, then the partial image containing that face is the target object's initial face image.

[0032] In one specific application scenario, a trained object detection model can be used to detect target objects in a video image to determine the initial image of the target object. In another specific application scenario, a trained image recognition model can be used to recognize images in each frame of a video image to determine the initial image of the target object. In yet another specific application scenario, object detection algorithms such as Two-Stage and One-Stage algorithms can be used to first detect target objects in the video image, and then track the bounding boxes to determine the initial image of the target object. The specific method for obtaining the initial image is not limited here.

[0033] In particular, when this embodiment is applied to scenarios with dense pedestrian activity, where there is occlusion between people and a high degree of overlap between human bodies, such as waiting halls, subway stations, shopping malls, and other important event venues, the method for determining the initial image of the target object may easily misjudge the initial image of a non-target object into the initial image of the target object due to the large flow of people and dense crowds. Therefore, this embodiment also requires adjustment of the initial image.

[0034] Step S12: In response to the tracking stability of the tracking trajectory not being greater than the stability threshold, determine the target face image and target body image of the target object based on the face similarity and body similarity among multiple initial images of the target object, and establish the association relationship between the target face image and the target body image.

[0035] In response to the tracking stability of the tracking trajectory not exceeding a stability threshold, the target face image and target body image of the target object are determined based on the face similarity and body similarity corresponding to multiple initial images of the target object. The specific value of the stability threshold can be set based on actual needs and is not limited here.

[0036] In one specific implementation, in response to the tracking trajectory exceeding a stability threshold, each face and each body of the target object in the initial face image, initial body image, and initial face and body image of the initial image can be used as the target face image and target body image of the target object. In another specific implementation, in response to the tracking trajectory exceeding a stability threshold, the face and body with the best feature quality in the initial face image, initial body image, and initial face and body image of the initial image can be used as the target face image and target body image of the target object. In this embodiment, "greater than" includes both "greater than" and "equal to".

[0037] The tracking stability can be determined by the target tracking confidence when the tracking trajectory of the target is in the deletion state. The target tracking confidence is determined by the tracking algorithm by judging the number of consecutive frames the target is occluded or the number of consecutive frames the target is lost. The confidence is initially 1 and is reduced to 0, with a value range of [0, 1].

[0038] One method to determine the facial and human body similarity among multiple initial images is to calculate the Euclidean distance, cosine distance, Manhattan distance, Pearson correlation coefficient, Minkowski distance, etc.

[0039] In a specific application scenario, the similarity between each initial face image and each face in each initial face and body image can be calculated, and the target face image of the target object can be determined based on the similarity between each face. Similarly, the similarity between each initial body image and each body in each initial face and body image can be calculated, and the target body image of the target object can be determined based on the similarity between each body.

[0040] In another specific application scenario, the quality of each initial image can be selected first, and the target face image and target body image of the target object can be determined based on the face similarity and body similarity between the initial images that meet the quality requirements.

[0041] The target face image and target body image of the target object are identified, and the association between the target face image and the target body image is established, that is, the target face image and the target body image are identified as images of the same target object, thus completing the association between the face and body of the target object.

[0042] Through the above steps, the face and body association method of this embodiment obtains video images, determines the tracking trajectory of the target object in the video images and multiple initial images of the target object, and then, in response to the tracking stability of the tracking trajectory not being greater than the stability threshold, determines the target face image and target body image of the target object based on the face similarity and body similarity between the multiple initial images of the target object, and establishes the association relationship between the target face image and the target body image. Based on the initial images, it can effectively reduce the misassociation of faces and bodies through similarity comparison, thereby improving the accuracy of the face and body association of the target object.

[0043] In other embodiments, the method for associating a face with a human body may further include:

[0044] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the method for associating a face with a human body provided by the present invention.

[0045] Step S21: Obtain the video image and determine the tracking trajectory of the target object in the video image and multiple initial images of the target object.

[0046] Once a video image is acquired, target tracking can be performed on the video image to determine the tracking trajectory of the target object in the video image; and target detection can be performed on each image frame of the video image to determine multiple initial images of the target object.

[0047] In a specific application scenario, target detection can be performed on each frame of a video image to obtain the faces and bodies of all people in the video image. Then, the tracking identifiers and tracking trajectories of each face and body can be obtained through target tracking methods. Based on the tracking identifiers and tracking trajectories of each face and body, the tracking trajectory of the target object and multiple initial images of the target object can be determined.

[0048] The initial image of the target object includes an initial face image, an initial human body image, and an initial face and human body image.

[0049] In a specific application scenario, this embodiment can divide the motion state of a target object in a video image into four states: target creation, target update, target loss, and target deletion. Each target has a tracking identifier that distinguishes it from other targets. When it is determined that the target object is in the target deletion state, the tracking trajectory of the target object and multiple initial images of the target object are determined during the time period from target creation to deletion in the video image.

[0050] Step S22: Perform quality scoring on each initial image of the target object to determine the quality score of each initial image.

[0051] Each initial image of the target object is scored to determine the quality score of each face image, each body image, and each face and body image of the target object in each initial image.

[0052] In a specific application scenario, a deep learning multi-label classification algorithm can be used. Using preset attributes of each face image, each body image, and each face / body image as training labels, a quality scoring model is trained based on a convolutional neural network. This model is then used to score the quality of each face image, each body image, and each face / body image based on its preset attributes. These preset attributes may include the integrity of the face or body in the image, whether the face or body is occluded, image resolution, the size of the face or body in the image, the sharpness of the face / body target image, and the presence of noise interference, etc., without further limitation.

[0053] In another specific application scenario, it is also possible to manually score the quality of each face image, each body image, and each face and body image to obtain the quality scores of each face image, each body image, and each face and body image of the target object in each initial image.

[0054] Step S23: In response to the tracking stability of the tracking trajectory not being greater than the stability threshold, determine multiple quality images from multiple initial images based on the quality scores of each initial image.

[0055] It determines whether the tracking trajectory of the target object exceeds a stability threshold. Specifically, the tracking stability can be determined based on the tracking trajectory. The specific value of the stability threshold can be set based on actual needs and is not limited here.

[0056] The tracking stability can be determined by the target tracking confidence when the tracking trajectory of the target is in the deletion state. The target tracking confidence is determined by the tracking algorithm by judging the number of consecutive frames the target is occluded or the number of consecutive frames the target is lost. The confidence is initially 1 and is reduced to 0, with a value range of [0, 1].

[0057] In response to the tracking stability of the tracking trajectory not exceeding a stability threshold, multiple quality images are determined from multiple initial images based on the quality scores of each initial image. Conversely, in response to the tracking trajectory exceeding the stability threshold, the faces and bodies of the target object within the initial face image, initial body image, and initial face / body image of the initial image can be used as the target face image and target body image of the target object. In another specific implementation, in response to the tracking trajectory exceeding the stability threshold, the face and body images with the best quality score or exceeding a preset quality score among the initial face image, initial body image, and initial face / body image of the initial image can be used as the target face image and target body image of the target object. The preset quality score can be set based on actual conditions and is not limited here.

[0058] In one specific implementation, in response to the tracking stability of the tracking trajectory not being greater than a stability threshold, it can be determined whether the quality scores of each initial face image, each initial human body image, and each initial face and human body image exceed a preset quality score, and the images exceeding the preset quality score are determined as quality images.

[0059] In another specific implementation, the multiple quality images may include the optimal face quality image F. best The human body image B corresponding to the best quality face image F-best The face image F in the best overall image of face and body 2-best and human body image B 2-best Among them, the image with the best face quality is F. best The image with the highest quality score among the initial face images is F, which is the face image with the best overall face and body quality. 2-best and human body image B 2-best This is the image of the face and body portion from the image with the highest quality score in the initial face and body image set.

[0060] That is, in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, the optimal face image F of the target object is determined based on the quality scores of each initial face image, each initial human body image, and each initial face and human body image. best The human body image B corresponding to the best quality face image F-best The face image F in the best overall image of face and body 2-best and human body image B 2-best In a specific application scenario, multiple quality images also include multiple human body quality images B. best That is, based on the quality scores of each initial human body image in each initial image, multiple human body quality images B of the target object can be determined from multiple initial images. best Among them, human body quality image B best The image with the highest quality score among the initial human body images.

[0061] Specifically, determine the optimal face image F. best The method can be as follows: Sort each initial face image, each initial human body image, and each initial face and human body image in descending order of quality score to obtain a face image sequence, a human body image sequence, and a face and human body image sequence; extract face features from the first preset number of initial face images in the face image sequence to obtain a face image feature set including multiple face images; cluster the face image feature set to obtain at least one face image cluster; and determine the face image with the highest quality score in the largest face image cluster as the face quality optimal image F. best And obtain the human body image B corresponding to the best face quality image. F-best Among them, the human body image B corresponding to the best face quality image. F-best This refers to the image F with the best face quality. best Human images that coexist in the same image frame and correspond to the same target object, wherein the correspondence is obtained by target detection in step S21, and there may be errors.

[0062] The above determines the optimal face quality image F best The method first selects a preset number of initial face images from the face image sequence, which can improve the accuracy of the source of the target face image of the target object, thereby improving the accuracy of the final target face image of the target object. In addition, face recognition and clustering are performed on the preset number of face images to exclude other faces that are misidentified in the target detection and target tracking in step S11, improve the targeting of the face image clusters to the target object, and further improve the accuracy of the final target face image of the target object.

[0063] Determine the face image F in the image with the best overall quality of face and body.2-best and human body image B 2-best and human body quality image B best One possible method is to determine the initial human image with the highest quality score in the human image sequence as human quality image B. best ; and the initial face and body image with the highest quality score in the face and body image sequence is determined as the face and body image with the best overall quality, thus obtaining the face image F in the face and body image with the best overall quality. 2-best and human body image B 2-best .

[0064] In one specific implementation, human body quality image B best It includes multi-pose human body images, including: frontal human body quality images, side human body quality images, and back human body quality images.

[0065] Specifically, a frontal human image that meets the quality criteria in a human image sequence can be defined as a frontal human quality image; a side human image that meets the quality criteria in a human image sequence can be defined as a side human quality image; and a back human image that meets the quality criteria in a human image sequence can be defined as a back human quality image. The quality criteria include having the highest quality score or a quality score exceeding a preset quality score. A frontal human image refers to an image of a human body facing the camera of the video image from the front; a side human quality image refers to an image of a human body facing the camera of the video image from the side; and a back human quality image refers to an image of a human body facing the camera of the video image from the back.

[0066] Among them, the optimal face image F of the target object obtained in this step best The human body image B corresponding to the best quality face image F-best The face image F in the best overall image of face and body 2-best and human body image B 2-best And multiple human body quality images B best The correspondence between the target object and the target object is obtained by target detection in step S21, which may contain errors. This embodiment eliminates these errors through step S24.

[0067] Step S24: Based on the face similarity and body similarity among multiple quality images of the target object, determine the target face image and the target body image of the target object, and establish the association between the target face image and the target body image.

[0068] When comparing similarity, feature vectors of each quality image can be obtained through feature extraction. The similarity between multiple quality images can be determined by using the feature vectors of each quality image and calculating similarity in step S11.

[0069] In a specific application scenario, after obtaining multiple quality images of a target object, the target face image and target body image of the target object can be determined based on the face similarity and body similarity among these multiple quality images. Furthermore, a correlation can be established between the target face images and the target body images. For example, the similarity of each face image and each body image among quality images exceeding a preset quality score can be compared. At least two face images with a similarity exceeding a preset face similarity score are identified as the target face images of the target object, and at least two body images with a similarity exceeding a preset body similarity score are identified as the target body images of the target object. The specific values ​​of the preset face similarity and preset body similarity scores can be set based on actual conditions and are not limited here.

[0070] In a specific application scenario, the optimal image F of the target object's face can be used as a basis. best The human body image B corresponding to the best quality face image F-best The face image F in the best overall image of face and body 2-best and human body image B 2-best The similarity between faces and bodies is used to determine the target face image and the target body image of the target object, and to establish the association between the target face image and the target body image.

[0071] In one specific implementation, in response to the optimal face quality image F best Face image F in the best overall quality image of face and body 2-best The facial similarity between the two images is no less than the preset facial similarity. The image with the best facial quality, F, is selected. best Face image F in the best overall quality image of face and body 2-best The target face image is identified as the target object; the response is the human body image B corresponding to the face image with the best quality. F-best Human image B in the best overall quality image of face and body 2-best The similarity between the human bodies is less than the preset similarity, so the human body image B in the image with the best overall quality of face and human body is selected. 2-best The target human image is identified as the target object. The human image B corresponding to the highest quality face image is selected. F-best The human image identified as another target object generates a new label. This is in response to the human image B corresponding to the highest quality face image. F-best Human image B in the best overall quality image of face and body 2-best The similarity between the human bodies is no less than the preset similarity between human bodies, and the human body image B corresponding to the best quality face image is selected. F-best Human image B in the best overall quality image of face and body2-best The target human image that has been identified as the target object.

[0072] In one specific implementation, in response to the optimal face quality image F best Face image F in the best overall quality image of face and body 2-best If the facial similarity between two images is less than the preset facial similarity, then the facial image F from the image with the best overall facial and human body quality will be selected. 2-best The target face image is identified as the primary target. The image with the highest face quality, F... best The face image identified as another target object generates a new label. This is in response to the human body image B corresponding to the best-quality face image. F-best Human image B in the best overall quality image of face and body 2-best The similarity between the human bodies is no less than the preset similarity between human bodies, and the human body image B corresponding to the best quality face image is selected. F-best Human image B in the best overall quality image of face and body 2-best All are identified as target human images of the same object. F-best Human image B in the best overall quality image of face and body 2-best The similarity between the human bodies is less than the preset similarity, so the human body image B in the image with the best overall quality of face and human body is selected. 2-best The target human image that has been identified as the target object.

[0073] In one specific implementation, when acquiring quality images, multiple human body quality images B of the target object can also be determined. best When the similarity between the human body quality image and the target human body image of the target object is greater than the preset similarity, the human body quality image is determined as the target human body image of the target object.

[0074] In a specific application scenario, individual body quality images B can be used. best Human body images B corresponding to the best quality face image F-best Human image B in the best overall quality image of face and body 2-best Perform similarity comparison when the human body quality image B best If the similarity of the human body to any of the human body images exceeds a preset similarity threshold, then the human body quality image B will be... best Assign the same label to human images whose similarity to human bodies exceeds the preset similarity.

[0075] In a specific application scenario, B can be any human body quality image whose similarity to the target human body image is no greater than a preset human body similarity.best Once a set of unclassified human images is identified, the human image with the highest quality score in this set is compared with the remaining human images in the set. If the similarity between the remaining human image and the highest quality score human image is higher than a certain threshold, then both images are classified as representing the same target object and are assigned the same label. If the similarity is not higher than the threshold, then both images are classified as representing different target objects and are assigned different labels. These labels can be the tracking labels acquired during target tracking, i.e., the initial labels. This process is repeated for the remaining unclassified human images until the last human image B is identified. best To be labeled.

[0076] After identifying the target face image and the target body image of the target object, establish the association between the target face image and the target body image, that is, determine that the target face image and the target body image are the same body and face images of the same target object, and complete the association between the face and body of the target object.

[0077] By using the aforementioned method of associating faces and bodies, erroneously associated images from multiple initial images of the target object in step S21 can be filtered out, thus determining the true target face and body images of the target object in the video image, thereby improving the accuracy of face-body association for the same target object. Furthermore, this method can associate the faces and bodies of all target objects in the video image, enabling the association of faces and bodies among target objects in dense scenes.

[0078] Through the above steps, this embodiment obtains video images for face and body association, determines the tracking trajectory of the target object in the video images and multiple initial images of the target object, performs quality scoring on each initial image of the target object, determines the quality score of each initial image, and, in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, determines multiple quality images from the multiple initial images based on the quality scores of each initial image. Finally, based on the face similarity and body similarity among the multiple quality images of the target object, the target face image and target body image of the target object are determined, and the association relationship between the target face image and the target body image is established. This approach, based on the initial images, first performs quality screening through quality scoring to obtain multiple quality images, and then compares the similarity of the quality images, effectively reducing falsely associated faces and bodies, thereby improving the accuracy of face and body association of the target object. Furthermore, this embodiment determines the true target face image and target body image of the target object through logical judgment of similarity, without introducing additional complex algorithms, which can improve the accuracy of face and body association while maintaining the efficiency of face and body association.

[0079] Please see Figure 3 , Figure 3 This is a schematic diagram of the framework of an embodiment of the image recognition device of the present invention. The face and body association device 30 includes an acquisition module 31 and a determination module 32. The acquisition module 31 is used to acquire a video image and determine the tracking trajectory of the target object in the video image and multiple initial images of the target object, including an initial face image, an initial body image, and an initial face and body image; the determination module 32 is used to determine the target face image and the target body image of the target object based on the face similarity and body similarity between the multiple initial images of the target object in response to the tracking stability of the tracking trajectory not being greater than a stability threshold, and to establish the association relationship between the target face image and the target body image.

[0080] The determination module 32 is also used to perform quality scoring on each initial image of the target object and determine the quality score of each initial image; in response to the tracking stability of the tracking trajectory not being greater than the stability threshold, it determines multiple quality images from multiple initial images based on the quality scores of each initial image; it determines the target face image and target body image of the target object based on the face similarity and body similarity between the multiple quality images of the target object, and establishes the association relationship between the target face image and the target body image.

[0081] The determination module 32 is also used to perform quality scoring on each initial image of the target object, and determine the quality scores of each initial face image, each initial human body image, and each initial face and human body image in each initial image; in response to the tracking stability of the tracking trajectory not being greater than the stability threshold, based on the quality scores of each initial face image, each initial human body image, and each initial face and human body image, determine the optimal face image of the target object, the human body image corresponding to the optimal face image, and the face and human body images in the optimal face and human body combined image; based on the face similarity and human body similarity between the optimal face image of the target object, the human body image corresponding to the optimal face image, and the face and human body images in the optimal face and human body combined image, determine the target face image and the target human body image of the target object, and establish the association relationship between the target face image and the target human body image.

[0082] The determining module 32 is further configured to, in response to the fact that the face similarity between the face image in the optimal face quality image and the optimal face-body composite image is not less than a preset face similarity, determine the face image in the optimal face quality image and the optimal face-body composite image as the target face image of the target object, and determine whether the body similarity between the body image corresponding to the optimal face quality image and the body image in the optimal face-body composite image is less than a preset body similarity; in response to the fact that the body similarity between the body image corresponding to the optimal face quality image and the body image in the optimal face-body composite image is less than a preset body similarity, determine the body image in the optimal face-body composite image as the target body image of the target object, and determine the body image corresponding to the optimal face quality image as the body image of other target objects.

[0083] The determining module 32 is further configured to, in response to a face similarity less than a preset face similarity between the face image in the optimal face quality image and the face image in the optimal face-body composite quality image, determine the face image in the optimal face-body composite quality image as the target face image of the target object, and determine whether the body similarity between the body image corresponding to the optimal face quality image and the body image in the optimal face-body composite quality image is less than a preset body similarity; in response to a body similarity less than a preset body similarity between the body image corresponding to the optimal face quality image and the body image in the optimal face-body composite quality image, determine the body image in the optimal face-body composite quality image as the target body image of the target object; and in response to a body similarity not less than a preset body similarity between the body image corresponding to the optimal face quality image and the body image in the optimal face-body composite quality image, determine the body image in the optimal face-body composite quality image and the body image corresponding to the optimal face quality image as the target body image of the target object.

[0084] The determining module 32 is further configured to determine the human body quality image of the target object from multiple initial images based on the quality scores of each initial human body image in each initial image; and to determine the human body quality image as the target human body image of the target object in response to the human body similarity between the human body quality image and the target human body image of the target object being greater than the human body preset similarity.

[0085] The determination module 32 is also used to determine the tracking stability of the tracking trajectory based on the feature quality of the target object in the video image corresponding to the tracking trajectory; in response to the tracking stability not being greater than the stability threshold, it determines the optimal face quality image of the target object, the human body image corresponding to the optimal face quality image, and the face image and human body image in the optimal combined face and human body image based on the quality scores of each initial face image, each initial human body image, and each initial face and human body image.

[0086] The determining module 32 is further configured to sort each initial face image, each initial human body image, and each initial face and human body image in descending order of quality score, thereby obtaining a face image sequence, a human body image sequence, and a face and human body image sequence; extract face features from the first preset number of initial face images in the face image sequence, thereby obtaining a face image feature set including multiple face images; cluster the face image feature set to obtain at least one face image cluster; determine the face image with the highest quality score in the largest face image cluster as the face quality optimal image, and obtain the human body image corresponding to the face quality optimal image; determine the initial human body image in the human body image sequence that meets the quality conditions as the human body quality image; and determine the initial face and human body image with the highest quality score in the face and human body image sequence as the face and human body comprehensive quality optimal image, thereby obtaining the face image and human body image in the face and human body comprehensive quality optimal image.

[0087] The determining module 32 is further configured to determine frontal human images that meet quality conditions in the human image sequence as frontal human quality images; and to determine side human images that meet quality conditions in the human image sequence as side human quality images; and to determine back human images that meet quality conditions in the human image sequence as back human quality images; wherein, the quality conditions include the highest quality score or a quality score exceeding a preset quality score. The human quality images include: frontal human quality images, side human quality images, and back human quality images.

[0088] Acquisition module 31: performs target tracking on video images to determine the tracking trajectory of the target object in the video images; and performs target detection on each image frame of the video images to determine multiple initial images of the target object.

[0089] The above method can improve the accuracy of the association between faces and bodies.

[0090] Based on the same inventive concept, the present invention also proposes an electronic device capable of executing the face-body association method of any of the above embodiments. Please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of an embodiment of the electronic device provided by the present invention. The electronic device includes a processor 41 and a memory 42.

[0091] The processor 41 is used to execute the program instructions stored in the memory 42 to implement the steps of any of the above-described methods for associating a face with a human body. In a specific implementation scenario, the electronic device may include, but is not limited to, a microcomputer or a server. In addition, the electronic device may also include mobile devices such as laptops and tablets, which are not limited here.

[0092] Specifically, processor 41 controls itself and memory 42 to implement the steps of any of the above embodiments. Processor 41 may also be referred to as a CPU (Central Processing Unit). Processor 41 may be an integrated circuit chip with signal processing capabilities. Processor 41 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 41 may be implemented using integrated circuit chips.

[0093] The above method can improve the accuracy of the association between faces and bodies.

[0094] Based on the same inventive concept, the present invention also proposes a computer-readable storage medium, please refer to [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of an embodiment of a computer-readable storage medium provided by the present invention. The computer-readable storage medium 50 stores at least one program data 51, which is used to implement any of the methods described above. In one embodiment, the computer-readable storage medium 50 includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0095] In the several embodiments provided by this invention, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0097] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium.

[0099] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

[0100] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A method for associating a human face with a human body, characterized in that, include: The video image is acquired, and the tracking trajectory of the target object in the video image and multiple initial images of the target object are determined. The initial images include an initial face image, an initial human body image, and an initial face and human body image. In response to the tracking stability of the tracking trajectory not being greater than a stability threshold, the target face image and target body image of the target object are determined based on the face similarity and body similarity among multiple initial images of the target object, and the association between the target face image and the target body image is established. Specifically, each initial image of the target object is scored to determine the quality score of each initial face image, each initial human body image, and each initial face-human body image in each initial image; in response to the tracking stability of the tracking trajectory not exceeding a stability threshold, based on the quality scores of each initial face image, each initial human body image, and each initial face-human body image, the optimal face quality image of the target object, the human body image corresponding to the optimal face quality image, and the face image and human body image in the optimal overall face-human body quality image of the target object are determined; based on the face similarity and human body similarity between the optimal face quality image, the human body image corresponding to the optimal face quality image, and the face image and human body image in the optimal overall face-human body quality image of the target object, the target face image and the target human body image are determined, and the association between the target face image and the target human body image is established.

2. The method for associating a face with a human body according to claim 1, characterized in that, The determination of the target face image and target body image of the target object based on the optimal face quality image of the target object, the corresponding body image of the optimal face quality image, and the face similarity and body similarity between the face image and body image in the optimal combined face and body quality image of the target object includes: In response to the fact that the face similarity between the face image with the best face quality and the face image with the best overall face and body quality is not less than the preset face similarity, the face image with the best face quality and the face image with the best overall face and body quality are determined as the target face image of the target object, and it is determined whether the human body similarity between the human body image corresponding to the face image with the best face quality and the human body image with the best overall face and body quality is less than the preset human body similarity. In response to a situation where the similarity between the human body image corresponding to the optimal face quality image and the human body image in the optimal face-human body composite image is less than a preset similarity, the human body image in the optimal face-human body composite image is determined as the target human body image of the target object, and the human body image corresponding to the optimal face quality image is determined as the human body image of other target objects.

3. The method for associating a face with a human body according to claim 1, characterized in that, The determination of the target face image and target body image of the target object based on the optimal face quality image of the target object, the corresponding body image of the optimal face quality image, and the face similarity and body similarity between the face image and body image in the optimal combined face and body quality image of the target object includes: In response to the fact that the face similarity between the face image with the best face quality and the face image with the best overall face and body quality is less than the preset face similarity, the face image in the best overall face and body quality image is determined as the target face image of the target object, and it is determined whether the human body similarity between the human body image corresponding to the best face quality image and the human body image in the best overall face and body quality image is less than the preset human body similarity. In response to the fact that the similarity between the human body image corresponding to the optimal face quality image and the human body image in the optimal face and human body combined image is less than a preset similarity, the human body image in the optimal face and human body combined image is determined as the target human body image of the target object. In response to the condition that the similarity between the human body image corresponding to the optimal face quality image and the human body image in the optimal overall face and human body quality image is not less than a preset similarity, the human body image in the optimal overall face and human body quality image and the human body image corresponding to the optimal face quality image are determined as the target human body image of the target object.

4. The method for associating a face with a human body according to claim 2 or 3, characterized in that, The step of determining the optimal face image, the corresponding body image, and the face and body image in the optimal combined face and body image of the target object based on the quality scores of each of the initial face images, each of the initial body images, and each of the initial face and body images, further includes: Based on the quality scores of each initial human body image in each of the initial images, a human body quality image of the target object is determined from the plurality of initial images; The method of determining the target face image and target body image of the target object based on the optimal face image of the target object, the corresponding body image of the optimal face image, and the face and body similarity between the face image and body image in the optimal combined face and body image of the target object, further includes: In response to the fact that the similarity between the human body quality image and the target human body image of the target object is greater than a preset similarity, the human body quality image is determined as the target human body image of the target object.

5. The method for associating a face with a human body according to claim 1, characterized in that, In response to the tracking stability of the tracking trajectory not exceeding a stability threshold, a plurality of quality images are determined from the plurality of initial images based on the quality scores of each initial image, including: The tracking stability of the target object is determined based on the tracking trajectory; In response to the tracking stability not being greater than the stability threshold, the optimal face image, the corresponding body image, and the face and body images in the optimal combined face and body image of the target object are determined based on the quality scores of each initial face image, each initial body image, and each initial face and body image.

6. The method for associating a face with a human body according to claim 5, characterized in that, The step of determining multiple quality images from the plurality of initial images based on the quality scores of each initial image includes: Based on the quality scores from high to low, the initial face images, initial human body images, and initial face and human body images are sorted to obtain face image sequences, human body image sequences, and face and human body image sequences. Facial features are extracted from the first preset number of initial facial images in the facial image sequence to obtain a set of facial image features including multiple facial images; Cluster the set of facial image features to obtain at least one facial image cluster; The face image with the highest quality score in the largest cluster of face images is identified as the optimal face image, and the corresponding human image is obtained; and The initial human images in the human image sequence that meet the quality conditions are determined as human quality images; and The initial face and body image with the highest quality score in the face and body image sequence is determined as the face and body image with the best overall quality, thus obtaining the face image and body image in the face and body image with the best overall quality.

7. The method for associating a face with a human body according to claim 6, characterized in that, The human body quality images include: a frontal human body quality image, a side human body quality image, and a back human body quality image; The step of determining the initial human images in the human image sequence that meet the quality conditions as human quality images includes: The frontal human images in the human image sequence that meet the quality conditions are identified as the frontal human quality images; and The side-view human image in the human image sequence that meets the quality condition is determined as the side-view human quality image; and The back-facing human body image that meets the quality conditions in the human body image sequence is determined as the back-facing human body quality image; The quality conditions include having the highest quality score or having a quality score that exceeds a preset quality score.

8. The method for associating a face with a human body according to claim 1, characterized in that, The process involves acquiring video images, determining the tracking trajectory of the target object within the video images, and identifying multiple initial images of the target object. These initial images include an initial face image, an initial body image, and an initial face-body image, comprising: Target tracking is performed on the video image to determine the tracking trajectory of the target object in the video image; and Target detection is performed on each frame of the video image to determine multiple initial images of the target object.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the face-body association method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data that can be executed to implement the face-body association method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Human face and human body association method and system, and computer readable storage medium

    CN113657434A