Re-identification method and device for lost facial images after capture in complex environments, and electronic equipment

By using the first frame features of the target video to confirm the target in a complex environment, and then using auxiliary features to compare subsequent frames, the problems of low efficiency and high computational complexity in cross-regional face tracking are solved, and efficient and accurate cross-regional face recognition is achieved.

CN114022806BActive Publication Date: 2025-09-16BEIJING BESCO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111212387.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-09-16
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

In existing technologies, face tracking across acquisition areas is inefficient and computationally intensive. Especially in complex environments, when the target object moves, changes in facial and body features lead to decreased recognition accuracy, making continuous monitoring across areas impossible.

Method used

By extracting features from the first frame of the target video and comparing them with the preset target features, after confirming the existence of the target, the auxiliary features of the target object are used to compare with the features of subsequent frames, reducing the number of comparisons with the database and improving recognition efficiency and accuracy.

Benefits of technology

It reduces the amount of calculation, improves the accuracy and efficiency of cross-region face recognition, and solves the problem of loss of recognition of moving target objects in different regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022806B_ABST
    Figure CN114022806B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, and electronic device for re-identifying a face image that is lost after capture in a complex environment. The method includes: performing feature extraction on a first video frame in a target video to obtain a plurality of first extracted features; comparing the plurality of first extracted features with a preset target feature for identifying a preset target object; when it is determined that the first video frame contains a preset target object, performing feature extraction on a second video frame in the target video to obtain a plurality of second extracted features; comparing the plurality of second extracted features with the preset target feature; when it is determined that the second video frame does not contain the preset target feature, comparing the plurality of second extracted features with the auxiliary target feature of the preset target object; and determining whether the preset target object exists in the second video frame based on the comparison result of the plurality of second extracted features with the auxiliary target feature of the preset target object. The embodiment of the present application reduces the computational complexity of processing each frame after the first frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and device, and an electronic device for re-identifying facial images lost after capture in a complex environment. Background Art

[0002] With the advancement of video acquisition technology, high-resolution image capture services are now available for users at various locations, allowing them to view real-time images of multiple target areas without leaving their homes. Based on these real-time images, image processing techniques can be applied to identify specific targets, for example. However, because existing image acquisition devices often only cover a limited area, they can only identify or even track specific targets within that area. However, as people's living standards improve, their range and areas of activity are also expanding. Therefore, existing technologies have proposed using multiple acquisition devices at multiple locations or angles to achieve multi-area coverage monitoring.

[0003] However, existing image processing techniques only support the recognition of target objects within a single frame, meaning continuous recognition and tracking of target objects within a captured area against a single background. When a user leaves the capture area of ​​one terminal and enters the capture area of ​​another, the image recognition methods for these two capture areas can only identify the target within their respective capture areas and cannot achieve continuous monitoring across these areas. This is particularly true in target tracking scenarios based on facial recognition, where the target appears in different capture areas or at different angles, significantly impacting facial recognition accuracy. Summary of the Invention

[0004] The embodiments of the present application provide a method and device, as well as an electronic device, for re-identifying facial images lost after capture in a complex environment, to address the shortcomings of the prior art of low efficiency and high computational complexity in face tracking across acquisition areas.

[0005] To achieve the above objectives, the present invention provides a method for re-identifying a facial image that is lost after capture in a complex environment, comprising:

[0006] Performing feature extraction on a first video frame in a target video to obtain a plurality of first extracted features;

[0007] comparing the plurality of first extracted features with preset target features for identifying a preset target object;

[0008] When it is determined that the first video frame contains the preset target object, performing feature extraction on a second video frame in the target video to obtain a plurality of second extracted features, wherein the acquisition time of the first video frame is earlier than the acquisition time of the second video frame;

[0009] comparing the plurality of second extracted features with the preset target features;

[0010] When it is determined that the second video frame does not contain the preset target feature, comparing the plurality of second extracted features with auxiliary target features of the preset target object, wherein the auxiliary target feature is other features of the preset target object other than the preset target feature;

[0011] It is determined whether the preset target object exists in the second video frame according to a comparison result of the plurality of second extracted features and the auxiliary target features of the preset target object.

[0012] The present application also provides a device for re-recognizing a facial image lost after capture in a complex environment, comprising:

[0013] A first extraction module is used to perform feature extraction on a first video frame in a target video to obtain a plurality of first extracted features;

[0014] A first comparison module, configured to compare the plurality of first extracted features with a preset target feature for identifying a preset target object;

[0015] a second extraction module configured to, when determining that the first video frame contains the preset target object, perform feature extraction on a second video frame in the target video to obtain a plurality of second extracted features, wherein the acquisition time of the first video frame is earlier than the acquisition time of the second video frame;

[0016] a second comparison module, configured to compare the plurality of second extracted features with the preset target feature; and when it is determined that the second video frame does not contain the preset target feature, compare the plurality of second extracted features with auxiliary target features of the preset target object, wherein the auxiliary target feature is other features of the preset target object other than the preset target feature;

[0017] A determination module is used to determine whether the preset target object exists in the second video frame based on a comparison result of the multiple second extracted features and the auxiliary target features of the preset target object.

[0018] An embodiment of the present application further provides an electronic device, including:

[0019] Memory, used to store programs;

[0020] The processor is used to run the program stored in the memory, and when the program is run, the method for re-identifying facial images lost after capture in a complex environment provided by an embodiment of the present application is executed.

[0021] The embodiments of the present application provide a method and device, and an electronic device for re-identifying facial images that are lost after capture in a complex environment. When determining that a target object exists in the first video frame based on a comparison of a first extracted feature extracted from a first video frame in the target video with a target feature, a second extracted feature is further extracted from a second video frame of the target video to compare with the target feature, and an auxiliary target feature of the target object is further compared with the second extracted feature to determine whether the target object exists in the second video frame. Therefore, compared with the solution in the prior art that performs feature comparison one by one in the database for each frame, the present application starts from finding the first frame containing the target object, and directly uses the target features of the target object to compare with the features extracted from the subsequent frames, and further uses other auxiliary features for comparison when the comparison fails, thereby eliminating the need in the prior art to compare each extracted feature with a large number of different target features in the database, and thus greatly reduces the computational complexity of processing each frame after the first frame, and by using direct comparison of the first target feature and supplementary comparison using auxiliary features, it also greatly improves the accuracy and efficiency of re-identification, thereby being able to solve the defect in the prior art that the facial and / or body features of the moving target object change when moving, resulting in the loss of the tracking object.

[0022] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0024] Figure 1a A schematic diagram of an application scenario of the solution for re-identifying facial images lost after capture in a complex environment provided by an embodiment of the present application;

[0025] Figure 1b A schematic diagram of a facial image recognition solution according to the prior art;

[0026] Figure 1cA schematic diagram of a solution for re-identifying facial images lost after capture in a complex environment provided by an embodiment of the present application;

[0027] Figure 2 A flowchart of an embodiment of a method for re-identifying a facial image lost after capture in a complex environment provided by this application;

[0028] Figure 3 A flowchart of another embodiment of the method for re-identifying a facial image lost after capture in a complex environment provided by the present application;

[0029] Figure 4 This is a schematic diagram of the structure of an embodiment of a device for re-identifying a facial image lost after capture in a complex environment provided by the present application;

[0030] Figure 5 This is a schematic structural diagram of an electronic device embodiment provided in this application. DETAILED DESCRIPTION

[0031] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0032] Example 1

[0033] The solution provided in the embodiments of the present application can be applied to any system with image recognition capabilities, such as an image recognition terminal, etc. Figure 1a A schematic diagram of an application scenario of the solution for re-identifying facial images lost after capture in a complex environment provided by an embodiment of the present application; Figure 1b A schematic diagram of a facial image recognition solution according to the prior art; Figure 1c A schematic diagram of a solution for re-identifying facial images that are lost after capture in a complex environment provided by an embodiment of the present application. Figure 1c The principle shown is only one example of the principle of the technical solution of this application.

[0034] With the advancement of video acquisition technology, high-resolution image capture services are now available for users at various locations, allowing them to view real-time images of multiple target areas without leaving their homes. Based on these real-time images, image processing techniques can be applied to identify specific targets, for example. However, because existing image acquisition devices often only cover a limited area, they can only identify or even track specific targets within that area. However, as people's living standards improve, their range and areas of activity are also expanding. Therefore, existing technologies have proposed using multiple acquisition devices at multiple locations or angles to achieve multi-area coverage monitoring.

[0035] However, existing image processing techniques only support the recognition of target objects within a single frame, meaning continuous recognition and tracking of target objects within a captured area against a single background. When a user leaves the capture area of ​​one terminal and enters the capture area of ​​another, the image recognition methods for these two capture areas can only identify the target within their respective capture areas and cannot achieve continuous monitoring across these areas. This is particularly true in target tracking scenarios based on facial recognition, where the target appears in different capture areas or at different angles, significantly impacting facial recognition accuracy.

[0036] For example, in the application scenario of face recognition, a video or image acquisition device can capture a video or image of a target area and generate multiple image frames by decoding, and then transmit them to a processing device such as the one provided by the present application for processing. Figure 1a In the scenario shown in , the three video frames output after decoding by the acquisition device can be respectively passed through the comparison module to identify the objects contained in the three video frames to confirm the objects contained in the three video frames. Figure 1a In the scene shown in , after processing by the comparison module, it can be further confirmed whether the three video frames contain a specified target object, such as object 1.

[0037] For this reason, in the prior art, Figure 1bAs shown in , typically, multiple features are first extracted from, for example, a first frame output by an acquisition device. A comparison module can then compare these features one by one with the features of multiple objects already stored in, for example, an object database, and determine which objects are included in the first frame based on the comparison results. For example, in the prior art, the similarity between each feature extracted from the first frame and the features of each object stored in the database can be calculated, and objects in the database whose similarity to the features extracted from the first frame is greater than a preset threshold are determined to be objects included in the first frame. Subsequently, a determination can be made based on a specified target object identifier whether the one or more objects thus determined in the first frame include a specified target object.

[0038] Afterwards, the same operation is performed on the second frame output by the acquisition device, that is, each feature extracted from the second frame is compared with the features of each object in the database one by one to determine which objects are contained in the second frame, and then the specified target object identifier is used to determine whether the one or more objects thus determined in the second frame contain the specified target object. Figure 1a In the example shown in , the comparison module can identify, through such feature-by-feature comparison, that the first frame contains Object 1 and Object 2, the second frame contains Object 1, Object 3, and Object 4, and the third frame contains Object 2 and Object 3. Therefore, the comparison module can further output the video frame containing Object 1, designated as the target object, as needed for tracking, for example, the first and second frames. However, in many cases, the objects in the image frames captured by the capture device are not actually stationary but are constantly moving. Furthermore, the captured area contains not only the objects but also various objects and backgrounds, such as trees and buildings. Therefore, in practical applications, when a target object, such as Object 1, is continuously moving in the captured area, it is likely that its front view may be captured in the first frame. Thus, the comparison module can identify Object 1 in the first frame through facial feature comparison. However, in the second frame, Object 1 may move, with its face obscured by trees or buildings, or the capture device may only capture a portion of its face due to body movements such as turning. Therefore, in this case, the prior art is likely to fail to identify such an object by simply comparing the features extracted from the video frame with the object features in the database. In other words, the prior art is likely to mistakenly identify the second frame as not containing the target object, i.e., Object 1.

[0039] Therefore, in the prior art, it is usually necessary to first compare each feature extracted from each frame with the features of multiple objects in the database one by one to identify the objects contained in each frame, and then determine whether the specified target object is contained in the identified objects, so that the video frames determined to contain the specified target object, especially the second frame after the first frame and other video frames identified to contain the specified target object, are used as tracking video frames for the target object corresponding to the first frame. However, such a recognition process requires that each frame be compared with the features of each object in the database one by one. Especially when the number of objects in the database currently storing objects is very large, such one-to-one comparison is not only inefficient but also consumes a large amount of computing resources, resulting in an inability to meet the current requirements of high recognition speed and high accuracy required for tracking specified target objects in real-time captured video frames.

[0040] In this regard, the recognition method according to the embodiment of the present application provides a solution that is particularly suitable for re-identifying a captured target after it is lost in a complex environment. Figure 1a In the application scenarios of the above face recognition shown in Figure 1c As shown in , according to the recognition scheme of the embodiment of the present application, the comparison module can directly compare the features extracted from the first frame output by the acquisition device with the target object, that is, the specified features of object 1. After confirming that the target object is contained in the first frame, on the one hand, the specified features of the target object can be continued to be compared with the features extracted from the second frame acquired after the first frame. On the other hand, the first frame confirmed to contain the target object can be stored as a reference frame, and when it is confirmed that the second frame does not contain the target object when comparing the specified features of the target object with the features extracted from the second frame, the auxiliary features related to the target object extracted from the reference frame can be used for further comparison and identification to confirm whether the second frame contains the target object. For example, in the case of Figure 1aIn the scenario shown in , although only a portion of the target object's face is captured in the second frame, the comparison module is unable to identify the target object when comparing the features extracted from the second frame with the features of the target object because the feature similarity is lower than the threshold. However, in this embodiment of the application, the first frame that has been identified to contain the target object can be used as a reference frame to further extract auxiliary features related to the target object from the frame. For example, the target object, i.e., the clothing style, color, height, hairstyle, and other features of the target object, i.e., object 1, can be extracted from the first frame and used as auxiliary features when comparing the features extracted from the second frame with the target features. In this way, even if the facial features have a low degree of match, for example, the face only matches 40% of the features of the target object, the other features of the candidate object in the second frame that match 40% of the facial features of the target object can be further compared with these auxiliary features extracted from the first frame. Therefore, even if the candidate object turns around so that the facial features only partially match the facial features of the target object, since the auxiliary features such as clothing style, color, and height hardly change in these frames, it is possible to confirm that the candidate object is the target object by matching the auxiliary features with a very high degree of match, even if the facial match is below the threshold. Therefore, it is possible to well identify that the target object, i.e., object 1, is actually contained in the second frame.

[0041] Therefore, compared to the prior art solutions that require comparing the features extracted from each frame with the features of each object in the database one by one, the recognition solution proposed in the present application can fully utilize the reference frame that has been previously confirmed to contain the target object. In particular, it can use auxiliary features related to the target object in the reference frame to assist in comparing the extracted features in the second frame with the features of the target object. That is, if the target object is captured from the front in the first frame, but in the subsequent second frame, the target object's face is obscured or its facial features are missing due to body movements such as turning, then in the prior art solutions, the features of the target object contained in the second frame are insufficient to match the features of the target object, resulting in the second frame being incorrectly identified as not containing the target object. However, the target object is actually present in the second frame, only its facial features are missing. Therefore, using the recognition solution proposed in the embodiments of the present application, in this case, auxiliary features related to the target object extracted from the reference frame, such as the first frame, that has been confirmed to contain the target object can be introduced for use in object recognition in the second frame, thereby avoiding the second frame being incorrectly identified as a video frame that does not contain the target object simply due to the missing facial features of the target object.

[0042] The present invention provides a re-identification solution for facial images lost after capture in complex environments. When a target object is determined to exist in the first video frame by comparing a first extracted feature extracted from a first video frame of the target video with a target feature, a second extracted feature is extracted from a second video frame of the target video and compared with the target feature. The second extracted feature is then compared with the target feature using an auxiliary target feature of the target object to determine whether the target object exists in the second video frame. Therefore, compared to the prior art solution that performs feature comparisons for each frame in a database, the present invention directly uses the target feature of the target object to compare with features extracted from subsequent frames starting from the first frame containing the target object. If the comparison fails, further comparison is performed using other auxiliary features. This eliminates the need in the prior art to compare each extracted feature with a large number of different target features in the database. This significantly reduces the computational complexity of processing each frame after the first frame. Furthermore, by using direct comparison of the first target feature and supplementary comparison using auxiliary features, the accuracy and efficiency of re-identification are significantly improved, thereby resolving the prior art defect of losing track of moving target objects due to changes in facial and / or body features during movement.

[0043] The above embodiments are illustrations of the technical principles and exemplary application frameworks of the embodiments of the present application. The specific technical solutions of the embodiments of the present application are further described in detail below through multiple embodiments.

[0044] Example 2

[0045] Figure 2 This is a flowchart of an embodiment of the method for re-identifying a face image lost after capture in a complex environment provided by this application. The execution subject of this method can be various terminals or server devices with image recognition capabilities, or devices or chips integrated on these devices. Figure 2 As shown, the method for re-identifying a face image lost after capture in this complex environment includes the following steps:

[0046] S201: Perform feature extraction on a first video frame in a target video to obtain a plurality of first extracted features.

[0047] In an embodiment of the present application, a target video may be acquired from a video source such as a capture device, and multiple video frames may be obtained by decoding the acquired target video. In an embodiment of the present application, the multiple video frames may be video frames that are sequential in time. For example, in Figure 1aIn the scene shown in , the first frame can be a video frame captured before the second frame, and the second frame can be a video frame captured before the third frame. Moreover, the first frame, the second frame, and the third frame only need to be in a temporal order and are not necessarily continuous. In other words, in the embodiment of the present application, the first frame, the second frame, and the third frame can be separated by other video frames, but they can also be continuous in time, that is, there are no other video frames between them.

[0048] Therefore, in step S201, feature extraction can be performed on the received video frame, such as the first video frame, to obtain a plurality of first extracted features. In the embodiment of the present application, the first extracted features can be features extracted using various feature extraction algorithms in the art, such as facial feature extraction algorithms.

[0049] S202: Compare the plurality of first extracted features with preset target features for identifying a preset target object.

[0050] After extracting the first extracted features of the first video frame in step S201, the extracted first extracted features can be compared with the preset target features in step S202. In particular, in the embodiment of the present application, the preset target features can be features for identifying the preset target object. For example, in the example Figure 1a In the scenario shown in , the preset target features used in step S202 may be facial features of object 1. Of course, in this embodiment of the application, other features of the object may also be used for comparison, as long as the features can identify the target object.

[0051] S203: When it is determined that the first video frame contains the preset target object, feature extraction is performed on the second video frame in the target video to obtain a plurality of second extracted features.

[0052] After the comparison in step S202, when it is determined that the first video frame contains the preset target object according to the comparison result in step S202, feature extraction can be performed on the second video frame collected after the first video frame in step S203 to obtain a plurality of second extracted features. In the embodiment of the present application, the feature extraction algorithm in step S203 can be the same as or different from that in step S201 to perform feature extraction on the second video frame collected after the first video frame. For example, in the example Figure 1aIn the scenario shown in , a facial feature extraction algorithm is used in step S201 to extract facial features of each object in the first video frame. The same facial feature extraction algorithm can be used to extract facial features of each object in the second video frame captured after the first video frame as second extracted features. Various other algorithms can also be used to extract various other features together with facial features as second extracted features. In other words, in this embodiment of the present application, the second extracted features extracted in step S203 are not limited to facial features, but rather all possible features that can be extracted from the second video frame can be extracted as second extracted features.

[0053] S204: Compare the plurality of second extracted features with preset target features.

[0054] After extracting a plurality of second extracted features in step S203, these second extracted features can be compared with the preset target features used in step S202 to preliminarily determine whether the second video frame contains the target features. Figure 1a As described above, since the target object may continuously move or change its body posture, its facial features in the first video frame, for example, may be blocked or not captured in the second video frame. However, it is still possible that the similarity between part of the facial features captured in the second video frame and the target features is still high enough, for example, higher than a preset similarity threshold. In this case, it is still possible to directly determine that the target object is contained in the second video frame based on the comparison result in step S204.

[0055] S205 : When it is determined that the second video frame does not include the preset target feature, the plurality of second extracted features are compared with the auxiliary target features of the preset target object.

[0056] When it is determined based on the result of the comparison in step S204 that the target object is not contained in the second video frame captured after the first video frame, for example, when the similarity between the facial features of the object captured in the second video frame and the target features is calculated to be lower than a preset threshold in step S204, contrary to the prior art method of directly determining that the second video frame does not contain the target object, in an embodiment of the present application, the auxiliary features of the target object can be further used in step S205 to compare with the second extracted features. As described above, the second extracted features include not only the facial features of the object, but also various features extracted using various other algorithms. For example, style features of clothes, color features, hair features of the object, etc., and therefore can be compared with other features other than the preset facial features of the target object.

[0057] S206 , determining whether a preset target object exists in the second video frame according to a comparison result of the plurality of second extracted features and the auxiliary target features of the preset target object.

[0058] Therefore, in step S206, it can be determined whether the second video frame contains the preset target object according to the comparison result in step S205. Figure 1a In the scenario shown in , although only a portion of the target object's face is captured in the second frame, the comparison module is unable to identify the target object when comparing the features extracted from the second frame with the features of the target object because the feature similarity is lower than the threshold. However, in this embodiment of the application, the first frame that has been identified to contain the target object can be used as a reference frame to further extract auxiliary features related to the target object from the frame. For example, the target object, i.e., the clothing style, color, height, hairstyle, and other features of the target object, i.e., object 1, can be extracted from the first frame and used as auxiliary features when comparing the features extracted from the second frame with the target features. In this way, even if the facial features have a low degree of match, for example, the face only matches 40% of the features of the target object, the other features of the candidate object in the second frame that match 40% of the facial features of the target object can be further compared with these auxiliary features extracted from the first frame. Therefore, even if the candidate object turns around so that the facial features only partially match the facial features of the target object, since the auxiliary features such as clothing style, color, and height hardly change in these frames, it is possible to confirm that the candidate object is the target object by matching the auxiliary features with a very high degree of match, even if the facial match is below the threshold. Therefore, it is possible to well identify that the target object, i.e., object 1, is actually contained in the second frame.

[0059] The present invention provides a method for re-identifying a face image lost after capture in a complex environment. When a target object is detected in the first video frame by comparing a first extracted feature extracted from a first video frame of the target video with a target feature, the method further extracts a second extracted feature from a second video frame of the target video and compares it with the target feature. The method then uses an auxiliary target feature of the target object to further compare the second extracted feature to determine whether the target object is detected in the second video frame. Therefore, compared to the prior art method of performing feature comparisons for each frame in a database, the present invention, starting from the first frame containing the target object, directly uses the target feature of the target object to compare with features extracted from subsequent frames. If the comparison fails, the method further uses other auxiliary features for comparison. This eliminates the need in the prior art to compare each extracted feature with a large number of different target features in the database. This significantly reduces the computational complexity of processing each frame after the first frame. Furthermore, by using direct comparison of the first target feature and supplementary comparison using auxiliary features, the accuracy and efficiency of re-identification are significantly improved, thereby resolving the prior art defect of losing track of moving target objects due to changes in facial and / or body features during movement.

[0060] Example 3

[0061] Figure 3 This is a flowchart of another embodiment of the method for re-identifying a face image lost after capture in a complex environment provided by this application. The execution subject of this method can be various terminals or server devices with image recognition capabilities, or devices or chips integrated on these devices. Figure 3 As shown in FIG, the re-recognition method for face images lost after capture in this complex environment includes the following steps:

[0062] S301, obtaining a target video.

[0063] In an embodiment of the present application, a target video may be acquired from a video source such as a capture device, and multiple video frames may be obtained by decoding the acquired target video. In an embodiment of the present application, the multiple video frames may be video frames that are sequential in time. For example, in Figure 1a In the scene shown in , the first frame can be a video frame captured before the second frame, and the second frame can be a video frame captured before the third frame. Moreover, the first frame, the second frame, and the third frame only need to be in a temporal order and are not necessarily continuous. In other words, in the embodiment of the present application, the first frame, the second frame, and the third frame can be separated by other video frames, but they can also be continuous in time, that is, there are no other video frames between them.

[0064] S302: Perform feature extraction on a first video frame in a target video to obtain a plurality of first extracted features.

[0065] In step S302, feature extraction may be performed on the video frame received in step S301, for example, the first video frame, to obtain a plurality of first extracted features. In an embodiment of the present application, the first extracted features may be features extracted using various feature extraction algorithms in the art, such as a facial feature extraction algorithm.

[0066] S303: Compare the plurality of first extracted features with preset target features for identifying a preset target object.

[0067] After the first extracted features of the first video frame are extracted in step S302, the extracted first extracted features can be compared with the preset target features in step S303. In particular, in the embodiment of the present application, the preset target features can be features for identifying the preset target object. For example, in the example Figure 1a In the scenario shown in , the preset target features used in step S202 may be facial features of object 1. Of course, in the embodiment of the present application, other features of the object may also be used for comparison, as long as the features can identify the target object. In particular, in the embodiment of the present application, the preset target features may be facial recognition features.

[0068] S304: When it is determined that the first video frame contains the preset target object, feature extraction is performed on the second video frame in the target video to obtain a plurality of second extracted features.

[0069] After the comparison in step S303, when it is determined that the first video frame contains the preset target object according to the comparison result in step S303, feature extraction can be performed on the second video frame collected after the first video frame in step S304 to obtain a plurality of second extracted features. In the embodiment of the present application, the feature extraction algorithm in step S304 can be the same as or different from that in step S302 to perform feature extraction on the second video frame collected after the first video frame. For example, in the example Figure 1a In the scenario shown in , a facial feature extraction algorithm is used in step S302 to extract facial features of each object in the first video frame. The same facial feature extraction algorithm can be used to extract facial features of each object in the second video frame captured after the first video frame as second extracted features. Various other algorithms can also be used to extract various other features together with facial features as second extracted features. In other words, in this embodiment of the present application, the second extracted features extracted in step S304 are not limited to facial features, but rather all possible features that can be extracted from the second video frame can be extracted as second extracted features.

[0070] S305 : Using the first video frame containing the recognized preset target object as a reference video frame.

[0071] Simultaneously with, before, or after step S304 , when it is determined according to the comparison result of step S303 that the first video frame contains the preset target object, the first video frame may be determined as a reference video frame and may be temporarily stored.

[0072] S306 , extracting other features of the preset target object except the preset target features from the reference video frame as auxiliary target features of the target object.

[0073] In step S306, auxiliary target features can be extracted from the first video frame determined as the reference video frame. Figure 1a In the scenario shown in , other features besides facial features can be further extracted from the first video frame, such as the clothing style, color, height, hairstyle, and other features of subject 1. These features can be used as auxiliary features when comparing the features extracted from the second frame with the target features. For example, the auxiliary target features can be one or more of color features, texture features, shape features, and spatial relationship features.

[0074] In addition, in the embodiment of the present application, in addition to extracting auxiliary target features from the first video frame determined as the reference video frame, auxiliary target features related to the target object can also be obtained from the candidate tracking list. For example, various features of the target object can be set in a list of pre-specified target objects as auxiliary features of the target object. For example, in the example Figure 1a In the scenario shown in , it is known in advance that subject 1 is bald. Therefore, the hairstyle features of subject 1 can be pre-entered into the candidate tracking list and used when recognizing the second video frame. For example, the auxiliary target features can be one or more of color features, texture features, shape features, and spatial relationship features.

[0075] S307: Compare the plurality of second extracted features with the preset target features.

[0076] After extracting a plurality of second extracted features in step S304, these second extracted features can be compared with the preset target features used in step S303 to preliminarily determine whether the second video frame contains the target features. Figure 1aAs described above, since the target object may continuously move or change its body posture, its facial features in the first video frame, for example, may be blocked or not captured in the second video frame. However, it is still possible that the similarity between part of the facial features captured in the second video frame and the target features is still high enough, for example, higher than a preset similarity threshold. In this case, it can still be directly determined that the target object is contained in the second video frame based on the comparison result in step S307.

[0077] S308: When it is determined that the second video frame does not contain the preset target feature, the plurality of second extracted features are compared with the auxiliary target features of the preset target object.

[0078] When it is determined based on the result of the comparison in step S307 that the second video frame captured after the first video frame does not contain the target object, for example, when the similarity between the facial features of the object captured in the second video frame and the target features is calculated to be lower than a preset threshold in step S307, contrary to the prior art method of directly determining that the second video frame does not contain the target object, in an embodiment of the present application, the auxiliary features of the target object can be further used in step S308 to compare with the second extracted features. As described above, the second extracted features include not only the facial features of the object, but also various features extracted using various other algorithms. For example, style features of clothes, color features, hair features of the object, etc., and therefore can be compared with other features other than the preset facial features of the target object.

[0079] S309 , determining whether a preset target object exists in the second video frame according to a comparison result of the plurality of second extracted features and the auxiliary target features of the preset target object.

[0080] Therefore, in step S309, it can be determined whether the second video frame contains the preset target object according to the comparison result in step S308. Figure 1aIn the scenario shown in , although only a portion of the target object's face is captured in the second frame, the comparison module is unable to identify the target object when comparing the features extracted from the second frame with the features of the target object because the feature similarity is lower than the threshold. However, in this embodiment of the application, the first frame that has been identified to contain the target object can be used as a reference frame to further extract auxiliary features related to the target object from the frame. For example, the target object, i.e., the clothing style, color, height, hairstyle, and other features of the target object, i.e., object 1, can be extracted from the first frame and used as auxiliary features when comparing the features extracted from the second frame with the target features. In this way, even if the facial features have a low degree of match, for example, the face only matches 40% of the features of the target object, the other features of the candidate object in the second frame that match 40% of the facial features of the target object can be further compared with these auxiliary features extracted from the first frame. Therefore, even if the candidate object turns around so that the facial features only partially match the facial features of the target object, since the auxiliary features such as clothing style, color, and height hardly change in these frames, it is possible to confirm that the candidate object is the target object by matching the auxiliary features with a very high degree of match, even if the facial match is below the threshold. Therefore, it is possible to well identify that the target object, i.e., object 1, is actually contained in the second frame.

[0081] The present invention provides a method for re-identifying a face image lost after capture in a complex environment. When a target object is detected in the first video frame by comparing a first extracted feature extracted from a first video frame of the target video with a target feature, the method further extracts a second extracted feature from a second video frame of the target video and compares it with the target feature. The method then uses an auxiliary target feature of the target object to further compare the second extracted feature to determine whether the target object is detected in the second video frame. Therefore, compared to the prior art method of performing feature comparisons for each frame in a database, the present invention, starting from the first frame containing the target object, directly uses the target feature of the target object to compare with features extracted from subsequent frames. If the comparison fails, the method further uses other auxiliary features for comparison. This eliminates the need in the prior art to compare each extracted feature with a large number of different target features in the database. This significantly reduces the computational complexity of processing each frame after the first frame. Furthermore, by using direct comparison of the first target feature and supplementary comparison using auxiliary features, the accuracy and efficiency of re-identification are significantly improved, thereby resolving the prior art defect of losing track of moving target objects due to changes in facial and / or body features during movement.

[0082] Example 4

[0083] Figure 4 This is a schematic diagram of an embodiment of a device for re-recognizing a face image lost after capture in a complex environment provided by the present application, which can be used to perform the following operations: Figure 2 and Figure 3 The method steps shown are as follows. Figure 4 As shown, the apparatus for re-identifying a facial image lost after capture in a complex environment may include: an acquisition module 46 , a first extraction module 41 , a first comparison module 42 , a second extraction module 43 , a second comparison module 44 , and a determination module 45 .

[0084] The acquisition module 46 may be configured to acquire a target video.

[0085] In the embodiment of the present application, the acquisition module 46 can acquire the target video from a video source such as an acquisition device, and can obtain multiple video frames by decoding the acquired target video. In the embodiment of the present application, the multiple video frames can be video frames that have a temporal order. For example, in the example Figure 1a In the scene shown in , the first frame can be a video frame captured before the second frame, and the second frame can be a video frame captured before the third frame. Moreover, the first frame, the second frame, and the third frame only need to be in a temporal order and are not necessarily continuous. In other words, in the embodiment of the present application, the first frame, the second frame, and the third frame can be separated by other video frames, but they can also be continuous in time, that is, there are no other video frames between them.

[0086] The first extraction module 41 may be configured to perform feature extraction on a first video frame in a target video to obtain a plurality of first extracted features.

[0087] The first extraction module 41 can perform feature extraction on the video frame received by the acquisition module 46, such as the first video frame, to obtain a plurality of first extracted features. In the embodiment of the present application, the first extracted features can be features extracted using various feature extraction algorithms in the art, such as a facial feature extraction algorithm.

[0088] The first comparison module 42 may be configured to compare the plurality of first extracted features with a preset target feature for identifying a preset target object.

[0089] After the first extraction module 41 extracts the first extraction features of the first video frame, the first comparison module 42 can compare the extracted first extraction features with the preset target features. In particular, in the embodiment of the present application, these preset target features can be features used to identify the preset target object. For example, in Figure 1aIn the scenario shown in , the preset target features used by the first comparison module 42 may be facial features of the object 1. Of course, in the embodiment of the present application, other features of the object may also be used for comparison, as long as the features can identify the target object. In particular, in the embodiment of the present application, the preset target features may be facial recognition features.

[0090] The second extraction module 43 may be configured to perform feature extraction on the second video frame in the target video to obtain a plurality of second extracted features when it is determined that the first video frame contains a preset target object.

[0091] After the first comparison module 42 performs the comparison, when it is determined that the first video frame contains the preset target object according to the comparison result of the first comparison module 42, the second extraction module 43 can perform feature extraction on the second video frame collected after the first video frame to obtain multiple second extracted features. In the embodiment of the present application, the second extraction module 43 can use the same or different feature extraction algorithm as the first extraction module 41 to perform feature extraction on the second video frame collected after the first video frame. For example, in the example Figure 1a In the scenario shown in , the first extraction module 41 uses a facial feature extraction algorithm for the first video frame to extract facial features of each object in the first video frame, and the second extraction module 43 can use the same facial feature extraction algorithm for the second video frame collected after the first video frame to extract facial features of each object in the second video frame as second extracted features. Various other algorithms can also be used to extract various other features together with facial features as second extracted features. In other words, in this embodiment of the present application, the second extracted features extracted by the second extraction module 43 are not limited to facial features, but can extract all possible features from the second video frame as second extracted features.

[0092] In addition, the re-identification device of the present application may further include a first processing module 47, which can be used to use the first video frame containing the recognition of the preset target object as a reference video frame, and extract other features of the preset target object other than the preset target features from the reference video frame as auxiliary target features of the target object.

[0093] When the first processing module 47 determines that the first video frame contains the preset target object according to the comparison result of the first comparison module 42, the first video frame can be determined as a reference video frame, and can be temporarily stored, and auxiliary target features can be extracted from the first video frame determined as the reference video frame. Figure 1aIn the scenario shown in , other features besides facial features can be further extracted from the first video frame, such as the clothing style, color, height, hairstyle, and other features of subject 1. These features can be used as auxiliary features when comparing the features extracted from the second frame with the target features. For example, the auxiliary target features can be one or more of color features, texture features, shape features, and spatial relationship features.

[0094] In addition, in the embodiment of the present application, in addition to extracting auxiliary target features from the first video frame determined as the reference video frame, auxiliary target features related to the target object can also be obtained from the candidate tracking list. For example, various features of the target object can be set in a list of pre-specified target objects as auxiliary features of the target object. For example, in the example Figure 1a In the scenario shown in , it is known in advance that subject 1 is bald. Therefore, the hairstyle features of subject 1 can be pre-entered into the candidate tracking list and used when recognizing the second video frame. For example, the auxiliary target features can be one or more of color features, texture features, shape features, and spatial relationship features.

[0095] The second comparison module 44 may be configured to compare the plurality of second extracted features with a preset target feature.

[0096] After the second extraction module 43 extracts a plurality of second extraction features, the second comparison module 44 can compare these second extraction features with the preset target features used by the first comparison module 42 to preliminarily determine whether the second video frame contains the target features. Figure 1a As described above, since the target object may continuously move or change its body posture, its facial features in the first video frame, for example, are likely to be blocked or not captured in the second video frame. However, it is still possible that the similarity between part of the facial features captured in the second video frame and the target features is still high enough, for example, higher than a preset similarity threshold, then it can still be directly determined that the second video frame contains the target object based on the comparison result.

[0097] The second comparison module 44 may also compare the plurality of second extracted features with the auxiliary target features of the preset target object when it is determined that the second video frame does not contain the preset target feature.

[0098] When the second comparison module 44 determines that the second video frame captured after the first video frame does not contain the target object based on the comparison result, for example, when the similarity between the facial features of the object captured in the second video frame and the target features is calculated to be lower than a preset threshold, contrary to the prior art directly determining that the second video frame does not contain the target object, in an embodiment of the present application, the auxiliary features of the target object can be further used to compare with the second extracted features. As described above, the second extracted features include not only the facial features of the object, but also various features extracted using various other algorithms. For example, style features of clothes, color features, hair features of the object, etc., and therefore can be compared with other features other than the preset facial features of the target object.

[0099] The determination module 45 may be configured to determine whether the preset target object exists in the second video frame according to a comparison result of the plurality of second extracted features and the auxiliary target features of the preset target object.

[0100] Therefore, the determination module 45 can determine whether the second video frame contains the preset target object according to the comparison result of the second comparison module 44. Figure 1a In the scenario shown in , although only a portion of the target object's face is captured in the second frame, the comparison module is unable to identify the target object when comparing the features extracted from the second frame with the features of the target object because the feature similarity is lower than the threshold. However, in this embodiment of the application, the first frame that has been identified to contain the target object can be used as a reference frame to further extract auxiliary features related to the target object from the frame. For example, the target object, i.e., the clothing style, color, height, hairstyle, and other features of the target object, i.e., object 1, can be extracted from the first frame and used as auxiliary features when comparing the features extracted from the second frame with the target features. In this way, even if the facial features have a low degree of match, for example, the face only matches 40% of the features of the target object, the other features of the candidate object in the second frame that match 40% of the facial features of the target object can be further compared with these auxiliary features extracted from the first frame. Therefore, even if the candidate object turns around so that the facial features only partially match the facial features of the target object, since the auxiliary features such as clothing style, color, and height hardly change in these frames, it is possible to confirm that the candidate object is the target object by matching the auxiliary features with a very high degree of match, even if the facial match is below the threshold. Therefore, it is possible to well identify that the target object, i.e., object 1, is actually contained in the second frame.

[0101] The present invention provides a device for re-identifying a face image lost after capture in a complex environment. When a target object is detected in the first video frame by comparing a first extracted feature extracted from the first video frame of the target video with a target feature, the device further extracts a second extracted feature from the second video frame of the target video and compares it with the target feature. The device then uses an auxiliary target feature of the target object to further compare the second extracted feature to determine whether the target object is detected in the second video frame. Therefore, compared to the prior art approach of performing feature comparisons for each frame in a database, the present invention, starting with the first frame containing the target object, directly uses the target feature of the target object to compare with features extracted from subsequent frames. If the comparison fails, the device further uses other auxiliary features for comparison. This eliminates the need in the prior art to compare each extracted feature with a large number of different target features in the database. This significantly reduces the computational complexity of processing each frame after the first frame. Furthermore, by using direct comparison of the first target feature and supplementary comparison using the auxiliary features, the device significantly improves the accuracy and efficiency of re-identification, thereby resolving the prior art issue of losing track of a moving target object due to changes in facial and / or body features during movement.

[0102] Example 5

[0103] The above describes the internal functions and structure of the device for re-identifying a facial image lost after capture in a complex environment. The device can be implemented as an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device embodiment provided by this application. Figure 5 As shown, the electronic device includes a memory 51 and a processor 52 .

[0104] Memory 51 is used to store programs. In addition to the aforementioned programs, memory 51 may also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, images, videos, etc.

[0105] The memory 51 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0106] Processor 52 is not limited to a central processing unit (CPU) and may also be a processing chip such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. Processor 52 is coupled to memory 51 and executes a program stored in memory 51. When the program is executed, the method for re-recognizing a facial image lost after capture in a complex environment described in the second and third embodiments above is performed.

[0107] Further, if Figure 5 As shown, the electronic device may further include: a communication component 53, a power component 54, an audio component 55, a display 56 and other components. Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown.

[0108] The communication component 53 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 53 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 53 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0109] The power supply assembly 54 provides power to various components of the electronic device. The power supply assembly 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0110] The audio component 55 is configured to output and / or input audio signals. For example, the audio component 55 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 51 or transmitted via the communication component 53. In some embodiments, the audio component 55 also includes a speaker for outputting audio signals.

[0111] The display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0112] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for re-identifying a face image lost after capture in a complex environment, comprising: Performing feature extraction on a first video frame in a target video to obtain a plurality of first extracted features; comparing the plurality of first extracted features with preset target features for identifying a preset target object; When it is determined that the first video frame contains the preset target object, performing feature extraction on a second video frame in the target video to obtain a plurality of second extracted features, wherein the acquisition time of the first video frame is earlier than the acquisition time of the second video frame; comparing the plurality of second extracted features with the preset target features; When it is determined that the second video frame does not contain the preset target feature, comparing the plurality of second extracted features with auxiliary target features of the preset target object, wherein the auxiliary target feature is other features of the preset target object other than the preset target feature; It is determined whether the preset target object exists in the second video frame according to a comparison result of the plurality of second extracted features and the auxiliary target features of the preset target object.

2. The method for re-identifying a face image lost after capture in a complex environment according to claim 1, wherein: The method further comprises: Using the first video frame containing the identified preset target object as a reference video frame; Other features of the preset target object except the preset target feature are extracted from the reference video frame as auxiliary target features of the target object.

3. The method for re-identifying a face image lost after capture in a complex environment according to claim 1, wherein: The method further comprises: An auxiliary target feature related to the target object is obtained from the candidate tracking list.

4. The method for re-identifying a face image lost after capture in a complex environment according to any one of claims 1 to 3, wherein: The method further comprises: A target video is acquired, where the target video includes a first video frame and a second video frame.

5. The method for re-identifying a face image lost after capture in a complex environment according to claim 1, wherein: The preset target feature is a face recognition feature, and the auxiliary target feature is one or more of a color feature, a texture feature, a shape feature, and a spatial relationship feature.

6. A device for re-identifying a face image lost after capture in a complex environment, comprising: A first extraction module is used to perform feature extraction on a first video frame in a target video to obtain a plurality of first extracted features; A first comparison module, configured to compare the plurality of first extracted features with a preset target feature for identifying a preset target object; a second extraction module configured to, when determining that the first video frame contains the preset target object, perform feature extraction on a second video frame in the target video to obtain a plurality of second extracted features, wherein the acquisition time of the first video frame is earlier than the acquisition time of the second video frame; a second comparison module, configured to compare the plurality of second extracted features with the preset target feature; and when it is determined that the second video frame does not contain the preset target feature, compare the plurality of second extracted features with auxiliary target features of the preset target object, wherein the auxiliary target feature is other features of the preset target object other than the preset target feature; A determination module is used to determine whether the preset target object exists in the second video frame based on a comparison result of the multiple second extracted features and the auxiliary target features of the preset target object.

7. The apparatus for re-recognizing a face image lost after capture in a complex environment according to claim 6, wherein: The device further comprises: The first processing module is used to use the first video frame containing the preset target object as a reference video frame; and extract other features of the preset target object except the preset target features from the reference video frame as auxiliary target features of the target object.

8. The apparatus for re-recognizing a face image lost after capture in a complex environment according to claim 6, wherein: The device further comprises: The second processing module is configured to obtain auxiliary target features related to the target object from the candidate tracking list.

9. The apparatus for re-recognizing a face image lost after capture in a complex environment according to claim 6, wherein: The preset target feature is a face recognition feature, and the auxiliary target feature is one or more of a color feature, a texture feature, a shape feature, and a spatial relationship feature.

10. An electronic device comprising: Memory, used to store programs; A processor is configured to run the program stored in the memory to execute the method for re-identifying a facial image lost after capture in a complex environment as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Pedestrian re-identification device and method

    JP2021039741A