Target tracking identification method and apparatus, and electronic device
By using low-resolution images in the target tracking module and replacing them with high-resolution images during face recognition, and combining trajectory marking with face feature binding mechanism, the problem of target tracking and face recognition under limited computing power and bandwidth conditions is solved, achieving high recognition accuracy and resource utilization efficiency.
Patent Information
- Application Number
- CN202510893826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-21
AI Technical Summary
How can we simultaneously achieve continuous tracking of target objects and high-precision face recognition under limited computing power and bandwidth conditions, while avoiding increased system costs and reduced recognition accuracy?
By acquiring low-resolution images in the target tracking module for tracking, and using high-resolution images to generate synchronous images to replace the current frame for face recognition, and combining trajectory markers with face feature binding mechanisms, the consistency between image content and timestamps is achieved, reducing system load and improving recognition accuracy.
Without increasing the system load, this approach balances facial recognition accuracy with the continuity of target tracking, improving recognition accuracy and resource utilization efficiency, and enhancing the system's robustness and recognition integrity.
Smart Images

Figure CN120997887A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a target tracking and recognition method, apparatus and electronic device. Background Technology
[0002] With the development of artificial intelligence and computer vision technologies, multi-object tracking and facial recognition based on video analytics have been widely applied in scenarios such as smart classrooms, behavior analysis, and automated broadcasting. Multi-object tracking is primarily used to simultaneously identify and continuously track multiple individuals (such as students and teachers) in a video sequence, recording their movement trajectories and behavioral states over time. Facial recognition, on the other hand, focuses on determining individual identity and can be used for automated attendance tracking and targeted analysis of individual behavior. The two technologies typically work together, but their technical requirements for video input differ: multi-object tracking relies on high temporal resolution to ensure the continuity and relevance of behavioral information; while facial recognition requires higher spatial resolution to ensure clear facial features and accurate recognition.
[0003] Under ideal conditions, if the system has sufficient computing power and bandwidth, it can input video images with high spatial and temporal resolution (e.g., simultaneously meeting the requirements of ultra-high-definition image quality and multi-frame-per-second image acquisition) to achieve parallel processing of multi-target tracking and face recognition tasks, thereby taking into account both spatial and temporal information acquisition and achieving high-precision intelligent perception. However, such high-specification video input will significantly increase processor load, storage capacity requirements, and data transmission bandwidth, leading to an increase in the overall system cost. Summary of the Invention
[0004] One objective of this application is to provide a target tracking and recognition method, apparatus, and electronic device to solve the technical problem of how to simultaneously achieve continuous tracking of target objects and high-precision face recognition under limited computing power and bandwidth conditions.
[0005] To address the aforementioned technical problems, one technical solution adopted in this application is as follows: a target tracking and recognition method is provided, comprising: continuously acquiring a first image sequence of a target scene, and tracking a target object in the target scene based on the first image sequence; during the tracking of the target object, when a face recognition condition is triggered, acquiring a second image of the target scene, and scaling the second image to generate a synchronized image; the resolution of the second image is greater than the resolution of the first image, and the difference between the resolution of the synchronized image and the resolution of the first image is within a preset range; when a face recognition condition is triggered, replacing the first image with the synchronized image to add the synchronized image to the first image sequence, and acquiring the target identification information corresponding to the synchronized image, wherein the first image is a frame image in the first image sequence, and the acquisition time of the first image is the time when the face recognition condition is triggered; extracting features from the face region of the target object based on the second image to obtain face features; associating the target identification information corresponding to the synchronized image with the face features to obtain the identity and trajectory binding information of the target object.
[0006] This method uses low-resolution images for target tracking to reduce the system's computational and bandwidth burden. Simultaneously, upon triggering face recognition, a synchronous image is generated from the high-resolution image to replace the current frame's tracking input image, ensuring consistency between image content and timestamps. This guarantees accurate correspondence and binding between facial features in face recognition and target identification information in target tracking, thereby eliminating binding deviations caused by asynchronous sampling between different modules. This method balances the accuracy requirements of face recognition with the continuity of target tracking without requiring continuous high-resolution acquisition or significantly increasing processing frequency, effectively improving overall recognition accuracy and resource utilization efficiency. It possesses good practicality and scalability.
[0007] In some embodiments, during the tracking of a target object in a target scene based on a first image sequence, if the face recognition condition is triggered multiple times, the method further includes: each time the face recognition condition is triggered, extracting features from the face region of the target object based on the second image of the target scene obtained each time, to obtain face features; adding the face features obtained each time to a face feature set, and binding the face features obtained each time with the target identification information in the corresponding synchronized image to obtain multiple identity and trajectory binding information; wherein, the corresponding synchronized image refers to the image generated after scaling the second image acquired when the face recognition condition is triggered, and the image generated after scaling replaces the first image for tracking the target object.
[0008] In this way, when the target object meets the facial recognition conditions multiple times, its facial feature information can be accumulated and bound to the corresponding trajectory markers multiple times, thereby improving the integrity of identity recognition and the accuracy of trajectory association.
[0009] In some embodiments, the method further includes: matching facial features in a set of facial features with features in a preset facial database to obtain successfully matched facial features and unmatched facial features; assigning facial identification information to the successfully matched facial features; using the facial identification information to identify the identity of the target object; and obtaining the identity information corresponding to the unmatched facial features based on multiple identity and trajectory binding information and facial identification information.
[0010] This method is based on the continuous target identification information of the target object in multiple frames of images. Even if a valid facial feature cannot be identified in some frames, indirect identity completion can be achieved through the facial identification information associated with it in other frames. Therefore, it effectively improves the fault tolerance and coverage of face recognition.
[0011] In some embodiments, obtaining the identity information corresponding to a face feature that failed to match based on multiple identity and trajectory binding information and face identification information includes: obtaining the target identification information corresponding to the face feature that failed to match based on multiple identity and trajectory binding information; the target identification information corresponding to the face feature that failed to match is the first target identification information; searching for other binding information that is the same as the first target identification information among multiple identity and trajectory binding information; if the other binding information contains a binding relationship between a face feature that successfully matches and the first target identification information, then assigning the identity information corresponding to the face feature that successfully matches to the face feature that failed to match.
[0012] In this way, by leveraging the persistence of target identification information across multiple frames of images, the identity information of the identified target can be extended to the facial features of those that were not successfully matched, thereby achieving automatic completion of the target object's identity and improving the completeness of identity recognition and the robustness of the system.
[0013] In some embodiments, after matching facial features in a set of facial features with features in a preset facial database to obtain successfully matched facial features and unmatched facial features, the method further includes: clustering the unmatched facial features with the successfully matched facial features to obtain clustering results of the unmatched facial features and the successfully matched facial features; and obtaining the identity information corresponding to the unmatched facial features based on multiple identity and trajectory binding information and facial identification information, including: obtaining the identity information corresponding to the unmatched facial features based on multiple identity and trajectory binding information, facial identification information, and clustering results.
[0014] This method first clusters the facial features of successfully matched and unmatched faces, grouping similar targets together to infer the attribution of unmatched faces. By combining multiple identity and trajectory binding information, the method further accurately completes the identity information of unmatched faces. This effectively improves the system's recognition completeness and robustness in complex environments, reduces missed detections due to pose, occlusion, or poor image quality, and enhances the continuity and stability of target identity recognition.
[0015] In some embodiments, based on multiple identity and trajectory binding information, face identification information, and clustering results, the identity information corresponding to the unmatched face features is determined, including: obtaining the target identification information corresponding to the unmatched face features, denoted as the first target identification information; searching for other binding information that is the same as the first target identification information among the multiple identity and trajectory binding information; searching for clusters containing unmatched face features in the clustering results; determining whether there are successfully matched face features located in the cluster among the other binding information; if so, assigning the identity information corresponding to the successfully matched face features to the unmatched face features.
[0016] This method combines temporal continuity (through target identification information) with feature similarity (through clustering), effectively improving the system's accuracy in identity recognition under poor image quality conditions such as occlusion, head tilting, and blurriness.
[0017] In some embodiments, the method further includes: matching facial features in a set of facial features with features in a preset facial database to obtain successfully matched facial features and unmatched facial features; assigning facial identification information to the successfully matched facial features; the facial identification information is used to identify the identity of the target object; based on multiple identity and trajectory binding information, obtaining target identification information bound to the facial identification information, denoted as second target identification information; the second target identification information is the target identification information bound to the successfully matched facial features; among the multiple identity and trajectory binding information, detecting whether there are other target identification information that is the same as the facial identification information but different from the second target identification information; if so, updating the other target identification information that is different from the second target identification information to the second target identification information.
[0018] In this way, by merging multiple target identification information bound to the same facial identification information, the problem of target identification information switching caused by rapid target movement or occlusion can be solved, thereby improving the stability and consistency of target tracking results and ensuring the continuous identification of individual identity in the time series.
[0019] To address the aforementioned technical problems, one technical solution adopted in this application is as follows: a target tracking and recognition device is provided, comprising: a first target tracking module, configured to continuously acquire a first image sequence of a target scene and track a target object in the target scene based on the first image sequence; an image processing module, configured to acquire a second image of the target scene and scale the second image to generate a synchronous image when a face recognition condition is detected during the tracking of the target object; the resolution of the second image is greater than the resolution of the first image, and the difference between the resolution of the synchronous image and the resolution of the first image is within a preset range; a second target tracking module, configured to replace the first image with the synchronous image when the face recognition condition is triggered, thereby adding the synchronous image to the first image sequence, and acquiring target identification information corresponding to the synchronous image; the first image is a frame image in the first image sequence, and the acquisition time of the first image is the time when the face recognition condition is triggered; a face recognition module, configured to extract features from the face region of the target object based on the second image to obtain face features; and an information binding module, configured to associate the target identification information corresponding to the synchronous image with the face features to obtain the identity and trajectory binding information of the target object.
[0020] To address the aforementioned technical problems, one technical solution adopted in this application is to provide an electronic device, including a memory and a processor. The memory is connected to the processor, and the processor executes one or more computer programs stored in the memory. When the processor executes the one or more computer programs, it enables the electronic device to implement a target tracking and recognition method applicable to the electronic device. This electronic device possesses the beneficial effects corresponding to the aforementioned target tracking and recognition method applicable to the electronic device.
[0021] To address the aforementioned technical problems, one technical solution adopted in this application is to provide a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by an electronic device, the electronic device performs the aforementioned target tracking and identification method. This non-volatile computer-readable storage medium possesses the beneficial effects corresponding to the aforementioned target tracking and identification method. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a target tracking and recognition method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the process of executing a target tracking and identification method using a target object as an example, provided in an embodiment of this application. Figure 3 This is a flowchart of a target tracking and recognition method provided in another embodiment of this application; Figure 4 This is a flowchart of a target tracking and recognition method provided in another embodiment of this application; Figure 5 This is a schematic diagram illustrating the process of using Face ID to complete the identity of a target object, as provided in the embodiments of this application. Figure 6 This is a flowchart of a target tracking and recognition method provided in another embodiment of this application; Figure 7 This is a schematic diagram illustrating the process of identity information completion through the combination of facial feature clustering and target identification information, as provided in the embodiments of this application. Figure 8 This is a schematic diagram illustrating the process of merging multiple target identifiers bound to the same face identifier information, as provided in an embodiment of this application. Figure 9 This is a schematic diagram of the structure of a target tracking and recognition device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0025] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0026] Object tracking is a computer vision task that aims to simultaneously identify and continuously track the identity and location of target objects in a video sequence. For each target object, continuous tracking is performed to record its motion trajectory over time and manage identity consistency. For example, in a classroom setting, teachers and students can be tracked in real time to obtain each person's location, movement trajectory, and behavioral state at different points in time, supporting applications such as teaching behavior analysis and student attention detection. This process typically includes key steps such as object detection, feature extraction, object association, and trajectory maintenance to ensure that the system can accurately distinguish and continuously identify each individual even in dynamic multi-person scenarios.
[0027] Facial recognition is a technology that identifies or verifies an individual's identity by analyzing facial images. It's primarily used for identity verification, not just "tracking." It typically involves processes such as face detection, face alignment, feature extraction, and identity comparison. For example, in a classroom setting, facial recognition technology can be used to confirm the identity of students entering the classroom, enabling automatic attendance. It can also continuously identify individuals during class, allowing for targeted analysis of specific students' behavior in conjunction with multi-target tracking technology.
[0028] In practical applications, multi-target tracking and face recognition have different requirements for video input resolution and frame rate. Multi-target tracking (such as classroom observation or intelligent broadcasting) typically relies on high temporal resolution to continuously track the behavioral states of targets (such as students and teachers), such as raising hands, walking, turning around, etc. Therefore, a video stream of at least 1080P resolution at 3fps (i.e., full HD (1080P) at 3 frames per second) is required to ensure the continuity of target motion and the stability of data association between different frames. If the frame rate is too low, key actions may be lost or the same target may not be accurately associated between frames, especially when the target is occluded, such as when turning or looking down, potentially leading to incorrect identification as two different individuals.
[0029] In contrast, facial recognition (such as for classroom attendance tracking) is more dependent on spatial resolution, requiring that facial features in the image be clearly distinguishable. Therefore, it typically requires at least 4K resolution video input at 1 / 30fps (ultra-high definition (4K) images, captured at a frequency of 1 frame per 30 seconds or lower) to capture the faces of all targets in the same frame and extract effective features. Lower resolution will cause facial features to be blurred, thus affecting the accuracy of feature vector extraction and recognition.
[0030] Therefore, face recognition requires high resolution (at least 4K) to acquire enough facial pixels (>40px) within a certain distance to ensure accurate recognition. Multi-object tracking, on the other hand, is limited by frame rate requirements (at least 2 to 5fps), but practical systems often cannot meet this frame rate due to insufficient computing power or gimbal lifespan limitations. Related technologies primarily achieve face recognition by sacrificing frame rate and capturing high-resolution images at low frame rates, but this approach cannot simultaneously handle multi-object tracking tasks involving behavioral trajectories.
[0031] Under ideal conditions, if the system has sufficient computing power and bandwidth, for example, it can directly use 4K@3fps (acquiring 4K ultra-high-definition images at a frequency of 3 frames per second) video input, uniformly supporting the parallel execution of multi-target tracking and face recognition, thereby achieving intelligent perception of target scenes while maintaining both spatial and temporal resolution. However, this will significantly increase processor load, storage, and transmission pressure, leading to increased system costs. Therefore, in practical deployments, a balance needs to be struck between recognition accuracy, computing resources, and deployment costs.
[0032] Based on this, this application provides a target tracking and recognition method that can simultaneously handle face attendance and target tracking and observation in an embedded environment with limited computing power and bandwidth. The target tracking and recognition method in this application is primarily based on an optimization mechanism that integrates task collaboration and data reuse.
[0033] One proposed mechanism is an image replacement and synchronous analysis mechanism. When the multi-target tracking module operates at a low frame rate (e.g., 3fps), the system acquires images in real-time from the camera or cache for target tracking. These real-time images can be low-resolution (e.g., 1080P). When the system detects a face recognition trigger (e.g., entering the recognition area or taking a timed photo), it scales down an acquired ultra-high-resolution image (e.g., 4K image) to generate a low-resolution image (e.g., 1080P image) and uses it to replace the input image of the target tracking module in the current frame. This replacement image has the same timestamp and image source as the image used for face recognition, ensuring that the face recognition result is directly aligned with the trajectory information of the current frame. When face recognition ends, the multi-target tracking module continues to sample and run independently, only temporarily accessing a frame with the same time as the recognition image at the key frame triggered by face recognition through the image replacement mechanism, achieving information alignment and image reuse between the two modules.
[0034] It also provides a mechanism for binding trajectory identifiers with facial features. To establish a correspondence between individual identity and behavioral trajectory, the system assigns the target identifier information (Track ID) of the corresponding target in the current multi-target tracking module to the feature value extracted by the face recognition module while performing face recognition. This establishes a temporary binding relationship between "identity features and target identifier information," i.e., identity and trajectory binding information. This binding relationship can be used for subsequent trajectory backtracking and identity tracking, and also provides structured data support for cluster analysis, avoiding the problem of traditional face recognition operating independently and being separated from trajectory information.
[0035] A face recognition compensation mechanism based on temporal clustering is also provided. Considering that in the target scene, factors such as pose changes, occlusion, and poor lighting may cause some face images to fail to match the face feature database during the recognition process (e.g., failing to reach a preset similarity threshold), resulting in individual recognition failure, a unified clustering analysis is performed on all face feature vectors collected during the period (including successfully matched and unmatched face features) after the complete activity cycle of the target scene (e.g., after the end of a class in a classroom scenario). This clustering analysis can employ unsupervised learning algorithms, such as K-Means, DBSCAN, or hierarchical clustering methods based on cosine distance, to group multiple highly similar but not clearly identified face features into the same category. Furthermore, by leveraging the binding relationship between the target identifier information corresponding to the synchronized image and the face feature value, this mechanism can further associate the unidentified features in the clustering results with previously identified samples, thereby achieving retrospective identity compensation for unidentified individuals and assigning them unified face identifier information (Face ID). This process significantly improves the recall rate of facial recognition and the overall attendance coverage, especially for individuals who occasionally look down or have poor viewing angles.
[0036] A trajectory consistency enhancement mechanism based on identity-based reverse association is also provided. Considering that during target tracking, especially in scenarios with rapid movement, occlusion, or drastic changes in posture (e.g., teachers walking, students standing, or large movements), the system may assign the same target different Track IDs, resulting in erroneous trajectory splitting for the same person. To address this issue, after matching facial features with the face database, this mechanism searches for and compares the Track ID corresponding to the identified face ID in the identity and trajectory binding information. If different Track IDs are detected bound to the same identity (i.e., the same Face ID), an update operation merges these different Track IDs into a unified identifier, thereby unifying and correcting the trajectory of the same target at different time periods. This mechanism effectively improves the trajectory coherence of target tracking and reduces Track ID switching caused by target occlusion or visual interference, making it particularly suitable for joint analysis of identity and behavior in dynamic multi-person scenarios such as classrooms and meetings.
[0037] Based on the above concept, the target tracking and recognition method will be described below through specific embodiments.
[0038] Please see Figure 1 , Figure 1 This is a flowchart of a target tracking and recognition method provided in an embodiment of this application. The method includes: S101. Continuously acquire the first image sequence of the target scene, and track the target object in the target scene based on the first image sequence.
[0039] The target scenario can be a classroom, a meeting room, a waiting area, an office environment, a factory workshop, or other specific scenarios that require continuous observation and behavioral analysis of people or objects.
[0040] The first image sequence refers to a series of image frames continuously acquired in the target scene using an image acquisition device (such as a camera). These image frames are arranged in chronological order, reflecting the changes in the visual state of the target scene at different times. Each frame in this image sequence is collectively referred to as the "first image." To meet the requirements of target object tracking, the acquisition frequency of the first image must meet certain requirements, such as a resolution of no less than 1080P and a frame rate of no less than 3 frames per second (3fps), to ensure image clarity and temporal continuity, and to ensure that the tracking algorithm has sufficient visual information for effective recognition and association.
[0041] A target object refers to an entity existing within a target scene that is of interest to the image processing system and requires detection, identification, tracking, or behavior analysis. Examples include teachers and students in a classroom setting, and attendees and speakers in a meeting setting. Target objects possess identifiable features such as appearance, posture, and actions; they also exhibit spatial and temporal trajectories; and they can be used for processing purposes such as identity verification and behavior recognition.
[0042] When tracking a target object in a target scene based on a first image sequence, the target object in each frame of the first image can be detected and its features extracted. This extraction process captures the object's appearance features (such as clothing color and human outline) and spatiotemporal features (such as position and motion trajectory). The detection results of the current frame are then correlated and matched with the target features in the previous or historical frames. This allows for the preservation of the target object's identity and trajectory tracking within consecutive images. Multi-Object Tracking (MOT) algorithms can be used for target tracking.
[0043] When tracking a target object in a target scene based on the first image sequence, the target tracking result of each frame of the first image can include corresponding target identification information, such as the target ID, bounding box position, motion state, and target category of the target object, which is used to identify the consistency and continuity of the target object between different frames.
[0044] S102. During the tracking of the target object, when the face recognition condition is triggered, a second image of the target scene is acquired, and the second image is scaled to generate a synchronized image.
[0045] In this system, the resolution of the second image is greater than that of the first image, and the difference between the resolution of the synchronized image and the resolution of the first image is within a preset range. The synchronized image refers to the image obtained by scaling the second image, whose spatial resolution differs from the resolution of the first image in the current frame within a set tolerance range, thus ensuring visual scale alignment or near-consistency. If the resolution of the first image is 1080P, then the resolution of the synchronized image can also be 1080P. The second image is used for face recognition, and its resolution is greater than that of the first image, for example, at least 4K. The scaling process aims to generate an image with a resolution close to that of the first image. The scaling process can be based on the resolution of the first image in the current frame, scaling the second image (e.g., a high-resolution image, 4K) proportionally or with bilateral constraints.
[0046] Face recognition conditions refer to the conditional judgment logic that triggers face recognition when a state meeting specific face information extraction requirements is detected during target object tracking. For example, if the target object's face region is clearly visible, facing forward, and the confidence level exceeds a set threshold, or if the target object enters a specific recognition area and remains there for more than a set number of frames, then the face recognition condition is considered triggered, and the current frame is acquired as a second image for face recognition processing. Another example is if the current image frame comes from a 4K ultra-high-definition input channel and is within an acquireable timeframe (e.g., one frame every 30 seconds), then the face recognition condition is triggered.
[0047] S103. When the face recognition condition is triggered, the first image is replaced with a synchronized image to add the synchronized image to the first image sequence and obtain the target identification information corresponding to the synchronized image.
[0048] Specifically, the target object in the target scene is tracked based on the replaced synchronized image, and target identification information corresponding to the synchronized image is generated. The first image is a frame from a first image sequence, and the first image is acquired at the moment when the face recognition condition is triggered.
[0049] When face recognition is triggered, the first image is replaced with a scaled-down synchronized image. Then, target tracking is performed based on this synchronized image, and corresponding target identification information is generated. This target identification information may include the target object's ID, location coordinates, tracking status, etc.
[0050] S104. Based on the second image, extract features from the face region of the target object to obtain face features.
[0051] A pre-trained face detection model can be used to accurately locate the face region of the target object from the second image. Then, a deep convolutional neural network (such as ResNet, MobileNet, or a dedicated face recognition network ArcFace, FaceNet, etc.) is used to encode the features of the face region and extract a discriminative face feature vector. This face feature vector can effectively represent the unique identity information of the face and support subsequent face comparison, identity verification, and classification tasks.
[0052] S105. Associate the target identification information corresponding to the synchronized image with the facial features to obtain the identity and trajectory binding information of the target object.
[0053] When the target identification information corresponding to the synchronized image is associated with the facial features, the synchronized image is generated by scaling up the second image when the facial features are extracted. In this way, visual tracking information with consistent timestamps and matching spatial scales is accurately bound to facial identity features, effectively eliminating the deviation in identity and trajectory binding caused by asynchronous sampling of multi-source images.
[0054] The identity and trajectory binding information of the target object combines the target object's identity with its location information, thereby revealing who the target object is, where it is, and how it moves, thus forming a complete and traceable record of identity recognition and behavioral trajectory. Specifically, the identity and trajectory binding information may include target identity information (such as a unique identifier ID, facial feature vector, name, student ID, etc.), trajectory information (such as timestamps, spatial location information, etc.), and other auxiliary information.
[0055] The target tracking and recognition method in this embodiment can be applied to a single target object in a target scene or to multiple target objects at the same time, thereby supporting parallel tracking and recognition of one or more targets in the scene.
[0056] For example, such as Figure 2 As shown, taking a target object as an example, a first image (e.g., a 1080P original image) corresponding to the target object is continuously acquired, and the target identification information corresponding to the target object is obtained from the first image and set as Track ID=1. If a face recognition condition is triggered, a second image (e.g., a 4K image) is acquired, then scaled to a synchronous image (e.g., scaled to 1080P), and the target identification information corresponding to the synchronous image is acquired and also set as Track ID=1. Simultaneously, the facial features of the target object are identified based on the second image, and these facial features are represented as Face ID=a. Furthermore, Face ID=a is bound to Track ID=1 corresponding to the synchronous image to obtain identity and trajectory binding information.
[0057] The method in this application reduces the system's computational and bandwidth burden by acquiring low-resolution images for target tracking. Simultaneously, upon triggering face recognition, a synchronous image is generated from the high-resolution image to replace the current frame's tracking input image, ensuring consistency between image content and timestamps. This guarantees accurate correspondence and binding between facial features in face recognition and target identification information in target tracking, thereby eliminating binding deviations caused by asynchronous sampling between different modules. This method eliminates the need for continuous high-resolution acquisition and significantly increases processing frequency, balancing the accuracy requirements of face recognition with the continuity of target tracking. It effectively improves overall recognition accuracy and resource utilization efficiency, demonstrating good practicality and scalability.
[0058] Please see Figure 3 , Figure 3 This is a flowchart of a target tracking and recognition method provided in another embodiment of this application. The method includes: S201. Continuously acquire the first image sequence of the target scene, and track the target object in the target scene based on the first image sequence.
[0059] S202. During the target object tracking process, determine whether the face recognition conditions are met.
[0060] If the conditions are met, proceed to steps S203 to S205; otherwise, continue with step S201.
[0061] S203. Obtain a second image of the target scene and scale the second image to generate a synchronized image.
[0062] The resolution of the second image is greater than that of the first image, and the difference between the resolution of the synchronized image and the resolution of the first image is within a preset range.
[0063] S204. When the face recognition condition is triggered, the first image is replaced with a synchronized image to add the synchronized image to the first image sequence, and the target identification information corresponding to the synchronized image is obtained.
[0064] Specifically, the target object in the target scene is tracked based on the replaced synchronized image, and target identification information corresponding to the synchronized image is generated. The first image is a frame from a first image sequence, and the first image is acquired at the moment when the face recognition condition is triggered.
[0065] S205. Based on the second image, extract features from the face region of the target object to obtain face features, and associate the face features with the target identification information corresponding to the synchronized image to obtain an identity and trajectory binding information.
[0066] S206. Determine whether the target tracking and recognition task has ended.
[0067] If the process does not end, return to step S202; otherwise, proceed to step S207.
[0068] S207. Add the facial features obtained each time to the facial feature set, and add the identity and trajectory binding information obtained each time to the binding information set.
[0069] The facial feature set includes facial features extracted after facial recognition of each second image. The binding information set includes multiple identity and trajectory binding information. If there is only one tracked target object in the target scene, the binding information set only contains the identity and trajectory binding information corresponding to that target object. If there are multiple target objects, the above steps can be executed in parallel, processing each target object independently, thereby obtaining multiple identity and trajectory binding information simultaneously within the same processing cycle; each identity and trajectory binding information corresponds to one target object and can be uniformly stored in the binding information set to support multi-target recognition and trajectory management.
[0070] The method in this embodiment repeatedly executes the steps of facial feature extraction and target identification information binding when the facial recognition condition is triggered multiple times. For detailed process information, please refer to [reference needed]. Figure 1 The corresponding implementation allows for multiple acquisitions of facial feature information from the target object at different times, angles, or states (such as looking up or turning the head), and multiple bindings with corresponding trajectory identifiers. This effectively compensates for single-recognition failures caused by factors such as posture, lighting, or occlusion, improving overall recognition coverage. Multiple bindings ensure that the trajectory and identity of the same target object remain consistent throughout the entire activity cycle, reducing the risk of Track ID breakage or identity drift. Furthermore, the accumulated facial feature set and binding information can be used for further clustering analysis, improving the ability to backtrack and identify unidentified objects.
[0071] In some embodiments, please refer to Figure 4 It also provides a target tracking and recognition method, which Figure 4 With the above Figure 3 The difference is that the method also includes: S208. Match the facial features in the facial feature set with the features in the preset facial database to obtain the successfully matched facial features and the unmatched facial features. For example, calculate the similarity between each facial feature in the facial feature set and the feature vector in the preset facial database. If the similarity exceeds a set matching threshold, the match is considered successful, and the corresponding identity information is obtained; otherwise, it is considered unmatched.
[0072] S209. Assign facial identification information to the successfully matched facial features. This facial identification information is used to identify the identity of the target object.
[0073] For example, such as Figure 5As shown, successfully matched facial features are assigned face identifier information, which is Face ID=1. According to the above embodiment, all facial features in the facial feature set are bound to the target identifier information corresponding to the synchronized image and stored in the aforementioned binding information set. Unmatched facial features can be assigned face identifier information, which is Face ID=0.
[0074] S210. Based on the above-mentioned binding information set and the face identification information, obtain the identity information corresponding to the unmatched face features.
[0075] First, based on multiple identity and trajectory binding information, the target identifier information corresponding to the unmatched facial features is obtained and set as the first target identifier information. For example, such as Figure 5 As shown, the target identification information corresponding to the unmatched face feature (Face ID=0) is the first target identification information, i.e., Track ID=1.
[0076] Next, among multiple identity and trajectory binding information, search for other binding information that matches the first target identifier information (Track ID=1). If other binding information contains a successfully matched face feature (Face ID=1) binding relationship with the first target identifier information (Track ID=1), then the identity information corresponding to the successfully matched face feature (Face ID=1) is assigned to the unmatched face feature. For example... Figure 5 As shown, changing "Track ID=1; Face ID=0" to "Track ID=1; Face ID=1" allows for indirect identity completion even when the preset face database fails to match a corresponding facial feature, thanks to the Face ID associated with that feature in other frames. This effectively improves the fault tolerance and coverage of face recognition. This process leverages the persistence of target identification information across multiple frames to extend the identity information of the identified target to unmatched facial features, thereby achieving automatic identity completion for the target object.
[0077] In some embodiments, please refer to Figure 6 It also provides a target tracking and recognition method, which Figure 6 With the above Figure 3 The difference is that the method also includes: S211. Match the facial features in the facial feature set with the features in the preset facial database to obtain the successfully matched facial features and the unmatched facial features.
[0078] S212. Cluster the unmatched face features with the matched face features to obtain the clustering results of the unmatched face features and the matched face features.
[0079] S213. Assign facial identification information to the successfully matched facial features. This facial identification information is used to identify the identity of the target object.
[0080] S214. Based on the above binding information set, the face identification information and the clustering result, obtain the identity information corresponding to the unmatched face features.
[0081] Specifically, the process involves: obtaining the target identifier information corresponding to the unmatched facial feature, denoted as the first target identifier information; searching for other binding information that is the same as the first target identifier information among the multiple identity and trajectory binding information; searching for a cluster containing the unmatched facial feature in the clustering results; determining whether the successfully matched facial feature exists in the other binding information within the cluster; and assigning the identity information corresponding to the successfully matched facial feature to the unmatched facial feature if it does.
[0082] The difference between this embodiment and the above-described method embodiment is that it proposes a face recognition enhancement method that integrates matching and clustering. This method is used in scenarios combining multi-target tracking and face recognition to complete the identity information of some unmatched faces, thereby improving the overall recognition coverage and accuracy.
[0083] Specifically, firstly, the features in the facial feature set are matched with a preset facial database to distinguish between successfully matched and unmatched facial features (S211). Then, a clustering algorithm (such as K-means) is used to jointly cluster the unmatched and successfully matched facial features to obtain their clustering relationship in the feature space (S212). For successfully matched facial features, their corresponding identity identifier is directly assigned (S213). Next, through a dual association between the clustering results and target tracking identifier information (S214), an attempt is made to infer the identity of unmatched facial features: specifically, the first target identifier information corresponding to the unmatched feature is obtained, and a record matching the target identifier is searched in the existing binding information set; if the successfully matched facial feature in that record is in the same cluster as the current unmatched feature, the identity information corresponding to the successfully matched feature is assigned to the unmatched feature, achieving identity propagation-based completion.
[0084] For example, such as Figure 7As shown, during target tracking, both the first image and the synchronized image can generate corresponding target identification information (e.g., Track ID=1) to identify the trajectory continuity of the same target object. Before performing face feature clustering, the system has obtained several face features extracted from the second image, including unmatched face features (e.g., Face ID=0) and successfully matched face features (e.g., Face ID=1). Subsequently, cluster analysis is performed on the unmatched and successfully matched face features. Through similarity calculation in the feature space, some unmatched face features (Face ID=0) and successfully matched face features (Face ID=1) can be grouped into the same cluster. Since these features are similar in feature distribution and their corresponding target objects have the same trajectory identifier (Track ID=1), the identity information corresponding to the successfully matched face features can be assigned to the unmatched features, thereby completing the identity completion. However, due to the possibility of not capturing clear faces or other interference factors in some second images, some unmatched face features (Face ID=0) may not be grouped into any known identity cluster through clustering. For these unidentified facial features, the system can further search for successfully matched facial features under the same trajectory in the existing clustering results by using their corresponding target identification information (Track ID=1). If they exist, the identity information corresponding to the successfully matched feature can also be assigned to the currently unmatched feature.
[0085] This method, without relying on a comprehensive face database, utilizes the continuity of target tracking and the clustering consistency of facial features to effectively compensate for some blind spots in recognition, and improves the ability to identify individuals in situations with weak features or occlusion. It is suitable for multi-target, low-frequency recognition scenarios such as classroom observation and group behavior analysis. The method in this embodiment is the same as described above. Figure 4 The corresponding embodiments share the same goal: to complete identity information for "unmatched facial features." However, this embodiment utilizes a joint analysis of facial feature clustering and target identification information. Before completing the identity information, it first clusters unmatched and matched facial features, grouping them into the same cluster based on similarity in the feature space. This is then combined with target identification information (Track ID) to further confirm identity. This approach is suitable for identifying and completing weak or blurry images. Figure 4 The corresponding implementation relies entirely on existing identity and trajectory binding information, and inherits identity through the consistency of trajectory identifiers.
[0086] In some embodiments, based on the association between Track ID and Face ID, for Track IDs that are not yet associated with Face ID, the identity information of the target object corresponding to the Track ID can be inferred through the identity and trajectory binding information. In this way, during the tracking of the target object, continuous identity labeling and trajectory identity binding of the target object can be achieved, thereby improving the performance of the multi-target tracking system in terms of the continuity and completeness of identity recognition.
[0087] In some embodiments, the method further includes: matching facial features in a facial feature set with features in a preset facial database to obtain successfully matched facial features and unmatched facial features; assigning facial identification information to the successfully matched facial features, the facial identification information being used to identify the identity of the target object; obtaining target identification information bound to the facial identification information based on multiple identity and trajectory binding information, denoted as second target identification information, the second target identification information being the target identification information bound to the successfully matched facial features; detecting, among the multiple identity and trajectory binding information, whether there exists other target identification information that is the same as the facial identification information but different from the second target identification information; if so, updating the other target identification information different from the second target identification information to the second target identification information. For example, such as Figure 8 As shown, the target identifier information bound to the successfully matched face feature (Face ID=1) is the second target identifier information (TrackID=1). However, there is other target identifier information (Track ID=0) in the target identifier information. In this case, Track ID=0 is modified to Track ID=1 based on the successfully matched face feature (Face ID=1).
[0088] This embodiment utilizes Face ID information to merge multiple different Track IDs, which is suitable for solving the problem of Track ID switching. By merging multiple target identification information bound to the same Face ID information, the problem of target identification information (Track ID) switching caused by rapid target movement or occlusion can be solved, thereby improving the stability and consistency of target tracking results and ensuring the continuous identification of individual identities in a time series.
[0089] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a target tracking and identification device provided in an embodiment of this application. The device 30 includes: The first target tracking module 31 is used to continuously acquire a first image sequence of the target scene and track the target object in the target scene based on the first image sequence. The image processing module 32, during the tracking of the target object, when a face recognition condition is triggered, acquires a second image of the target scene and scales the second image to generate a synchronized image; the resolution of the second image is greater than the resolution of the first image, and the difference between the resolution of the synchronized image and the resolution of the first image is within a preset range. The second target tracking module 33, when a face recognition condition is triggered, replaces the first image with the synchronized image to add the synchronized image to the first image sequence and acquires the target identification information corresponding to the synchronized image. The first image is a frame in the first image sequence, and the acquisition time of the first image is the time when the face recognition condition is triggered. The face recognition module 34 is used to extract features from the face region of the target object based on the second image to obtain face features. The information binding module 35 is used to associate the target identification information corresponding to the synchronized image with the face features to obtain the identity and trajectory binding information of the target object.
[0090] The target tracking and recognition device 30 described above can be a software module. The software module includes several instructions, which are stored in a memory. The processor can access the memory and call the instructions to execute them in order to complete the target tracking and recognition methods described in the above embodiments.
[0091] In some embodiments, the target tracking and identification device 30 can also be constructed from hardware devices. For example, the target tracking and identification device 30 can be constructed from one or more chips, and the chips can work in coordination to complete the target tracking and identification method described in the various embodiments. As another example, the target tracking and identification device 30 can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0092] It should be noted that the target tracking and identification device 30 described above can execute the target tracking and identification method for electronic devices provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the embodiments of the target tracking and identification device 30 can be found in the target tracking and identification method for electronic devices provided in the embodiments of this application.
[0093] Please see Figure 10 , Figure 10This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 40 includes one or more processors 41 and a memory 42. The memory 42 is connected to one or more processors 41, for example, via a bus.
[0094] Processor 41 is configured to support the electronic device 40 in performing the corresponding functions in the methods described in the above method embodiments. Processor 41 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0095] Memory 42 is used to store program code, etc. Memory 42 may include volatile memory (VM), such as random access memory (RAM); memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 42 may also include combinations of the above types of memory.
[0096] The memory 42 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the target tracking and identification method in the embodiments of this application. The processor 41 executes various functional applications and data processing of the target tracking and identification method and the target tracking and identification device by running the non-volatile software programs, instructions, and modules stored in the memory 42, that is, it realizes the functions of each module or unit of the target tracking and identification method and the target tracking and identification device provided in the above method embodiments.
[0097] The memory 42 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the target tracking and identification device. In some embodiments, the memory 42 may include memory remotely located relative to the processor, which can be connected to the target tracking and identification device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0098] The one or more modules are stored in the memory 42. When executed by the one or more processors 41, they perform the target tracking and recognition method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.
[0099] The electronic device 40 in this application embodiment refers to an embedded vision processing terminal with camera acquisition, face recognition, and multi-target tracking functions, specifically including but not limited to: Smart cameras: They have built-in image sensors, processors, and network interfaces, supporting high-resolution video capture and real-time analysis, and are widely used in scenarios such as smart classroom attendance and security monitoring; Edge computing devices integrate dedicated vision processing chips (such as AI accelerators), storage and communication modules, enabling video preprocessing, multi-target tracking and face recognition to be completed locally, reducing dependence on the cloud and making them suitable for application environments with limited computing power and bandwidth. Intelligent terminal devices, such as embedded controllers, industrial computers, or robot vision systems equipped with cameras, can work with cloud services to achieve video analysis and identity verification; Camera systems with cloud-based collaborative capabilities: Some computational tasks are completed locally, and after multi-target tracking and preliminary identification, key frames or feature data are uploaded to the cloud for further processing.
[0100] This electronic device integrates a multi-resolution video acquisition module and an intelligent algorithm processing module to achieve simultaneous replacement of routine tracking of low-resolution images with face recognition triggered by high-resolution images in the above method, thereby effectively balancing the needs of target tracking and face recognition under limited computing power and bandwidth conditions.
[0101] This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 10One of the processors 41 can enable the above one or more processors 41 to execute the target tracking and recognition method in any of the above method embodiments, for example, to execute the method steps described in the above method embodiments and to realize the function of the module described in the above device embodiments.
[0102] This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by the electronic device 41, the electronic device 41 is able to perform the target tracking and recognition method in any of the above method embodiments. For example, it can perform the method steps described in the above method embodiments and realize the functions of the modules described in the above device embodiments.
[0103] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0104] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A target tracking and identification method, characterized by, The method comprises: continuously collecting a first image sequence of a target scene, and tracking a target object in the target scene according to the first image sequence; in the process of tracking the target object, when a face recognition condition is triggered, a second image of the target scene is acquired, and the second image is scaled to generate a synchronization image; the resolution of the second image is greater than the resolution of the first image, and the difference between the resolution of the synchronization image and the resolution of the first image is within a preset range; when the face recognition condition is triggered, the synchronization image is used to replace the first image, so that the synchronization image is added to the first image sequence, and target identification information corresponding to the synchronization image is acquired; the first image is a certain frame of image in the first image sequence, and the first image is collected at the time when the face recognition condition is triggered; feature extraction is performed on a face region of the target object based on the second image to obtain face features; the target identification information corresponding to the synchronization image is associated with the face features to obtain identity and trajectory binding information of the target object.
2. The method of claim 1, wherein, In the process of tracking the target object in the target scene according to the first image sequence, if the face recognition condition is triggered multiple times, the method further comprises: when the face recognition condition is triggered each time, feature extraction is performed on the face region of the target object based on the second image of the target scene obtained each time to obtain face features; the face features obtained each time are added to a face feature set, and the face features obtained each time are bound to target identification information in the corresponding synchronization image to obtain multiple identity and trajectory binding information; wherein the corresponding synchronization image refers to the image generated after scaling the second image collected when the face recognition condition is triggered, and the image generated after scaling is used to replace the first image for tracking the target object.
3. The method of claim 2, wherein, The method further comprises: matching the face features in the face feature set with features in a preset face library to obtain face features that match successfully and face features that do not match successfully; assigning the face features that match successfully to face identification information; the face identification information is used to identify the identity of the target object; based on the multiple identity and trajectory binding information and the face identification information, acquiring identity information corresponding to the face features that do not match successfully.
4. The method of claim 3, wherein, The acquiring of the identity information corresponding to the face features that do not match successfully based on the multiple identity and trajectory binding information and the face identification information comprises: based on the multiple identity and trajectory binding information, acquiring target identification information corresponding to the face features that do not match successfully; the target identification information corresponding to the face features that do not match successfully is first target identification information; in the multiple identity and trajectory binding information, searching for other binding information that is the same as the first target identification information; and based on the other binding information that is the same as the first target identification information, acquiring identity information corresponding to the face features that do not match successfully. If the other binding information contains the binding relationship between the matched face feature and the first target identification information, the identity information corresponding to the matched face feature is assigned to the unmatched face feature.
5. The method of claim 3, wherein, After the step of matching the face features in the face feature set with the features in the preset face library to obtain the matched face features and the unmatched face features, the method further comprises: clustering the unmatched face features and the matched face features to obtain a clustering result of the unmatched face features and the matched face features; The method further comprises: obtaining the identity information corresponding to the unmatched face feature based on the plurality of identity and trajectory binding information and the face identification information.
6. The method of claim 5, wherein, The method further comprises: obtaining the identity information corresponding to the unmatched face feature based on the plurality of identity and trajectory binding information, the face identification information, and the clustering result. The method further comprises: obtaining the identity information corresponding to the unmatched face feature based on the plurality of identity and trajectory binding information, the face identification information, and the clustering result. The method further comprises: obtaining the target identification information corresponding to the unmatched face feature, denoted as first target identification information; 7. The method of claim 2, wherein, searching for other binding information identical to the first target identification information in the plurality of identity and trajectory binding information; searching for a clustering cluster containing the unmatched face feature in the clustering result; determining whether the matched face feature in the clustering cluster exists in the other binding information; if the matched face feature exists, assigning the identity information corresponding to the matched face feature to the unmatched face feature. The method further comprises: matching the face features in the face feature set with the features in the preset face library to obtain the matched face features and the unmatched face features; 8. A target tracking and identification apparatus, characterized by, assigning the matched face features to face identification information; the face identification information is used to identify the identity of the target object; obtaining target identification information bound to the face identification information based on the plurality of identity and trajectory binding information, denoted as second target identification information; the second target identification information is the target identification information bound to the matched face feature; detecting whether other target identification information identical to the face identification information but different from the second target identification information exists in the plurality of identity and trajectory binding information; if the other target identification information exists, updating the other target identification information different from the second target identification information to the second target identification information. The method further comprises: a first target tracking module configured to continuously collect a first image sequence of a target scene and track a target object in the target scene according to the first image sequence; an image processing module configured to, when detecting that a face recognition condition is triggered in the process of tracking the target object, acquire a second image of the target scene, and scale the second image to generate a synchronous image. The resolution of the second image is greater than the resolution of the first image, and the resolution of the synchronization image is within a preset range of the resolution of the first image. The second target tracking module is configured to replace a first image with the synchronization image to add the synchronization image into the first image sequence and obtain target identification information corresponding to the synchronization image when the face recognition condition is triggered, the first image being a certain frame of image in the first image sequence, and the first image being captured at a time when the face recognition condition is triggered; The face recognition module is configured to perform feature extraction on a face region of the target object based on the second image to obtain a face feature. The information binding module is configured to associate the target identification information corresponding to the synchronization image with the face feature to obtain identity and trajectory binding information of the target object.
9. An electronic device, comprising: The memory and the processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, and the processor being configured to cause the electronic device to implement the method according to any one of claims 1 to 7 when executing the one or more computer programs. The non-volatile computer readable storage medium stores computer executable instructions, and when the computer executable instructions are executed by the electronic device, the electronic device executes the method according to any one of claims 1 to 7.
10. A non-transitory computer readable storage medium, comprising:
Citation Information
Patent Citations
Target tracking method and target tracking device
CN102867311A
Clustering face recognition method and device based on tracking
CN119274223A
Target Object Tracking Method and Apparatus, and Storage Medium
US20200327678A1