Method of tracking person based on user interaction and system for the same
By extracting features and clustering individuals in video content while incorporating user interaction, the method addresses limitations of conventional person recognition, achieving improved accuracy and adaptability in recognizing individuals across varying appearances.
Patent Information
- Application Number
- PCT/KR2024/018036
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-14
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-22
AI Technical Summary
Conventional person recognition methods in video content face challenges such as difficulty in recognizing individuals without clear photos, issues with varying facial expressions and directions, and impracticality of obtaining photos from multiple angles.
The method involves extracting features of persons in video content, performing similarity calculations to cluster the same individuals, and constructing a character database. This approach includes user interaction to improve accuracy and allows for updating person clusters based on user input.
This solution enhances person recognition accuracy in video content, facilitates easy addition of new persons to the database, and allows for effective handling of changes in appearance, such as makeup or aging, thereby supporting video editing and search functionalities.
Smart Images

Figure KR2024018036_22052025_PF_FP_ABST
Abstract
Description
METHOD OF TRACKING PERSON BASED ON USER INTERACTION AND SYSTEM FOR THE SAME
[0001] The present disclosure relates to a method of tracking a person based on user interaction and a system for the same, and more particularly, to a method of constructing a character database for a single video content item by extracting features of persons appearing in video content and tracking and clustering respective characters, and a system for the same.
[0002] Video content items are produced by shooting and editing them at various scales from a number of locations and directions using multiple cameras, and video content produced through this production method consists of various shots, making it very difficult to find all the sections where characters appear using a conventional method.
[0003] Conventional person recognition methods have been carried out to compare a photo of a person to be recognized with a person (persons) appearing in video content so as to determine whether the person in the photo appears in the video content. However, the method has limitations in that it cannot recognize a person at all when it is difficult to obtain a photo of the person or a person who has previously appeared is wearing makeup, and moreover, the method of comparing with a photo of a person also has a problem in that the recognition performance is high when it is close to the direction and expression of the face in the photo, and the recognition performance is low otherwise. In order to solve the foregoing problem, it is necessary to secure photos captured from various directions while making various facial expressions, but it is practically impossible to secure those photos every time for each person to be recognized.
[0004] As such, the present disclosure is proposed in consideration of the problems of the conventional person recognition methods, and is intended to achieve better person recognition in video content than conventional methods.
[0005] An aspect of the present disclosure is to extract features of persons detected in video content, and then perform a similarity calculation so as to cluster a same person, thereby ultimately constructing a database of characters in the video content.
[0006] In particular, an aspect of the present disclosure is to prevent an error in an output due to a method of automatically detecting a person, that is, a result of incorrectly identifying a character. In addition, an aspect of the present disclosure is to enable a user to accurately identify a character by allowing him or her to participate in a person identification process.
[0007] In some example embodiments, the method of tracking a person in a video by a computing device having a central processing unit and a memory comprises the steps of : (a) receiving an arbitrary video; (b) tracking a specific person in the video, and storing information on the tracked person to generate a person cluster; (c) generating a user page provided to a user by utilizing information stored in the person cluster; and (d) reflecting a user input received from the user to update the person cluster.
[0008] In some example embodiments, the information on the person comprises: a face image of the person, frame information for identifying a frame containing the person, and area information in which the person is located within a specific frame.
[0009] In some example embodiments, the step (b) comprises: (b-1) detecting a face image of a specific person within arbitrary frames constituting the video; and (b-2) tracking a person determined to be similar to the specific person based on the detected face image.
[0010] In some example embodiments, the steps (b-1) and (b-2) are executed in units of a plurality of shots constituting the video.
[0011] In some example embodiments, the method further comprises: (b-3) extracting landmark information from face images of a tracked person; (b-4) aligning the face images using the extracted landmark information; and (b-5) generating a person cluster for the tracked person.
[0012] In some example embodiments, the step (b-4) aligns face images with reference to a face direction template, wherein the face direction template comprises a plurality of face direction images in which the face is three-dimensionally rotated by a unit value in upward, downward, leftward, and rightward directions based on a frontal direction thereof.
[0013] In some example embodiments, the method further comprises, subsequent to the step (b-5), the step of performing a comparison operation on a person cluster of the tracked person and another person cluster to determine whether to combine the person cluster with the other person cluster.
[0014] In some example embodiments, the comparison operation calculates similarities of at least some of representative face images selected by direction in a preset number, each corresponding to the person cluster and the other person cluster.
[0015] In some example embodiments, the method further comprises the step of extracting a facial feature vector of the person based on the aligned face images; wherein the comparison operation calculates a similarity between facial feature vectors of persons corresponding to the person cluster and other person clusters, respectively.
[0016] In some example embodiments, the method further comprises, subsequent to the step (b-5), the step of performing a comparison operation on a face image of an arbitrary person detected from the video and at least one of the face images stored in the person cluster to determine whether to cluster the arbitrary person into the person cluster.
[0017] In some example embodiments, the step (b-1) further detects a body image of the specific person, and the step (b-2) tracks a person based on the detected face image or body image.
[0018] In some example embodiments, the step (b) further comprises:(bb-1) detecting a body image of a third person within arbitrary frames constituting the video; and (bb-2) extracting a body feature vector from the detected body image.
[0019] In some example embodiments, the method further comprises the step of (bb-3) generating a person cluster for the third person.
[0020] In some example embodiments, the method further comprises the step of performing a comparison operation on a person cluster for the third person and the person cluster to determine whether to combine the person clusters with each other.
[0021] In some example embodiments, the step (c) comprises the steps of : extracting a face image of at least one person from information stored in the person cluster; and generating a user page including the extracted face image.
[0022] In some example embodiments, the user page comprises a cluster editing function that allows merging or splitting of person clusters by a user input.
[0023] In another embodiments, the person tracking system comprises a central processing unit and memory, wherein the central processing unit is configured to execute instructions for executing a person tracking method stored in the memory, where the person tracking method comprises: (a) receiving an arbitrary video; (b) tracking a specific person in the video, and storing information on the tracked person to generate a person cluster; (c) generating a user page provided to a user by utilizing information stored in the person cluster; and (d) reflecting a user input received from the user to update the person cluster.
[0024] The present disclosure has an effect of improving accuracy compared to conventional methods of tracking a same person.
[0025] In addition, according to the present disclosure, when a DB for a video search is constructed, there is an effect of easily adding a new person.
[0026] In addition, according to the present disclosure, there is an effect of making it possible to easily set a person as the same person even when there is a change in the style, aging, makeup or the like of a person in a video.
[0027] In addition, according to the present disclosure, there is an effect capable of supporting the work of editing and reproducing a video (short video, data screen, etc.) by person.
[0028] In addition, according to the present disclosure, there is an effect of easily acquiring necessary person and action appearance section information to blur or remove a controversial person or his or her action in video content.
[0029] In addition, according to the present disclosure, as a demand for a video search is rapidly increasing, there is an effect of capable providing a video search result such as user-friendly video classification, summary, recommendation or the like by performing an analysis on an accurate person in advance.
[0030] FIG. 1 is a diagram for easily understanding the concept of a person tracking method according to the present disclosure.
[0031] FIG. 2 is a flowchart sequentially showing a person tracking method according to one embodiment of the present disclosure.
[0032] FIG. 3 shows detailed steps in a person tracking and person cluster generation step.
[0033] FIG. 4 conceptually shows an embodiment of tracking a person in units of shots.
[0034] FIG. 5 shows an example of a face direction template.
[0035] FIG. 6 shows a comparison operation process between person clusters.
[0036] FIG. 7 shows an embodiment of comparing a person cluster and a newly detected face image from a video.
[0037] FIG. 8 shows an example in which a body image, in addition to a person's face, is detected and used to generate a person cluster.
[0038] FIG. 9 shows a process of detecting only a person's body image to generate a separate person cluster.
[0039] FIG. 10 shows an example of a user page provided for user interaction.
[0040] FIG. 11 shows another example of a user page.
[0041] The details of the objects and technical configurations of the present disclosure and operational effects thereof will be more clearly understood from the following detailed description based on the accompanying drawings appended hereto. Hereinafter, embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings.
[0042] Embodiments disclosed herein should not be interpreted as limiting or used to limit the scope of the present disclosure. It is apparent for those skilled in the art that a description including embodiments herein has various applications. Therefore, any embodiments described in the detailed description of the present disclosure are illustrative for better understanding of the present disclosure and are not intended to limit the scope of the present disclosure to the embodiments.
[0043] Functional blocks illustrated in the drawings and described hereunder are only examples of possible implementations. In other implementations, other functional blocks may be used without departing from the concept and scope of the detailed description. Furthermore, one or more functional blocks of the present disclosure are illustrated as separate blocks, but one or more of the functional blocks of the present disclosure may be a combination of various hardware and software elements that execute the same function.
[0044] In addition, an expression that some elements are "included" is an expression of an "open type", and the expression simply denotes that the corresponding elements are present, but should not be construed as excluding additional elements.
[0045] Moreover, in case where it is mentioned that one element is "connected" or "coupled" to the other element, it should be understood that one element may be directly connected to the other element, but another element may be present therebetween.
[0046] First, a person tracking method according to the present disclosure will be briefly described with reference to FIG. 1.
[0047] With reference to the drawings, a person tracking method according to the present disclosure is characterized to detect a face of a person appearing in units of frames or shots constituting a video, track the detected face, and individually cluster a character(characters). Moreover, the method is further characterized to provide information on previously clustered persons to a human, that is, a user with thinking and cognitive abilities, and receive a feedback input for the persons from the user, thereby ultimately generating an accurate database of information on the persons appearing in a video. A left side of the diagram briefly shows detecting and tracking of persons in a video, and a right side of the diagram briefly shows an example of a user page provided to the user. As can be seen from the diagram, the user page may include at least images of persons previously identified in a video, and images of frames in which each person appears, thereby allowing the user to check whether characters in the video have been accurately identified, and if necessary, allowing the user to directly enter information so as to more accurately update information on the characters.
[0048] In this manner, a person tracking method according to the present disclosure proposes a technology for tracking the same person from video content based on user interaction while tracking the same person in a manner recognized by a human. A human may recognize persons in different shots even if he or she doesn't know who the persons are, and in particular, recognize the same persons regardless of sizes thereof or directions from which they are viewed. The present disclosure implements a method of tracking a person in video content using the foregoing concept. Meanwhile, there has been an increase in the use of artificial intelligence algorithms to identify persons in a video, but it is assessed that artificial intelligence algorithms have not yet reached a level where they can accurately identify all the persons in the video. The present disclosure relates to accurately tracking a person in video content by using the fact that a human with thinking abilities can identify, even when the person is not properly recognized by an artificial intelligence algorithm as described above, the person even though the person's face is not clearly visible.
[0049] Meanwhile, prior to going into a detailed description of the present disclosure, it is to be understood that a person tracking method according to the present disclosure can be executed by a system or a computing device having a central processing unit and a memory (in this detailed description, for convenience, a subject executing the person tracking method will be referred to as a 'system'). The type of such a system or computing device may include both a portable terminal such as a smartphone, a PDA, a tablet PC, and a terminal fixedly placed at a predetermined position, such as a desktop PC. The central processing unit may also be referred to as a controller, a microcontroller, a microprocessor, a microcomputer, or the like. Furthermore, the central processing unit may be implemented by hardware or firmware, software, or a combination thereof, and configured to include an appli1cation specific integrated circuit (ASIC) or a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), or a field programmable gate array (FPGA) when implemented using hardware, and configured with firmware or software to include a module, a procedure, a function or the like that performs the foregoing functions or operations when implemented using firmware or software. In addition, the memory may be implemented as Read Only Memory (ROM), Random Access Memory (RAM), Erasable Programmable Read Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory, Static RAM (SRAM), a hard disk drive (HDD), a solid state drive (SSD) or the like.
[0050] In some cases, the system may be a server, and in this case, the server may be a device that stores and executes a program, that is, a set of instructions, for actually implementing a person tracking method according to the present disclosure. The type of the server may be at least one server PC managed by a specific user or may be a type of cloud server provided by another company, that is, a type of cloud server that the user can sign up to use.
[0051] Additionally, in some cases, a person tracking method according to the present disclosure may be executed not on a single system but on a cluster system consisting of a plurality of systems or computing devices. Within the cluster system, a plurality of computing devices required to execute the person tracking method may be respectively set to perform different operations.
[0052] Hereinafter, various embodiments related thereto will be described with reference to the drawings.
[0053] FIG. 2 is a flowchart sequentially showing each step of a person tracking method according to the present disclosure.
[0054] Referring to the diagram, the person tracking method first includes receiving, by the system, an arbitrary video (S100). In this case, the type of the video may include any type of video without limitation in format and content, such as a movie, drama, entertainment, news, documentary, CCTV video, or the like, and such a video may be transmitted to the system through various routes, such as the user's local storage device, network drive, external storage device, online streaming, and broadcast signal reception. In this detailed description, in order to facilitate understanding of the description, a description will be made on the assumption that the video content is a drama.
[0055] Subsequent to receiving a video, the system may execute tracking of a person in the video, and generating a person cluster for the tracked person. Tracking a person means more than simply finding a person appearing in a video, which may also mean identifying and recording the movement(s) and change(s) of a specific person (specific persons) over time. More specifically, a process of tracking a person may include a person identification process that recognizes a person in a video and distinguishes him or her from other objects, a same person recognition process that determines whether a person appearing across a number of frames is a same person (identifies a same person even if the person moves from one shot to another shot or disappears from the screen for a moment and then reappears), a person movement and change tracking process that continuously identifies changes in a person's location, size, posture, expression, clothing, and the like over time, and an information linking process that links information acquired through tracking in a chronological order to identify the appearance sections, actions, or the like, of a specific person.
[0056] FIG. 3 shows step S200 in more detail, and when referring to the diagram, the step S200 may include detecting a face image of a specific person within a frame constituting a video (S210), and tracking a person determined to be similar to the specific person based on the detected face image in the video (S212).
[0057] The detecting of a face image (S210) may be carried out on some or all of frames constituting the video, and for each extracted frame, an operation may be executed to find a face area of a specific person by utilizing a predetermined face detection algorithm. Other known algorithms including deep learning-based object detection algorithms may be used for the face detection algorithm.
[0058] The tracking of a person (S212) may be understood as tracking a same person by connecting faces previously detected in the video in a chronological order. At this step, a similarity of a face detected in each frame may be calculated, and information on a path a person moved along as well as the face may be utilized for person tracking. Additionally, information on a location or area occupied by a person within a frame may also be utilized in a person tracking step. For example, when an area occupied by a specific person within a frame suddenly changes, it may be identified that a shot (scene) is changing or that a different person has appeared at that point. In addition, a person's posture, a color of clothing a person is wearing, a skin color may be used as a reference to track a specific person in a video.
[0059] Meanwhile, the detecting of a face image and tracking of a specific person based thereon may be executed for each of shot. A shot is a unit of video filmed continuously without movement of a camera or scene changes during filming, and a single shot may be easily understood as a single scene. A single video consists of a very large number of frames, and performing face detection and person tracking on all frames may place a significant burden on the system's computational power, and in order to efficiently perform a person tracking method, and ultimately to accurately obtain person information appearing in a video, it is preferable to split the video into a plurality of shot units and then perform person tracking on each shot. FIG. 4 conceptually shows how steps S210 and S212 are performed in units of shots, and it is seen that a first person and a second person are identified as a result of performing person tracking for shot #1 consisting of a plurality of frames, and a second person and a third person are identified as a result of performing person tracking for shot #2.
[0060] Subsequent to performing person tracking for each shot, a process of determining whether persons appearing in different shots are the same person may be additionally carried out. For example, in FIG. 4, it is described that the second person in shot #1 and the second person in shot #2 are the same person, but described a result after a process of determining whether persons in respective shots are identical has already been performed. In a real system, the same person determination operation may include a process of calculating a degree of facial similarity and a degree of physical similarity (body shape, clothing, race, etc.) of the persons identified in each shot.
[0061] Returning again to the description of FIG. 3, subsequent to the step S212, extracting landmark information from face images of persons tracked in a video (S214) and aligning the face images using the extracted landmark information (S216) may be carried out. Landmark information may be understood as location information of major facial features such as eyes, nose, mouth, eyebrows, and jawline, and the information may preferably be expressed as coordinate values. When landmark information is extracted from a plurality of face images, the landmark information may be used to align the face images, and the direction of a face may be adjusted by rotating or resizing a face image based on landmark information.
[0062] For reference, the present disclosure may be implemented such that a face direction template is referenced when aligning face images. FIG. 5 shows an example of a face direction template, according to which the face direction template may include a plurality of face direction images in which the face is three-dimensionally rotated by a unit value in upward, downward, leftward, and rightward directions based on a frontal direction thereof. As mentioned above, face images may be aligned based on landmark information, and if the direction of the face is calculated based on the extracted landmark information (e.g., the direction of the face is estimated using a line connecting centers of two eyes and a direction of a nose), then a face direction image having a high similarity to the direction of the face may be selected from among the face direction templates, and a face image of a specific person may be adjusted to the face direction template by rotating or transforming the face image to match the selected face direction image. When a plurality of face images of a specific person are acquired from a video, they may be rotated or transformed to match all directions of face direction images included in the face direction template, and as a result, the face images of the specific person, that is, aligned face images facing in all directions, may be obtained.
[0063] Subsequent to step S216, extracting a facial feature vector based on the aligned face images may be carried out. A facial feature vector is a vector representation of unique features extracted from a face image, which may be used to distinguish each person's face or calculate similarity. In the present disclosure, the extracting of facial feature vectors is placed after the aligning of face images, which because more accurate and consistent results can be acquired when facial feature vectors are extracted by using aligned face images as input.
[0064] For reference, although this detailed description describes that the extracting of facial feature vectors is carried out subsequent o aligning all face images, the timing of extracting facial feature vectors is not necessarily limited thereto, and it is to be understood that it may also be executed individually when individual face images are acquired, that is, when extraction of facial feature vectors is possible.
[0065] Then, the system may generate a person cluster for the tracked person (S218). A person cluster may refer to a data set that stores information on each person appearing in a video. A single person cluster may typically store information on one specific person. For example, a face image of a person, a frame (frame information) in which a person appears, area information where the person is located within a specific frame, landmark information, a facial feature vector, or the like may be stored.
[0066] Meanwhile, the foregoing person tracking and person cluster generation steps may also be carried out individually for a plurality of other persons appearing in the video.
[0067] In addition, two or more of the plurality of generated person clusters may be merged into a single person cluster depending on the comparison operation results, or when it is determined that heterogeneous persons are included in a single person cluster, the person cluster may be split into two.
[0068] FIG. 6 is a diagram showing that a comparison operation is carried out between two person clusters in the system, and that the two person clusters can be merged into one based on a result thereof. That is, when it is determined that persons appearing in different frames or shots are the same person, the system may merge the person clusters corresponding to that person into one. In order to determine whether to merge, a comparison operation may be carried out, for example, the system may calculate similarity by comparing representative image(s) extracted from each person cluster, or calculate similarity by comparing facial feature vectors stored in each person cluster and merge them if the value is greater than a preset value. For another comparison operation method, the system may infer that two persons corresponding to each person cluster are the same person if they appear in frames or shots that are temporally close, and may also infer that two persons are the same person if they appear in spatially close locations. Additionally, it may be determined whether the person clusters are of the same person by utilizing a person's physical or external appearance information, such as not only his or her face but also clothing, hairstyle, clothes, skin color, and tattoos. Merging between person clusters may be carried out by transferring information on a person included in any one person cluster to another person cluster and storing the information therein.
[0069] For reference, FIG. 7 shows an example in which a new person (person D) appearing in a video is detected and compared with the existing person cluster(s) when a single person cluster is present. This embodiment may be carried out when a new person appears in the process of tracking the same person to determine whether the person belongs to an existing person cluster or a new person cluster should be generated.
[0070] Meanwhile, so far, only embodiments of detecting and tracking a person's face when tracking the person have been described, but in a person tracking method according to the present disclosure, as shown in FIG. 8, a body image of a person may be detected (S220), a body feature vector may be extracted based on the detected body image(s) (S222), and the extracted body feature vector may be stored in a previously generated person cluster, or a new person cluster may be generated and stored therein. The body image may include all visual information that can be acquired from a person's appearance excluding his or her face, such as the person's body shape, clothing, posture, clothing color, and skin color. The diagram shows that step S220 can be carried out as a step subsequent to steps S210 and S212, which means that an operation of detecting a body image of a person can be carried out simultaneously with or subsequent to detecting a face image of a specific person, and can also be carried out simultaneously with or subsequent to tracking a person in a video (S212).
[0071] On the other hand, the system of the present disclosure may be implemented to simultaneously detect and track a face and a body image of a specific person as in FIG. 8, and may also be implemented to detect only a body image of a person (S250), track the person within a frame or shot based on the detected body image only (S252), extract a body feature vector based on a plurality of body images acquired in the process (S254), and generate a person cluster consisting of information on the tracked person, that is, person information based on body features (S256). An embodiment as illustrated in FIG. 9 may be particularly useful when it is difficult to detect a person's face in an image. Difficulties in detecting faces in a video may occur for a variety of reasons. In the case of videos produced to be provided to consumers, such as dramas or movies, the quality of the videos is high, so there may not be many such cases, but in videos with low resolution, such as CCTV videos, videos shot in places with drastic lighting changes, and videos where a person is obscured by an object and cannot be seen clearly, it is often difficult to detect the person's face. In this case, an embodiment as illustrated in FIG. 9 may generate a person cluster based on a body image of a person, and allow the user to later determine which person it is related to, thereby allowing a person in a video to be detected as accurately as possible.
[0072] In the above, various embodiments of tracking a person in a video and generating a person cluster have been described with reference to FIGS. 3 to 9.
[0073] Returning to FIG. 2 again, a person tracking method according to the present disclosure may further include, subsequent to step S200, generating a user page provided to a user by utilizing information stored in the person cluster (S300), and reflecting a user input received from the user to update the person cluster (S400).
[0074] As mentioned in the introduction of the detailed description, the present disclosure is characterized in that a person appearing in a video is identified, and then such a result is provided to the user, and an input is received from the user to accurately construct a person database, wherein step S300 may be understood as generating a page to be provided to the user for such user interaction.
[0075] FIG. 10 shows an example of a user page 1000, which may include a list 1011 of persons tracked in preceding steps, a scene 1012 in which a person appears, an area 1013 in which the appearing person is present within a frame, or a timeline 1014 in which the person appears in a video. In addition, the user page 1000 may be provided with an editing environment that allows the user to directly edit information on a person, and for example, a person name editing function 1015 that allows editing a name of a tracked person, and a cluster editing function (not shown) that allows merging or splitting person clusters may be provided. The cluster editing function may be implemented to merge clusters of two overlapping persons by a drag, for example, when the user makes a mouse input to drag a specific person from the list to overlap another person, or implemented conversely to add the dragged person to a separate person cluster when making a mouse input to drag and drop a specific person on an "add new person" icon.
[0076] On the other hand, FIG. 11 shows an example of another user page 1100. The user page 1100 is characterized in that it is difficult to detect a face of a person, so a person (person E) tracked based on a body image is included in the list, and is characterized in that scenes 1112 in which body images appear in relation to person E, and a timeline 1114 in which sections in which the person E appears are displayed. For reference, the timeline 1114 may also show appearance sections of other persons in addition to person E, but in this diagram, only the appearance sections of the person E are shown to emphasize the person E. Meanwhile, compared to FIG. 10, it is characterized in that a person description 1106 is further included in a text format. The character description 1106 may describe facts that can be determined from the detected body image, and for example, information that can be acquired through an image analysis, such as a name of body part, a skin color, a clothing color, whether a hat is worn, and a hairstyle, may be described. Additionally, the character description 1106 may include an estimated opinion as to which person among persons tracked based on a face is most similar to the person tracked based on the body image. The system may track a specific person by detecting a body image at the same time as detecting a face, and in this manner, when a body image is used together to track a person, it may be possible to calculate which person a person cluster generated based only on the body image is likely to match, and the system may include an estimated opinion thereon on the user page 1100 to help the user edit the person. In addition, since the user determines whether a person is identical based on human cognitive abilities, when displaying information on a person based on a body image (e.g., a pattern of clothing captured in a specific frame, a shape of a hat, etc.) on the user page 1100, and further displaying an estimated opinion on which person it is estimated to resemble, there is an effect of allowing the user to more accurately identify who the body image is.
[0077] Meanwhile, although not shown separately in the diagram, the user page may further include functions for performing image editing or image processing on specific tracked persons in addition to functions for accurately distinguishing persons appearing in a video. For example, when a specific performer becomes controversial after a variety show is broadcast and the person needs to be blurred, the user page provided by the system may display information on the specific performer for which tracking has been completed, and when the performer is selected, image editing or image processing options such as mosaic processing or blurring may be provided to hide the performer within the frame. In addition, a timeline may be displayed on the user page such that scenes or frames in which a specific performer appears can be visually viewed at a glance, and sections in which the specific performer appears can be marked with a color that is easily recognizable to the user, such as red, thereby supporting the user to quickly edit sections in which the performer appears.
[0078] In the above, a person tracking method according to the present disclosure and a system for the same have been described. Meanwhile, the present disclosure is not limited to the foregoing specific embodiments and application examples, it will be of course understood by those skilled in the art that various modifications may be made without departing from the gist of the present disclosure as defined in the following claims, and it is to be noted that those modifications should not be understood individually from the technical concept and prospect of the present disclosure.
Claims
1.A method of tracking a person in a video by a computing device having a central processing unit and a memory, the method comprising:(a) receiving an arbitrary video;(b) tracking a specific person in the video, and storing information on the tracked person to generate a person cluster;(c) generating a user page provided to a user by utilizing information stored in the person cluster; and(d) reflecting a user input received from the user to update the person cluster.2.The method of claim 1, wherein information on the person comprises:a face image of the person, frame information for identifying a frame containing the person, and area information in which the person is located within a specific frame.3.The method of claim 1, wherein the step (b) comprises:(b-1) detecting a face image of a specific person within arbitrary frames constituting the video; and(b-2) tracking a person determined to be similar to the specific person based on the detected face image.4.The method of claim 3, wherein the steps (b-1) and (b-2) are executed in units of a plurality of shots constituting the video.5.The method of claim 3, further comprising:(b-3) extracting landmark information from face images of a tracked person;(b-4) aligning the face images using the extracted landmark information; and(b-5) generating a person cluster for the tracked person.6.The method of claim 5, wherein the step (b-4) aligns face images with reference to a face direction template, wherein the face direction template comprises a plurality of face direction images in which the face is three-dimensionally rotated by a unit value in upward, downward, leftward, and rightward directions based on a frontal direction thereof.7.The method of claim 5, further comprising:subsequent to the step (b-5), performing a comparison operation on a person cluster of the tracked person and another person cluster to determine whether to combine the person cluster with the other person cluster.8.The method of claim 7, wherein the comparison operation calculates similarities of at least some of representative face images selected by direction in a preset number, each corresponding to the person cluster and the other person cluster.9.The method of claim 7, further comprising:extracting a facial feature vector of the person based on the aligned face images;wherein the comparison operation calculates a similarity between facial feature vectors of persons corresponding to the person cluster and other person clusters, respectively.10.The method of claim 5, further comprising:subsequent to the step (b-5), performing a comparison operation on a face image of an arbitrary person detected from the video and at least one of the face images stored in the person cluster to determine whether to cluster the arbitrary person into the person cluster.11.The method of claim 7, wherein the step (b-1) further detects a body image of the specific person, andwherein the step (b-2) tracks a person based on the detected face image or body image.12.The method of claim 11, wherein the step (b) further comprises:(bb-1) detecting a body image of a third person within arbitrary frames constituting the video; and(bb-2) extracting a body feature vector from the detected body image.13.The method of claim 12, further comprising:(bb-3) generating a person cluster for the third person.14.The method of claim 13, further comprising:performing a comparison operation on a person cluster for the third person and the person cluster to determine whether to combine the person clusters with each other.15.The method of claim 1, wherein the step (c) comprises:extracting a face image of at least one person from information stored in the person cluster; andgenerating a user page including the extracted face image.16.The method of claim 15, wherein the user page comprises a cluster editing function that allows merging or splitting of person clusters by a user input.17.A person tracking system including a central processing unit and memory, wherein the central processing unit executes instructions for executing a person tracking method stored in the memory, andwherein the person tracking method comprises:(a) receiving an arbitrary video;(b) tracking a specific person in the video, and storing information on the tracked person to generate a person cluster;(c) generating a user page provided to a user by utilizing information stored in the person cluster; and(d) reflecting a user input received from the user to update the person cluster.
Citation Information
Patent Citations
Information processing apparatus, method of controlling the same, and storage medium
US10762133B2
Person tracking across video instances
US11048919B1
Retrieving Visual Media
US20140193048A1
Method, medium, and apparatus for person-based photo clustering in digital photo album, and person-based digital photo albuming method, medium, and apparatus
US7756334B2
Video retrieval system for human face content
US7881505B2