Method and system for object tracking using online learning

By using an online learning model of global pattern matching in computer systems, combining the motion factor of local patterns and the appearance factor of global patterns, the problem of tracking accuracy degradation caused by target occlusion and fast movement in the prior art is solved, and more accurate and efficient object tracking is achieved.

CN113454640BActive Publication Date: 2025-05-02NAVER CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080014716.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-28
Filing Date
2020-02-11
Publication Date
2025-05-02
Estimated Expiration
2040-02-11

AI Technical Summary

Technical Problem

When existing object pose estimation technology deals with occlusion, fast movement and other situations, it is difficult to accurately track the target, resulting in a decrease in ID offset and tracking accuracy.

Method used

By using an online learning model of global pattern matching in a computer system, the global pattern of each goal is learned and tracked in combination with the motion factor of the local pattern and the appearance factor of the global pattern.

Benefits of technology

More accurate target tracking is achieved, ID offset and tracking error are reduced, and the accuracy and efficiency of object pose estimation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113454640B_ABST
    Figure CN113454640B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for object tracking using online learning. The object tracking method comprises the following steps: learning a classifier model using global pattern matching; and classifying and tracking each target through online learning including the classifier model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following description relates to an object tracking technology. Background Art

[0002] Object pose estimation is an important part of computer vision, human-computer interaction and other related fields. For example, when the user's head is regarded as the object to be estimated, the rich personalized information that the user wants to express can be obtained by estimating the user's continuous head pose. In addition, the estimated result of the object (such as the head) pose can be used for human-computer interaction. For example, by estimating the head pose, the user's visual focus can be obtained, and more effective human-computer interaction can be performed.

[0003] As an example of the object pose estimation technology, Korean Patent Publication No. 10-2008-0073933 (published on Aug. 12, 2008) discloses a technology for automatically tracking the motion of an object in real time from an input video image and determining the pose of the object.

[0004] The object pose estimation methods currently used are generally divided into tracking-based methods and learning-based methods.

[0005] The tracking-based method estimates the pose of an object by matching the current frame (Current Frame) and the previous frame (Previous Frame) in a video sequence into a paired method.

[0006] Learning-based methods generally define object pose estimation as a classification method or a regression method, train with labeled samples and use the obtained training model to estimate the pose of the object. Summary of the invention

[0007] 1. Technical issues to be resolved

[0008] The global pattern of each object can be learned by an online learning model to which a classifier that classifies the ID (identification code) of each object is added.

[0009] You can create learning data for each target accumulated over time and use the learning data to learn a classifier model.

[0010] A motion factor based on a local pattern and an appearance factor based on a global pattern can be used together for tracking.

[0011] (II) Technical solution

[0012] The present invention provides an object tracking method executed in a computer system, wherein the computer system comprises: at least one processor configured to execute computer-readable instructions included in a memory, and the present invention provides an object tracking method, wherein the object tracking method comprises the following steps: at least one of the processors learns a classifier model using global pattern matching; and at least one of the processors classifies and tracks each target through online learning including the classifier model.

[0013] According to one aspect, the learning step may include the following steps: learning a global pattern of each target through a learning model to which a classifier for classifying each target is added.

[0014] According to another aspect, the learning step may include the following steps: creating learning data of each target accumulated on a time axis by sample mining, and repeatedly learning the classifier model using the accumulated learning data.

[0015] According to another aspect, the learning step may include the following steps: distinguishing valid periods where the target exists within the entire continuous interval of the input video; after marking any one of the valid periods, creating learning data to learn the classifier model; and after marking the next valid period, creating learning data and merging it with the previously created learning data to create accumulated learning data and repeatedly learn the classifier model.

[0016] According to yet another aspect, the annotation utilizes a similarity matrix of the classifier model calculated based on an appearance factor according to a global pattern of the target.

[0017] According to yet another aspect, the learning step may further include the following step: marking invalid periods other than the valid period by using the classifier model learned using the valid period.

[0018] According to another aspect, the according step may include the following steps: finding the position of the target for all frames of the input video and calculating the coordinates of the key points of each target; calculating the matching score between the targets in adjacent frames using the coordinates of the key points of each target; and performing posture matching between frames based on the matching scores between the targets.

[0019] According to yet another aspect, the step of performing pose matching may include performing the pose matching using a similarity matrix calculated based on a motion factor of a box representing the position of the target.

[0020] According to yet another aspect, the matching score may represent a degree of proximity between an object in a previous frame and an object in a next frame.

[0021] According to another aspect, the tracking step may further include: performing at least one post-processing process of removing the error of the posture matching by measuring the error based on the bounding box representing the position of the target, correcting the error of the posture matching by interpolation, and smoothing the posture matching based on a moving average.

[0022] The present invention provides a computer-readable recording medium, characterized in that a program for executing the object tracking method in a computer is recorded therein.

[0023] The present invention provides a computer system, comprising: a memory; and at least one processor, configured to be connected to the memory and execute computer-readable instructions included in the memory, at least one of the processors processing the following processes: learning a classifier model using global pattern matching; and classifying and tracking each target through online learning including the classifier model.

[0024] (III) Beneficial effects

[0025] According to an embodiment of the present invention, the global pattern of each target may be learned through an online learning model to which a classifier for classifying the ID of each target is added.

[0026] According to an embodiment of the present invention, it is possible to create learning data for each target accumulated on a time axis, and learn a classifier model using the learning data.

[0027] According to an embodiment of the present invention, a motion factor according to a local mode and an appearance factor according to a global mode may be used together for tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a block diagram for explaining an example of the internal configuration of a computer system according to one embodiment of the present invention.

[0029] Figure 2 is a diagram showing an example of components that a processor of a computer system according to one embodiment of the present invention may include.

[0030] Figure 3is a flow chart illustrating an example of an object tracking method that can be performed by a computer system according to an embodiment of the present invention.

[0031] Figure 4 An example of a process of calculating key point coordinates of a target according to an embodiment of the present invention is shown.

[0032] Figure 5 An example of measuring the Intersection over Union (IoU) ratio representing the degree of overlap between regions according to an embodiment of the present invention is shown.

[0033] Figure 6 to Figure 7 An example of a process of learning a global pattern of an objective according to an embodiment of the present invention is shown.

[0034] Best Mode for Carrying Out the Invention

[0035] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0036] Embodiments of the present invention relate to techniques for tracking the position of an object through an online learning model.

[0037] In an embodiment including the contents specifically disclosed in this specification, the global patterns of each pattern can be learned by adding an online learning model with a classifier that classifies the ID of each target, thereby achieving great advantages in terms of accuracy, efficiency, and cost reduction.

[0038] Figure 1 is a block diagram for explaining an example of the internal configuration of a computer system according to an embodiment of the present invention. Figure 1 The computer system 100 is implemented.

[0039] like Figure 1 As shown, the computer system 100 may include components for executing the object tracking method, such as a processor 110 , a memory 120 , a permanent storage device 130 , a bus 140 , an input and output interface 150 , and a network interface 160 .

[0040] The processor 110 is a component for object tracking, which may include any device capable of processing a sequence of instructions or be a part of such a device. The processor 110 may include, for example, a computer processor, a processor within a mobile device or other electronic device, and / or a digital processor. The processor 110 may be included in, for example, a server computing device, a server computer, a series of server computers, a server farm, a cloud computer, a content platform, etc. The processor 110 may be connected to the memory 120 via a bus 140.

[0041] The memory 120 may include volatile memory, permanent memory, virtual memory, or other memory for storing information used or output by the computer system 100. The memory 120 may include, for example, random access memory (RAM) and / or dynamic RAM (DRAM). The memory 120 may be used to store any information such as state information of the computer system 100. The memory 120 may also be used to store instructions of the computer system 100 including, for example, instructions for tracking objects. The computer system 100 may include one or more processors 110 as needed or appropriate.

[0042] The bus 140 may include a communication infrastructure that enables interaction between the various components of the computer system 100. The bus 140 may transmit data, for example, between components of the computer system 100, such as between the processor 110 and the memory 120. The bus 140 may include wireless and / or wired communication media between the components of the computer system 100 and may include parallel, serial, or other topological arrangements.

[0043] Persistent storage 130 may include components such as memory or other persistent storage used by computer system 100 to store data for specified extended intervals (e.g., compared to memory 120). Persistent storage 130 may include non-volatile main memory such as used by processor 110 within computer system 100. Persistent storage 130 may include, for example, flash memory, a hard disk, an optical disk, or other computer-readable media.

[0044] The input-output interface 150 may include an interface of a keyboard, a mouse, a voice command input, a display, or other input or output devices. Input for configuration commands and / or object tracking may be received through the input-output interface 150 .

[0045] The network interface 160 may include one or more interfaces to a network such as a local area network or the Internet. The network interface 160 may include an interface for a wired or wireless connection. Input for configuration instructions and / or object tracking may be received through the network interface 160.

[0046] In addition, in another embodiment, the computer system 100 may further include Figure 1However, it is not necessary to clearly show most of the prior art components. For example, the computer system 100 may include at least a portion of input and output devices connected to the above-mentioned input and output interface 150, or may also include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, a database, etc.

[0047] When object tracking is performed in an actual image, there may be a problem that the comparison cannot be made correctly or even the same object is recognized as a different object due to occlusion of the object by another object or the object moving quickly and appearing blurred.

[0048] For these reasons, pose estimation for existing object tracking is not 100% accurate and has the limitation of estimating through similar positions with local patterns. Therefore, there may be a problem that the ID of the target is shifted, and the accumulation of these small errors will lead to results far away from the target object.

[0049] In the present invention, the target object can be tracked more accurately by utilizing an online learning model of global pattern matching.

[0050] In this specification, although human tracking is taken as a representative example, it is not limited thereto, and various things or other types of objects other than humans may be applied.

[0051] Figure 2 is a diagram showing an example of components that a processor of a computer system according to one embodiment of the present invention may include, Figure 3 is a flow chart illustrating an example of an object tracking method that can be performed by a computer system according to an embodiment of the present invention.

[0052] like Figure 2 As shown, the processor 110 may include an estimation unit 210, a similarity calculation unit 220, a matching unit 230, a post-processing unit 240, and a position providing unit 250. Such components of the processor 110 may be representations of different functions performed by the processor 110 according to control instructions provided by at least one program code. For example, the estimation unit 210 may be used as a functional representation of operating to control the computer system 100 so that the processor 110 performs posture estimation.

[0053] The processor 110 and the components of the processor 110 may execute Figure 3The object tracking method includes steps S310 to S350. For example, the processor 110 and the components of the processor 110 may execute the code of the operating system included in the memory 120 and the instructions according to the at least one program code. Here, the at least one program code may correspond to the code of the program for processing the object tracking method.

[0054] The object tracking method may not be performed in the order shown, and some steps may be omitted or may further include additional processes.

[0055] The processor 110 may load the program code stored in the program file for the object tracking method into the memory 120. For example, the program file for the object tracking method may be stored in a program file for the object tracking method. Figure 1 The processor 110 may control the computer system 110 to load the program code from the program file stored in the permanent storage device 130 to the memory 120 through the bus. At this time, the processor 110 and the estimation unit 210, the similarity calculation unit 220, the matching unit 230, the post-processing unit 240 and the position providing unit 250 included in the processor 110 may be instructions for executing corresponding parts of the program code loaded into the memory 120 to perform different functional performances of the processor 110 of the subsequent steps S310 to S350. In order to perform steps S310 to S350, the processor 110 and the components of the processor 110 may directly process the operation according to the control instruction or control the computer system 100.

[0056] In step S310, when a video file is input, the estimation unit 210 can perform posture estimation with the input video as the object. At this time, the estimation unit 210 can find the position of the person corresponding to the target object for all frames of the input video and calculate the coordinates of the key points of each person.

[0057] For example, refer to Figure 4 After finding the position of the target person in all frames constituting the input video, the coordinates of 17 positions of the found person, such as the head, left and right shoulders, left and right elbows, left and right hands, left and right knees, and left and right feet, can be used as key points. For example, the estimation unit 210 can find people in the frame by using a human detection algorithm based on you only look once (YOLO), and calculate the coordinates of the key points of each person in a top-down manner.

[0058] Refer again Figure 3In step S320, the similarity calculation unit 220 can calculate the pose similarity between adjacent frames based on the key point coordinates of each person in each frame. In other words, the similarity calculation unit 220 can calculate a matching score representing the pose similarity between the characters in two adjacent frames. At this time, the matching score can represent an indicator of the degree of proximity between the K persons in the nth frame and the K‵ persons in the n+1th frame.

[0059] In particular, in the present invention, the matching score representing the similarity of posture may include a motion factor according to a local pattern and an appearance factor according to a global pattern. The model for calculating the matching score may be implemented as an online learning model to which a classifier for classifying the ID of each target is added, and the global pattern of each target may be learned through the online learning model.

[0060] The classifier model according to the present invention can accumulate learning data of each target in time axis, and as an example of learning data, all key points of the target can be included. In other words, the global pattern of each target can be learned by the classifier model. At this time, the classifier for learning the global pattern can apply all network models that can be classified.

[0061] The motion factor can be calculated based on the bounding box IoU (Intersection Over Union) and the pose IoU of the target location area. Figure 5 As shown, IoU represents the degree of overlap between two regions, thereby measuring the accuracy of the predicted value when detecting an object with a ground truth (actual object boundary). In addition, the appearance factor can be calculated by utilizing sample mining for judging objective probability and global pattern matching based on online learning.

[0062] Refer again Figure 3 In step S330, the matching unit 230 may perform posture matching between the frames using the result of step S320. In other words, the matching unit 230 may actually match the i-th frame (ie, the target position) of the n-th frame with the j-th frame of the n+1-th frame based on the matching score indicating the similarity of the posture.

[0063] The matching unit 230 may perform pose matching using a matching algorithm such as the Hungarian method. The matching unit 230 may match each frame by first calculating a similarity matrix between adjacent frames and then optimizing it using the Hungarian method. At this time, a motion factor representing IoU may be used to calculate a similarity matrix for pose matching.

[0064] In step S340, the post-processing unit 240 may perform a post-processing process including eliminating false detections on the gesture matching result of step S330. For example, the post-processing unit 240 may remove the matching error by measuring the error based on the bounding box IoU. In addition, the post-processing unit 240 may correct the matching error using interpolation, and may further perform smoothing for the gesture matching based on a moving average, etc.

[0065] In step S350, the position providing unit 250 can provide the position of each target matched according to the posture as the tracking result. The position providing unit 250 can provide the coordinate value of each target as an output. The area showing the target position is called a frame, and at this time, the position of the target can be provided by the position coordinates within the frame. The position coordinates of the target can be marked in the form of [left line X coordinate, upper line Y coordinate, right line X coordinate, lower line Y coordinate], [left line X coordinate, upper line Y coordinate, rectangle width, rectangle height], etc.

[0066] Figure 6 to Figure 7 An example of a process of learning a global pattern of an objective according to an embodiment of the present invention is shown.

[0067] Figure 6 to Figure 7 The sample mining process is shown.

[0068] Reference Figure 6 , 1. The model result value is the result of applying the existing tracking technology using the motion factor. In the present invention, the appearance factor can be calculated for the second time to track the object after the existing tracking is applied for the first time.

[0069] 2. You can distinguish between valid periods and invalid periods by defining the valid period and invalid period in the entire video. The valid period refers to the period where all targets exist. Figure 6 The shaded area in the figure indicates the valid range.

[0070] Reference Figure 7 ,3. Learning examples can be added by repeatedly training the model and using the model to specify the label for the next valid interval.

[0071] The learning data uses the entire continuous interval consisting of multiple frames. At this time, the input unit of the learning model can be a mini-batch sampled in the entire continuous interval, and the size of the mini-batch can be determined as a predetermined default value or determined by the user.

[0072] The learning data includes a frame image including the target position and the ID of the corresponding target. The frame image refers to an image in which only the area representing the position of each person is cut out from the entire image.

[0073] When a frame image including an arbitrary person is given, the output of the learning model (network) is the probability value of each target ID of the frame image.

[0074] like Figure 7 As shown, in the first step (1st) of learning, the longest valid interval 710 is used to create learning data of the first interval, and the learning data of the first interval is used to learn the model. At this time, the learning data can be directly annotated with the results obtained by using the existing object tracking technology, or the frame image and the target ID can be used as the learning data.

[0075] In the second step (2nd), after labeling the next target interval, i.e., the second longest valid interval 720, using the model learned in the first interval, the learning data of the second interval is created. Then, the learning data of the first interval and the second interval are merged to create cumulative learning data, and the model is learned again using it.

[0076] After learning the valid intervals by repeating this method, the invalid intervals are predicted (labeled) by the model learned from the valid intervals.

[0077] In the above-mentioned annotation process, after calculating the similarity matrix for the classifier model, the similarity matrix can be used to match each frame. At this time, the similarity of the classifier model can be calculated using the appearance factor instead of the motion factor.

[0078] As described above, according to an embodiment of the present invention, the global pattern of each target can be learned by adding an online learning model to which a classifier for classifying the ID of each target is added, and learning data of each target accumulated along a time axis is created, and the classifier model is learned using the learning data, thereby enabling the motion factor according to the local pattern and the appearance factor according to the global pattern to be used together for object tracking.

[0079] The above-described device can be implemented as a combination of hardware components, software components and / or hardware components and software components. For example, the device and components described in the embodiment can be implemented using one or more general-purpose computers or special-purpose computers, such as processors, controllers, arithmetic logic units (arithmetic logic unit, ALU), digital signal processors (digitl signal processors), microcomputers, field programmable gate arrays (field programmable gate array, FPGA), programmable logic units (programmable logic unit, PLU), microprocessors or any device capable of executing and responding to instructions (instruction). The processing device can execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device can also access, store, operate, process and generate data in response to the execution of the software. For ease of understanding, although there is a description of the use of a processing device, it will be recognized by those of ordinary skill in the art that the processing device may include multiple processing elements (processing element) and / or multiple types of processing elements. For example, the processing device may include multiple processors or a processor and a controller. In addition, other processing configurations (processing configurations) such as parallel processors (parallel processors) can also be used.

[0080] Software may include a computer program, code, instructions, or a combination of one or more thereof, and may configure a processing device or independently or collectively instruct a processing device to operate the software as required. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device to interpret or provide instructions or data to a processing device through a processing device. Software may be distributed on networked computer systems and stored or executed in a distributed manner. Software and data may be stored in one or more computer-readable recording media.

[0081] The method according to the embodiment can be implemented in the form of program instructions that can be executed by various computer devices and recorded in a computer-readable medium. At this point, the medium can continue to store the computer executable program, or can be temporarily stored for execution or downloading. In addition, the medium can be various recording devices or storage devices in the form of a single or multiple hardware combinations, but is not limited to the medium directly connected to the computer system, and can also be distributed on the network. Examples of media include magnetic media such as hard disks, floppy disks and tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as optical disks (magneto-optical medium), and ROM, RAM, flash memory, etc., so that they can be configured to store program instructions. In addition, examples of other media can include recording media or storage media managed by application stores that distribute applications or sites that provide or distribute various other software and servers. DETAILED DESCRIPTION

[0082] As described above, although the embodiments are described through limited embodiments and drawings, a person skilled in the art of the present invention can make various modifications and variations based on the above description. For example, even if the described techniques are performed in a different order than the described method, and / or the described systems, structures, devices, circuits, etc., are combined or combined in a different form than the described method, or replaced or substituted by other components or equivalents, appropriate results can be achieved.

[0083] Therefore, other implementations, other examples, and equivalents to the claims also fall within the scope of the claims.

Claims

1. An object tracking method, which is an object tracking method executed in a computer system, wherein: The computer system comprises: at least one processor configured to execute computer-readable instructions contained in the memory, The object tracking method comprises the following steps: learning, by at least one of the processors, a classifier model using global pattern matching; and At least one of the processors classifies and tracks each target by online learning including the classifier model, The learning step comprises the following steps: Distinguish the valid intervals where all targets exist in the entire continuous interval of the input video; After marking the longest valid interval among the valid intervals, creating learning data and learning the classifier model; After annotating the second longest valid interval using the classifier model learned in the longest valid interval, creating learning data and merging it with previously created learning data to create cumulative learning data, and using it to learn the classifier model again; and The effective interval is learned by repeating the steps, thereby completing the learning of the classifier model.

2. The object tracking method according to claim 1, wherein: The learning step comprises the following steps: The global pattern of each object is learned by adding a learning model to which a classifier for classifying each object is added.

3. The object tracking method according to claim 1, wherein: The learning step comprises the following steps: Learning data of each target accumulated on a time axis is created through sample mining, and the classifier model is repeatedly learned using the accumulated learning data.

4. The object tracking method according to claim 1, characterized in that: The annotation utilizes a similarity matrix of the classifier model calculated based on an appearance factor according to a global pattern of the target, wherein the appearance factor is used as a matching score representing posture similarity, and is calculated by utilizing sample mining for judging objective probability and global pattern matching based on online learning.

5. The object tracking method according to claim 1, wherein: The learning step further comprises the following steps: The intervals outside the valid interval are labeled by using the classifier model learned using the valid interval.

6. The object tracking method according to claim 1, wherein: The tracking step includes the following steps: Find the location of the target for all frames of the input video and calculate the coordinates of the key points of each target; Calculating the matching scores between objects in adjacent frames using the coordinates of the key points of each object; and Pose matching between frames is performed based on the matching scores between the objects.

7. A computer-readable recording medium, characterized in that: A program for executing the object tracking method according to any one of claims 1 to 6 in a computer is recorded.

8. A computer system comprising: Memory; as well as at least one processor configured to communicate with the memory and execute computer-readable instructions included in the memory, At least one of the processors processes the following: learning a classifier model using global pattern matching; and Classifying and tracking each target by online learning including the classifier model, The learning process includes the following steps: Distinguish the valid intervals where all targets exist in the entire continuous interval of the input video; After marking the longest valid interval among the valid intervals, creating learning data and learning the classifier model; After labeling the second longest valid interval using the classifier model learned in the longest valid interval, creating learning data and merging it with previously created learning data to create cumulative learning data, and using it to learn the classifier model again; and The effective interval is learned by repeating the process, thereby completing the learning of the classifier model.

9. The computer system according to claim 8, characterized in that: The learning process learns the global pattern of each target through a learning model added with a classifier for classifying each target.

10. The computer system according to claim 8, characterized in that The learning process creates learning data of each target accumulated on a time axis through sample mining, and repeatedly learns the classifier model using the accumulated learning data.

11. The computer system according to claim 8, characterized in that: The annotation utilizes a similarity matrix of the classifier model calculated based on an appearance factor according to a global pattern of the target, wherein the appearance factor is used as a matching score representing posture similarity, and is calculated by utilizing sample mining for judging objective probability and global pattern matching based on online learning.

12. The computer system according to claim 8, wherein: The learning process further includes the following process: The intervals outside the valid interval are labeled by using the classifier model learned using the valid interval.

13. The computer system according to claim 8, wherein: The tracking process includes the following steps: Find the location of the target for all frames of the input video and calculate the coordinates of the key points of each target; The coordinates of the key points of each target are used to calculate the matching scores between targets in adjacent frames; as well as Pose matching between frames is performed based on the matching scores between the objects.

Citation Information

Patent Citations

  • Object tracking method and apparatus, and object pose information calculating method and apparatus

    KR1020080073933A

  • Information processing unit, control method, program

    JP2017117139A