Image processing methods, image processing devices, programs

The image processing method addresses tracking shifts by switching to individual matching and stable features like the head or face, enhancing tracking stability and accuracy.

JP2026063632APending Publication Date: 2026-04-13CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing image tracking systems are prone to shifting from the intended tracking target to another object due to changes in feature amounts caused by hiding or orientation changes, leading to continuous tracking errors.

Method used

An image processing method and apparatus that includes detecting a first region for tracking, identifying an individual within a second region, and switching to a second tracking method based on individual matching when a mismatch is detected, using features like the upper body for robust tracking and switching to more stable features like the head or face for continued tracking.

Benefits of technology

This approach effectively suppresses further tracking shifts by switching to more stable features, ensuring continuous and accurate tracking even in situations where the target is obscured or changes orientation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063632000001_ABST
    Figure 2026063632000001_ABST
Patent Text Reader

Abstract

When attempting to return a target that has switched places with another target during tracking back to the correct target, if the conditions that made the target prone to switching places have not been resolved, another switching will occur. [Solution] An image processing method comprising: a first detection step of detecting a first region including a subject to be tracked from a first image; a first tracking step of tracking a subject in a second image taken after the first image based on information obtained from the first region; a second detection step of detecting a second region in which an individual of the subject can be identified; a matching step of matching the subject and the individual information of the subject to be tracked based on information obtained from the second region; and a second tracking step of tracking the subject to be tracked based on information obtained from the second region, wherein if the matching determines that a subject other than the subject to be tracked is being tracked in the first tracking step, the method switches from the first tracking step to the second tracking step, starting from the second region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] There is a technique for tracking an object in images sequentially acquired in time series. Generally, an object to be tracked in an image is detected, a feature amount of the detected object is extracted, and an area having a feature amount similar to the feature amount in images acquired thereafter is sequentially specified to track the object. At this time, due to a change in the feature amount caused by the hiding or orientation change of the object, the tracking target may shift from the object to be tracked to another object. In contrast, Patent Document 1 discloses a method of tracking an object by using in combination a technique for specifying an individual to be tracked, such as face authentication.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Even when the subject can be identified as an individual, it is considered that the situation in which tracking is likely to shift continues immediately after the shift. Therefore, even if the tracking target is reset, there is a possibility that another shift will occur. The present invention has been made in view of the above problems, and an object thereof is to suppress another shift after determining a shift in tracking.

Means for Solving the Problems

[0005] One aspect of the present invention is an image processing method comprising: a first detection step of detecting a first region including a subject to be tracked from a first image; a first tracking step of tracking a subject in a second image taken after the first image based on information obtained from the first region; a second detection step of detecting a second region in which an individual of the subject can be identified; a matching step of matching the subject and the individual information of the subject to be tracked based on information obtained from the second region; and a second tracking step of tracking the subject to be tracked based on information obtained from the second region, wherein if the matching determines that a subject other than the subject to be tracked is being tracked in the first tracking step, the method switches from the first tracking step to the second tracking step, starting from the second region.

[0006] Another aspect of the present invention is an image processing apparatus comprising: a first detection means for detecting a first region including a subject to be tracked from a first image; a first tracking means for tracking a subject in a second image taken after the first image based on information obtained from the first region; a second detection means for detecting a second region in which an individual of the subject can be identified; a matching means for matching the individual information of the subject and the subject to be tracked based on information obtained from the second region; and a second tracking means for tracking the subject to be tracked based on information obtained from the second region, wherein during tracking by the first tracking means, it is determined by the individual matching whether a transfer of tracking has occurred, and if the matching determines that the first tracking means is tracking a subject other than the subject to be tracked, the apparatus switches from the first tracking means to the second tracking means, starting from the second region.

[0007] A further aspect of the present invention is a program comprising: a first detection means for detecting a first region containing a subject to be tracked from a first image; a first tracking means for tracking a subject in a second image taken after the first image based on information obtained from the first region; a second detection means for detecting a second region in which an individual of the subject can be identified; a matching means for matching the subject with individual information of the subject to be tracked based on information obtained from the second region; and a second tracking means for tracking the subject to be tracked based on information obtained from the second region, wherein during tracking by the first tracking means, the program determines whether a transfer of tracking has occurred by matching the individual, and if the matching determines that the first tracking means is tracking a subject other than the subject to be tracked, the program causes a computer to switch from the first tracking means to the second tracking means, starting from the second region. [Effects of the Invention]

[0008] According to the present invention, it is possible to suppress a second transfer of tracking after a transfer of tracking has been detected. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of an image processing apparatus according to the first embodiment. [Figure 2] This is a block diagram showing the internal configuration of an image processing apparatus according to the first embodiment. [Figure 3] This diagram shows the internal configuration of the imaging unit of the image processing apparatus according to the first embodiment. [Figure 4] This is a block diagram showing the basic configuration of the process for switching the tracking area when matching an individual in the image processing apparatus according to the first embodiment. [Figure 5] This is a block diagram showing the basic configuration of the process for returning the switched tracking area in the image processing apparatus according to the first embodiment. [Figure 6] This is a flowchart showing the tracking procedure in the image processing apparatus according to the first embodiment. [Figure 7]This is a flowchart showing the procedure for extracting matching features to be used as the matching source for the individual to be tracked in the image processing apparatus according to the first embodiment. [Figure 8] This is a flowchart showing the procedure for switching the tracking area when an individual is identified in the image processing device according to the first embodiment. [Figure 9] This flowchart shows the procedure for returning the tracked portion in the image processing apparatus according to the first embodiment. [Figure 10] This is a conceptual example illustrating a series of processing steps in the image processing apparatus according to the first embodiment. [Modes for carrying out the invention]

[0010] Hereinafter, preferred embodiments of the present invention will be described with reference to the attached drawings. The embodiments described below are examples of how the present invention can be specifically implemented and are one of the specific embodiments of the configuration described in the claims. Here, we will describe a case in which a person captured by a camera lens is continuously tracked by half-pressing the capture button of the imaging device or by instructing on the touchscreen.

[0011] (First Embodiment) The first embodiment will be described using Figures 1 to 10.

[0012] FIG. 1 is a diagram showing an imaging device which is an example of an image processing device in the present embodiment. Reference numeral 100 denotes the housing of a camera which is an imaging device, and it includes an imaging unit 101 for acquiring an image, a shooting button 102, a display unit 103 for displaying an image, and a cross key 104. When the user presses the shooting button 102, the imaging unit 101 receives the light information of the subject with a sensor (imaging element), and an image (digital data) is acquired. The acquired image can be displayed on the display unit 103. The display unit 103 can also display a predetermined related information superimposed on the captured image, or display character information such as a menu. The cross key 104 can be used for the user to switch or select and instruct the display of the content displayed on the display unit 103 by input in the up, down, left, and right directions. The positions, types, and numbers of the buttons, display unit, and imaging unit of the camera are not limited to this, and for example, an infrared sensor, a mode dial, etc. may be further provided.

[0013] FIG. is a block diagram showing the internal configuration of the imaging device 100 according to the present embodiment. The imaging device 100 includes a CPU 201, a RAM 202, a ROM 203, a bus 204, an operation unit 205, a display control unit 206, an imaging unit control unit 207, a digital signal processing unit 208, an encoder unit 209, a memory control unit 210, a communication control unit 211, and an image processing unit 212.

[0014] Reference numeral 201 denotes a central processing unit (CPU) in the imaging device 100. The CPU 201 comprehensively controls each component of the imaging device 100 described below.

[0015] Reference numeral 202 denotes a RAM (Random Access Memory) in the imaging device 100. The RAM 202 functions as the main memory and work area of the CPU 201.

[0016] Reference numeral 203 denotes a ROM (Read Only Memory) in the imaging device 100. The ROM 203 stores a control program and the like executed by the CPU 201.

[0017] Bus 204 is a bus that serves as a transfer path for various types of data in imaging device 100. For example, the digital data acquired by imaging unit 101 is sent to image processing unit 212 via this bus 204.

[0018] Operation unit 205 is an operation unit through which imaging device 100 receives user instructions. In the example of FIG. 1, operation unit 205 includes a shooting button 102, a cross key 104, and the like.

[0019] Display control unit 206 is a display control unit in imaging device 100. In the example of FIG. 1, it performs display control of the captured images and characters displayed on display unit 103. For example, a liquid crystal display is used as display unit 103. Display unit 103 may have a touch screen function, and user instructions using the touch screen may also be used as input signals for operation unit 205.

[0020] Imaging unit control unit 207 is an imaging unit control unit in imaging device 100. Imaging unit control unit 207 performs control of the imaging system based on instructions from CPU 201, such as focusing, opening and closing the shutter, and adjusting the aperture.

[0021] Digital signal processing unit 208 is a digital signal processing unit in imaging device 100. Digital signal processing unit 208 performs various processes on the digital data received via bus 204, such as white balance processing, gamma processing, and noise reduction processing.

[0022] Encoder unit 209 is an encoder unit in imaging device 100, and performs processing to convert digital data into a file format such as JPEG or MPEG.

[0023] Memory control unit 210 is a memory control unit in imaging device 100 and is an interface for connecting imaging device 100 to various media. Here, the various media are, for example, a hard disk, a memory card, a CF card, an SD card, a USB memory, and the like.

[0024] The communication control unit 211 is the communication control unit in the imaging device 100, and is an interface for connecting the imaging device 100 to external devices via wireless LAN or short-range wireless communication to send and receive data. External devices include, for example, PCs and smartphones. If the external device can provide input / output functions to the user, the communication control unit 211 may be used as an input source to the operation unit 205 or as an output destination from the display control unit 206.

[0025] The image processing unit 212 is an image processing unit that detects the upper body, face, and head of a person from images acquired by the imaging unit 101, images output from the digital signal processing unit 208, or images acquired from the media via the memory control unit 210 or the communication control unit 211. Details of the image processing unit 212 will be described later.

[0026] Figure 3 shows the internal configuration of the imaging unit 101 of the imaging device according to this embodiment.

[0027] The imaging unit 101 is composed of a lens group 301, an aperture 302, a shutter 303, a filter group 304, a sensor 305, and an A / D conversion unit 306. The lens group 301 is, for example, a zoom lens, a focus lens, an image stabilization lens, etc. The filter group 304 is, for example, a low-pass filter, an iR cut filter, a color filter, etc. The sensor 305 is, for example, a CMOS sensor, a CCD sensor, or a SPAD sensor. When the sensor 305 detects light from a subject, the detected amount of light is converted into a digital value by the A / D conversion unit 306 and output as digital data to the bus 204. Note that if the sensor 305 is a SPAD sensor, the A / D conversion unit 306 is not a required component.

[0028] The basic configuration of the image processing unit 212 of the imaging device according to this embodiment will be described with reference to Figures 4 and 5.

[0029] In this embodiment, the target of tracking is a person, and the tracking is performed for the purpose of photographing the person with a camera. In this case, the upper body of the person is basically tracked, as the change in feature quantities due to the orientation of the body is minimal. However, the present invention is not limited to tracking the upper body, and the entire body, including the lower body, may also be tracked. Depending on the field of view and subject during shooting, the part that makes it easier to distinguish the target of tracking may be used for feature extraction. In addition, in this embodiment, personal identification by face is performed in conjunction with upper body tracking, but it is assumed that the shooting is performed at a resolution that allows for the acquisition of the feature quantities necessary for such identification.

[0030] The tracking target in this embodiment is not limited to people; it can be any object that has features that allow for individual identification and from which such identifying features can be extracted from the captured image. For example, if an individual can be identified by differences in face or pattern, the tracking target may be an animal or an inanimate object.

[0031] First, the process of determining when the tracking target moves and switching the tracking area in this embodiment will be explained with reference to Figure 4. In Figure 4, 401 is an image acquisition unit that acquires an image including the tracking target. The image including the tracking target is the image data currently input via the imaging unit 101, and is, for example, an image of a person as shown in Figure 10.

[0032] The upper body detection unit 402 detects the upper body of a person from the image acquired by the image acquisition unit 401 and determines its position in the image. In this embodiment, a model that estimates the location of the upper body in an image, composed of a multi-layered neural network, is pre-trained and used. This model can be any known network that is a multi-layered Convolutional Neural Network (CNN) such as VGG or ResNet. However, the implementation method of upper body detection is not limited to CNNs; the architecture is not limited as long as it can estimate the position of the upper body of a person using an image as input.

[0033] In this embodiment, the following explanation will use a model trained to estimate a map in which the value of the pixel at the center of the upper body position is close to 1, and the values ​​of all other pixels are close to 0. For model training, a training image and a ground truth map are prepared in advance. The ground truth map is a map in which the value of the pixel at the center of the upper body position in the training image is set to 1, and the values ​​of all other pixels are set to 0. Training is repeated to minimize the error between the estimated map and the ground truth map. As a loss function to evaluate this error, a commonly known cross-entropy loss is used. Note that it is not necessary to estimate so that only the pixel at the center of the upper body position has a value; the ground truth map may be a region that is blurred to follow a Gaussian distribution, for example, where the value increases towards the center coordinate of the upper body position.

[0034] Furthermore, in order to determine the extent of the upper body, the model is trained to estimate a map that outputs values ​​that can be converted from the pixels surrounding the center position of the upper body to the vertical and horizontal dimensions (number of pixels) of the upper body. In other words, in addition to the map of the center position of the upper body mentioned above, a total of three maps are output: one for the vertical dimension of the upper body and one for the horizontal dimension of the upper body. When training the model to estimate the pixel value maps that represent the vertical and horizontal dimensions, the error between the correct value and the output value can be evaluated, for example, by evaluating the squared error. That is, the Mean Square Error, which is the average of the errors between the correct pixel value and the output value across all pixels, can be used as the loss function, and the model is trained to optimize the output of the vertical and horizontal dimensions maps.

[0035] The face / head detection unit 403, like the upper body detection unit 402, detects the face and head of a person from the image acquired by the image acquisition unit 401. In this embodiment, the face / head detection unit 403 performs face detection and head detection separately. For both face detection and head detection, the position, vertical size, and horizontal size are output. Since both detection methods can be trained in the same way as described in 402, a detailed explanation is omitted.

[0036] In this embodiment, the face is defined as the area including the facial features such as the eyes, nose, and mouth, while the head is defined as the area including the face, hair, neck, and shoulders. As will be described later, when tracking using the face and head occurs when a transfer of information occurs during upper body tracking, head detection is performed in conjunction with face detection in order to obtain tracking features from the head, which contains more information than the face.

[0037] In this embodiment, it is necessary to associate the face and head of the same person. In this case, if the detection tendencies for the face and head differ, it becomes difficult to determine the position of the head of the same person from the position of the face. Therefore, it is possible to train the system to estimate both simultaneously. By making the positional relationship of the face and head provided as training data such that the face and head of the same person have similar positional relationship tendencies, simultaneous estimation of the face and head can be expected. This makes it easier to link parts output from the same person, even if the acquisition site of the feature quantities used for matching with individual information is different from the acquisition site of the feature quantities used for tracking.

[0038] The matching unit 404, when a face is detected by the face / head detection unit 403, compares whether the detected face belongs to the same person as a face previously registered. In this embodiment, the face detected at the start of tracking is registered as a face and its features are acquired. This allows face matching to be used even when the target being tracked changes each time a tracking operation is performed. On the other hand, in shooting scenes where the target being tracked does not change frequently, such as when photographing one's child at a sports day, the face image of the target being tracked may be acquired in advance, and the face features obtained from that face may be registered as the face. The method of acquiring features for matching is not limited to these methods.

[0039] In this embodiment, a multi-layer neural network is trained to perform the personal identification task. For the personal identification task, a large number of face images are assigned an ID to identify an individual, and the network is trained to estimate the ID from the input face. This task can be considered a type of classification task that estimates which class an image belongs to. For classification tasks, a method is known in which the output layer is trained to output the likelihood distribution of each class. This method, which uses cross-entropy loss as the loss function, is a common method for multi-class classification using CNNs. This embodiment also adopts a configuration that follows this method.

[0040] In a model that has been sufficiently trained and can estimate IDs for personal identification with high accuracy, it is known that the output vectors in the intermediate layer preceding the output layer will be similar for the same person. By utilizing this similarity between vectors, personal matching can be performed even for unknown faces (people) that have not been trained. In this embodiment, the output vectors of the intermediate layer of such a model are used as matching features, and the similarity between the feature vectors is used as the matching score. Whether or not two faces are the same is determined by whether or not this score exceeds a certain threshold.

[0041] The tracking unit 405 performs tracking of a person by extracting tracking features from the detected area. The tracking features are obtained from a rectangular area in the image, which is derived from the center position, height, and width of the detected area. As tracking features, it is desirable to use features that are robust to changes in posture and orientation caused by the person's movement and movement. For this reason, features obtained from areas such as the upper body or the whole body are used. For the same reason, features are extracted from the head, which can obtain more information from a wider area, rather than the face, from which information can only be obtained from the front.

[0042] In this embodiment, upper body detection by the upper body detection unit 402 and face / head detection by the face / head detection unit 403 are performed for each image. The tracking features of the target to be tracked obtained from the previous image are acquired from the tracking result holding unit 408, which will be described later, and the detection region having feature quantities similar to those tracking features is set as the tracking target position in the current image. It is desirable that the tracking feature quantities be features that can be extracted at a high frame rate. For example, template features using histograms of the RGB values ​​of each color of each pixel within the detection region can be used as feature quantities.

[0043] However, the method of tracking the target is not limited to the method described above. For example, when acquiring images at a high frame rate, the feature detection process may not keep up, making it difficult to detect the target in each image. In such cases, tracking using a general template matching method may be performed. That is, starting from the upper body region detected first, the target may be tracked by repeatedly greedily searching the area around that region in the next image for regions with feature quantities similar to the tracking features of the upper body region. In this case, the face / head detection unit 403 performs detection processing from the image when it is operational and performs the transfer determination described above.

[0044] The possession determination unit 406 determines whether the tracked target has possessed another person. The matching unit 404 detects the face of the same person as the tracked target, and if the upper body with a high correlation to that face is not being tracked as the tracked target, it is determined that possession has occurred. The correlation between the upper body and the face can be determined, for example, from the distance and size between the center of the upper body and the center of the face, but it is not limited to this, as long as it is possible to determine which is the face of the person whose upper body is being tracked by some method. For example, the joint points of the person can be estimated in the same way as the center of the upper body and face, and the correlation can be determined using the joint points that are close to the face, such as the neck and head.

[0045] The tracking area switching unit 407 switches the tracking method to tracking using the area used for matching, according to the judgment result of the possession determination unit 406. In a situation where possession is occurring, it is difficult to continue tracking with the upper body, but there are areas that are visible enough to be matched. Therefore, the tracking process switches to using the matching area. Note that the purpose of this switch is to continue tracking, and it is sufficient to acquire features for tracking, so the area used for tracking does not have to be the matching area itself, but rather an area with a range equivalent to the matching area. Here, tracking is performed using the head, which is more robust to changes in orientation than the face used for matching. Since the face and head are in close proximity, it is sufficient to select the head whose head center is closest to the face center of the person being tracked. The head corresponding to the face may be selected more accurately by using face size, etc., in combination.

[0046] The tracking result holding unit 408 holds information such as the tracking features, position, and size of the target being tracked in the current image, which are obtained as outputs from the tracking unit 405, the transfer determination unit 406, and the tracking part switching unit 407, as well as the switching state of the tracking part.

[0047] The configuration and processing procedure for switching the tracking unit back to the upper body after switching from upper body tracking to face / head tracking will be explained with reference to Figure 5. In Figure 5, it is assumed that the tracking unit has switched to head tracking at the time of image acquisition. The image acquisition unit 401, upper body detection unit 402, face / head detection unit 403, tracking unit switching unit 407, tracking unit 405, and tracking result holding unit 408 shown in Figure 5 may be the same as those shown in Figure 4, so their explanation will be omitted.

[0048] In Figure 5, the return determination unit 501 uses the position of the head being tracked by the tracking unit 405, the upper body detection result by the upper body detection unit 402, and the head detection result by the face / head detection unit 403 to determine whether or not to return the tracking area from the head to the upper body. Immediately after a transfer of the tracked target occurs, it is considered that the target is in a state where transfers are likely to occur again. Therefore, even if the tracking is returned from head tracking to upper body tracking, there is a possibility that the target will immediately transfer to another person. In this embodiment, a process is performed to associate the head and upper body near the tracked target position in the acquired image with each person. As a result, if the number of associated heads and upper bodies matches, and the overlap between the upper body region of the tracked target and the upper body region of the target where a transfer may occur is sufficiently small, the tracking area is considered to be in a state where another transfer is unlikely to occur, and the tracking area is switched from the head to the upper body.

[0049] The method for determining whether the tracking part is restored is not limited to this; any process that restores the tracking method in situations where possession is unlikely to occur may be used, either through other processes or in combination with other processes. For example, more simply, the conditions for possession may improve after a certain amount of time has passed since possession occurred. Therefore, it is also acceptable to determine whether to restore the tracking part based on whether a predetermined amount of time has passed since the tracking part was switched.

[0050] Next, the processing procedure in the image processing apparatus of this embodiment will be explained using flowcharts from Figures 6 to 9. Figure 6 is a flowchart illustrating the procedure for starting the tracking process using the upper body.

[0051] First, in step S601, the image acquisition unit 401 acquires the initial image captured by the imaging unit 101.

[0052] In step S602, the upper body detection unit 402 detects the upper body of a person, which is the first region, from the first image acquired in step S601 (first detection step).

[0053] In step S603, the tracking target is determined from the upper body detected in step S602. If shooting with a camera, the photographer can determine the tracking target by selecting the detected upper body area displayed on the display unit 103 using touch operation or the directional keys. Alternatively, the upper body of the target can be determined as the tracking target by deciding on the focus target during shooting and half-pressing the shutter button.

[0054] In step S604, the tracking unit 405 extracts tracking features from the upper body image of the target to be tracked, which was determined in step S603. The extracted features, along with the position of the target to be tracked, are stored in the tracking result holding unit 408.

[0055] In step S605, the image acquisition unit 401 acquires the image of the frame following the image from which tracking features were acquired in step S604. If the image is captured by a camera, images are usually acquired sequentially at a frame rate of 30fps or 60fps. However, if processing cannot keep up and it is difficult to process all frames, images may be acquired every few frames to achieve a processing-capable frame rate.

[0056] In step S606, the tracking unit 405 obtains the tracking position in the previous frame's image from the tracking result holding unit 408. Then, from the vicinity of the position of the previously tracked target, it searches for a region similar to the tracking feature quantity obtained in the previous image from the second image obtained in step S605 (first tracking step). In this embodiment, when the image is obtained in step S605, the upper body detection unit 402 detects the upper body and obtains tracking features, and the region of the upper body of the subject that is most similar to the previously tracked feature quantity is set as the tracking position in the current image. Note that the tracking method is not limited to detecting the upper body; tracking features may be obtained for the entire screen and the location with the highest correlation may be selected.

[0057] In step S607, the tracking position information in the tracking result holding unit 408 is updated with the tracking position discovered in step S606.

[0058] In step S608, the tracking feature information in the tracking result storage unit 408 is updated with the tracking features used when the tracking position was determined in step S606. Alternatively, the tracking features may not be updated, and the tracking features of the target acquired at the start of tracking may be searched continuously. Or, the corresponding tracking features of the previous image and the tracking features at the start of tracking may be used in combination, and only the features of the previous image may be updated.

[0059] After that, the system returns to Step S605 and updates the tracking position to continue tracking the target.

[0060] Figure 7 is a flowchart illustrating the procedure for registering facial feature quantities for matching the person to be tracked at the start of the tracking process. This process is performed when the target to be tracked is determined in step S603 as described in the flowchart of Figure 6.

[0061] In step S701, the same image as the one obtained in the flowchart of Figure 6 described above is acquired.

[0062] In step S702, the face / head detection unit 403 detects a face, which is a second region in the image acquired in step S701 that can identify an individual subject (second detection step).

[0063] In step S703, the system identifies the face with a high correlation to the tracking target from among the faces detected in step S703. If a face is identified, the system proceeds to step S704. If a face is not identified, the system returns to step S701 to acquire the next image and repeats image acquisition and face detection until a face with a high correlation to the tracking target is obtained.

[0064] In step S704, the matching unit 404 extracts matching features from the face to be tracked, which was acquired in step S703.

[0065] In step S705, the matching unit 404 registers the matching features extracted in step S704 as the features of the target to be tracked and keeps them therein during tracking. From this point onward, the unit determines whether the face in the sequentially acquired images is that of the person being tracked by looking at the similarity to these registered features of the target to be tracked.

[0066] Figure 8 is a flowchart illustrating the procedure for determining when the target being tracked has moved and switching the tracking area.

[0067] Here, we assume that the upper body tracking process (first tracking step) shown in Figure 6 is underway, and that the feature quantities of the matching source of the tracking target have already been acquired, as shown in Figure 7.

[0068] In step S801, the same image as the one obtained in step S605 in the flowchart of Figure 6 described above is acquired.

[0069] In step S802, the face / head detection unit 403 detects faces in the image acquired in step S801.

[0070] In step S803, the matching unit 404 identifies the face of the person in question from the upper body position being tracked, and obtains a face that is not the target of tracking from the faces acquired in step S802.

[0071] In step S804, matching features are extracted sequentially from faces other than the target being tracked, as acquired in step S803, and compared with the registered matching features of the target being tracked, as in Figure 7, to calculate a matching score. Note that extracting matching features, such as those from human faces, is generally a costly process. Therefore, if it is difficult to extract matching features from all faces before proceeding to the tracking process for the next image, matching should be performed sequentially, starting with faces that are geographically close to the person being tracked, as long as time allows.

[0072] In step S805, the transfer detection unit 406 refers to the matching scores calculated in step S804. If any of the matching scores exceed the threshold, it is determined that the tracking of the person currently being tracked in the image has transferred to a person who is not the target of tracking, and the process proceeds to step S806. If none of the matching scores exceed the threshold, it is determined that no transfer has occurred, and the next image is acquired (step S605).

[0073] In step S806, the tracking area switching unit 407 switches the tracking method to head tracking (second tracking process) starting from the face position that was above the threshold in step S805. Specifically, the tracking result holding unit 408 also holds that the current tracking area is the head.

[0074] In step S807, the tracking unit 405 extracts feature quantities for head tracking and updates the current tracking feature quantities held by the tracking result retention unit 408.

[0075] In step S808, the tracking position and size of the target in the current image held by the tracking result holding unit 408 are updated.

[0076] From this point onward, the same processing as the upper body tracking process shown in Figure 6 will be performed. The decision to switch back from head tracking to upper body tracking will be made in the process shown in Figure 9, which will be described later. Until then, the processing that targeted the upper body in the process in Figure 6 will be performed on the head.

[0077] In the flow chart of Figure 8, matching was performed on faces other than the person being tracked during the possession detection process. However, the matching process is not limited to this. For example, the face of the person being tracked may also be acquired, a matching score may be calculated for that person's face as well, and possession may be determined to have occurred if it is further confirmed that the matching score of the person being tracked's face is below a threshold. Alternatively, possession may be determined to have occurred if the matching scores of both the face of the person being tracked and the faces of other people exceed a threshold, and the matching score of the faces of other people is higher. The above detection methods allow for more accurate detection of possession by the person being tracked.

[0078] The situation in which a transfer is detected using the flow shown in Figure 8 can be described as a state where the upper body is prone to transfers and the face is able to identify the individual simultaneously. In this situation, continuing to track the upper body would cause another transfer to occur. Therefore, by switching to tracking the face and head, where at least individual identification is possible, it becomes possible to suppress transfers and sustain tracking for as long as possible.

[0079] Figure 9 is a flowchart illustrating the procedure for switching back from head tracking to upper body tracking. Here, we assume that the system has switched to head tracking and is currently tracking the head, as shown in Figure 8. We also assume that the tracking process shown in Figure 6 has been completed, that is, image acquisition, detection of the head and upper body from the image, and identification of the head position of the target to be tracked in the acquired image have already been performed.

[0080] In step S901, the recovery determination unit 501 acquires information on the position and size of heads that are present around the head position being tracked.

[0081] In step S902, the recovery determination unit 501 associates each head acquired in step S901 with the detected upper body.

[0082] Step S903 determines whether the correspondence between the head and upper body performed in Step S902 was successful. For example, if there are parts that cannot be matched after the correspondence in Step S902, it is determined that the correspondence in Step S902 is not possible, and the flow is terminated while maintaining head tracking. If it is determined that the correspondence in Step S902 was successful, the process proceeds to Step S904.

[0083] In step S904, the degree of overlap between the upper body associated with each head in step S902 and the other upper body parts is calculated. In this embodiment, IoU (Intersection over Union), a common index indicating the degree of overlap of rectangular regions, is used. At this time, the highest IoU value is obtained between the rectangular region of the upper body associated with the head being tracked and the rectangular regions of the other upper body parts.

[0084] In step S905, it is determined whether the IoU value obtained in step S904 is below a predetermined threshold. If the IoU is below the threshold, it is determined that the upper bodies are separated enough to be distinguishable for tracking purposes, and the process proceeds to step S906. If the overlap exceeds the threshold, it is considered that the overlap between the upper bodies has not been resolved, and this flow is terminated to maintain tracking by the head.

[0085] In step S906, the tracking state is switched from head tracking to upper body tracking. Specifically, the tracking result holding unit 408 also holds that the current tracking area is the upper body.

[0086] In step S907, the tracking unit 405 extracts upper body feature quantities that were associated with the head being tracked in step S902, and updates the current tracking feature quantities held by the tracking result holding unit 408.

[0087] In step S908, the tracking position and size of the target in the current image, which are held in the tracking result retention unit 408, are updated.

[0088] As described above, when the upper body and face of the subject being tracked are matched beyond a predetermined level of confidence, it is determined that it is sufficiently possible to switch back from facial recognition tracking to upper body tracking, and the tracking method is changed back to upper body tracking. This suppresses the occurrence of further transfers and maintains accurate tracking.

[0089] Finally, Figure 10 provides a supplementary explanation of the overall process of the present invention.

[0090] In Figure 10(a), 1001 represents the upper body frame of the target to be tracked, and tracking is performed using tracking features obtained from this region (flowchart in Figure 6). Furthermore, matching features of the face to be matched with the target are extracted from 1002, which is the face of the target to be tracked (flowchart in Figure 7).

[0091] Figure 10(b) is an image taken a few frames after Figure 10(a). The tracked object crosses paths with the person in the foreground and is hidden by the person in the foreground, indicating that the tracked object has taken over the person in the foreground. In other words, the upper body frame 1003 of the tracked object is capturing the person in the foreground.

[0092] Figure 10(c) is an image taken several frames later than Figure 10(b), and the face 1004, the target of tracking, is visible. Face matching can be performed with high accuracy as long as the positions of the facial organs are visible, so face 1004 is matched as the target of tracking. At this point, if the tracking frame is immediately moved to the upper body of the target, it will be confused with the upper body of the person in the foreground, 1003. Therefore, for several to several tens of frames from Figure 10(c) to Figure 10(d), tracking is switched to the head as shown in Figure 8.

[0093] Figure 10(d) is an image taken several to several tens of frames after Figure 10(c). As described above, tracking is performed using the head in the frames up to Figure 10(d). However, in the state of Figure 10(d), the head of the target being tracked and the head of the person in the foreground, as well as the upper body 1005 of the target being tracked and the upper body of the person in the foreground, can all be detected. There is little overlap between the upper body of the target being tracked and the upper body of the person in the foreground, so the possibility of confusion is low. After confirming that the situation that is likely to cause confusion has been resolved, the tracking method is changed back to tracking using the upper body region as shown in Figure 9.

[0094] In this way, by switching to localized tracking using individual matching, tracking becomes possible even in images between Figure 10(c) and Figure 10(d). Such cases are likely to occur in sports photography. For example, even when photographing scenes such as players vying for a ball in a ball game, if the subject to be photographed can be continuously tracked, the opportunity to take the picture will not be lost. Alternatively, in a foot race at a school sports day, the child you want to photograph may be temporarily hidden by a child in the foreground, causing a tracking switch. This is especially likely to occur at school sports days, as participants often wear the same uniform, making tracking switches more likely due to the visual similarity.

[0095] In such cases, individual matching is highly effective in preventing the tracking target from switching over. However, even if the switching can be temporarily resolved by individual matching, if the switching occurs again immediately, the effect on the user is limited. Therefore, as in the present invention, by using the individual matching results to suppress the switching of the tracking target again, and by confirming that the situation is such that the switching is unlikely to occur, and then returning to tracking the upper body, it is possible to reduce the loss of opportunity during shooting.

[0096] (Second Embodiment) As a second embodiment, we will describe cases in which the present invention is applied to purposes other than tracking people. Cell names common to the first embodiment will be omitted, and the differences from the first embodiment will be described in detail.

[0097] In the first embodiment, when a transfer of a person occurred during tracking of the upper body, and this was detected by matching the individual's face, the tracking was switched to the area around the face, such as the head, to prevent further transfer. Similarly, the object being tracked using the present invention does not have to be a person, as long as it satisfies the relationship of switching from tracking a large area such as the overall shape to identifying and tracking a smaller area.

[0098] For example, in the case of a vehicle, detection of the entire vehicle, which corresponds to detection of the upper body in the case of a person, and detection of the license plate, which corresponds to detection of the face and head in the case of a person, may be performed separately. In this case, if a change of vehicle during tracking is identified by the license plate, the system may switch to acquiring tracking features from the vicinity of that license plate and continuing the tracking.

[0099] In the case of horse racing, for example, the detection of the racehorse, which corresponds to the upper body detection unit in the case of a person, and the detection of the rider's hat and clothing color (detection of racing silks), which corresponds to the face and head detection unit in the case of a person, may be performed separately. In this case, if a change of mount during tracking of a racehorse is identified by the color of the rider's hat and racing silks, the system may switch to tracking using the rider's hat and upper body.

[0100] Local features that allow for the identification of tracking transitions are visible in the image at least at the time the transition is detected, making it possible to maintain tracking even in situations where transitions are likely to occur. This embodiment enhances the sustainability of tracking in various shooting scenes.

[0101] (Third embodiment) As a third embodiment, we will describe a case in which a part other than the face is used for personal identification. Cell names common to the first embodiment will be omitted, and the differences from the first embodiment will be described mainly.

[0102] In the first embodiment, only the face was used as the matching area for identifying the individual being tracked. However, multiple matching areas may be detected and used as switching points for matching and tracking. For example, if the subject has identifying parts other than the face, such as a name tag or bib, the name tag may be detected and matched with the individual information, and the system may switch from tracking the upper body to tracking the name tag. Alternatively, the system may switch from tracking the upper body to tracking the face, and then further switch to tracking the name tag.

[0103] Increasing the number of body parts used for individual matching and tracking will increase the consumption of computing resources, but if it is possible to select a body part to use for tracking from among multiple body parts that have been used for individual matching, tracking can be sustained for a longer period of time.

[0104] (Other embodiments) The object of the present invention can also be achieved in the following case: A recording medium containing program code for software that realizes the functions of the embodiments described above is supplied to a system or device, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored on the recording medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiments described above, and the storage medium containing the program code constitutes the present invention.

[0105] For storing program code, storage media such as flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and DVDs can be used.

[0106] Furthermore, this also includes cases where the operating system (OS) running on the computer performs some or all of the actual processing based on the instructions of the program code read by the computer, and the functions of the aforementioned embodiments are realized through that processing.

[0107] Furthermore, the program code read from the storage medium may be written to the memory of a function expansion board inserted into a computer or a function expansion unit connected to a computer. Subsequently, the CPU or other components of the function expansion board or function expansion unit perform some or all of the actual processing based on the instructions of the program code, and the functions of the embodiments described above are realized through this processing.

[0108] It should be noted that the above embodiments are merely examples of how the present invention can be implemented, and the technical scope of the present invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various forms without departing from its technical concept or its main features.

[0109] Furthermore, this disclosure includes the following components.

[0110] (Composition 1) A first detection step involves detecting a first region containing the target to be tracked from the first image, In a second image taken after the first image, a first tracking step is performed to track the subject based on information obtained from the first region, A second detection step involves detecting a second region in which an individual subject can be identified, A matching step of matching the subject and the individual information of the tracking target based on the information obtained from the second region, The process includes a second tracking step of tracking the target based on information obtained from the second region, An image processing method characterized in that, if the comparison determines that a subject other than the target subject is being tracked in the first tracking step, the process switches from the first tracking step to the second tracking step, starting from the second region.

[0111] (Configuration 2) The first detection step involves detecting a region including the outline of the object to be tracked. The image processing method according to configuration 1, characterized in that the second detection step detects a local area of ​​the subject.

[0112] (Composition 3) The subject being tracked is a person. The first region includes the upper body of the person, The matching process involves matching the person's face with the individual information of the target being tracked. The image processing method according to configuration 1 or 2, characterized in that the second tracking step performs tracking of the target to be tracked based on information obtained from the head of the person.

[0113] (Composition 4) The matching step involves matching the feature quantity of the tracking target with the feature quantity of the second region. An image processing method according to any one of configurations 1 to 3, characterized in that, if the similarity between the feature quantity of the tracking target and the feature quantity of the second region is less than a threshold, it is determined that the first tracking step is tracking a subject other than the tracking target.

[0114] (Composition 5) The matching step involves matching the feature quantities other than the tracking target with the feature quantities of the second region. An image processing method according to any one of configurations 1 to 3, characterized in that, if the similarity between the feature quantities other than the tracking target and the feature quantities of the second region is greater than or equal to a threshold, it is determined that the first tracking step is tracking a subject other than the tracking target.

[0115] (Composition 6) The image processing method according to any one of configurations 1 to 5, characterized in that, in the second tracking step, if the second region and the first region are associated with each other beyond a predetermined level of confidence, the method switches from the second tracking step to the first tracking step, starting from the first region.

[0116] (Composition 7) An image processing method according to any one of configurations 1 to 6, characterized in that when the first region and the second region do not overlap, the system switches from the second tracking step to the first tracking step.

[0117] (Composition 8) A first detection means for detecting a first region including the target to be tracked from a first image, In a second image taken after the first image, a first tracking means tracks the subject based on information obtained from the first region, A second detection means for detecting a second region capable of identifying an individual subject, A matching means for matching the subject and the individual information of the tracking target based on the information obtained from the second region, The system includes a second tracking means that tracks the target based on information obtained from the second region, During tracking by the first tracking means, it is determined by matching the individual whether or not a transfer of tracking has occurred. An image processing apparatus characterized in that, if the verification determines that the first tracking means is tracking a subject other than the target subject, it switches from the first tracking means to the second tracking means, starting from the second region.

[0118] (Composition 9) A first detection means for detecting a first region including the target to be tracked from a first image, In a second image taken after the first image, a first tracking means tracks the subject based on information obtained from the first region, A second detection means for detecting a second region capable of identifying an individual subject, A matching means for matching the subject and the individual information of the tracking target based on the information obtained from the second region, The system includes a second tracking means that tracks the target based on information obtained from the second region, During tracking by the first tracking means, it is determined by matching the individual whether or not a transfer of tracking has occurred. A program to cause a computer to switch from the first tracking means to the second tracking means, starting from the second region, if the verification determines that the first tracking means is tracking a subject other than the target subject. [Explanation of symbols]

[0119] 402 Upper body detection unit 403 Face / Head Detection Unit 404 Verification Unit 405 Tracking part 406 Transfer detection unit 407 Tracking section switching unit

Claims

1. A first detection step involves detecting a first region containing the target to be tracked from the first image, In a second image taken after the first image, a first tracking step is performed to track the subject based on information obtained from the first region, A second detection step involves detecting a second region in which an individual subject can be identified, A matching step of matching the subject and the individual information of the tracking target based on the information obtained from the second region, The process includes a second tracking step of tracking the target based on information obtained from the second region, An image processing method characterized in that, if the comparison determines that a subject other than the target subject is being tracked in the first tracking step, the process switches from the first tracking step to the second tracking step, starting from the second region.

2. The first detection step involves detecting a region including the outline of the object to be tracked. The image processing method according to claim 1, characterized in that the second detection step detects a local area of ​​the subject.

3. The subject being tracked is a person. The first region includes the upper body of the person, The matching process involves matching the person's face with the individual information of the target being tracked. The image processing method according to claim 1, characterized in that the second tracking step performs tracking of the target to be tracked based on information obtained from the head of the person.

4. The matching step involves matching the feature quantity of the tracking target with the feature quantity of the second region. The image processing method according to claim 1, characterized in that, if the similarity between the feature quantity of the tracking target and the feature quantity of the second region is less than a threshold, it is determined in the first tracking step that a subject other than the tracking target is being tracked.

5. The matching step involves matching the feature quantities other than the tracking target with the feature quantities of the second region. The image processing method according to claim 1, characterized in that, if the similarity between the feature quantities other than the tracking target and the feature quantities of the second region is greater than or equal to a threshold, it is determined that the first tracking step is tracking a subject other than the tracking target.

6. The image processing method according to claim 1, characterized in that, in the second tracking step, if the second region and the first region are associated with each other beyond a predetermined level of confidence, the method switches from the second tracking step to the first tracking step, starting from the first region.

7. The image processing method according to claim 6, characterized in that when the first region and the second region do not overlap, the system switches from the second tracking step to the first tracking step.

8. A first detection means for detecting a first region including the target to be tracked from a first image, In a second image taken after the first image, a first tracking means tracks the subject based on information obtained from the first region, A second detection means for detecting a second region capable of identifying an individual subject, A matching means for matching the subject and the individual information of the tracking target based on the information obtained from the second region, The system includes a second tracking means that tracks the target based on information obtained from the second region, During tracking by the first tracking means, it is determined by matching the individual whether or not a transfer of tracking has occurred. An image processing apparatus characterized in that, if the verification determines that the first tracking means is tracking a subject other than the target subject, it switches from the first tracking means to the second tracking means, starting from the second region.

9. A first detection means for detecting a first region including the target to be tracked from a first image, In a second image taken after the first image, a first tracking means tracks the subject based on information obtained from the first region, A second detection means for detecting a second region capable of identifying an individual subject, A matching means for matching the subject and the individual information of the tracking target based on the information obtained from the second region, The system includes a second tracking means that tracks the target based on information obtained from the second region, During tracking by the first tracking means, it is determined by matching the individual whether or not a transfer of tracking has occurred. A program to cause a computer to switch from the first tracking means to the second tracking means, starting from the second region, if the verification determines that the first tracking means is tracking a subject other than the target subject.

Citation Information

Patent Citations

  • Tracking device and method

    JP2020096262A