Image processing device, image processing method, and system

By employing a dual prediction method using frame groups to accurately predict subject positions, the system addresses the issue of takeover in PTZ cameras, ensuring accurate tracking even during subject intersections.

JP2025080595APending Publication Date: 2025-05-26CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023193852
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

The existing tracking techniques in PTZ cameras face issues with 'takeover,' where a subject is erroneously determined as the tracking subject when it intersects with another subject, leading to incorrect identification and disrupted shooting.

Method used

The system employs a dual prediction approach using first and second frame groups to calculate predicted positions of a subject, and based on these predictions, it selects the correct subject to be tracked, reducing misidentification even during intersections.

Benefits of technology

This configuration effectively reduces misjudgment of the tracking subject, even in situations where subjects are likely to intersect, thereby minimizing the occurrence of takeover errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080595000001_ABST
    Figure 2025080595000001_ABST
Patent Text Reader

Abstract

To reduce misidentification of a tracking subject even in situations where the tracking subject intersects with another subject and identity switching is likely to occur.SOLUTION: An image processing device is configured to: calculate a first predicted position of a subject using a first frame group; calculate a second predicted position of the subject using a second frame group; and select a tracking subject to be tracked on the basis of first identification results of the subject based on the detection position of the subject and the first predicted position and second identification results of the subject based on the detection position and the second predicted position.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for selecting a subject to be tracked.

Background Art

[0002] Generally, in a camera capable of adjusting pan, tilt, and zoom, which is generally called a PTZ camera, a technique is known in which a subject to be tracked (hereinafter referred to as a tracking subject) specified by a user is detected from a captured image and tracking shooting is performed. In such a tracking technique, the pan, tilt, and zoom are automatically controlled so as to continuously capture the tracking subject in the target shooting composition. At this time, even when the tracking subject moves, tracking shooting can be performed by continuously determining that the tracking subject after the movement is the same subject as the tracking subject before the movement. Patent Document 1 discloses a method of identifying the same tracking subject over a plurality of consecutive frames and continuing to track it.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The tracking technique has a problem called so-called takeover, in which when the tracking subject intersects with another subject, the other subject is erroneously determined as the tracking subject. When takeover occurs, a subject different from the original tracking subject is photographed, and thus shooting cannot be continued. The present invention provides a technique for reducing the erroneous determination of the tracking subject even in a situation where the tracking subject intersects with another subject and takeover is likely to occur.

Means for Solving the Problems

[0005] One aspect of the present invention includes: a first calculation means for calculating a first predicted position of a subject using a first frame group; a second calculation means for calculating a second predicted position of the subject using a second frame group; a first identification result of the subject based on the detected position of the subject and the first predicted position; and a second identification result of the subject based on the detected position and the second predicted position, and a selection means for selecting a following subject to be tracked based on the first and second identification results.

Advantages of the Invention

[0006] According to the configuration of the present invention, it is possible to reduce the misjudgment of the following subject even in a situation where the following subject is likely to cross with other subjects and transfer occurs.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0008] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0009] [First Embodiment] As shown in FIG. 1, the system according to this embodiment includes a camera 100 and a controller 200 which is a control device of the camera 200, and the camera 100 and the controller 200 are connected to a network 400. Thereby, the system according to this embodiment is configured such that the camera 100 and the controller 200 can perform data communication with each other via the network 400. The network 400 includes networks such as a LAN (Local Area Network) and the Internet.

[0010] <Description of Each Device> Next, an example of the hardware configuration of each of the camera 100 and the controller 200 will be described with reference to the block diagram of FIG. 2. Note that the configuration shown in FIG. 2 is merely an example of the hardware configuration of the camera 100 and the controller 200, and can be appropriately changed / modified.

[0011] First, an example of the hardware configuration of camera 100 will be described. Camera 100 has a mechanism capable of performing pan-tilt operations for changing the shooting direction by rotating the device itself. Further, camera 100 detects a subject from the captured image and changes the shooting direction based on the result of the detection.

[0012] CPU 101 executes various processes using the computer programs and data stored in RAM 102. Thereby, CPU 101 controls the overall operation of camera 100 and executes or controls various processes described as the processes performed by camera 100.

[0013] RAM 102 is a high-speed storage device such as a DRAM, and has an area for storing computer programs and data loaded from ROM 103 and storage device 119, and an area for storing the captured image output from image processing unit 106. Further, RAM 102 has an area for storing various information received from controller 200 via network I / F 105, and a work area used when CPU 101 and inference unit 110 execute various processes. Thus, RAM 102 can appropriately provide various areas.

[0014] ROM 103 stores setting data of camera 100, computer programs and data related to the startup of camera 100, computer programs and data related to the basic operation of camera 100, and the like. Further, ROM 103 also stores computer programs and data for causing CPU 101 and inference unit 110 to execute or control various processes described as the processes performed by camera 100.

[0015] Network I / F 105 is an interface for connecting to the above network 400, and is responsible for communication with external devices such as controller 200 via a communication medium such as Ethernet (registered trademark). Note that serial communication I / F may be separately prepared and used for communication.

[0016] The image processing unit 106 converts the video signal output from the image sensor 107 into a captured image which is data in a predetermined format, compresses the captured image generated by the conversion as necessary, and then outputs it to the RAM 102. Note that the image processing unit 106 may perform various processes such as image quality adjustment including color correction, exposure correction, and sharpness correction on the video represented by the video signal acquired from the image sensor 107, and crop processing for cutting out only a predetermined area. Further, these processes may be carried out according to instructions received from the controller 200 via the network I / F 105.

[0017] The image sensor 107 receives the light reflected from the subject, converts the brightness and color of the received light into electric charges, and outputs a video signal based on the result of the conversion. For the image sensor 107, for example, a photodiode, a CCD (Charge Coupled Device) sensor, a CMOS (Complementary Metal Oxide Semiconductor) sensor, etc. can be used.

[0018] The drive I / F 108 is an interface for transmitting and receiving instruction signals such as control signals to and from the drive unit 109. The drive unit 109 is a rotation mechanism for changing the shooting direction of the camera 100, and has a mechanical drive system and a motor as a drive source. The drive unit 109 performs a pan-tilt operation for changing the shooting direction horizontally and vertically and a zoom operation for optically changing the shooting angle of view according to an instruction received from the CPU 101 via the drive I / F 108.

[0019] The inference unit 110 performs inference processing for estimating the presence and position of a subject from the captured image. The inference unit 110 is, for example, an arithmetic device specialized for image processing and inference processing such as a GPU (Graphics Processing Unit). Although it is generally effective to use a GPU for the inference processing, the same function may be realized by a reconfigurable logic circuit such as an FPGA (Field Programmable Gate Array). Further, the CPU 101 may be responsible for the processing of the inference unit 110.

[0020] The storage device 119 is a non-volatile storage device such as a flash memory, HDD, SSD, SD card, etc., and stores computer programs such as an OS and data. Further, the storage device 119 is also used as a storage area for various short-term data. Also, part or all of the computer programs and data described as being stored in the ROM 103 may be stored in the storage device 119.

[0021] The CPU 101, RAM 102, ROM 103, network I / F 105, image processing unit 106, drive I / F 108, inference unit 110, and storage device 119 are all connected to the system bus 111.

[0022] Next, the controller 200 will be described. The controller 200 can receive the captured image and detection result transmitted from the camera 100 via the local network 400, and can also transmit the selection result of the tracking subject based on the user operation to the camera 100. Using this system, the user can select a tracking subject using the controller 200 and perform tracking shooting of the selected tracking subject with the camera 100.

[0023] The CPU 201 executes various processes using the computer programs and data stored in the RAM 202. As a result, the CPU 201 controls the overall operation of the controller 200 and executes or controls various processes described as processes performed by the controller 200.

[0024] The RAM 202 is a high-speed storage device such as a DRAM. The RAM 202 has an area for storing computer programs and data loaded from the ROM 203 and the storage device 219, and an area for storing various data received from the camera 100 via the network I / F 204. Further, the RAM 202 has a work area used when the CPU 201 and the inference unit 210 execute various processes. In this way, the RAM 202 can appropriately provide various areas.

[0025] The ROM 203 stores the setting data of the controller 200, computer programs and data related to the startup of the controller 200, computer programs and data related to the basic operations of the controller 200, and so on.

[0026] The inference unit 210 performs inference processing for estimating the presence or absence, position, etc. of a subject from a captured image. The inference unit 210 is, for example, an arithmetic device specialized for image processing and inference processing such as a GPU (Graphics Processing Unit). Generally, a GPU is effective for use in inference processing, but a reconfigurable logic circuit such as an FPGA (Field Programmable Gate Array) may be used to achieve an equivalent function. Also, the CPU 201 may perform the processing of the inference unit 210.

[0027] The network I / F 204 is an interface for connecting to the network 400 and is responsible for communication with external devices such as the camera 100 via a communication medium such as Ethernet. For example, communication with the camera 100 includes transmitting a control command to the camera 100 and receiving a captured image from the camera 100.

[0028] The display unit 205 is a display unit having a screen such as a liquid crystal screen or a touch panel screen, and can display a captured image received from the camera 100, a detection result, a setting screen of the controller 200, and so on. In the present embodiment, a case where the display unit 205 is a display unit having a touch panel screen will be described.

[0029] Note that the controller 200 does not necessarily have the display unit 205. For example, the display unit 205 may be omitted from the controller 200, and a display device may be connected to the controller 200 to display a captured image, a detection result, a setting screen of the controller 200, and so on on the display device.

[0030] The user input I / F 206 is an interface for receiving operations from the user to the controller 200, and includes, for example, buttons, dials, joysticks, touch panels, and the like.

[0031] The storage device 219 is a non-volatile storage device such as a flash memory, HDD, SSD, SD card, etc. The storage device 219 stores an OS, computer programs and data for causing the CPU 201 and the inference unit 210 to execute or control various processes described as processes performed by the controller 200. Further, the storage device 219 is also used as a storage area for various short-term data.

[0032] The CPU 201, RAM 202, ROM 203, inference unit 210, storage device 219, network I / F 204, display unit 205, and user input I / F 206 are all connected to the system bus 207. Note that the controller 200 may be a PC (personal computer) having a mouse, keyboard, etc. as the user input I / F 206.

[0033] Next, a software configuration example in each of the camera 100 and the controller 200 will be described using the block diagram of FIG. 3. Note that in FIG. 3, illustration of general-purpose software such as an operating system is omitted.

[0034] In the ROM 103 of the camera 100, as software, a photographing unit 301, an inference unit 302, a drive control unit 303, a communication unit 304, and an arithmetic unit 309 are stored, and the CPU 101 appropriately expands these software from the ROM 103 to the RAM 102 and uses them.

[0035] The photographing unit 301 has a software function for acquiring a photographed image including a subject by causing the CPU 101 to control the image processing unit 106. The inference unit 302 has a software function for detecting a subject from the photographed image by causing the CPU 101 to control the inference unit 110.

[0036] The drive control unit 303 has a software function for rotating the camera 100 so that the direction thereof faces the subject by causing the CPU 101 to control the drive unit 109. The communication unit 304 has a software function for causing the CPU 101 to perform data communication with the controller 200.

[0037] The arithmetic unit 309 has a software function for causing the CPU 101 to perform various arithmetic operations such as motion prediction processing, arithmetic operations associated with control command calculation, and logical operations for branch processing.

[0038] In the storage device 219 of the controller 200, the user interface unit 305, the inference unit 306, and the communication unit 308 are stored as software. The CPU 201 appropriately expands these software from the storage device 219 to the RAM 202 and uses them.

[0039] The user interface unit 305 has a software function for displaying necessary information to the user and receiving user operations by causing the CPU 201 to control the display unit 205 and the user input I / F 206.

[0040] The inference unit 306 has a software function for detecting a subject from a captured image received from the camera 100 by causing the CPU 201 to control the inference unit 210. The communication unit 308 has a software function for causing the CPU 201 to perform data communication with the camera 100.

[0041] The arithmetic unit 310 has a software function for causing the CPU 201 to perform various arithmetic operations such as motion prediction processing, arithmetic operations associated with control command calculation, and logical operations for branch processing.

[0042] Note that the software configuration shown in FIG. 3 is an example. For example, one functional unit may be divided into a plurality of functional units according to functions, or a plurality of functional units may be integrated into one functional unit. Also, one or more of the functional units shown in FIG. 3 may be implemented by hardware.

[0043] <Explanation of the Operations of Each Device> Next, the operations of the camera 100 and the controller 200 in the system according to this embodiment will be described. First, the operation of the camera 100 will be described according to the flowchart in Fig. 4(a).

[0044] In step S101, the CPU 101 reads the imaging unit 301 from the ROM 103, expands it in the RAM 102, and executes the expanded imaging unit 301. As a result, the CPU 101 acquires a captured image from the image processing unit 106 and stores the captured image in the RAM 102.

[0045] In step S102, the CPU 101 reads the inference unit 302 from the ROM 103, expands it in the RAM 102, and executes the expanded inference unit 302. As a result, the CPU 101 inputs the captured image stored in the RAM 102 in step S101 to the inference unit 110, controls the inference unit 110 to detect all subjects in the captured image, and stores the result of the detection (detection result) in the RAM 102.

[0046] At this time, the inference unit 110 reads the learned model created using a machine learning method such as deep learning from the ROM 103 and expands it in the RAM 102. Then, the inference unit 110 inputs the captured image to the learned model and performs the arithmetic processing of the learned model to detect the subject in the captured image, and outputs the attribute information such as the position information, size information, and orientation information of the subject as the result of the detection (detection result). Note that the CPU 101 may reduce the captured image, and the inference unit 110 may input the reduced captured image to the learned model. Thereby, the processing amount of the inference unit 110 can be reduced, and the inference processing can be made faster.

[0047] Here, the detection result of the subject in the inference unit 110 will be described. When the inference unit 110 inputs the captured image into the learned model and performs the arithmetic processing of the learned model, it outputs rectangle information (for example, the coordinates of the upper left vertex and the lower right vertex of the rectangle) that defines a rectangle including the entire subject as the position information of the subject in the captured image. Note that the rectangle information is not limited to the entire subject, and may also be information indicating a part of the subject, for example, the position of the head or face of a human subject. In that case, it is necessary to change the learned model used to a learned model with the desired input and output. Also, the position information of the subject is not limited to the coordinates of the upper left vertex and the lower right vertex of the rectangle including the entire subject, and any information that can represent the position of the subject in the captured image, such as the center coordinates, width, and height of the rectangle, may be used. Further, the inference unit 110 outputs one of the four directions of front, right, back, and left as the orientation information of the subject. Note that the orientation of the subject is not limited to such discontinuous directions, and may also be a continuous angle such as 0 degrees or 90 degrees.

[0048] Note that the method for detecting the subject from the captured image by the inference unit 110 is not limited to a specific method. For example, the inference unit 110 may use a template matching method in which a template image of the subject is registered in advance, and a region with a high degree of similarity to the template image in the captured image is detected as the subject region.

[0049] In step S103, the CPU 101 reads the arithmetic unit 309 from the ROM 103 and expands it in the RAM 102, and executes the expanded arithmetic unit 309. As a result, based on the motion prediction using the detection results (past detection results) stored in the RAM 102, the CPU 101 assigns the same identification information as the identification information of the past subject detected from the captured image of the past frame to the same subject as the past subject detected from the captured image of the current frame (current frame) among the subjects detected from the captured image of the current frame. The CPU 101 writes the identification information assigned to the subject detected from the captured image of the current frame into the RAM 102. Details of the processing in step S103 will be described later.

[0050] In step S104, the CPU 101 reads the communication unit 304 from the ROM 103, expands it in the RAM 102, and executes the expanded communication unit 304. Thereby, the CPU 101 reads the captured image, the detection result, and the identification information from the RAM 102, and transmits the read captured image, detection result, and identification information to the controller 200 via the network I / F 105. In the present embodiment, when the detection result and the captured image are not synchronized due to the execution time of the inference process, the CPU 101 transmits the past detection result as the current detection result.

[0051] In step S105, the CPU 101 executes the arithmetic unit 309 expanded in the RAM 102. Thereby, the CPU 101 determines whether it has received the identification information of the tracking subject from the controller 200 via the network I / F 105. As a result of this determination, if the identification information has been received, the CPU 101 transfers the process to step S106, and if the identification information has not been received, the CPU 101 transfers the process to step S107.

[0052] In step S106, the CPU 101 executes the arithmetic unit 309 expanded in the RAM 102. Thereby, the CPU 101 stores the identification information received from the controller 200 via the network I / F 105 in the RAM 102.

[0053] In step S107, the CPU 101 executes the arithmetic unit 309 expanded in the RAM 102. Thereby, the CPU 101 selects the tracking subject based on the detection result and the identification information of the subject stored in the RAM 102, and writes the identification information of the tracking subject to the RAM 102. Details of the process in step S107 will be described later.

[0054] In step S108, the CPU 101 executes the arithmetic unit 309 expanded in the RAM 102. As a result, the CPU 101 reads out from the RAM 102 "the position information of the tracking subject in the captured image of the current frame (the subject corresponding to the identification information written to the RAM 102 in step S106 or step S107)" and "the position information of the tracking subject in the target shooting composition", and calculates the difference between the respective position information on the captured image of the current frame. Subsequently, the CPU 101 converts this difference into an angular difference as seen from the camera 100. For example, the CPU 101 approximately calculates the angle per pixel of the captured image using the information on the shooting resolution and shooting angle of view of the camera 100, and calculates the result of multiplying the angle by the above difference as the angular difference. Then, the CPU 101 calculates the angular velocity in the pan / tilt direction according to the calculated angle difference. For example, if the angle of the tracking subject in the current frame is a_1, the angle of the tracking subject in the target shooting composition is a_2, and an arbitrary speed coefficient is g, the CPU 101 obtains the angular velocity ω according to the following (Equation 1).

[0055] ω=(a_2 - a_1)×g … (Equation 1) Note that the speed coefficient may be determined experimentally or may be arbitrarily set by the user via the controller 200. Also, the method for calculating the angular velocity in the pan / tilt direction is not limited to the above-described calculation method. For example, if the difference in the horizontal direction is large, the angular velocity of the pan may be increased, and if the difference in the vertical direction is large, the angular velocity of the tilt may be increased. Subsequently, the CPU 101 writes the calculated angular velocity in the pan / tilt direction to the RAM 102.

[0056] In step S109, the CPU 101 reads the drive control unit 303 from the ROM 103, expands it in the RAM 102, and executes the expanded drive control unit 303. As a result, the CPU 101 derives drive parameters for panning and tilting at a desired speed in a desired direction from the angular velocity in the pan-tilt direction read from the RAM 102. Here, the drive parameters refer to parameters for controlling motors (not shown) in the pan direction and the tilt direction included in the drive unit 109. Subsequently, the CPU 101 controls the drive unit 109 via the drive I / F 108 based on the derived drive parameters. When the drive unit 109 rotates based on the drive parameters, the camera 100 changes the shooting direction, that is, a panning and tilting operation is performed.

[0057] In step S110, the arithmetic unit 309 expanded in the RAM 102 is executed. As a result, the CPU 101 writes the position information of the subject in the captured image of the current frame to the RAM 102 for use as the position information of the subject in the past frames hereafter.

[0058] Note that in this embodiment, in the motion prediction process, at least the RAM 102 holds the detection results up to five frames before in order to refer to the detection results up to five frames before.

[0059] In step S111, the CPU 101 executes the arithmetic unit 309 expanded in the RAM 102. As a result, the CPU 101 determines whether or not an end condition for ending the tracking is satisfied. Various conditions can be applied to the end condition, and it is not limited to specific conditions. Examples of the end condition include "receiving an end instruction for tracking from the controller 200", "the current date and time reaching the specified date and time", "a specified time having elapsed since the start of tracking", and the like.

[0060] When such a judgment result indicates that the end condition is satisfied, the CPU 101 ends the process according to the flowchart in FIG. 4(a). On the other hand, when the end condition is not satisfied, the CPU 101 advances the process to step S101.

[0061] Next, the operation of the controller 200 will be described according to the flowchart of FIG. 4(b). In step S201, the CPU 201 reads the arithmetic unit 310 from the ROM 203, expands it in the RAM 202, and executes the arithmetic unit 310. As a result, the CPU 201 determines whether it has received a captured image, a detection result, and identification information from the camera 100 via the network I / F 204.

[0062] If, as a result of this determination, the CPU 201 has received a captured image, a detection result, and identification information from the camera 100, the CPU 201 stores the received captured image, detection result, and identification information in the RAM 202 and advances the process to step S202. On the other hand, if, as a result of this determination, the CPU 201 has not received a captured image, a detection result, and identification information from the camera 100, the CPU 201 advances the process to step S201.

[0063] In step S202, the CPU 201 reads the user interface unit 305 from the ROM 203, expands it in the RAM 202, and executes the expanded user interface unit 305. As a result, the CPU 201 reads the captured image and the detection result from the RAM 202 and causes the display unit 205 to display the read captured image and detection result.

[0064] An example of the display of the captured image and the detection result on the display unit 205 is shown in FIG. 5(a). As shown in FIG. 5(a), a captured image including the subject 700 is displayed on the display screen of the display unit 205, and a rectangular frame 701 defined by the rectangular information of the subject 700 is superimposed and displayed on the captured image as the detection result of the subject 700. The user can confirm the captured image and the detection result by the camera 100 by looking at the screen of the display unit 205 like this.

[0065] In step S203, the CPU 201 executes the user interface unit 305 expanded in the RAM 202. As a result, the CPU 201 receives a touch operation by the user for the "tracking subject selection operation" on the display unit 205. The state where the user selects the central subject as the tracking subject in the display example of FIG. 5(a) is shown in FIG. 5(b). In FIG. 5(b), the user touches the central subject 700 with his / her finger to select the central subject 700 as the tracking subject. Note that the method for selecting the tracking subject is not limited to a specific method. For example, the user may operate the user I / F 206 to select the tracking subject.

[0066] Then, the CPU 201 determines whether a touch operation for the "tracking subject selection operation" has been input. As a result of this determination, if a touch operation for the "tracking subject selection operation" has been input, the CPU 201 transfers the process to step S204. On the other hand, if a touch operation for the "tracking subject selection operation" has not been input, the CPU 201 advances the process to step S201.

[0067] In step S204, the CPU 201 reads the communication unit 308 from the ROM 203, expands it in the RAM 202, and executes the communication unit 308. As a result, the CPU 201 reads the identification information of the subject selected by the user as the tracking subject from the identification information group stored in the RAM 202, and transmits the read identification information to the camera 100 via the network I / F 204.

[0068] Next, the details of the processing in step S103 above will be described according to the flowchart in FIG. 6(a). In step S301, the CPU 101 reads, from the RAM 102, the detection result one frame before (in the past) and the detection result of the current frame as "detection results with a short acquisition period (hereinafter, short-term detection results)". Then, the CPU 101 calculates the displacement vector of the subject from one frame before to the current frame using the detection result one frame before (in the past) and the detection result of the current frame. Then, the CPU 101 predicts the position of the subject one frame later based on the calculated displacement vector.

[0069] The processing in step S301 will be described taking FIG. 7(a) as an example. The horizontal axis of the graph in FIG. 7(a) is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the state of the subject 500 and the subject 501 shown in the captured image moving over time is shown.

[0070] The subject 500a represents the subject 500 in frame (t - 1), the subject 500b represents the subject 500 in frame t, and the subject 502 represents the subject 500 in frame (t + 1).

[0071] The following describes the process for predicting the position of the subject 500 (the position of the subject 502) in the frame (t + 1). To predict the position of the subject 500 in the frame (t + 1), first, the CPU 101 calculates a displacement vector x_t from the center position of the subject 500a to the center position of the subject 500b. Next, the CPU 101 derives a predicted displacement vector x_t+1 that has the same direction and magnitude as the displacement vector x_t with the center position of the subject 500b as the starting point. Then, the CPU 101 obtains (predicts) the position of the end point of the derived predicted displacement vector x_t+1 as the position of the subject 500 (the position of the subject 502) in the frame (t + 1). And the CPU 101 writes the obtained position into the RAM 102 as a predicted position. Note that when predicting the position of the subject 501 in the frame (t + 1), it can also be predicted in the same way.

[0072] In this way, the CPU 101 can predict the position of the subject after one frame from short-term motion prediction. Since only the immediately previous detection result is reflected in the prediction for short-term motion prediction, local changes in the motion of the subject can be observed, and it is possible to predict a subject with a large speed change.

[0073] Note that the prediction is not limited to the position of the subject in the next instance. In addition to predicting the position of the subject, the rectangular area may be calculated from the rectangular information of the subject, and the rectangular area of the rectangle including the subject in the next instance may be predicted from the transition of the rectangular area. Thereby, since the difference in the rectangular areas of the two intersecting subjects shown in the captured image can be discriminated, even when subjects at different positions in the depth direction intersect, matching between the prediction result described later and the current detection result becomes possible.

[0074] In addition to predicting the position of the subject, the attribute information of the subject (for example, the orientation of the face) in the next instance may be predicted from the transition of the attribute information. By referring to the moving direction of the position of the subject and the attribute information, the CPU 101 can match the prediction result described later and the current detection result based on the orientation of the face even when the subject moves in the same direction and intersects.

[0075] Also, in this embodiment, the face orientation is described in four directions. However, if a model that can be acquired in more detail is adopted, it becomes possible to recognize the difference in face orientation in more detail, and robust matching processing can be realized. Furthermore, the CPU 101 may cooperate these prediction processes to realize more robust prediction processing.

[0076] Next, in step S302, the CPU 101 reads out from the RAM 102 the detection results in each frame from 5 frames ago (past) to the current frame as "detection results with a long acquisition period (hereinafter, long-term detection results)". Here, let the current frame be frame t. Then, the CPU 101 uses the read detection results to calculate the displacement vector of the subject from frame (t - i) to frame (t - i + 1) for i = 5 to 1. Then, the CPU 101 calculates a predicted displacement vector whose direction and magnitude are the average of the directions of the calculated displacement vectors and the average of the magnitudes of the calculated displacement vectors, respectively, and predicts the position of the subject one frame later based on the calculated predicted displacement vector.

[0077] The processing in step S302 will be described by taking FIG. 7(b) as an example. The horizontal axis of the graph in FIG. 7(b) is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the state where the subjects 503 and 504 shown in the captured image are moving over time is shown. The subject 503a is the subject 503 in frame (t - 5), the subject 503b is the subject 503 in frame (t - 4), the subject 503c is the subject 503 in frame (t - 3), and the subject 503d is the subject 503 in frame (t - 2). Also, the subject 503e is the subject 503 in frame (t - 1), the subject 503f is the subject 503 in frame t, and the subject 505 is the subject 503 in frame (t + 1). Hereinafter, the processing for predicting the position of the subject 503 in frame (t + 1) (the position of the subject 505) will be described.

[0078] To predict the position of the subject 503 in the frame (t + 1), first, the CPU 101 calculates the displacement vector x_t-4 from the center position of the subject 503a to the center position of the subject 503b. Also, the CPU 101 calculates the displacement vector x_t-3 from the center position of the subject 503b to the center position of the subject 503c. Also, the CPU 101 calculates the displacement vector x_t-2 from the center position of the subject 503c to the center position of the subject 503d. Also, the CPU 101 calculates the displacement vector x_t-1 from the center position of the subject 503d to the center position of the subject 503e. Also, the CPU 101 calculates the displacement vector x_t from the center position of the subject 503e to the center position of the subject 503f.

[0079] Next, the CPU 101 calculates a displacement vector x_avg whose direction and magnitude are the average of the directions of the displacement vectors x_t to x_t-4 and the average of the magnitudes of the displacement vectors x_t to x_t-4, respectively.

[0080] For the directions of the respective displacement vectors x_t-4 to x_t, when the inclination angles in the vector direction with the horizontal being 0 degrees are a_t-4 to a_t, respectively, the direction (inclination angle) a_avg of the displacement vector x_avg is obtained by the following (Equation 2).

[0081] a_avg = ((a_t-4 + a_t-3 + a_t-2 + a_t-1 + a_t)) / 5 … (Equation 2) Also, when the magnitudes of the respective displacement vectors x_t-4 to x_t are v_t-4 to v_t, respectively, the magnitude v_avg of the displacement vector x_avg is obtained by the following (Equation 3).

[0082] v_avg = ((v_t-4 + v_t-3 + v_t-2 + v_t-1 + v_t)) / 5 … (Equation 3) Next, the CPU 101 derives a predicted displacement vector x_t+1 with the same direction and magnitude as the displacement vector x_avg, starting from the center position of the subject 503f. Then, the CPU 101 obtains (predicts) the position of the end point of the derived predicted displacement vector x_t+1 as the position of the subject 503 (the position of the subject 505) in the frame (t + 1). Then, the CPU 101 writes the obtained position into the RAM 102 as a predicted position. Note that when predicting the position of the subject 504 in the frame (t + 1), it can also be predicted in the same way.

[0083] Also, in this embodiment, the predicted displacement vector is calculated from the average of the direction and magnitude of the displacement vectors between frames, but the calculation method of the predicted displacement vector is not limited to this. For example, the direction and magnitude of the predicted displacement vector may be the average of the directions of the first and last displacement vectors (x_t-4, x_t) and the average of the magnitudes of the first and last displacement vectors (x_t-4, x_t) during the period of referring to the detection results.

[0084] Furthermore, for each of the direction and magnitude of the displacement vectors between frames, the optimal value of the least squares method may be obtained, and the optimal value of the direction may be the direction of the predicted displacement vector, and the optimal value of the magnitude may be the magnitude of the predicted displacement vector. That is, any method that can predict the average movement from the detection results over a long period may be used.

[0085] In this way, the CPU 101 can predict the position of the subject one frame later from the long-term movement prediction. By referring to the detection results of the past over a long period, the overall change in the movement of the subject can be seen. Therefore, even if an error occurs in the detection position of the subject due to partial occlusion of the intersecting subjects or the subjects overlapping and being visible, the relative error is small with respect to the overall movement of the subject. Thus, a prediction with a small influence of the error can be made, and it becomes possible to re-identify the subject after intersection.

[0086] That is, compared with a conventional system that performs tracking using one type of motion prediction, the present embodiment realizes a system that performs more robust tracking by selectively using two different types of motion predictions according to the condition of whether or not the subjects cross each other.

[0087] In the present embodiment, in each of steps S301 and S302, motion prediction is performed for a 1-frame period and a 5-frame period, but the number of frames in each period is not limited to this. The number of frames in each period may be experimentally determined according to the movement of the subject, and it is sufficient that there are a plurality of motion predictions with different numbers of frames in each period.

[0088] In step S303, the CPU 101 reads from the RAM 102 the detection result of the subject detected from the captured image of the current frame, the position of the subject predicted in step S301 for the captured image of the current frame (first predicted position), and the position of the subject predicted in step S302 for the captured image of the current frame (second predicted position). Then, the CPU 101 matches (associates) the closest combination among the detection result, the first predicted position, and the second predicted position.

[0089] The process in step S303 will be described taking FIG. 8 as an example. FIG. 8(a) is a diagram schematically explaining the matching method between the detection result of the subject and the short-term motion prediction result (prediction result in step S301) and the process of matching when the subjects cross each other.

[0090] The horizontal axis of the graph is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the states of the subjects 506 and 507 included in the captured image moving over time are shown. The position 506i indicates the position of the subject 506 in frame i, and the position 507i indicates the position of the subject 507 in frame i. Furthermore, the subject 506 and the subject 507 cross each other during movement, and on the way, the subject 506 is blocked by the subject 507.

[0091] Regarding the matching process, first, at step S301, the CPU 101 performs motion prediction of the subject 506, and derives the predicted position 508c of the subject 506 from the detected positions 506a and 506b of the subject 506. Similarly, for the subject 507, the CPU 101 also derives the predicted position 509c of the subject 507 from the detected position in frame a and the detected position in frame b of the subject 507.

[0092] Next, the CPU 101 acquires from the RAM 102 the detected position 506c of the subject 506 detected from the captured image of frame c and the detected position 507c of the subject 507 detected from the captured image of frame c.

[0093] Then, the CPU 101 matches the detected positions 506c and 507c with the predicted positions 508c and 509c according to the similarity of the positions. The similarity of the positions is calculated by a calculation formula such that the value is higher when the detected position and the predicted position are closer, and the value is lower when the detected position and the predicted position are farther. For example, it is calculated by a calculation formula such that it becomes 1 when the detected position and the predicted position match, and becomes 0 when the detected position and the predicted position are separated by the distance of one side of the rectangle of the subject. Specifically, when the detected position is p1, the predicted position is p2, and the distance of one side of the rectangle is L, the "position similarity P" between the detected position p1 and the predicted position p2 can be obtained by calculating the following (Equation 4).

[0094] P = MAX(1 - |((p2 - p1) / L)|, 0) … (Equation 4) In the example of Fig. 8(a), the detected position closest to the predicted position 508c is the detected position 506c, and the detected position closest to the predicted position 509c is the detected position 507c. However, in this case, the CPU 101 matches (associates) the predicted position 508c with the detected position 506c, and matches (associates) the predicted position 509c with the detected position 507c.

[0095] In this embodiment, the similarity of positions is used as a condition for matching the subject of the detection result and the subject of the prediction result. However, the matching condition is not limited to this. For example, instead of or in addition to the similarity of positions, the area of the subject may be calculated from the rectangular information of the detection result, and the similarity of the areas of the respective subjects may be used as a matching condition. Alternatively, the orientation information of the subject may be extracted from the attribute information of the detection result, and the similarity of the orientations may be used as a matching condition.

[0096] For example, the similarity of areas can be calculated by a calculation formula such that it is 1 when the areas match and 0 when the area is twice as far from the subject of the detection result. Specifically, when the area of the subject in the detection result is s1 and the area of the subject in the prediction result is s2, the similarity of areas S is obtained by the following (Equation 5).

[0097] S = MAX(1 - |((s2 - s1) / s1)|, 0) … (Equation 5) Also, the similarity of orientations can be calculated by a calculation formula such that it is 1 when the orientations match and 0 when the orientations do not match. Note that the calculation processes of the respective similarities may be coordinated. For example, the weighted average value of the similarities of positions, areas, and orientations may be used. Thereby, robust matching processing based on a plurality of pieces of information can be realized.

[0098] Subsequently, the process of matching while the subjects are intersecting will be described. First, the CPU 101 performs motion prediction of the subject 506 in step S301, and derives the predicted position 508e of the subject 506 from the detection positions 506c and 506d of the subject 506. Similarly, the CPU 101 also derives the predicted position 509e of the subject 507 from the detection positions 507c and 507d of the subject 507.

[0099] Next, the CPU 101 acquires from the RAM 102 the detection position 507e of the subject 507 detected from the captured image of frame e. Here, in frame e, since the subject 506 is shielded by the subject 507, the subject 506 is not detected from the captured image of frame e, and thus the detection position of the subject 506 in frame e is not obtained.

[0100] Then, the CPU 101 matches the detection position 507e with the predicted positions 508e and 509e according to the similarity of positions. In the example of Fig. 8(a), since the subjects 506 and 507 are in the vicinity due to intersection, the detection position 507e, the predicted position 508e, and the predicted position 509e are all located in the vicinity, but the predicted position 508e is closer to the detection position 507e than the predicted position 509e. Therefore, the CPU 101 matches (associates) the predicted position 508e with the detection position 507e.

[0101] In this way, during the intersection of subjects, the predicted position of a subject may match the detection position of another subject different from the subject. That is, it is the cause of the tracked subject switching in the tracking shooting.

[0102] In the example of Fig. 8(a), the predicted position 509e does not match any detection position, but the CPU 101 writes the predicted position 509e to the RAM 102 as a temporary detection position of the subject 507 in frame e, and uses this for the subsequent movement prediction of the subject 507. Thereby, even when the subject 507 cannot be detected temporarily due to shielding or the like, or when the subject 507 cannot be detected temporarily due to the performance of the inference unit, it becomes possible to continue the position prediction of the subject. However, since the error of the predicted position increases when the prediction based on the prediction result is repeated, it is effective to set an upper limit on the number of times of performing the prediction using the prediction result. For simplicity in this embodiment, the use of the prediction result as a temporary detection result is described with an upper limit of 1 time.

[0103] Next, at step S301, the CPU 101 performs motion prediction of the subject 506, and derives a predicted position 508f of the subject 506 from the detected position 506d of the subject 506 and the detected position 507e which is the detected position of the subject 507 in the frame e. Similarly, for the subject 507, the CPU 101 also derives a predicted position 509f of the subject 507 from the detected position 507d and the predicted position 509e of the subject 507.

[0104] Next, the CPU 101 acquires, from the RAM 102, the detected position 507f of the subject 507 detected from the captured image of the frame f. Here, also in the frame f, since the subject 506 is blocked by the subject 507, the subject 506 is not detected from the captured image of the frame f, and thus, the detected position of the subject 506 in the frame f is not obtained.

[0105] Then, the CPU 101 matches the detected position 507f with the predicted positions 508f and 509f based on the similarity of the positions. In the example of FIG. 8(a), since the subjects 506 and 507 are present in the vicinity due to intersection, the detected position 507f, the predicted position 508f, and the predicted position 509f are all located in the vicinity, but the predicted position 508f is closer to the detected position 507f than the predicted position 509f. Therefore, the CPU 101 matches (associates) the predicted position 508f with the detected position 507f.

[0106] In this way, in short-term motion prediction, since the previous motion is reflected in the prediction, it becomes difficult to return to the original subject once the matching is incorrect. Also, the predicted position 509f does not match any of the detected positions, and since the number of times of "performing prediction using the prediction result" has reached the upper limit, it is not written out as a detected position either, and this prediction ends.

[0107] Next, at step S301, the CPU 101 predicts the movement of the subject 507, and derives the predicted position 508g of the subject 507 from the detection position 507e of the subject 507 and the detection position 507f of the subject 507. At this point, since the detection positions in frame e and the detection positions in frame f, which are necessary for predicting the movement of the subject 506, have not been obtained, the predicted position of the subject 506 cannot be determined.

[0108] Next, the CPU 101 acquires from the RAM 102 the detection position 506g of the subject 506 and the detection position 507g of the subject 507 detected from the captured image of frame g. At this time, since the intersection of the subject 506 and the subject 507 has ended, the detection positions of the respective subjects can be obtained.

[0109] Then, the CPU 101 matches the detection position 506g and the detection position 507g with the predicted position 508g according to the similarity of the positions. In the example of Fig. 8(a), the predicted position 508g is closer to the detection position 507g than to the detection position 506g. Therefore, the CPU 101 matches (associates) the predicted position 508g with the detection position 507g.

[0110] Fig. 8(b) is a diagram schematically explaining a method of matching the detection result of the subject and the long-term movement prediction result (the prediction result in step S302), and the process of matching when the subjects intersect.

[0111] The horizontal axis of the graph is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the states of the subjects 506 and 507 included in the captured image moving over time are shown. The position 506i indicates the position of the subject 506 in frame i, and the position 507i indicates the position of the subject 507 in frame i. Further, the subject 506 and the subject 507 intersect during movement, and on the way, the subject 506 is blocked by the subject 507.

[0112] Regarding the matching process, the basic operation is the same as the method of matching the detected subject in FIG. 8(a) with the predicted subject obtained from short-term motion prediction. However, long-term motion prediction as described in step S302 shall be used for motion prediction.

[0113] Subsequently, the process of matching during the intersection of subjects will be described. First, the CPU 101 performs motion prediction of the subject 506 in step S302, and derives the predicted position 510e of the subject 506 using the detection positions in each of frames a, b, c, and d of the subject 506. Note that in the long-term motion prediction at this time, the number of detection positions is insufficient, but if the number of detection positions is insufficient, motion prediction may be performed using only the existing detection positions.

[0114] Similarly, the CPU 101 performs motion prediction of the subject 507 in step S302, and derives the predicted position 511e of the subject 507 using the detection positions in each of frames a, b, c, and d of the subject 507.

[0115] Then, the CPU 101 acquires the detection position 507e of the subject 507 detected from the captured image of frame e from the RAM 102. Here, in frame e, since the subject 506 is blocked by the subject 507, the subject 506 is not detected from the captured image of frame e, and thus the detection position of the subject 506 in frame e is not obtained.

[0116] Then, the CPU 101 matches the detection position 507e with the predicted position 510e and the predicted position 511e based on the similarity of the positions. In the example of FIG. 8(b), since the subjects 506 and 507 are in the vicinity due to intersection, the detection position 507e, the predicted position 510e, and the predicted position 511e are all in the vicinity, but the predicted position 510e is closer to the detection position 507e than the predicted position 511e. Therefore, the CPU 101 matches (associates) the predicted position 510e with the detection position 507e. That is, similar to the case of performing short-term prediction, the predicted position of the subject 506 transfers to the subject 507.

[0117] Next, the CPU 101 predicts the movement of the subject 506 in step S302, and derives the predicted position 510f of the subject 506 using the detection positions in each of the frames a, b, c, and d of the subject 506 and the detection position 507e.

[0118] Similarly, for the subject 507 as well, the CPU 101 derives the predicted position 511f of the subject 507 using the detection positions in each of the frames a, b, c, and d of the subject 507 and the predicted position 511e.

[0119] Then, the CPU 101 acquires the detection position 507f of the subject 507 detected from the captured image of frame f from the RAM 102. Here, in frame f, since the subject 506 is shielded by the subject 507, the subject 506 is not detected from the captured image of frame f, and thus the detection position of the subject 506 in frame f is not obtained.

[0120] Then, the CPU 101 matches the detection position 507f with the predicted positions 510f and 511f based on the similarity of the positions. In the example of Fig. 8(b), unlike the case of making a prediction in a short period, the predicted position 511f is closer to the detection position 507f than the predicted position 510f. Therefore, the CPU 101 matches (associates) the predicted position 511f with the detection position 507f. In this way, in long-term movement prediction, since past movements are reflected in the prediction, even if a matching error occurs once, it is possible to return to the original subject.

[0121] Then, the CPU 101 writes the matching result (association result) between the detection position of the subject and the result of short-term movement prediction (predicted position), and the matching result (association result) between the detection result of the subject and the result of long-term movement prediction (predicted position) to the RAM 102.

[0122] In step S304, if there is a matching (association) result between the detection position of the subject detected from the captured image of the current frame and the result of short-term motion prediction (predicted position), the CPU 101 assigns the same identification information as the identification information of the subject of the "result of short-term motion prediction (predicted position)" associated by the matching to the detection position to the subject. Also, if there is a matching (association) result between the detection position of the subject detected from the captured image of the current frame and the result of long-term motion prediction (predicted position), the CPU 101 assigns the same identification information as the identification information of the subject of the "result of long-term motion prediction (predicted position)" associated by the matching to the detection position to the subject. The identification information may be information unique to each subject. For example, an ID is applicable.

[0123] Note that if both the matching result between the detection position of the subject detected from the captured image of the current frame and the result of short-term motion prediction and the matching result between the detection position of the subject detected from the captured image of the current frame and the result of long-term motion prediction exist, two pieces of identification information will be assigned to the subject detected from the captured image of the current frame. Also, if there is a new subject, uniquely determined identification information is newly assigned. Then, the CPU 101 writes out the identification information of all the subjects for each frame to the RAM 102.

[0124] Next, the details of the processing in step S107 above will be described according to the flowchart in FIG. 6(b). In step S305, the CPU 101 reads out from the RAM 102 the detection results of each subject detected from the captured image of the current frame and the identification information written to the RAM 102 in step S304 above.

[0125] In step S306, for each subject detected from the captured image of the current frame, the CPU 101 calculates the center position of the rectangle defined by the rectangle information included in the detection result of the subject as the position of the subject. Then, the CPU 101 uses the positions of the respective subjects to obtain the distance between the subjects. Then, the CPU 101 determines whether there is a distance less than the threshold among the obtained distances. The threshold is based on the shorter of the lengths of one side in the horizontal direction calculated from the rectangle information included in the detection result among the subjects for which the distance is determined, and is set to twice that value. Note that the threshold is not limited to this, and it may be based on the longer of the lengths of one side in the horizontal direction of the rectangle, or may be a value determined experimentally in advance. That is, a value determined to indicate that the subjects are close to each other may be used as the threshold.

[0126] As a result of such determination, if there is a distance less than the threshold among the obtained distances, the CPU 101 proceeds to step S307, and if there is no distance less than the threshold among the obtained distances, the CPU 101 proceeds to step S308.

[0127] In this embodiment, the distance between the center coordinates of the subjects is used for distance determination. However, the distance to be compared is not limited to this. Instead, the prediction result may be read from the RAM 102, and the distance between the predicted positions of the subjects may be used. Thereby, even when detection results that cannot be compared due to occlusion by interleaving are obtained, the distance between the interleaving subjects can be determined using the prediction result as an alternative.

[0128] In step S307, the CPU 101 selects the identification information obtained by matching the identification information read in step S305 between the detection result of the subject and the long-term motion prediction result (the prediction result in step S302). Then, the CPU 101 overwrites the "identification information obtained by matching the detection result of the subject and the short-term motion prediction result (the prediction result in step S301)" with the "identification information obtained by matching the detection result of the subject and the long-term motion prediction result".

[0129] As a result, the identification information of the subject whose distance between subjects is less than the threshold value, that is, the subject with a high possibility of intersection, is corrected with the identification information based on the prediction result for a long period effective for intersection. Then, the CPU 101 writes the selected identification information to the RAM 102.

[0130] In step S308, the CPU 101 selects the identification information obtained by matching the detection result of the subject and the short-term motion prediction result (the prediction result in step S301) among the identification information read in step S305. Then, the CPU 101 overwrites the "identification information obtained by matching the detection result of the subject and the long-term motion prediction result (the prediction result in step S302)" with the "identification information obtained by matching the detection result of the subject and the short-term motion prediction result". Then, the CPU 101 writes the selected identification information to the RAM 102.

[0131] Note that in steps S307 and S308, the CPU 101 does not select from the identification information based on the prediction results for each period for a newly emerged subject, but selects the newly assigned identification information.

[0132] In step S309, the CPU 101 selects, as the current tracking subject, the subject to which the same identification information as the identification information of the previous tracking subject is assigned among the identification information selected in step S307 or step S308. Then, the CPU 101 writes the identification information of the tracking subject to the RAM 102 and shifts the process to step S108.

[0133] It is a feature of this embodiment to utilize the characteristics of short-term motion prediction when tracking a subject in a non-intersecting scene or a subject with a large speed change, and to utilize the characteristics of long-term motion prediction when tracking a subject in an intersecting scene.

[0134] In this way, according to the tendency of the occurrence of handover due to intersection, it is possible to reduce the handover by performing a plurality of motion predictions from subject information with different acquisition periods and selecting the identification information of the tracking subject based on the prediction results.

[0135] In this embodiment, a first predicted position of a subject is calculated using a first group of frames, a second predicted position of the subject is calculated using a second group of frames, and a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position are used to select a subject to be tracked. An example of the process is described below.

[0136] Note that when the distance between the subject to be tracked and other subjects other than the subject to be tracked is equal to or greater than a threshold value, the camera 100 selects the subject to be tracked based on identification information based on short-term motion prediction. When the distance is less than the threshold value, the camera 100 may select the subject to be tracked based on identification information based on long-term motion prediction.

[0137] In this embodiment, the camera 100 calculates the driving amount for detecting and tracking the subject inside the camera. However, part or all of these processes may be executed by the controller 200. In that case, first, the camera 100 transmits the captured image to the controller 200. Next, the controller 200 detects the subject from the received captured image, derives the driving parameters for tracking the subject to be tracked, and transmits the driving parameters to the camera 100. Then, the camera 100 captures the subject to be tracked according to the received driving parameters. At this time, the process performed by the inference unit 110 of the camera may be performed by the inference unit 210 of the controller 200, and the software operations of the inference unit 302 and the arithmetic unit 309 of the camera 100 may be respectively corresponding to the inference unit 306 and the arithmetic unit 310 of the controller 200.

[0138] That is, it is also possible to have a system including a camera that controls tracking shooting of a subject to be tracked, a control device that calculates a first predicted position of the subject using a first group of frames by the camera, calculates a second predicted position of the subject using a second group of frames by the camera, and selects the subject to be tracked based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position.

[0139] [Second Embodiment] Hereinafter, differences from the first embodiment will be described, and unless otherwise specified below, it is assumed to be the same as the first embodiment. In this embodiment, a subject is detected from a captured image by the camera 100, the above matching is performed using a plurality of motion predictions with different periods, and reliable identification information is selected according to the presence or absence of identification information based on each period.

[0140] Details of the process of step S103 according to this embodiment will be described according to the flowchart of FIG. 9(a). In step S401, the CPU 101 reads from the RAM 102 the detection result one frame before (past) and the detection result of the current frame as "detection results with a short acquisition period (hereinafter, short period)". The "detection result one frame before (past) and the detection result of the current frame" are the detection results of frames arranged at a period of "1". Then, in the same manner as in step S301 above, the CPU 101 uses the detection result one frame before (past) and the detection result of the current frame to calculate the displacement vector of the subject from one frame before to the current frame. Then, in the same manner as in step S301 above, the CPU 101 predicts the position of the subject one frame later based on the calculated displacement vector. Since short-period motion prediction reflects detection results with a short time interval in the prediction, local changes in the motion of the subject can be observed, and prediction of a subject with a large speed change is possible.

[0141] In step S402, the CPU 101 reads from the RAM 102 the detection result three frames before (past) and the detection result of the current frame as "detection results with a long acquisition period (hereinafter, long period)". The "detection result three frames before (past) and the detection result of the current frame" are the detection results of frames arranged at a period of "3". Then, in the same manner as in step S301 above, the CPU 101 uses the detection result three frames before (past) and the detection result of the current frame to calculate the displacement vector of the subject from three frames before to the current frame. Then, in the same manner as in step S301 above, the CPU 101 predicts the position of the subject three frames later based on the calculated displacement vector.

[0142] Regarding the process in step S402, an example will be taken with reference to FIG. 10 for explanation. The horizontal axis of the graph in FIG. 10 is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the states of the subjects 600 and 601 shown in the captured image moving over time are shown.

[0143] Subject 600a represents subject 600 in frame (t - 3), subject 600b represents subject 600 in frame t, and subject 602 represents subject 600 in frame (t + 3).

[0144] Hereinafter, the process for predicting the position of subject 600 in frame (t + 3) will be described. To predict the position of subject 600 in frame (t + 3), first, the CPU 101 calculates a displacement vector x_t from the center position of subject 600a to the center position of subject 600b. Next, the CPU 101 derives a predicted displacement vector x_t+3 with the same direction and magnitude as the displacement vector x_t starting from the center position of subject 600b. Then, the CPU 101 obtains (predicts) the position of the end point of the derived predicted displacement vector x_t+3 as the position of subject 600 (the position of subject 602) in frame (t + 3). And the CPU 101 writes the obtained position into the RAM 102 as a predicted position. Note that when predicting the position of subject 601 in frame (t + 3), it can also be predicted in the same way.

[0145] In this way, the CPU 101 can predict the position of the subject three frames later next time from the detection results of a long period. Referring to the detection results in the past with a long period, the changes in the overall movement of the subject can be seen. Therefore, even if an error occurs in the detection position of the subject due to some subjects being partially blocked or subjects overlapping and being seen, the relative error is small with respect to the overall movement of the subject. Thus, since a prediction with a small influence of the error can be made, it becomes possible to re-identify the subject after interleaving.

[0146] In this embodiment, motion prediction is performed for one frame period and three frame periods in steps S401 and S402, respectively. However, the number of periods is not limited to this. The number of periods may be experimentally determined according to the movement of the subject, and any form with a plurality of motion predictions having different numbers of periods may be used.

[0147] Next, in step S403, the CPU 101 reads from the RAM 102 the detection result of the subject detected from the captured image of the current frame, the position of the subject predicted in step S401 for the captured image of the current frame (the first predicted position), and the position of the subject predicted in step S402 for the captured image of the current frame (the second predicted position). Then, the CPU 101 matches (associates) the closest combination among the detection result, the first predicted position, and the second predicted position in the same manner as in step S303 above.

[0148] The process in step S403 will be described by taking FIG. 11 as an example. FIG. 11 is a diagram schematically explaining the matching method between the detection result of the subject and the long-period motion prediction result (the prediction result in step S402) and the process of matching when the subjects cross each other.

[0149] The horizontal axis of the graph is the shooting time (frame) of the captured image, and the vertical axis is the horizontal position (it may be the vertical position or both) of each subject in the captured image. Also, on the graph, the states of the subjects 603 and 604 included in the captured image moving over time are shown. The position 603i indicates the position of the subject 603 in frame i, and the position 604i indicates the position of the subject 604 in frame i. Further, the subjects 603 and 604 cross each other during movement, and on the way, the subject 603 is blocked by the subject 604.

[0150] Regarding the matching process, first, the CPU 101 predicts the movement of the subject 603 in step S402, and derives the predicted position 605g from the detected positions 603a and 603d of the subject 603. Similarly, the CPU 101 also derives the predicted position 606g from the detected positions 604a and 604d for the subject 604.

[0151] Next, the CPU 101 obtains from the RAM 102 the detected position 603g of the subject 603 detected from the captured image of frame g, and the detected position 604g of the subject 604 detected from the captured image of frame g.

[0152] Then, the CPU 101 matches (associates) each of the detected positions 603g and 604g with the predicted positions 605g and 606g, respectively, based on the similarity of the positions. In the example of FIG. 11, the predicted position 605g is close to the detected position 603g, and the predicted position 606g is close to the detected position 604g. Therefore, the CPU 101 matches (associates) the predicted position 605g with the detected position 603g, and the predicted position 606g with the detected position 604g.

[0153] Subsequently, the matching process during interleaving will be described. First, the CPU 101 predicts the movement of the subject 603 in step S402, and derives the predicted position 605j from the detected positions 603d and 603g of the subject 603. Similarly, the CPU 101 derives the predicted position 606j from the detected positions 604d and 604g of the subject 604.

[0154] Next, the CPU 101 obtains from the RAM 102 the detected position 604j of the subject 604 detected from the captured image of frame j. Here, in frame j, since the subject 603 is blocked by the subject 604, the subject 603 is not detected from the captured image of frame j, and thus the detected position of the subject 603 in frame j is not obtained.

[0155] Next, the CPU 101 matches the detected position 604j with the predicted positions 605j and 606j based on the similarity of the positions. In the example of FIG. 11, since the subject 603 and the subject 604 are in the vicinity due to intersection, the detected position 604j, the predicted position 605j, and the predicted position 606j are all located in the vicinity. However, the predicted position 605j is closer to the detected position 604j than the predicted position 606j. Therefore, the CPU 101 matches (associates) the predicted position 605j with the detected position 604j. Thus, during the intersection, a transfer may occur in the same manner as in the first embodiment.

[0156] At this time, although the predicted position 606j does not match any of the detection results, the CPU 101 writes the predicted position 606j to the RAM 102 as a provisional detected position of the subject 604 and uses this for predicting the movement of the next subject 604.

[0157] Next, the CPU 101 predicts the movement of the subject 603 in step S402 and derives the predicted position 605m from the detected position 603g of the subject 603 and the detected position 604j of the subject 604. Similarly, the CPU 101 derives the predicted position 606m from the detected position 604g of the subject 604 and the predicted position 606j.

[0158] Next, the CPU 101 acquires from the RAM 102 the detected position 603m of the subject 603 detected from the captured image of frame m and the detected position 604m of the subject 604. Next, the CPU 101 matches the detected position 603m and the detected position 604m with the predicted position 605m and the predicted position 606m based on the similarity of the positions. In the example of FIG. 11, the predicted position 605m is close to the detected position 603m, and the predicted position 606m is close to the detected position 604m. Therefore, the CPU 101 matches (associates) the predicted position 605m with the detected position 603m and the predicted position 606m with the detected position 604m. That is, similar to the first embodiment, even if the matching is mistaken once, it is possible to return to the original subject.

[0159] Then, the CPU 101 writes the matching result (association result) between the detected position of the subject and the short-term motion prediction result (predicted position), and the matching result (association result) between the detection result of the subject and the long-term motion prediction result (predicted position) to the RAM 102.

[0160] In step S404, when there is a matching (association) result between the detected position of the subject detected from the captured image of the current frame and the short-term motion prediction result (predicted position), the CPU 101 assigns the same identification information as the identification information of the subject of the "short-term motion prediction result (predicted position)" associated by the matching to the detection position to the subject. Also, when there is a matching (association) result between the detected position of the subject detected from the captured image of the current frame and the long-term motion prediction result (predicted position), the CPU 101 assigns the same identification information as the identification information of the subject of the "long-term motion prediction result (predicted position)" associated by the matching to the detection position to the subject. Also, when there is a newly emerged subject, newly assign uniquely determined identification information. Then, the CPU 101 writes the identification information of all subjects for each frame to the RAM 102.

[0161] Next, the details of the process in step S107 above will be described according to the flowchart of FIG. 9(b). In step S405, the CPU 101 reads from the RAM 102 the detection result of each subject detected from the captured image of the current frame and the identification information written to the RAM 102 in step S404 above.

[0162] In step S406, the CPU 101 determines whether or not it has read the identification information obtained by matching the detection result of the subject and the long-term motion prediction result (prediction result in step S402).

[0163] As a result of this determination, when the CPU 101 reads out "identification information obtained by matching the subject detection result with the long-term motion prediction result (the prediction result in step S402)" in step S405, the CPU 101 advances the process to step S407. On the other hand, when the CPU 101 does not read out "identification information obtained by matching the subject detection result with the long-term motion prediction result (the prediction result in step S402)" in step S405, the CPU 101 advances the process to step S409.

[0164] In step S407, the CPU 101 determines whether "identification information obtained by matching the subject detection result with the long-term motion prediction result" and "identification information obtained by matching the subject detection result with the short-term motion prediction result" are the same or different.

[0165] As a result of this determination, if it is determined that they are the same, the CPU 101 advances the process to step S409, and if it is determined that they are different, the CPU 101 advances the process to step S408.

[0166] In step S408, the CPU 101 selects the identification information obtained by matching the subject detection result with the long-term motion prediction result (the prediction result in step S402) from among the identification information read out in step S405. Then, the CPU 101 overwrites "identification information obtained by matching the subject detection result with the short-term motion prediction result (the prediction result in step S401)" with "identification information obtained by matching the subject detection result with the long-term motion prediction result". After that, the CPU 101 writes out the selected identification information to the RAM 102.

[0167] In step S409, the CPU 101 selects the identification information obtained by matching the detection result of the subject among the identification information read in step S405 with the short-term motion prediction result (the prediction result in step S401). Then, the CPU 101 overwrites the "identification information obtained by matching the detection result of the subject with the long-term motion prediction result (the prediction result in step S402)" with the "identification information obtained by matching the detection result of the subject with the short-term motion prediction result". After that, the CPU 101 writes the selected identification information to the RAM 102.

[0168] Note that in steps S408 and S409, for a newly emerged subject, the CPU 101 does not select from the identification information based on the prediction results of each cycle, but selects the newly assigned identification information.

[0169] In step S410, the CPU 101 selects, as the current tracking subject, the subject to which the same identification information as the identification information of the previous tracking subject is assigned among the identification information selected in step S408 or step S409. After that, the CPU 101 writes the identification information of the tracking subject to the RAM 102 and transfers the process to step S108.

[0170] It is a feature of this embodiment that when tracking a non-overlapping scene or a subject with a large speed change, the characteristics of short-term motion prediction are utilized, and when tracking a subject in an overlapping scene, the characteristics of long-term motion prediction are utilized.

[0171] In this way, according to the tendency of the occurrence of transfer due to overlapping, multiple motion predictions are performed from subject information with different acquisition cycles, and by selecting the identification information of the tracking subject based on the prediction results, the same effect as in the first embodiment can be obtained.

[0172] The numerical values, processing timings, processing orders, processing entities, methods of acquiring / sending to / sending from / storing data (information), etc. used in the above embodiments are given as examples for specific explanations and are not intended to be limited to such examples.

[0173] Also, some or all of the above-described embodiments may be used in appropriate combination. Also, some or all of the above-described embodiments may be selectively used.

[0174] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0175] The invention described in this specification includes the following image processing apparatus, image processing method, system, and computer program. (Item 1) A first calculation means for calculating a first predicted position of a subject using a first frame group; A second calculation means for calculating a second predicted position of a subject using a second frame group; A selection means for selecting a tracking subject to be tracked based on a first identification result of the subject based on the detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position; An image processing apparatus comprising the same. (Item 2) The first calculation means calculates a first predicted position of a subject using a first frame group in a first frame period, The second calculation means calculates a second predicted position of a subject using a second frame group in a second frame period longer than the first frame period The image processing apparatus according to Item 1, characterized in that. (Item 3) The selection means assigns identification information of a subject at a first predicted position closest to the detection position of the subject in the current frame among the first predicted positions calculated for the current frame to the subject in the current frame. The image processing apparatus according to Item 1 or 2, characterized in that. (Item 4) The selection means assigns, to the subject in the current frame, the identification information of the subject at the second predicted position closest to the detection position of the subject in the current frame among the respective second predicted positions calculated for the current frame, to the image processing apparatus according to any one of Items 1 to 3. (Item 5) The selection means identifies the subject based on the detected area and the predicted area of the subject, to the image processing apparatus according to any one of Items 1 to 4. (Item 6) The selection means identifies the subject based on the detected orientation and the predicted orientation of the subject, to the image processing apparatus according to any one of Items 1 to 5. (Item 7) When there is no distance between subjects less than the threshold value in the current frame, the selection means selects a tracking subject from the subjects in the current frame based on the first identification result, to the image processing apparatus according to any one of Items 1 to 6. (Item 8) When there is a distance between subjects less than the threshold value in the current frame, the selection means selects a tracking subject from the subjects in the current frame based on the second identification result, to the image processing apparatus according to any one of Items 1 to 7. (Item 9) When the distance between the tracking subject and other subjects other than the tracking subject is equal to or greater than the threshold value, the selection means selects the tracking subject based on the first identification result, and when the distance is less than the threshold value, the selection means selects the tracking subject based on the second identification result, to the image processing apparatus according to any one of Items 1 to 6. (Item 10) The first calculation means calculates the first predicted position of the subject in each frame for each first period using a group of frames arranged in the first period. The second calculation means calculates the second predicted position of the subject in each frame for each second period using a group of frames arranged in a second period longer than the first period. The image processing apparatus according to item 1, characterized in that... (Item 11) The image processing apparatus according to item 10, characterized in that the selection means selects a tracking subject based on the first identification result when the second identification result cannot be obtained. (Item 12) The image processing apparatus according to item 10 or 11, characterized in that the selection means selects a tracking subject based on the first identification result when the first identification result and the second identification result are the same. (Item 13) The image processing apparatus according to any one of items 10 to 12, characterized in that the selection means selects a tracking subject based on the second identification result when the first identification result and the second identification result are different. (Item 14) The image processing apparatus according to any one of items 1 to 13, characterized in that the selection means selects a subject corresponding to the identification information as the tracking subject when receiving the identification information of the tracking subject from an external device. (Item 15) Furthermore, The image processing apparatus according to any one of items 1 to 14, characterized in that it comprises control means for controlling to track and photograph the tracking subject selected by the selection means. (Item 16) An image processing method performed by an image processing apparatus, comprising: a first calculation step in which a first calculation means of the image processing apparatus calculates a first predicted position of a subject using a first frame group; a second calculation step in which a second calculation means of the image processing apparatus calculates a second predicted position of a subject using a second frame group; a selection step in which a selection means of the image processing apparatus selects a tracking subject to be tracked based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position; An image processing method, characterized by comprising the above steps. (Item 17) A computer program for causing a computer to function as each means of the image processing apparatus according to any one of Items 1 to 15. (Item 18) A system comprising a camera and a control device, wherein the camera comprises photographing means, and control means for controlling so as to perform sequential shooting of a subject to be tracked, and the control device comprises first calculation means for calculating a first predicted position of a subject using a first frame group obtained by the photographing means, second calculation means for calculating a second predicted position of the subject using a second frame group obtained by the photographing means, and selection means for selecting the subject to be tracked based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position. A system characterized by the above.

[0176] The invention is not limited to the above embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Signs

[0177] 301: Photographing unit 302: Inference unit 303: Drive control unit 304: Communication unit 305: User interface unit 306: Inference unit 308: Communication unit 309: Calculation unit 310: Calculation unit

Claims

1. First calculation means for calculating a first predicted position of a subject using a first frame group; Second calculation means for calculating a second predicted position of the subject using a second frame group; Selection means for selecting a tracking subject to be tracked based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position An image processing apparatus comprising the same.

2. The first calculation means calculates a first predicted position of the subject using the first frame group in a first frame period, The second calculation means calculates a second predicted position of the subject using the second frame group in a second frame period longer than the first frame period The image processing apparatus according to claim 1, characterized in that.

3. The selection means assigns, to the subject in the current frame, identification information of the subject at the first predicted position closest to the detection position of the subject in the current frame among the first predicted positions calculated for the current frame. The image processing apparatus according to claim 1, characterized in that.

4. The selection means assigns, to the subject in the current frame, identification information of the subject at the second predicted position closest to the detection position of the subject in the current frame among the second predicted positions calculated for the current frame. The image processing apparatus according to claim 1, characterized in that.

5. The selection means identifies the subject based on the detected area and the predicted area of the subject. The image processing apparatus according to claim 1, characterized in that.

6. The selection means identifies the subject based on the detected orientation and the predicted orientation of the subject. The image processing apparatus according to claim 1, characterized in that.

7. When there is no distance between subjects less than a threshold value in the current frame, the selection means selects a tracking subject from the subjects in the current frame based on the first identification result. The image processing apparatus according to claim 1, characterized in that.

8. When there is a distance between subjects less than a threshold value in the current frame, the selection means selects a tracking subject from the subjects in the current frame based on the second identification result. The image processing apparatus according to claim 1, characterized in that.

9. When the distance between the following subject and another subject other than the following subject is equal to or greater than a threshold value, the selection means selects the following subject based on the first identification result, and when the distance is less than the threshold value, the following subject is selected based on the second identification result. The image processing apparatus according to claim 1, characterized in that

10. The first calculation means calculates a first predicted position of a subject in each frame for each first period using a group of frames arranged in the first period. The second calculation means calculates a second predicted position of a subject in each frame for each second period using a group of frames arranged in a second period longer than the first period. The image processing apparatus according to claim 1, characterized in that

11. When the second identification result cannot be obtained, the selection means selects the following subject based on the first identification result. The image processing apparatus according to claim 10, characterized in that

12. When the first identification result and the second identification result are the same, the selection means selects the following subject based on the first identification result. The image processing apparatus according to claim 10, characterized in that

13. When the first identification result and the second identification result are different, the selection means selects the following subject based on the second identification result. The image processing apparatus according to claim 10, characterized in that

14. When the selection means receives identification information of the following subject from an external device, the selection means selects the subject corresponding to the identification information as the following subject. The image processing apparatus according to claim 1, characterized in that

15. Furthermore, The image processing apparatus according to claim 1, further comprising control means for controlling so as to track and photograph the following subject selected by the selection means.

16. An image processing method performed by an image processing apparatus, A first calculation step in which the first calculation means of the image processing apparatus calculates a first predicted position of a subject using a first group of frames; A second calculation step in which the second calculation means of the image processing apparatus calculates a second predicted position of a subject using a second group of frames; A selection step in which the selection means of the image processing apparatus selects a following subject to be tracked based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position. An image processing method, characterized by comprising

17. A computer program for causing a computer to function as each means of the image processing apparatus according to any one of claims 1 to 15.

18. A system comprising a camera and a control device, wherein the camera comprises photographing means and control means for controlling so as to perform sequential tracking photography of a tracking subject, and the control device comprises first calculation means for calculating a first predicted position of a subject using a first frame group obtained by the photographing means, second calculation means for calculating a second predicted position of the subject using a second frame group obtained by the photographing means, and selection means for selecting the tracking subject based on a first identification result of the subject based on a detection position of the subject and the first predicted position, and a second identification result of the subject based on the detection position and the second predicted position. A system characterized by the above.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2023028908A