Information processing device, method and program
The information processing device controls the zoom of the imaging means to keep multiple tracking targets within the angle of view, addressing the issue of target detection loss by ensuring they remain visible.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-12
AI Technical Summary
Existing technologies fail to effectively control the imaging means to include multiple tracking targets within the angle of view while tracking each of the targets, risking detection loss due to targets becoming too small in the image.
An information processing device and method that sets multiple tracking targets, detects their sizes, and controls the zoom of the imaging means to keep each target within a predetermined range, ensuring they remain in the angle of view.
Enables the imaging means to track and include multiple targets by controlling zoom, preventing targets from becoming too small to be detected, thus maintaining effective tracking.
Smart Images

Figure 2026044263000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, method, and program. [Background technology]
[0002] Conventionally, there have been technologies that detect a tracking target, which is a subject designated as a target for tracking and photographing, from an image generated based on photographing by an imaging means, and track and photograph the detected tracking target. There are also technologies that, when zooming the imaging means to photograph the tracking target, control the zoom based on the size of the tracking target in the image, thereby maintaining the size of the tracking target in the image within a predetermined range even when the tracking target moves in the direction of the imaging optical axis of the imaging means. Patent Document 1 discloses a method for continuing to track and photograph the tracking target while maintaining an appropriate size in the image by measuring the size of the tracking target in the image and changing the zoom magnification so that the measured size does not exceed a desired size. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-112648 Summary of the Invention [Problem to be solved by the invention]
[0004] However, Patent Document 1 does not mention the operation for capturing multiple subjects as tracking targets within the angle of view. An object of the present invention is to control the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means while causing the imaging means to track each of the plurality of tracking targets that have been detected. [Means for solving the problem]
[0005] In order to solve the above problem, the information processing device of the present invention comprises a setting means for setting multiple tracking targets to be tracked in an image captured by an imaging means, a size detection means for detecting the size of the tracking targets in an image captured by the imaging means, and a control means for controlling the zoom of the imaging means so that the multiple tracking targets are within the angle of view of the imaging means, wherein the control means controls the zoom so that the size of each tracking target falls within a predetermined range for at least a portion of the multiple tracking targets. [Effects of the Invention]
[0006] According to the present invention, it is possible to control the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means while causing the imaging means to track each of the plurality of tracking targets that have been detected. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram illustrating the overall configuration of an imaging system. [Figure 2] FIG. 2 is a diagram illustrating the hardware configuration of a camera and the hardware configuration of a controller. [Figure 3] FIG. 1A is a diagram showing the functional configuration of the camera, and FIG. 1B is a diagram showing the functional configuration of the controller. [Figure 4] 10A to 10D are diagrams showing the relationship between a captured image and an index calculated by a calculation unit of the camera. [Figure 5] FIG. 10 is a sequence diagram showing a flow from when a camera transmits a captured image to a controller until a tracking target is set in the camera. [Figure 6] FIG. 10 is a flowchart showing the flow of a tracking process. [Figure 7] FIG. 10 is a flowchart showing the flow of a determination process. [Figure 8] FIG. 10 is a flowchart showing the flow of a tracking process. [Figure 9] FIG. 10 is a flowchart showing the flow of a determination process. [Figure 10]10A and 10B are diagrams for explaining the relationship between an area in a captured image where a tracking target object is located and restrictions on zooming in by a camera. [Figure 11] 10A and 10B are diagrams for explaining the relationship between whether or not a tracking target object is detected by a detection unit in a captured image and the zoom restrictions imposed by the camera. [Figure 12] FIG. 10 is a flowchart showing the flow of a tracking process. DETAILED DESCRIPTION OF THE INVENTION
[0008] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is an overall configuration diagram of an imaging system 1. The imaging system 1 detects multiple tracking targets and tracks and captures the detected tracking targets. The tracking targets are subjects predetermined as targets for tracking by the imaging means. Examples of tracking targets include humans. In this embodiment, a case is assumed in which it is desirable to set multiple subjects as tracking targets. In this case, it is necessary to operate the imaging means, such as zooming, so that the multiple tracking targets are within the field of view of the imaging means. However, if only consideration is given to ensuring that the multiple tracking targets are within the field of view of the imaging means, there is a risk that tracking each of the multiple tracking targets will be difficult, for example, because some tracking targets will not be detected due to their excessively small size in the image. Therefore, in this embodiment, tracking of multiple tracking targets is achieved by restricting imaging control when tracking conditions are not met. The imaging system 1 includes a camera 100 and a controller 200. The camera 100 and the controller 200 are connected via a network 300.
[0009] Camera 100, which is an example of an imaging means, operates to generate an image by capturing an image, detect multiple tracking targets from the generated image, and track and capture the detected multiple tracking targets. Operations performed by camera 100 include panning, tilting, zooming, and the like. Zooming, which is an operation performed by camera 100, includes zooming in and zooming out. Note that an image generated by capturing an image with camera 100 may be referred to as a captured image below. Camera 100 can also be considered as an information processing device. Controller 200, which is an example of an information processing device, controls camera 100. Controller 200 of this embodiment acquires from camera 100 captured images and the results of detection by camera 100 of subjects captured in the captured images, and displays the acquired information to accept a user's selection of which subject to set as a tracking target. Controller 200 also sets the selected subject as a tracking target, and transmits information indicating the set tracking target to camera 100.
[0010] The camera 100 and the controller 200 may be configured by a single computer, or may be realized by distributed processing using multiple computers. The network 300 is realized by, for example, a LAN (Local Area Network) such as the Internet, a WAN (Wide Area Network), etc. Furthermore, the network 300 may be realized not only by the Internet, but also by any one or a combination of telephone lines, dedicated digital lines, ATM (Asynchronous Transfer Mode), frame relay lines, cable television lines, wireless lines for data broadcasting, etc. Examples include:
[0011] FIG. 2 is a diagram showing the hardware configuration of the camera 100 and the hardware configuration of the controller 200. As shown in FIG. The camera 100 includes a CPU 101, a RAM 102, a ROM 103, a GPU 104, a network I / F 105, a sensor I / F 106, an image sensor 107, a drive I / F 108, and a drive unit 109. The CPU 101, the RAM 102, the ROM 103, the GPU 104, the network I / F 105, the sensor I / F 106, the image sensor 107, the drive I / F 108, and the drive unit 109 are connected to each other via a bus 110.
[0012] The CPU 101 controls the entire camera 100 by executing various processes using computer programs and data stored in the RAM 102. The RAM 102 is a high-speed storage device such as a DRAM, and stores various types of information, such as computer programs loaded from the ROM 103, captured images, and information acquired from the controller 200. The RAM 102 also has a working area used by the CPU 101 and the GPU 104 when executing various processes. The ROM 103 is a non-volatile storage device such as a flash memory, HDD, SSD, or SD card, and stores setting data for the camera 100, computer programs and data related to the startup and basic operation of the camera 100, and the like. The ROM 103 also stores computer programs and data for causing the CPU 101 and the GPU 104 to execute or control various processes described as processes performed by the camera 100. The GPU 104 performs inference processing to estimate the presence or absence of a subject and the area of the subject from a captured image. The GPU 104 is a computing device specialized for image processing and inference processing, such as a GPU (Graphics Processing Unit). Note that an arithmetic device such as an FPGA (Field Programmable Gate Array) may be used instead of the GPU 104. Furthermore, the processing of the GPU 104 may be performed by the CPU 101. The network I / F 105 is an interface for connecting to the network 300 and communicates with external devices such as the controller 200 via a communication medium such as Ethernet (registered trademark). The sensor I / F 106 converts the video signal output from the image sensor 107 into a captured image, which is data in a predetermined format, and outputs the converted captured image to the RAM 102 after compressing it as necessary. Note that the sensor I / F 106 may perform various types of processing on the image represented by the video signal acquired from the image sensor 107, such as image quality adjustments such as color correction, exposure correction, and sharpness correction, and cropping to cut out only a predetermined area. These processings by the sensor I / F 106 may be performed in accordance with instructions received from the controller 200 via the network I / F 105. The image sensor 107 receives light reflected from a subject, converts the brightness and color of the received light into electric charges, and outputs a video signal based on the result of the conversion.Examples of image sensor 107 include a photodiode, a CCD (Charge Coupled Device) sensor, and a CMOS (Complementary Metal Oxide Semiconductor) sensor. Drive I / F 108 is an interface for transmitting and receiving instruction signals such as control signals to and from drive unit 109. Drive unit 109 is a drive mechanism for changing the shooting direction of camera 100, and includes a mechanical drive system and a drive motor. Drive unit 109 performs panning and tilting to change the shooting direction horizontally and vertically, and zooming to optically change the shooting angle of view, in accordance with instructions received from CPU 101 via drive I / F 108.
[0013] The controller 200 also includes a CPU 201, a RAM 202, a ROM 203, a GPU 204, a network I / F 205, a display unit 206, and an operation unit 207. The CPU 201, the RAM 202, the ROM 203, the GPU 204, the network I / F 205, the display unit 206, and the operation unit 207 are connected to each other via a bus 208.
[0014] The CPU 201 controls the entire controller 200 by executing various processes using computer programs and data stored in the RAM 202. The RAM 202 is a high-speed storage device such as a DRAM. The RAM 202 stores computer programs and data loaded from the ROM 203 and various data acquired from the camera 100. The RAM 202 also has a working area used by the CPU 201 and the GPU 204 when executing various processes. The ROM 203 is a non-volatile storage device such as a flash memory, HDD, SSD, or SD card, and stores setting data for the controller 200, computer programs and data related to the startup and basic operations of the controller 200, etc. The ROM 203 also stores computer programs and data for causing the CPU 201 and the GPU 204 to control various processes. The GPU 204 performs inference processing to estimate the presence or absence of a subject and the area of the subject from a captured image. The GPU 204 is, for example, a computing device specialized for image processing and inference processing, such as a GPU. Note that an arithmetic device such as an FPGA may be used instead of the GPU 204. Furthermore, the processing of the GPU 204 may be performed by the CPU 201. The network I / F 205 is an interface for connecting to the network 300, and communicates with external devices such as the camera 100 via a communication medium such as Ethernet. The display unit 206 has a screen such as an LCD screen or a touch panel screen, and displays captured images acquired from the camera 100, a setting screen for the controller 200, and the like. The following describes a configuration in which the display unit 206 has a touch panel screen. Note that in the imaging system 1, the controller 200 may not be provided with the display unit 206, but instead may be connected to a display device (not shown) that displays information on the controller 200, and the captured images and the setting screen for the controller 200 may be displayed on the display device. The operation unit 207 is a user interface that accepts user operations on the controller 200, and may be, for example, a button, a dial, a joystick, a touch panel, or the like. The controller 200 may be a PC (personal computer) having a mouse, keyboard, and the like as the operation unit 207.
[0015] 3A is a diagram showing the functional configuration of camera 100. Camera 100 has an acquisition unit 111, a storage unit 112, a detection unit 113, an output unit 114, an extraction unit 115, a calculation unit 116, and a control unit 117. The acquisition unit 111 acquires information from the controller 200. Examples of the information acquired by the acquisition unit 111 include information indicating a subject set as a tracking target. The storage unit 112 stores information acquired by the acquisition unit 111 and information generated by the camera 100, such as captured images.
[0016] The detection unit 113 detects a subject, such as a person, from a captured image and detects the area of the subject in the captured image. Examples of the area of the subject in the captured image include the size of the subject in the captured image, the length of the subject in the captured image, and the position of the subject in the captured image. The detection unit 113 estimates the subject from the input captured image by performing inference processing using a trained model created using a machine learning method such as deep learning, and outputs information indicating coordinates corresponding to the area of the subject as a result. Note that examples of coordinates corresponding to the area of the subject include coordinates corresponding to a partial or entire area of the subject in the captured image. Examples of coordinates corresponding to a portion of the subject include coordinates corresponding to the outline of the subject, coordinates corresponding to the head or face of the subject, etc. Examples of coordinates corresponding to a portion of the subject include, when the detection unit 113 detects the subject as a rectangle, coordinates of the upper left and lower right vertices of the rectangle, coordinates of the center of the rectangle, coordinates corresponding to the width direction of the rectangle, and coordinates corresponding to the height direction of the rectangle. Here, the detection unit 113 can also be considered as a size detection unit that detects the size of the tracking target in the image captured by the camera 100. The detection unit 113 can also be regarded as a subject detection unit that detects a specific subject from a captured image. Examples of the specific subject include a person or other type of subject that is predetermined as a candidate for a tracking target. The detection unit 113 can also be regarded as a position detection unit that detects the position of the tracking target in the captured image. The method of detecting a subject by the detection unit 113 is not limited to a machine learning method. The detection unit 113 may use a template matching method in which a template image in which a subject to be detected appears is compared with the captured image, and an area in the captured image that has a high similarity to the subject appearing in the template image is detected as an area in which the subject is indicated. The detection unit 113 may use any method for detecting a subject. The subject to be detected by the detection unit 113 may be an object other than a person.
[0017] The detection unit 113 also identifies each of the detected subjects by detecting their features and generates information for identifying the subject for each identified subject. The detection unit 113 performs inference processing using a trained model created using a machine learning technique such as deep learning, thereby outputting information indicating the feature amounts of the subject based on the input photographed image and the area of the subject in the photographed image. The feature amounts extracted by the detection unit 113 may be the feature amounts of the entire subject, or the feature amounts of a part of the subject, such as a person's head or face. The feature amounts of the subject may include a feature vector of the image. In this case, a machine learning model may be used in which photographed images generated by photographing each subject from various angles are used as training images, images showing the same subject are labeled with the same ID, input to the training model, and feature vectors are output. The training model used by the detection unit 113 to identify the subject may be any training model. The method of identifying a subject by the detection unit 113 is not limited to a machine learning method. The detection unit 113 may predict the area of the subject in the latest captured image from the transition of the area of the subject in consecutive captured images in the past using a Kalman filter or the like, and identify the subject closest to the predicted area as the same subject. The method of identifying a subject by the detection unit 113 may be any method.
[0018] Upon detecting the subject, the detection unit 113 generates information indicating coordinates corresponding to the area of the subject and information indicating the feature amounts of the subject, and stores the generated information in the storage unit 112. Note that the information indicating coordinates corresponding to the area of the subject and the information indicating the feature amounts of the subject generated by the detection unit 113 are both regarded as the results of detection by the detection unit 113 of the subject appearing in the captured image. The output unit 114 outputs information to the controller 200. Examples of information output to the output unit 114 include information indicating the captured image and the result of detection by the detection unit 113 of the subject appearing in the captured image.
[0019] The extraction unit 115 extracts a tracking target from a captured image. The extraction unit 115 extracts the tracking target from information transmitted from the controller 200 as information indicating the subject set as the tracking target and the result of detection by the detection unit 113 on the captured image. Furthermore, when multiple tracking targets appear in the captured image, the extraction unit 115 extracts all of the tracking targets appearing in the captured image. When the extraction unit 115 extracts the tracking target, it stores in the storage unit 112 information indicating which of the subjects detected by the detection unit 113 is the tracking target. Furthermore, the extraction unit 115 determines, from the extracted tracking targets, a tracking target that satisfies a predetermined condition regarding size or length. The extraction unit 115 determines the tracking target that satisfies the predetermined condition from the detection result by the detection unit 113, such as an area in the captured image of the subject as the extracted tracking target.
[0020] The calculation unit 116 calculates the reference positions of the tracking target group in the captured image and the sizes of the tracking target group in the captured image, with the multiple tracking target objects extracted by the extraction unit 115 as a tracking target group. The calculation method used by the calculation unit 116 will be described in detail later. Control unit 117, which is an example of a control means, controls operations such as panning, tilting, and zooming by driving unit 109 (see FIG. 2) of camera 100. When causing driving unit 109 to zoom, control unit 117 of this embodiment controls zooming by driving unit 109 so that camera 100 can continue tracking each of the multiple tracking targets extracted by extraction unit 115. In other words, when causing driving unit 109 to zoom, control unit 117 limits zooming (for example, limits the range of the angle of view obtainable by zoom control) if causing camera 100 to zoom would make it impossible for camera 100 to continue tracking any of the multiple tracking targets extracted by extraction unit 115.
[0021] The processes performed by acquisition unit 111, detection unit 113, output unit 114, extraction unit 115, calculation unit 116, and control unit 117 of camera 100 are implemented by CPU 101 or GPU 104 loading programs stored in ROM 103 into RAM 102 and executing them. Also, storage unit 112 of camera 100 is implemented by RAM 102 or ROM 103.
[0022] 3B is a diagram showing the functional configuration of the controller 200. The controller 200 includes an acquisition unit 211, a storage unit 212, a target setting unit 213, and an output unit 214. The acquisition unit 211 acquires information from the camera 100. Examples of the information acquired by the acquisition unit 211 include a captured image and information indicating the result of detection by the detection unit 113 of the camera 100 on the captured image. The storage unit 212 stores the information acquired by the acquisition unit 211 and the information generated by the controller 200.
[0023] The target setting unit 213 sets a tracking target. The target setting unit 213 displays, on the display unit 206 of the controller 200 (see FIG. 2 ), information indicating the captured image transmitted from the camera 100 and the subject detected in the captured image by the detection unit 113. The target setting unit 213 then accepts a selection of a tracking target by the user based on the information displayed on the display unit 206, and sets the subject selected by the user as the tracking target. The target setting unit 213 also generates information indicating which subject has been set as the tracking target, and stores the generated information in the storage unit 212. Output unit 214 outputs information to camera 100 or controller 200. Examples of information output from output unit 214 to camera 100 include information generated by target setting unit 213 as information indicating which subject has been set as the tracking target. Examples of information output by output unit 214 to controller 200 include displaying an image on display unit 206.
[0024] The processes performed by the acquisition unit 211, the target setting unit 213, and the output unit 214 of the controller 200 are implemented by the CPU 201 or the GPU 204 loading programs stored in the ROM 203 into the RAM 202 and executing the programs. The storage unit 212 of the controller 200 is implemented by the RAM 202 or the ROM 203.
[0025] FIG. 4 is a diagram showing the relationship between a captured image and the indices calculated by the calculation unit 116 of the camera 100. As shown in FIG. 4(A) shows a captured image 401. Also shown in the captured image 401 are a plurality of tracking targets 402 consisting of tracking targets 402a and 402b, and a plurality of non-tracking targets 403, none of which are tracking targets. Here, the plurality of tracking targets 402 consisting of tracking targets 402a and 402b constitute a tracking target group. 4(A), a rectangular image 413 is shown in each of an area overlapping the tracking target object 402a and an area overlapping the tracking target object 402b. The rectangular image 413 is a rectangular image showing the area of the tracking target object 402 detected by the detection unit 113.
[0026] 4A, the calculation unit 116 calculates the reference position of the group of tracking targets from the captured image 401. More specifically, the calculation unit 116 calculates a rectangular area 404 that is a rectangular area that includes a rectangular image 413 for each of the tracking targets 402, and calculates the central area of the calculated rectangular area 404 as the reference position 405 of the group of tracking targets. Furthermore, as shown in FIG. 4B, the calculation unit 116 calculates a rectangular area 408, which is the area of the calculated rectangular area 404, as the size of the group of tracking targets.
[0027] The method used by calculation unit 116 to calculate the reference positions of the tracking target group and the size of the tracking target group is not limited to the above example. 4C shows a captured image 401 including a plurality of tracking targets 402, namely, a tracking target group, including tracking target 402a, tracking target 402b, and tracking target 402c. Also, FIG. 4C shows a rectangular image 413 for each tracking target 402. In this case, the calculation unit 116 may calculate, for each rectangular image 413, a central region 411 that is the center of the region surrounded by the rectangular image 413, and calculate, as the reference position 407 of the tracking target group, coordinates representing the average value of the coordinates corresponding to each calculated central region 411. In this case, the coordinates representing the average value of the coordinates corresponding to each central region 411 may be calculated as a weighted average of the coordinates corresponding to each central region 411. In this way, the reference position of the tracking target group becomes closer to a region in the tracking target group where more tracking targets are gathered. Therefore, when the driving unit 109 of the camera 100 operates using the reference position of the group of tracking targets as the target position, it becomes possible to operate in a region where more tracking targets are concentrated in the group of tracking targets.
[0028] 4(C), when the number of tracking targets 402 constituting the tracking target group is three or more, the coordinates corresponding to the area of the center of gravity of polygon 406 having each central region 411 as a vertex coincide with the coordinates as the average value of the coordinates corresponding to each central region 411. Therefore, the coordinates corresponding to the area of the center of gravity of polygon 406 may be calculated as reference position 407 of the tracking target group.
[0029] Furthermore, when the calculation unit 116 calculates each central region 411, it calculates a rectangular region 412 that includes each central region 411, as shown in FIG. 4(D). Then, the calculation unit 116 may calculate the larger of the ratio of the width of the rectangular region 412 to the width of the captured image 401 and the ratio of the height of the rectangular region 412 to the height of the captured image 401 as the size of the tracking target group. Here, the width of the captured image 401 and the rectangular region 412 refers to the length in the horizontal direction in the drawing. Also, the height of the captured image 401 and the rectangular region 412 refers to the length in the vertical direction in the drawing. In the illustrated example, the ratio of the width of the rectangular region 412 to the width of the captured image 401 is larger than the ratio of the height of the rectangular region 412 to the height of the captured image 401. Therefore, the ratio of the width of the rectangular region 412 to the width of the captured image 401 becomes the size of the tracking target group. By determining the size of the group of tracking targets in this manner, it becomes possible to control the operation of the drive unit 109 of the camera 100 based on the size in the direction in which the group of tracking targets is wider, either left-right or up-down, in the captured image 401.
[0030] The calculation unit 116 may determine, as the size of the tracking target group, the ratio of the diagonal line 409 of the rectangular region 412 to the diagonal line of the outer frame of the captured image 401. The calculation unit 116 may also determine, as the size of the tracking target group, the width, height, or length of the diagonal line 409 of the rectangular region 412. Furthermore, the calculation unit 116 stores information indicating the calculation result in the storage unit 112.
[0031] FIG. 5 is a sequence diagram showing the flow from when camera 100 transmits a captured image to controller 200 until a tracking target object is set in camera 100. In FIG. First, the acquisition unit 111 of the camera 100 acquires a captured image generated by photography (step (hereinafter sometimes referred to as “S”) 101). The captured image acquired by the acquisition unit 111 is stored in the storage unit 112. The detection unit 113 of the camera 100 detects the presence of a subject from the captured image acquired by the acquisition unit 111, and detects the area of the subject in the captured image (S102). The detection unit 113 also identifies the subject by detecting the characteristics of the subject, and generates information for identifying the subject. The detection unit 113 stores information indicating the detection result, including the information for identifying the subject, in the storage unit 112. The output unit 114 of the camera 100 transmits the photographed image acquired by the acquisition unit 111 and information indicating the result of detection by the detection unit 113 to the controller 200 (S103).
[0032] The output unit 214 of the controller 200 causes the captured image and the information identifying the subject transmitted from the camera 100 to be displayed on the display unit 206. More specifically, the output unit 214 causes the display unit 206 to display the captured image on which the information identifying the subject is superimposed (S104). The target setting unit 213 receives a user's selection of which subject in the captured image displayed on the display unit 206 is to be set as the tracking target (S105). The target setting unit 213 transmits information indicating the tracking target selected by the user to the camera 100 via the output unit 214 (S106). The extraction unit 115 of the camera 100 sets a tracking target object in accordance with the information transmitted from the controller 200 in step 106 (S107). Therefore, the extraction unit 115 can also be regarded as a setting unit that sets a plurality of tracking targets to be tracked in image capture by the camera 100. The extraction unit 115 also stores information indicating the set tracking targets in the storage unit 112.
[0033] 5 may be performed each time the camera 100 captures a new image and generates a captured image. Each time a tracking target object is selected by the user, the settings of the tracking target object may be updated in the camera 100 and the controller 200. Furthermore, in the above example, the subject selected by the user is set as the tracking target, but examples of setting the tracking target are not limited to the above example. For example, the target setting unit 213 may set all of the subjects of a predetermined type detected from the captured image as the tracking target. Examples of the predetermined type include humans. After the user selects one tracking target, the target setting unit 213 may set a subject located near the one tracking target selected by the user in the captured image as the tracking target. Furthermore, the target setting unit 213 may set a subject having characteristics similar to the one tracking target selected by the user in the captured image as the tracking target. This reduces the burden on the user for setting the tracking target compared to when the user is required to select all of the tracking targets to be set.
[0034] 6 is a flowchart showing the flow of the tracking process. The tracking process is a process in which the camera 100 tracks a tracking target group, which is a plurality of tracking targets. In this embodiment, when the camera 100 is set to a mode in which the tracking target group is tracked, the tracking process is started when a captured image is generated by capturing an image with the camera 100. The acquisition unit 111 of the camera 100 acquires a captured image generated by photographing (S301). The detection unit 113 of the camera 100 detects, from the captured image acquired by the acquisition unit 111, a subject of the same type as the tracking target set in the extraction unit 115 (S302). When the tracking target set in the extraction unit 115 is a human, the detection unit 113 detects the human as a subject that is a candidate for the tracking target from the captured image. In this case, the detection unit 113 detects the area of the subject and also identifies the subject by detecting the features of the subject, and generates information for identifying the subject.
[0035] The processing in steps 301 and 302 may be the same as the processing in steps 101 and 102 shown in Fig. 5. Furthermore, if the processing in steps 301 and 302 and the processing in steps 101 and 102 are the same processing for the same captured image, the processing in steps 301 and 302 may be omitted.
[0036] The extraction unit 115 determines whether or not a plurality of tracking targets have been extracted from the captured image (S303). The extraction unit 115 extracts tracking targets from the captured image, and performs the determination in step 303 depending on whether or not a plurality of tracking targets have been extracted. If the extraction unit 115 has not extracted a tracking target object from the captured image, or if the number of tracking targets extracted by the extraction unit 115 is 1 (No in S303), the tracking process ends. Here, if the extraction unit 115 has not extracted a tracking target object from the captured image, the control unit 117 does not control the driving unit 109. Furthermore, if the number of tracking targets extracted by the extraction unit 115 is 1, the control unit 117 causes the driving unit 109 to perform pan, tilt, and zoom operations so as to track the single extracted tracking target object.
[0037] Furthermore, if the extraction unit 115 extracts a plurality of tracking targets (Yes in S303), the calculation unit 116 calculates the reference position and size of the group of tracking targets (S304). The extraction unit 115 performs a determination process (S305). As will be described in detail later, in this determination process, the extraction unit 115 determines a tracking target that satisfies a predetermined condition for size or length from among the multiple tracking targets extracted by the extraction unit 115 as a representative tracking target among the multiple tracking targets. Note that the representative tracking target determined in the determination process may be referred to as a representative target hereinafter. Furthermore, in the determination process, the extraction unit 115 determines an index as the size or length of the representative target.
[0038] The control unit 117 determines whether the index determined in the determination process as the size or length of the representative object is equal to or greater than a predetermined threshold value as a lower limit (S306). The control unit 117 may, for example, determine whether the longitudinal length of the rectangular image 413 (see FIG. 4(A) etc.) generated for the representative object is equal to or greater than a threshold value as a lower limit. The control unit 117 may also, for example, determine whether the area of the region surrounded by the rectangular image 413 generated for the representative object is equal to or greater than a threshold value as a lower limit. The control unit 117 may also, for example, determine whether the length of the diagonal line connecting two vertices in the rectangular image 413 generated for the representative object is equal to or greater than a threshold value as a lower limit. In other words, the index for the representative object used as the criterion for determination in step 306 may be any index that allows comparison of the size or length of the representative object. The threshold value as the lower limit may be, for example, the minimum size or shortest length at which the detection unit 113 can detect a subject from a captured image. The threshold value as the lower limit may be, for example, a value obtained by adding a predetermined size or length to the minimum size or shortest length at which the detection unit 113 can detect a subject from a captured image. The predetermined size or length may be any value, but is preferably a value determined so that detection by the detection unit 113 is possible. The threshold value as the lower limit may be, for example, the minimum size or shortest length at which the detection unit 113 can detect a feature of the subject from a captured image. The threshold value as the lower limit may be, for example, a value obtained by adding a predetermined size or length to the minimum size or shortest length at which the detection unit 113 can detect a feature of the subject from a captured image. The threshold value as the lower limit may be a value set by a user's operation on the camera 100 or the controller 200. The threshold value as the lower limit may be a value set so as to ensure the image quality of the tracking target object.
[0039] If the index for the representative object is equal to or greater than the lower limit (Yes in S306), the process proceeds to the next step. The control unit 117 calculates the difference between the size of the tracking target group calculated by the calculation unit 116 and a value preset by the user as a target value for the size of the tracking target group in the captured image, and calculates the amount of zoom control according to the calculated difference (S307). In this case, the control unit 117 determines whether to cause the drive unit 109 to zoom in or zoom out so that the calculated difference becomes smaller. The control unit 117 also determines the zoom speed so that the larger the calculated difference, the faster the zoom speed. In this case, the zoom is controlled so that the calculated size of the tracking target group approaches the target value. Note that the control unit 117 may change the zoom speed even if the magnitude of the calculated difference is the same when causing the driving unit 109 to zoom in and when causing the driving unit 109 to zoom out. In particular, the control unit 117 may make the zoom speed faster when causing the driving unit 109 to zoom out than when causing the driving unit 109 to zoom in. In this case, the tracking responsiveness of the camera 100 to the spread of the group of tracking targets is improved.
[0040] Furthermore, when camera 100 tracks and photographs a group of tracking targets, the moving speed of the tracking targets may be faster than when camera 100 tracks and photographs a single tracking target, and a high-speed zoom may be required. Therefore, control unit 117 may set different zoom speeds when camera 100 tracks and photographs a group of tracking targets and when camera 100 tracks and photographs a single tracking target. In particular, control unit 117 may set the zoom speed to be faster when camera 100 tracks and photographs a group of tracking targets than when camera 100 tracks and photographs a single tracking target.
[0041] Furthermore, if the index for the representative object is less than the lower limit value (No in S306), the control unit 117 limits the control of the driving unit 109 to zoom (S308). As a zoom limit, the control unit 117 may, for example, not allow the driving unit 109 to zoom. As a zoom limit, the control unit 117 may slow down the zoom speed or shorten the zoom time compared to when zoom is not limited. Furthermore, the zoom limited in step 308 is specifically zoom out. That is, the control unit 117 does not need to limit zoom in by the driving unit 109 in step 308. However, the control unit 117 may limit both zoom in and zoom out by the driving unit 109 in step 308.
[0042] The control unit 117 compares the reference position calculated by the calculation unit 116 for the group of tracking targets with the position set by the user as the target for the position of the group of tracking targets in the captured image, and calculates the amount of control for panning and tilting according to the result of the comparison (S309). In this case, the control unit 117 determines the direction and speed of panning and tilting so that the reference position calculated for the group of tracking targets approaches the position set as the target. The control unit 117 also determines the speed of panning and tilting so that the greater the difference in distance between the reference position calculated for the group of tracking targets and the position set as the target, the faster the panning and tilting speeds become. In this case, panning and tilting are controlled so that the reference position calculated for the group of tracking targets approaches the target position.
[0043] The control unit 117 operates the driving unit 109 in accordance with the content determined in step 308 or step 309 (S310). More specifically, the control unit 117 gives the driving unit 109 an instruction for an operation indicated by the content determined in step 308 or step 309, to the driving unit 109. As a result, the driving unit 109 performs operations such as panning, tilting, and zooming in accordance with the instruction from the control unit 117.
[0044] The tracking process may be terminated when a termination condition is satisfied after the process of step 310. Examples of the termination condition include receiving an instruction to terminate the tracking process from camera 100 or controller 200, reaching a predetermined date and time, or a predetermined time having elapsed since the start of the tracking process. If the termination condition is not satisfied, the process from step 301 may be repeated for a newly generated captured image.
[0045] FIG. 7 is a flowchart showing the flow of the determination process (see step 305 in FIG. 6). The extraction unit 115 determines, as a representative object, a tracking object that satisfies a minimum condition among the multiple tracking objects extracted in step 303 of the tracking process (S501). The minimum condition is a predetermined condition regarding the size or length of the tracking object. In this embodiment, the minimum condition is determined to be the smallest size or the shortest length of the tracking object in the captured image. Note that the index of the size or length of the tracking object used by the extraction unit 115 to identify the tracking object that satisfies the minimum condition is the same index as the index used for the determination in step 306 of the tracking process.
[0046] The extraction unit 115 determines whether the index as the size or length of the representative object has changed from the index determined in the previous determination process (S502). If the index as the size or length of the representative object has not changed from the index determined in the previous determination process (No in S502), the determination process ends. In this case, the index determined in the previous determination process is taken over as the index used for the determination in step 306 of the tracking process.
[0047] In addition, there may be cases where the index of the size or length of the representative object has changed from the index determined in the previous determination process (Yes in S502), in which case the extraction unit 115 determines a new index of the size or length of the representative object as the index to be used for the determination in step 306 of the tracking process (S503). If the current determination process is the first determination process, the extraction unit 115 determines that the index of the size or length of the representative object has changed from the index determined in the previous determination process. Furthermore, information indicating the index determined in the determination process as the index to be used in the determination in step 306 of the tracking process is stored in the storage unit 112 of the camera 100.
[0048] As described above, in this embodiment, when the size of the smallest tracking target object in the captured image or the length of the shortest tracking target object in the captured image is equal to or greater than the threshold value as the lower limit, the camera 100 performs zooming in and out without any restrictions. When the size of the smallest tracking target object in the captured image or the length of the shortest tracking target object in the captured image is less than the threshold value as the lower limit, the camera 100 restricts zooming out. In this case, compared to when the zoom-out of camera 100 is not limited regardless of the size or length of the tracking target, it is possible to prevent the occurrence of a tracking target that becomes so small in the captured image that it cannot be detected by camera 100. Therefore, camera 100 can track each of the detected tracking targets and zoom out so that the tracking targets are included in the angle of view of the tracking target group. Furthermore, in this embodiment, an example of image capture control based on the size or length of all tracking targets has been described, but image capture control may also be performed based on the size of some of the tracking targets. For example, image capture control may be performed based only on high-priority tracking targets, such as a representative subject or a similar subject, for example, a subject that has undergone personal authentication. In other words, the zoom-out of the camera 100 is limited when the size or length of these relatively high-priority tracking targets is less than a threshold value as a lower limit, i.e., smaller than the reference value. In this case, other tracking targets with relatively lower priorities may be smaller than the reference value. Furthermore, the method of determining the multiple subjects to be tracked is not limited to the method introduced above, and other methods may be used. For example, a main subject determination process may be performed based on the position and size within the angle of view of the subject detected using detection of a specific subject such as a person (face, head) or an animal, personal authentication, or user selection, and the selected main subject may be set as the tracking target, and subjects that exist in the vicinity of the selected main subject within a predetermined distance range and that have also been detected in the detection process of the specific subject may be grouped together and set as the tracking target.
[0049] (Variation) Next, a modified example of the tracking process will be described. The tracking process of this embodiment is not limited to that shown in FIG. Fig. 8 is a flowchart showing the flow of a modified tracking process. Note that the processing of steps 601 to 604 in the tracking process shown in Fig. 8 is the same as the processing of steps 301 to 304 in the tracking process shown in Fig. 6. The extraction unit 115 performs a determination process (S605). Although details will be described later, the determination process performed in step 605 and the determination process performed in step 305 of the tracking process shown in Fig. 6 use different methods for determining a representative object by the extraction unit 115.
[0050] The control unit 117 determines whether the index determined in the determination process as the size or length of the representative object is equal to or less than a predetermined threshold value as an upper limit (S606). The control unit 117 may, for example, determine whether the longitudinal length of the rectangular image 413 (see FIG. 4(A) etc.) generated for the representative object is equal to or less than a threshold value as an upper limit. The control unit 117 may also, for example, determine whether the area of the region surrounded by the rectangular image 413 generated for the representative object is equal to or less than a threshold value as an upper limit. The control unit 117 may also, for example, determine whether the length of the diagonal line connecting two vertices in the rectangular image 413 generated for the representative object is equal to or less than a threshold value as an upper limit. In other words, the index for the representative object used as the criterion for determination in step 606 may be an index that allows comparison of the size or length of the representative object. The threshold value as the upper limit may be, for example, the maximum size or the longest length at which the detection unit 113 can detect a subject from a captured image. The threshold value as the upper limit may be, for example, a value obtained by subtracting a predetermined size or length from the maximum size or the longest length at which the detection unit 113 can detect a subject from a captured image. The predetermined size or length may be any value, but is preferably a value determined so that detection by the detection unit 113 is possible. The threshold value as the upper limit may be, for example, the maximum size or the longest length at which the detection unit 113 can detect a feature of a subject from a captured image. The threshold value as the upper limit may be, for example, a value obtained by subtracting a predetermined size or length from the maximum size or the longest length at which the detection unit 113 can detect a feature of a subject from a captured image. The threshold value as the upper limit may be a value set by a user's operation on the camera 100 or the controller 200. The threshold value as the upper limit may be a value set so as to ensure the image quality of the tracking target object.
[0051] If the index for the representative object is equal to or less than the upper limit, the process proceeds to step 607. The process in step 607 is the same as the process in step 307 in FIG. Furthermore, if the index for the representative object is greater than the upper limit, the control unit 117 limits the control of the driving unit 109 to zoom (S608). As a zoom limit, the control unit 117 may, for example, not allow the driving unit 109 to zoom. As a zoom limit, the control unit 117 may slow down the zoom speed or shorten the zoom time compared to when zoom is not limited. Furthermore, the zoom limited in step 608 is specifically zoom-in. That is, the control unit 117 may not limit zoom-out by the driving unit 109 in step 608. However, the control unit 117 may limit both zoom-in and zoom-out by the driving unit 109 in step 608. 6. The processes in steps 609 and 610 are the same as those in steps 309 and 310 in FIG.
[0052] FIG. 9 is a flowchart showing the flow of the determination process (see step 605 in FIG. 8) as a modified example. The extraction unit 115 determines the tracking target object that satisfies the maximum condition as the representative target object from among the multiple tracking target objects extracted in step 303 of the tracking processing (S701). The maximum condition is a condition determined in advance regarding the size or length of the tracking target object. In this embodiment, the maximum condition is determined to be that the size of the tracking target object in the captured image is the largest or the length of the tracking target object in the captured image is the longest. Note that the index of the size or length of the tracking target object used by the extraction unit 115 to identify the tracking target object that satisfies the maximum condition is the same index as the index used for the determination in step 606 of the tracking processing.
[0053] The extraction unit 115 determines whether the index of the size or length of the representative object has changed from the index determined in the previous determination process (S702). If the index of the size or length of the representative object has not changed from the index determined in the previous determination process (No in S702), the determination process ends. In this case, the index determined in the previous determination process is taken over as the index used for the determination in step 706 of the tracking process.
[0054] In addition, there may be cases where the index of the size or length of the representative object has changed from the index determined in the previous determination process (Yes in S702), in which case the extraction unit 115 determines the index of the size or length of the representative object as a new index to be used for the determination in step 606 of the tracking process (S703). If the current determination process is the first determination process, the extraction unit 115 determines that the index of the size or length of the representative object has changed from the index determined in the previous determination process. Furthermore, information indicating the index determined in the determination process as the index to be used in the determination in step 606 of the tracking process is stored in the storage unit 112 of the camera 100.
[0055] Thus, in this embodiment, when the size of the largest tracking target object in the captured image or the length of the longest tracking target object in the captured image is equal to or smaller than the threshold value as the upper limit, the zoom control of the camera 100 is freely performed for both zooming in and zooming out. When the size of the largest tracking target object in the captured image or the length of the longest tracking target object in the captured image is larger than the threshold value as the upper limit, the zooming in of the camera 100 is restricted. In this case, compared to when the zoom-in of camera 100 is not limited regardless of the size or length of the tracking target, it is possible to prevent the occurrence of a tracking target that becomes too large in the captured image to be detected by camera 100. Therefore, camera 100 can track each of the detected tracking targets, while zooming out so that the tracking targets are included in the angle of view of the tracking target group.
[0056] (Second embodiment) Next, an imaging system 1 according to a second embodiment will be described. Note that, for the imaging system 1 according to the second embodiment, configurations different from those of the imaging system 1 according to the first embodiment will be described, and descriptions of the same configurations as those of the imaging system 1 according to the first embodiment will be omitted. The control unit 117 of the camera 100 of this embodiment controls zooming by the driving unit 109 when the tracking conditions are satisfied, and limits zooming by the driving unit 109 when the tracking conditions are not satisfied. The tracking conditions are conditions used by the control unit 117 to determine whether or not the camera 100 can track individual tracking targets even when the driving unit 109 zooms. The tracking conditions will be described in detail later.
[0057] FIG. 10 is a diagram for explaining the relationship between the area in which the tracking target object 402 is located in the captured image 401 and the zoom-in restrictions imposed by the camera 100. In FIG. A captured image 401 shown in FIG. 10 shows a plurality of tracking targets 402, including a tracking target 402a and a tracking target 402b, and a plurality of non-tracking targets 403. In this embodiment, there is an area in the captured image 401 that is defined as a restricted area R. The restricted area R is an area in which zooming in by the camera 100 is restricted when a tracking target object 402 is located. The restricted area R may be any area, but in the illustrated example, it is an area indicated by diagonal lines as the periphery of the captured image 401.
[0058] When the tracking target 402 is located in the restricted area R, if the camera 100 zooms in, the tracking target 402 may be cut off from the captured image 401 or may not appear in the captured image 401. If the tracking target 402 is cut off from the captured image 401 or does not appear in the captured image 401, the tracking target 402 will not be detected by the detection unit 113, making it difficult for the camera 100 to track the tracking target 402. Therefore, in this embodiment, an area that may not be detected by the camera 100 when the camera 100 zooms in is defined as the restricted area R. One of the tracking conditions is that the areas detected by the detection unit 113 as the positions of the multiple tracking targets are not included in the restricted area R. In the example shown in the figure, a portion of the rectangular image 413 corresponding to the tracking target 402a is included in the restricted area R, and therefore the tracking condition is not satisfied.
[0059] Note that a tracking condition may be that no part of rectangular image 413 for each tracking target object 402 is included in restricted area R. Alternatively, a tracking condition may be that a part of rectangular image 413 for each tracking target object 402 that occupies a predetermined proportion or more is not included in restricted area R. The predetermined proportion may be any proportion, for example, half. Furthermore, the tracking condition determined regarding the relationship between the area detected by detection unit 113 as the position of each of the multiple tracking targets and restriction area R is the tracking condition applied when control unit 117 causes drive unit 109 to zoom in. In other words, when control unit 117 causes drive unit 109 to zoom out, control unit 117 does not need to apply the tracking condition determined regarding the relationship between the area detected by detection unit 113 as the position of each of the multiple tracking targets and restriction area R. In other words, control unit 117 may cause drive unit 109 to zoom out regardless of whether the tracking condition determined regarding the relationship between the area detected by detection unit 113 as the position of the tracking target and restriction area R is satisfied.
[0060] FIG. 11 is a diagram for explaining the relationship between whether or not the tracking target object 402 in the captured image 401 is detected by the detection unit 113 and the zoom restriction imposed by the camera 100. In FIG. 11(A) shows a plurality of tracking targets 402 consisting of tracking targets 402a and 402b. In this case, the number of tracking targets 402 detected by the detection unit 113 from the captured image 401 shown in FIG. 11(A) is assumed to be "2", i.e., tracking targets 402a and 402b.
[0061] Here, it is assumed that the captured image 401 shown in Fig. 11(B) is generated as the frame next to the captured image 401 shown in Fig. 11(A). The captured image 401 shown in Fig. 11(B) shows a tracking target object 402b, but does not show a tracking target object 402a (see Fig. 11(A)). In this embodiment, one of the tracking conditions is that the number of tracking targets extracted by the extraction unit 115 does not decrease. 11(B), the number of tracking targets 402 extracted by the extraction unit 115 from the captured image 401 is only one, i.e., tracking target 402b. In this case, the number of tracking targets extracted by the extraction unit 115 is reduced, and therefore the tracking condition is not satisfied.
[0062] Also, it is assumed that the captured image 401 shown in Fig. 11(C) is generated as the frame next to the captured image 401 shown in Fig. 11(A). The captured image 401 shown in Fig. 11(C) shows a tracking target object 402b, but does not show a tracking target object 402a (see Fig. 11(A)). Furthermore, the captured image 401 shown in Fig. 11(C) newly shows a tracking target object 402c. In this embodiment, one of the tracking conditions is that the tracking target object 402 detected by the extraction unit 115 in the first captured image 401 is also extracted by the extraction unit 115 in the second captured image 401, which was captured later than the first captured image 401. 11(C), the number of tracking targets 402 extracted by extraction unit 115 from captured image 401 is "2", i.e., tracking target 402b and tracking target 402c, and therefore there is no decrease in the number of tracking targets extracted by extraction unit 115. On the other hand, tracking target 402a, which was extracted in captured image 401 shown in Fig. 11(A), is not extracted in captured image 401 shown in Fig. 11(C), and therefore the tracking condition is not satisfied.
[0063] Although not shown in the figure, one of the tracking conditions is that the size or length of each of the multiple tracking targets 402 extracted by the extraction unit 115 is equal to or greater than a predetermined threshold value set as a lower limit. This threshold value set as a lower limit is a threshold value used for the determination in step 306 of the tracking process shown in Fig. 6. Another of the tracking conditions is that an index for the size or length of each of the multiple tracking targets 402 extracted by the extraction unit 115 is equal to or less than a predetermined threshold value set as an upper limit. This threshold value set as an upper limit is a threshold value used for the determination in step 606 of the tracking process shown in Fig. 8. In this manner, in this embodiment, a plurality of conditions are defined as tracking conditions. Then, when all of the plurality of conditions defined as tracking conditions are satisfied, the control unit 117 controls zooming by the driving unit 109. Furthermore, when at least one of the plurality of conditions defined as tracking conditions is not satisfied, the control unit 117 limits a portion of the zooming operation by the driving unit 109.
[0064] Fig. 12 is a flowchart showing the flow of the tracking process according to the second embodiment. Note that the processes of steps 1201 to 1203 in the tracking process shown in Fig. 12 are the same as the processes of steps 301 to 303 in the tracking process shown in Fig. 6. If multiple tracking targets have not been extracted from the captured image (No in S1203), the extraction unit 115 determines whether multiple tracking targets have been extracted from the captured image in the previous tracking process (S1204). Note that the captured image that is the subject of the determination as to whether multiple tracking targets have been extracted in the previous tracking process is a captured image generated by capturing at a time point earlier than the captured image that is the subject of the determination as to whether multiple tracking targets have been extracted in the current tracking process.
[0065] If multiple tracking targets were not extracted from the captured image in the previous tracking process (No in S1204), the tracking process ends. Note that if no tracking targets were extracted in the current tracking process, the control unit 117 does not control the driving unit 109. Also, if only a single tracking target was extracted in the current tracking process, the control unit 117 causes the driving unit 109 to perform pan, tilt, and zoom operations so as to track the extracted single tracking target. Furthermore, if multiple tracking targets are extracted from the captured image in the current tracking process (Yes in S1203), or if multiple tracking targets are extracted from the captured image in the previous tracking process (Yes in S1204), the process proceeds to step 1205. The process in step 1205 is the same as the process in step 304 in the tracking process shown in FIG.
[0066] The control unit 117 determines whether or not the tracking conditions are satisfied (S1206). More specifically, the control unit 117 determines whether or not all of the above-mentioned multiple conditions as the tracking conditions are satisfied. If the tracking condition is satisfied (Yes in S1206), the control unit 117 calculates a zoom control amount according to the difference between the size of the tracking target group calculated by the calculation unit 116 and a value preset by the user as a target value for the size of the tracking target group in the captured image. At this time, the control unit 117 also calculates a zoom control amount so that the tracking condition is satisfied even after zoom control by the driving unit 109 (S1207). More specifically, if the positions of the tracking targets constituting the tracking target group do not change, the control unit 117 calculates a zoom control amount so that the tracking condition is satisfied even after zooming by the driving unit 109 according to the control amount calculated in step 1207. If the positions of the tracking targets constituting the tracking target group do not change, this refers to a case where the positions of the tracking targets constituting the tracking target group do not change between before and after zooming by the driving unit 109 according to the control amount calculated in step 1207.
[0067] Furthermore, if the tracking condition is not satisfied (No in S1206), the control unit 117 determines whether or not the tracking target object that does not satisfy the tracking condition is a specific tracking target object (S1208). The specific tracking target object is a predetermined tracking target object among the tracking targets set in step 107 of FIG. 5. The specific tracking target object may be any tracking target object, but may be, for example, a tracking target object that has been determined to be of high importance for tracking. Furthermore, the specific tracking target object may be set by a user's operation of the camera 100 or the controller 200, or may be set by the detection unit 113 of the camera 100 or the target setting unit 213 of the controller 200.
[0068] If the tracking target object that does not satisfy the tracking condition is not a specific tracking target object (No in S1208), the control unit 117 determines whether or not the relaxation condition is satisfied (S1209). The relaxation condition is a condition used by the control unit 117 to determine whether or not to relax the restriction on zooming by the driving unit 109. An example of the relaxation condition is that a predetermined time has elapsed since zooming by the driving unit 109 was started in a state in which the tracking condition was not satisfied. Another example of the relaxation condition is that the user has not operated the camera 100 or the controller 200 to instruct the driving unit 109 to restrict zooming. Another example of the relaxation condition is that the distance from an area detected by the detection unit 113 as the position of a tracking target object that does not satisfy the tracking condition to a reference position of the tracking target group is equal to or greater than a predetermined distance. Another example of the relaxation condition is that the detection unit 113 has detected that the tracking target object that does not satisfy the tracking condition is not moving. The detection unit 113 may identify whether or not the tracking target is moving by comparing the captured image acquired in the currently running tracking process with the captured image acquired in the previous tracking process.
[0069] If the relaxed condition is satisfied (Yes in S1209), the process proceeds to step 1210. The process in step 1210 is the same as the process in step 307 in the tracking process shown in FIG. If the tracking target object that does not satisfy the tracking condition is a specific tracking target object (Yes in S1208), or if the relaxed condition is not satisfied (No in S1209), the process proceeds to step 1211. The process of step 1211 is the same as the process of step 308 in the tracking process shown in FIG. After step 1207, step 1210, or step 1211, the process proceeds to step 1212. The processes of step 1212 and step 1213 are the same as the processes of step 309 and step 310 in the tracking process shown in FIG.
[0070] As described above, the control unit 117 controls the zoom of the imaging unit so that the multiple tracking targets are included in the angle of view of the imaging unit. Furthermore, the control unit 117 controls the zoom so that the size of each of the multiple tracking targets falls within a predetermined range for at least some of the multiple tracking targets. The size of the tracking target may be the size or length of the tracking target detected by the detection unit 113. Furthermore, the predetermined range may be a range that is equal to or greater than a threshold value as a lower limit and equal to or less than a threshold value as an upper limit. In this case, compared to a configuration in which zooming is performed without controlling the size of each tracking target object, it is possible to prevent the detected multiple tracking targets from becoming undetectable after zooming. Therefore, it is possible to control the imaging means so that the multiple tracking targets are included in the angle of view of the imaging means while having the imaging means track each of the detected multiple tracking targets.
[0071] Furthermore, the process by which the imaging system 1 sets multiple tracking targets to be tracked in an image captured by the camera 100 can also be considered a setting process. Furthermore, the process by which the imaging system 1 detects the size of the tracking targets in an image captured by the camera 100 can also be considered a size detection process. Furthermore, the process by which the imaging system 1 controls the zoom of the camera 100 so that the multiple tracking targets are within the angle of view of the imaging means can also be considered a control process. In this control process, the zoom is controlled so that the size of each of the multiple tracking targets falls within a predetermined range for at least some of the multiple tracking targets. Furthermore, the function of the imaging system 1 to set multiple tracking targets to be tracked in an image captured by the camera 100 can also be considered a setting function. Furthermore, the function of the imaging system 1 to detect the size of the tracking targets in an image captured by the camera 100 can also be considered a size detection function. Furthermore, the function of the imaging system 1 to control the zoom of the imaging means so that the multiple tracking targets fall within the angle of view of the imaging means can also be considered a control function. This control function controls the zoom so that the size of each of the multiple tracking targets falls within a predetermined range for at least some of the multiple tracking targets.
[0072] It should be noted that the images used to detect the multiple tracking targets are not limited to captured images. For example, when the detection unit 113 detects a subject from a captured image, the detection unit 113 may generate a processed image based on the captured image, such as generating an image in which information about the detected subject is superimposed on the captured image. In this case, the extraction unit 115 may detect multiple tracking targets from the image generated by the detection unit 113. In this way, the image generated by processing the captured image is also included in the image generated based on the image captured by the imaging means.
[0073] Furthermore, when the size of at least some of the multiple tracking targets exceeds a threshold value as an upper limit, the control unit 117 limits zooming in. In this case, compared to a configuration in which zooming in is performed regardless of the size of each tracking target, it is possible to prevent the multiple tracking targets from becoming undetectable after zooming in. Furthermore, the control unit 117 limits zooming out when the size of at least some of the multiple tracking targets falls below a threshold value as a lower limit. In this case, the multiple tracking targets are prevented from becoming undetectable after zooming out, compared to a configuration in which zooming out is performed regardless of the size of each tracking target.
[0074] Furthermore, the detection unit 113 detects a specific subject from the captured image. Then, the extraction unit 115 sets a tracking target from among the subjects detected by the detection unit 113. In this case, the tracking target can be tracked from the time the tracking target is set. Furthermore, as described above, the above-mentioned predetermined range is determined from a size or length at which the detection unit 113 can detect a subject, or a value obtained by adding or subtracting a predetermined size to or from a size at which the detection unit 113 can detect a subject. In other words, the predetermined range is determined in relation to the detection capability of the detection unit 113. In this case, when a group of tracking targets is being tracked, zooming to a size at which each of the multiple tracking targets cannot be detected is suppressed. Furthermore, if the predetermined range is a range set by the user's operation of the camera 100 or the controller 200, the imaging means can be made to track and photograph the group of tracking targets while maintaining the size of the tracking targets desired by the user.
[0075] Furthermore, the control unit 117 limits the zoom when the number of tracking targets extracted by the extraction unit 115 is decreasing or when a tracking target that was extracted by the extraction unit 115 is no longer being extracted. In other words, the control unit 117 limits the zoom when each of the tracking targets among the multiple tracking targets is no longer detected by the detection unit 113. In this case, compared to a configuration in which zooming is performed regardless of whether each tracking target object is no longer detected by the detection unit 113, it is possible to prevent multiple tracking targets from becoming undetectable after zooming.
[0076] Furthermore, when the size of the smallest of the plurality of tracking targets falls below a threshold value as a lower limit, the control unit 117 limits zoom-out of the zoom. The smallest of the plurality of tracking targets may be a tracking target that satisfies a minimum condition. In this case, compared to a configuration in which zoom-out is performed regardless of the size of the smallest of the multiple tracking targets, it is possible to prevent the smallest of the multiple tracking targets from becoming undetectable after zoom-out.
[0077] Furthermore, when the size of the largest of the plurality of tracking targets exceeds a threshold value as an upper limit, the control unit 117 limits zooming in. The largest of the plurality of tracking targets may be a tracking target that satisfies a maximum condition. In this case, compared to a configuration in which zooming in is performed regardless of the size of the largest of the multiple tracking targets, it is possible to prevent the largest of the multiple tracking targets from becoming undetectable after zooming in.
[0078] Furthermore, when the position of at least some of the multiple tracking targets is not a predetermined position, the control unit 117 limits zooming in. An example of the predetermined position is a position that is not included in the restricted area R within the area captured in the captured image. In this case, compared to a configuration in which zooming in is performed regardless of the position of each tracking target object, it is possible to prevent a plurality of tracking targets from going undetected after zooming in.
[0079] Furthermore, when the size of at least some of the multiple tracking targets is not within a predetermined range, the control unit 117 restricts zooming by not allowing zooming, by reducing the zoom speed, or by shortening the zoom time. In this case, it is possible to prevent the cameras 100 from being unable to continue tracking and capturing the plurality of tracking targets.
[0080] Furthermore, the control unit 117 limits zooming when the size of at least some of the multiple tracking targets is not within a predetermined range. Even if the size of at least some of the multiple tracking targets is not within a predetermined range, the control unit 117 does not limit zooming if a predetermined condition is satisfied. Examples of the predetermined condition include relaxed conditions. In this case, if the predetermined condition is satisfied, the camera 100 can prioritize zooming over tracking the multiple tracking targets. Furthermore, the tracking targets include a specific tracking target. If the size of the specific tracking target is not within a predetermined range, the control unit 117 limits zooming even if predetermined conditions are met (see steps 1208, 1211, etc. in FIG. 12). In this case, if the predetermined conditions are met, the camera 100 can prioritize tracking of the specific tracking target over zooming.
[0081] Note that even if a negative result is obtained in step 306 of Fig. 6 (No in S306) or a negative result is obtained in step 606 of Fig. 8 (No in S606), if the alleviating condition is satisfied, the control unit 117 may not limit the zoom. Also, if a negative result is obtained in step 306 of Fig. 6 (No in S306) or a negative result is obtained in step 606 of Fig. 8 (No in S606), the result of detection by the detection means for a specific tracking target object may not satisfy the tracking condition. In this case, the control unit 117 may limit the zoom regardless of whether the alleviating condition is satisfied.
[0082] As described above, there are multiple tracking conditions, but the present invention is not limited to this. Any one or more of the multiple tracking conditions described above may be the tracking condition.
[0083] Furthermore, the zoom control or zoom limitation by the control unit 117 is not affected by non-tracked objects. For example, even if the detection result by the detection unit 113 for the non-tracked object does not satisfy the tracking condition, if the detection result by the detection means for each of the multiple tracking objects satisfies the tracking condition, the control unit 117 does not limit the zoom by the drive unit 109.
[0084] In addition, in the present disclosure, it has been described that the controller 200 sets which subject to be the tracking target, but this is not limiting. For example, the camera 100 may set which subject to be the tracking target. That is, the camera 100 may have the function of the target setting unit 213 of the controller 200. Furthermore, in the present disclosure, it has been described that camera 100 performs processes such as detecting an index related to a subject, extracting a tracking target, and determining whether to limit zooming by drive unit 109, but this is not limiting. For example, controller 200 may perform processes such as detecting an index related to a subject, extracting a tracking target, and determining whether to limit zooming by drive unit 109. In other words, controller 200 may have the functions of detection unit 113, extraction unit 115, calculation unit 116, and control unit 117 of camera 100.
[0085] The present invention also includes cases where a software program that realizes the functions of each of the above-mentioned embodiments is supplied to a system or device having a computer that can execute the program directly from a recording medium or via wired / wireless communication, and the program is executed. Therefore, the program code itself, supplied to and installed on a computer to implement the functional processes described above, also embodies the present invention. In other words, the computer program itself for implementing the functional processes of the present invention is also included in the present invention. In this case, the program may take any form, such as object code, a program executed by an interpreter, or script data supplied to an OS, as long as it has the program's functionality. Recording media for providing the program may include, for example, a hard disk, a magnetic recording medium such as a magnetic tape, an optical / magneto-optical storage medium, or a non-volatile semiconductor memory. Another possible method for providing the program is to store the computer program forming the present invention on a server on a computer network, and then download the computer program to a connected client computer. Furthermore, an OS running on a computer may perform some or all of the actual processing based on the instructions of the program code, and the functions of the above-described embodiments may be realized by this processing. Furthermore, program code read from a storage medium may be written to memory provided on a function expansion board inserted into a computer or a function expansion unit connected to the computer. Then, a CPU provided on the function expansion board or function expansion unit may perform some or all of the actual processing based on the instructions of the program code. Even in this case, the functions of the above-described embodiments are realized.
[0086] The disclosure of this embodiment includes the following configuration. (Configuration 1) a setting means for setting a plurality of tracking targets to be tracked in photographing by the imaging means; a size detection means for detecting a size of the tracking target object in an image captured by the imaging means; a control means for controlling the zoom of the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means; Equipped with The information processing device is characterized in that the control means controls the zoom so that the size of each of the plurality of tracking targets falls within a predetermined range for at least some of the tracking targets. (Configuration 2) The information processing device according to configuration 1, wherein the control means limits zoom-in of the zoom when the size of each of the plurality of tracking targets exceeds a threshold value as an upper limit value for at least some of the tracking targets. (Configuration 3) The information processing device according to configuration 1 or 2, wherein the control means limits zoom-out of the zoom when the size of each of the plurality of tracking targets falls below a threshold value as a lower limit value for at least some of the tracking targets. (Configuration 4) a subject detection means for detecting a specific subject from the photographed image, 4. The information processing device according to any one of configurations 1 to 3, wherein the setting means sets the tracking target from among the subjects detected by the subject detection means. (Configuration 5) 5. The information processing device according to configuration 4, wherein the range is determined in relation to the detection capability of the subject detection means. (Configuration 6) 5. The information processing device according to configuration 4, wherein the control means limits the zoom when each of the plurality of tracking targets is no longer detected by the subject detection means. (Configuration 7) The information processing device according to any one of configurations 1 to 6, wherein the control means limits zoom-out of the zoom when the size of the smallest of the plurality of tracking targets falls below a threshold value as a lower limit value. (Configuration 8) The information processing device according to any one of configurations 1 to 7, wherein the control means limits zoom-in of the zoom when the size of the largest of the plurality of tracking targets exceeds a threshold value as an upper limit value. (Configuration 9) further comprising a position detection means for detecting a position of the tracking target in the captured image; The information processing device according to any one of configurations 1 to 8, wherein the control means limits zoom-in of the zoom when the position of each of the plurality of tracking targets is not a predetermined position for at least some of the tracking targets. (Configuration 10) the control means limits the zoom when the size of each of at least some of the plurality of tracking targets is not within the range; The information processing device according to any one of configurations 1 to 9, wherein the restriction is that the control means does not allow the zoom, the control means reduces the speed of the zoom, or the control means shortens the time for the zoom. (Configuration 11) The control means limiting the zoom when the size of each of at least some of the plurality of tracking targets is not within the range; An information processing device described in any one of configurations 1 to 10, characterized in that even if the size of each of the plurality of tracking targets is not within the range for at least some of the tracking targets, the restriction is not imposed if a predetermined condition is met. (Configuration 12) The tracking target includes a specific tracking target, The information processing device according to configuration 11, characterized in that the control means imposes the restriction when the size of the specific tracking target object is not within the range, even if the predetermined condition is satisfied.
[0087] The present invention has been described in detail above based on preferred embodiments thereof, but the present invention is not limited to the above embodiments, and various modifications are possible based on the spirit of the present invention, and these modifications are not excluded from the scope of the present invention. [Explanation of symbols]
[0088] 1...imaging system, 100...camera, 200...controller, 300...network
Claims
1. a setting means for setting a plurality of tracking targets to be tracked in photographing by the imaging means; a size detection means for detecting a size of the tracking target object in an image captured by the imaging means; a control means for controlling the zoom of the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means; Equipped with The information processing device is characterized in that the control means controls the zoom so that the size of each of the plurality of tracking targets falls within a predetermined range for at least some of the tracking targets.
2. 2 . The information processing device according to claim 1 , wherein the control means limits zoom-in of the zoom when the size of at least some of the plurality of tracking targets exceeds a threshold value as an upper limit value.
3. 2 . The information processing device according to claim 1 , wherein the control means limits zoom-out of the zoom when the size of at least some of the plurality of tracking targets falls below a threshold value as a lower limit value.
4. a subject detection means for detecting a specific subject from the photographed image, 2. The information processing apparatus according to claim 1, wherein the setting means sets the tracking target from among the subjects detected by the subject detection means.
5. 5. The information processing apparatus according to claim 4, wherein the range is determined in relation to the detection capability of the subject detection means.
6. 5. The information processing apparatus according to claim 4, wherein the control means limits the zoom when each of the plurality of tracking targets is no longer detected by the subject detection means.
7. 2 . The information processing device according to claim 1 , wherein the control means limits the zoom-out of the zoom when the size of the smallest of the plurality of tracking targets falls below a threshold value as a lower limit value.
8. 2 . The information processing device according to claim 1 , wherein the control means limits zoom-in of the zoom when the size of the largest of the plurality of tracking targets exceeds a threshold value as an upper limit value.
9. further comprising a position detection means for detecting a position of the tracking target in the captured image; 2 . The information processing apparatus according to claim 1 , wherein the control means limits the zoom-in of the zoom when the positions of at least some of the plurality of tracking targets are not predetermined positions.
10. the control means limits the zoom when the size of each of at least some of the plurality of tracking targets is not within the range; 2. The information processing apparatus according to claim 1, wherein the restriction is one of the control means not allowing the zooming, the control means reducing the speed of the zooming, and the control means shortening the time period for the zooming.
11. The control means limiting the zoom when the size of each of at least some of the plurality of tracking targets is not within the range; 2. The information processing device according to claim 1, wherein even if the size of at least some of the plurality of tracking targets is not within the range, the restriction is not imposed if a predetermined condition is satisfied.
12. The tracking target includes a specific tracking target, The information processing device according to claim 11, characterized in that the control means imposes the restriction when the size of the specific tracking target object is not within the range, even if the predetermined condition is satisfied.
13. a setting step of setting a plurality of tracking targets to be tracked in image capture by the imaging means; a size detection step of detecting a size of the tracking target object in an image captured by the imaging means; a control step of controlling the zoom of the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means; and A method characterized in that, in the control step, the zoom is controlled so that the size of each of the plurality of tracking targets falls within a predetermined range for at least some of the tracking targets.
14. On the computer, a setting function for setting a plurality of tracking targets to be tracked in photographing by the imaging means; a size detection function for detecting the size of the tracking target object in the image captured by the imaging means; a control function for controlling the zoom of the imaging means so that the plurality of tracking targets are included in the angle of view of the imaging means; To achieve this, The control function controls the zoom so that the size of each of the plurality of tracking targets falls within a predetermined range for at least some of the tracking targets.
Citation Information
Patent Citations
Controller, imaging device, control method of controller, control program of controller, and storage medium
JP2020112648A