Image processing apparatus, method of controlling image processing apparatus, recording medium, and computer program product

By detecting the size of the tracked object and controlling the zoom of the camera unit, the viewing angle problem when multiple objects move along the camera's optical axis is solved, ensuring that all objects are within the viewing angle and achieving effective tracking of multiple objects.

CN121644971APending Publication Date: 2026-03-10CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously track multiple camera objects and maintain their appropriate size in an image, especially when the objects are moving along the camera's optical axis, making it difficult to ensure that all objects are within the field of view.

Method used

By setting components to detect the size of the tracked objects and using control components to control the zoom of the camera unit, multiple tracked objects are kept within the field of view of the camera unit, ensuring that at least some of the objects are within a predetermined size range.

Benefits of technology

It enables effective tracking of multiple objects, reduces undetected objects, and improves the completeness and accuracy of tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644971A_ABST
    Figure CN121644971A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing apparatus, a method of controlling the image processing apparatus, a recording medium, and a computer program product. The image processing apparatus includes: a setting means for setting a plurality of tracking objects to be tracked during imaging by an imaging means; a size detection means for detecting the size of the tracking object in the image captured by the image capturing means; and a control means for controlling zooming of the imaging means such that the plurality of tracking objects are included within a viewing angle of the imaging means. The control unit controls the zoom so that the size of each of at least a portion of the plurality of tracking objects falls within a predetermined range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image processing apparatus, methods, and recording media. Background Technology

[0002] In the prior art, there exists a technique for detecting a tracking object, identified as a camera subject to be tracked, from an image generated by a camera unit, tracking the detected tracking object, and capturing its image. Additionally, there is a technique where, when zooming the camera unit to capture an image of the tracking object, the zoom is controlled based on the size of the tracking object in the image, maintaining the size of the tracking object within a specified range in the image even if the tracking object moves along the camera's optical axis. Japanese Patent Application Laid-Open No. 2020-112648 discloses a method for continuing to track the camera while maintaining an appropriate size of the tracking object in the image by measuring the size of the tracking object in the image and changing the zoom ratio so that the measured size does not exceed a desired size. Summary of the Invention

[0003] This disclosure relates to controlling a camera unit such that, while the camera unit tracks each of a plurality of tracked objects detected, the plurality of tracked objects are included within the field of view of the camera unit.

[0004] According to one aspect of this disclosure, an image processing apparatus is provided, comprising: a setting component for setting a plurality of tracking objects to be tracked during imaging using a camera component; a size detection component for detecting the size of the tracking objects in an image captured by the camera component; and a control component for controlling the zoom of the camera component such that the plurality of tracking objects are included within the field of view of the camera component. The control component controls the zoom such that the size of each tracking object among at least a portion of the plurality of tracking objects falls within a predetermined range.

[0005] The features of this disclosure will become apparent from the following description of embodiments with reference to the accompanying drawings. The following description of the embodiments is given by way of example. Attached Figure Description

[0006] Figure 1 This is a diagram showing the overall configuration of the camera system.

[0007] Figure 2 This is a diagram showing the hardware configuration of the camera and the hardware configuration of the controller.

[0008] Figure 3A It is a diagram showing the functional configuration of a camera, and Figure 3B This is a diagram illustrating the functional structure of the controller.

[0009] Figures 4A to 4D It is a graph showing the relationship between captured images and metrics calculated by the camera's computing unit.

[0010] Figure 5 It is a sequence diagram showing the process from when the camera sends the captured image to the controller until the tracking object is set in the camera.

[0011] Figure 6 This is a flowchart illustrating the process of tracing.

[0012] Figure 7 This is a flowchart illustrating the process of determining the procedure.

[0013] Figure 8 This is a flowchart illustrating another process of the tracking procedure.

[0014] Figure 9 This is a flowchart illustrating another process for determining the procedure.

[0015] Figure 10 This is an illustration of the relationship between the area where the object is located in the captured image and the limitations of using camera magnification.

[0016] Figures 11A to 11C This diagram illustrates the relationship between the detection of the presence or absence of a tracked object in a captured image using a detection unit and the limitations imposed by the camera's zoom.

[0017] Figure 12 This is a flowchart illustrating another process of the tracking procedure. Detailed Implementation

[0018] <First Embodiment>

[0019] In the following description, embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0020] Figure 1This is an overall configuration diagram of camera system 1. Camera system 1 is a system that detects multiple tracking objects, tracks the detected multiple tracking objects, and captures images of them. Tracking objects are pre-determined subjects that are tracked using the camera unit. Tracking objects include humans, etc. In this embodiment, it is assumed that all multiple subjects are desired to be set as tracking objects. In this case, the camera unit needs to perform operations such as zooming to include multiple tracking objects within the camera unit's field of view. However, considering that only including multiple tracking objects within the camera unit's field of view presents the concern that it may be difficult to track each of the multiple tracking objects, for example, there may be undetected tracking objects due to their small size in the image. Therefore, in this embodiment, if the tracking conditions are not met, tracking of multiple tracking objects is achieved by restricting camera control.

[0021] The camera system 1 includes a camera 100 and a controller 200. The camera 100 and the controller 200 are connected via a network 300.

[0022] As an example of an imaging unit, camera 100 generates images by capturing images, detects multiple tracking objects from the generated images, and operates to track the detected multiple tracking objects and capture images of them. Operations performed by camera 100 include panning, tilting, and zooming. Furthermore, zooming, as an operation performed by camera 100, includes magnification and reduction. In the following text, the image generated by capturing images using camera 100 may be referred to as an captured image. Additionally, camera 100 may also be considered an image processing device.

[0023] As an example of an image processing device, the controller 200 controls the camera 100. In this embodiment, the controller 200 acquires captured images from the camera 100 and the detection results of subjects shown in the captured images using the camera 100, and displays the acquired information, thereby receiving the user's selection of which subjects should be set as tracking objects. In addition, the controller 200 sets the selected subjects as tracking objects and sends information indicating the set tracking objects to the camera 100.

[0024] The camera 100 and controller 200 can be comprised of a single computer, or they can be implemented using distributed processing across multiple computers.

[0025] For example, Network 300 is implemented using a Local Area Network (LAN) or Wide Area Network (WAN), such as the Internet. Furthermore, Network 300 can be implemented not only via the Internet, but also via any or a combination of telephone lines, dedicated digital lines, Asynchronous Transfer Mode (ATM), Frame Relay lines, cable TV lines, and data broadcast radio lines.

[0026] Figure 2 This is a diagram showing the hardware configuration of the camera 100 and the controller 200.

[0027] The camera 100 includes a CPU 101, RAM 102, ROM 103, GPU 104, network I / F 105, sensor I / F 106, image sensor 107, driver I / F 108, and driver unit 109. The CPU 101, RAM 102, ROM 103, GPU 104, network I / F 105, sensor I / F 106, image sensor 107, driver I / F 108, and driver unit 109 are interconnected via a bus 110.

[0028] The CPU 101 performs various processes to control the camera 100 as a whole by using computer programs and data stored in RAM 102. RAM 102 is a high-speed storage device such as DRAM, and stores computer programs loaded from ROM 103, captured images, and various information such as information obtained from the controller 200. Furthermore, RAM 102 has a working domain used by the CPU 101 or GPU 104 when performing various processes. ROM 103 is a non-volatile storage device such as flash memory, HDD, SSD, or SD card, and stores camera 100 setting data, as well as computer programs and data related to the startup and basic operation of the camera 100. Additionally, ROM 103 also stores computer programs and data for causing the CPU 101 or GPU 104 to perform or control various processes that will be described as being performed by the camera 100. GPU 104 performs inference processing for estimating the presence or absence of a subject, or the area of ​​a subject, from captured images. For example, GPU 104 is a computing device, such as a graphics processing unit (GPU) dedicated to image processing or inference processing. Instead of GPU 104, a computing device such as a field-programmable gate array (FPGA) can be used. Alternatively, CPU 101 can handle the processing of GPU 104. Network I / F 105 is an interface for connecting to network 300 and communicates with external devices such as controller 200 via a communication medium such as Ethernet (registered trademark). Sensor I / F 106 converts the video signal output from image sensor 107 into a captured image as data in a prescribed format, and outputs the converted captured image to RAM 102 after compression when needed. Sensor I / F 106 can perform various processing on the video image represented by the video signal acquired from image sensor 107, such as image quality adjustment, exposure correction, or sharpness correction, or cropping processing such as cropping only a prescribed area. Furthermore, sensor I / F 106 can perform processing according to instructions received from controller 200 via network I / F 105. Image sensor 107 receives light reflected from a subject, converts the brightness and color of the received light into electrical charges, and outputs a video signal based on the conversion result. Examples of image sensors 107 include photodiodes, charge-coupled device (CCD) sensors, and complementary metal-oxide-semiconductor (CMOS) sensors. Drive I / F 108 is an interface for sending and receiving signals such as control signals to drive unit 109. Drive unit 109 is a drive mechanism for changing the imaging direction of camera 100 and includes a mechanical drive system and a motor (drive source).The drive unit 109 performs panning and tilting to change the camera direction horizontally and vertically, and zooming to change the camera's viewing angle optically, according to instructions received from the CPU 101 via the drive I / F 108.

[0029] Additionally, the controller 200 includes a CPU 201, RAM 202, ROM 203, GPU 204, network I / F 205, display unit 206, and operation unit 207. The CPU 201, RAM 202, ROM 203, GPU 204, network I / F 205, display unit 206, and operation unit 207 are interconnected via a bus 208.

[0030] The CPU 201 performs various processes to control the controller 200 as a whole by using computer programs and data stored in RAM 202. RAM 202 is a high-speed storage device such as DRAM. RAM 202 stores computer programs and data loaded from ROM 203, as well as various data acquired from the camera 100. Furthermore, RAM 202 has a working domain used by the CPU 201 or GPU 204 when performing various processes. ROM 203 is a non-volatile storage device such as flash memory, HDD, SSD, or SD card, and stores setting data for the controller 200, as well as computer programs and data related to the startup and basic operation of the controller 200. Additionally, ROM 203 stores computer programs and data used to enable the CPU 201 or GPU 204 to control various processes. GPU 204 performs inference processing to estimate the presence or absence of a subject, or the area of ​​the subject, from the captured image. For example, GPU 204 is a computing device, such as a GPU dedicated to image processing or inference processing. Instead of GPU 204, a computing device such as an FPGA can be used. Alternatively, CPU 201 can handle the processing of GPU 204. Network I / F 205 is an interface for connecting to network 300 and communicating with external devices such as camera 100 via a communication medium such as Ethernet. Display unit 206 has a screen such as an LCD screen or a touch screen, and displays captured images from camera 100 and settings screens of controller 200. The configuration of display unit 206 with a touch screen will be described below. In camera system 1, instead of setting display unit 206 in controller 200, a display device (not shown) for displaying information can be connected to controller 200, and the display device can display captured images or settings screens of controller 200. Operation unit 207 is a user interface that receives user operations on controller 200, such as buttons, dials, joysticks, or touch panels.

[0031] The controller 200 may be a personal computer (PC) with a mouse and keyboard as operating units 207.

[0032] Figure 3A This diagram illustrates the functional configuration of a camera 100. The camera 100 includes an acquisition unit 111, a storage unit 112, a detection unit 113, an output unit 114, an extraction unit 115, a calculation unit 116, and a control unit 117.

[0033] The acquisition unit 111 acquires information from the controller 200. The information acquired by the acquisition unit 111 includes information such as information indicating the subject that is set as the tracking object.

[0034] Storage unit 112 stores the information acquired by acquisition unit 111 and information such as captured images generated by camera 100.

[0035] The detection unit 113 detects subjects, such as people, from the captured image and detects regions of the subjects within the captured image. The regions of the subjects in the captured image include the size, length, and position of the subjects. The detection unit 113 estimates the subjects from the input captured image by performing inference processing using a learned model created using machine learning techniques such as deep learning, and outputs information indicating the coordinates corresponding to the regions of the subjects as a result. The coordinates corresponding to the regions of the subjects include coordinates corresponding to parts or all of the subjects in the captured image. Additionally, the coordinates corresponding to parts of the subjects include coordinates corresponding to the outline of the subjects and coordinates corresponding to the head or face of the subjects (people). Furthermore, when the detection unit 113 detects the subject as a rectangle, the coordinates corresponding to each part of the subject include the coordinates of the top-left and bottom-right vertices of the rectangle, the coordinates of the center of the rectangle, the coordinates corresponding to the width direction of the rectangle, and the coordinates corresponding to the height direction of the rectangle. Here, the detection unit 113 can also be viewed as a size detection unit configured to detect the size of the tracked object in the image captured by the camera 100. Alternatively, the detection unit 113 can also be viewed as a subject detection unit configured to detect a specific subject from the captured image. Specific subjects include subjects of a predetermined type, such as people, which are candidates for tracking. Furthermore, the detection unit 113 can also be viewed as a position detection unit configured to detect the location of the tracked object in the captured image.

[0036] The technique for detecting a subject using detection unit 113 is not limited to machine learning-based techniques. Detection unit 113 can use a template matching method, in which a template image showing the subject as the detection object is compared with the captured image, and regions in the captured image that are highly similar to the subject shown in the template image are detected as regions showing the subject. Alternatively, any technique can be used to detect a subject using detection unit 113. Furthermore, the subject being detected using detection unit 113 can be an object other than a person.

[0037] Furthermore, the detection unit 113 identifies each detected subject by detecting its features and generates information for each identified subject. The detection unit 113 performs inference processing using a learned model created using machine learning techniques such as deep learning, and outputs information representing the feature quantities of the subject as a result, based on the input captured image and the region of the subject in the captured image. The feature quantities extracted by the detection unit 113 can be the feature quantities of all subjects, or they can be the feature quantities of individual parts of the subject (such as a person's head or face). Additionally, the feature quantities of the subject include feature vectors in the image. In this case, captured images generated from various angles for each subject can be used as images for learning, and for images showing the same subject, a machine learning model can be used, in which these images are labeled with the same ID and input into the learning model, and the machine learning model outputs feature vectors. Any learning model can be used as the learning model for identifying subjects using the detection unit 113.

[0038] The technique for identifying a subject using the detection unit 113 is not limited to machine learning-based techniques. The detection unit 113 can use a Kalman filter or similar method to predict the region of a subject in the latest captured image based on changes in the region of a subject in a series of past captured images, and identify the subject closest to the predicted region as the same subject. Alternatively, any technique can be employed to identify a subject using the detection unit 113.

[0039] After generating information indicating the coordinates corresponding to the area of ​​the subject and information indicating the feature quantities of the subject during subject detection, the detection unit 113 causes the storage unit 112 to store the generated information. The information generated by the detection unit 113, including the information indicating the coordinates corresponding to the area of ​​the subject and the information indicating the feature quantities of the subject, can also be considered as the detection result of the detection unit 113 on the subject shown in the captured image.

[0040] The output unit 114 outputs information to the controller 200. The information output to the output unit 114 includes information such as indicating the captured image and the detection results of the subject shown in the captured image by the detection unit 113.

[0041] Extraction unit 115 extracts the tracking objects from the captured image. Extraction unit 115 extracts the tracking objects from information sent by controller 200 as information indicating subjects set as tracking objects, and from the detection results of detection unit 113 on the captured image. Furthermore, if the captured image shows multiple tracking objects, extraction unit 115 extracts all tracking objects shown in the captured image. After extracting the tracking objects, extraction unit 115 causes storage unit 112 to store information indicating which subjects detected by detection unit 113 are tracking objects.

[0042] Furthermore, the extraction unit 115 determines the tracking objects that satisfy predetermined conditions related to size or length from the extracted tracking objects. The extraction unit 115 determines the tracking objects that satisfy the predetermined conditions based on the detection results of the detection unit 113 (such as the area of ​​the subject as the extracted tracking object in the captured image).

[0043] The calculation unit 116 treats the multiple tracking objects extracted by the extraction unit 115 as a tracking object group, and calculates the reference position of the tracking object group and the size of the tracking object group in the captured image. The calculation technique using the calculation unit 116 will be described in detail below.

[0044] As an example of a control unit, control unit 117 utilizes the drive unit 109 of camera 100 (see reference). Figure 2 The control unit 117 controls operations such as panning, tilting, and zooming. If the drive unit 109 zooms, the control unit 117 in this embodiment controls the zoom using the drive unit 109 so that the camera 100 can continue tracking each of the multiple tracking objects extracted by the extraction unit 115. In other words, the control unit 117 zooms the drive unit 109, and if the camera 100 can no longer continue tracking any of the multiple tracking objects extracted by the extraction unit 115, the zoom is limited (e.g., the range of the angle of view taken by the zoom control is limited).

[0045] Each step of the processing performed by the acquisition unit 111, detection unit 113, output unit 114, extraction unit 115, calculation unit 116, and control unit 117 of the camera 100 is implemented by loading the program stored in ROM 103 into RAM 102 and executing the program via CPU 101 or GPU 104. Furthermore, the storage unit 112 of the camera 100 is implemented by RAM 102 or ROM 103.

[0046] Figure 3B This is a diagram illustrating the functional configuration of the controller 200. The controller 200 includes an acquisition unit 211, a storage unit 212, an object setting unit 213, and an output unit 214.

[0047] The acquisition unit 211 acquires information from the camera 100. The information acquired by the acquisition unit 211 includes information such as instructions for capturing images and detection results of the captured images by the detection unit 113 of the camera 100.

[0048] Storage unit 212 stores the information acquired by acquisition unit 211 and the information generated by controller 200.

[0049] The object setting unit 213 sets the tracking object. The object setting unit 213 enables the display unit 206 of the controller 200 (see reference) to... Figure 2 The device displays captured images sent from the camera 100, as well as information indicating the subjects detected by the detection unit 113 in the captured images. Furthermore, the object setting unit 213 receives the user's selection of tracking objects based on the information displayed on the display unit 206, and sets the user-selected subjects as tracking objects. Additionally, the object setting unit 213 generates information indicating which subjects have been set as tracking objects, and stores the generated information in the storage unit 212.

[0050] Output unit 214 outputs information to camera 100 or controller 200. The information output from output unit 214 to camera 100 includes information generated by object setting unit 213 as information indicating which subjects have been set as tracking objects. Additionally, the information output from output unit 214 to controller 200 includes displaying an image to display unit 206.

[0051] Each step of the processing performed by the acquisition unit 211, object setting unit 213, and output unit 214 of the controller 200 is implemented by loading the program stored in ROM 203 into RAM 202 and executing the program via CPU 201 or GPU 204. Additionally, the storage unit 212 of the controller 200 is implemented using RAM 202 or ROM 203.

[0052] Figures 4A to 4D It is a graph showing the relationship between captured images and indicators calculated by the computing unit 116 of the camera 100.

[0053] Figure 4AImage 401 is shown. Furthermore, image 401 shows a plurality of tracking objects 402 consisting of tracking objects 402a and 402b, and a plurality of non-tracking objects 403, none of which are tracking objects. Here, the plurality of tracking objects 402 consisting of tracking objects 402a and 402b constitutes a group of tracking objects.

[0054] In addition, Figure 4A In the captured image 401 shown, rectangular images 413 are shown in the regions superimposed on the tracking object 402a and the tracking object 402b, respectively. The rectangular image 413 is an image with a rectangular shape used to indicate the region of the tracking object 402 detected by the detection unit 113.

[0055] Computing unit 116 from Figure 4A The captured image 401 shown calculates the reference position of the tracking object group. More specifically, the calculation unit 116 calculates a rectangular region 404 as a rectangular shape region including the rectangular image 413 of the corresponding tracking object 402, and calculates the center region of the calculated rectangular region 404 as the reference position 405 of the tracking object group.

[0056] In addition, such as Figure 4B As shown, the calculation unit 116 calculates the rectangular area 408, which is the area of ​​the calculated rectangular region 404, as the size of the tracking object group.

[0057] The technique of using the computing unit 116 to calculate the reference position and size of the group of tracked objects is not limited to the examples described above.

[0058] exist Figure 4C The captured image 401 shows multiple tracking objects 402 (in other words, a group of tracking objects) consisting of tracking object 402a, tracking object 402b, and tracking object 402c. Additionally, Figure 4C A rectangular image 413 is shown for the corresponding tracking object 402. In this case, the calculation unit 116 can calculate a center region 411, which is the center region of the area surrounded by the rectangular image 413, and calculate the coordinates as the average of the coordinates corresponding to each calculated center region 411 as the reference position 407 of the tracking object group. In this case, the coordinates as the average of the coordinates corresponding to each center region 411 can be calculated based on the weighted average of the coordinates corresponding to each center region 411. In this way, the reference position of the tracking object group becomes closer to the area where more tracking objects are clustered in the tracking object group. For this reason, when the driving unit 109 of the camera 100 operates using the reference position of the tracking object group as the target position, it can operate on the area where more tracking objects are clustered in the tracking object group.

[0059] like Figure 4C As shown in the example, if the number of tracking objects 402 constituting the tracking object group is three or more, the coordinates corresponding to the region of the centroid of the polygon 406 with each central region 411 as a vertex are consistent with the coordinates that are the average of the coordinates corresponding to each central region 411. For this reason, the coordinates corresponding to the region of the centroid of the polygon 406 can be calculated as the reference position 407 of the tracking object group.

[0060] Additionally, if each central region 411 has been calculated, then as follows Figure 4D As shown, the calculation unit 116 calculates a rectangular region 412 comprising each central region 411. Furthermore, the calculation unit 116 can calculate the size of the tracking object group as the larger of the following two: the ratio of the width of the rectangular region 412 to the width of the captured image 401, and the ratio of the height of the rectangular region 412 to the height of the captured image 401. Here, the width of the captured image 401 or the rectangular region 412 represents the length in the horizontal direction of the image. Additionally, the height of the captured image 401 or the rectangular region 412 represents the length in the vertical direction of the image. In the example shown, the ratio of the width of the rectangular region 412 to the width of the captured image 401 is greater than the ratio of the height of the rectangular region 412 to the height of the captured image 401. For this reason, the ratio of the width of the rectangular region 412 to the width of the captured image 401 becomes the size of the tracking object group. By determining the size of the tracking object group in this way, the operation of the drive unit 109 of the camera 100 can be controlled based on the size of the tracking object group in the horizontal and vertical directions of the captured image 401, where the tracking object group is more widely distributed.

[0061] The calculation unit 116 can determine the size of the tracking object group by the ratio of the diagonal 409 in the rectangular region 412 to the diagonal of the outer frame in the captured image 401. Additionally, the calculation unit 116 can determine the size of the tracking object group by the length of the width, the height, or the length of the diagonal 409 in the rectangular region 412.

[0062] In addition, the calculation unit 116 enables the storage unit 112 to store information used to indicate the calculation results.

[0063] Figure 5 This is a sequence diagram showing the process from when the camera 100 sends the captured image to the controller 200 until the tracking object is set in the camera 100.

[0064] First, the acquisition unit 111 of the camera 100 acquires the captured image generated by the camera (step (hereinafter referred to as "S") 101). The captured image acquired by the acquisition unit 111 is stored in the storage unit 112.

[0065] The detection unit 113 of the camera 100 detects the presence of a subject in the captured image acquired by the acquisition unit 111, and detects the region of the subject in the captured image (step S102). Furthermore, the detection unit 113 identifies the subject by detecting its features and generates information for identifying the subject. The detection unit 113 causes the storage unit 112 to store information indicating the detection result, including the information for identifying the subject.

[0066] The output unit 114 of the camera 100 sends the captured image acquired by the acquisition unit 111 and information indicating the detection result of the detection unit 113 to the controller 200 (step S103).

[0067] The output unit 214 of the controller 200 causes the display unit 206 to display the captured image sent from the camera 100 and information for identifying the subject. More specifically, the output unit 214 causes the display unit 206 to display the captured image superimposed with the information for identifying the subject (step S104).

[0068] The object setting unit 213 receives the user's selection of which subjects in the captured image displayed by the display unit 206 will be the tracking objects (step S105).

[0069] The object setting unit 213 sends information for instructing the user to track the object selected by the user to the camera 100 via the output unit 214 (step S106).

[0070] The extraction unit 115 of the camera 100 sets the tracking objects according to the information sent from the controller 200 in step S106 (step S107). For this reason, the extraction unit 115 can also be regarded as a setting unit, which is configured to set multiple tracking objects to be tracked during the recording of the camera 100. In addition, the extraction unit 115 causes the storage unit 112 to store information indicating the set tracking objects.

[0071] Whenever camera 100 captures a new image and generates the captured image, it can perform... Figure 5 The processing is shown. Furthermore, the settings for the tracked object in the camera 100 and controller 200 can be updated whenever the user selects a tracked object.

[0072] Furthermore, while the foregoing example describes a subject selected by the user being set as a tracking object, examples of setting tracking objects are not limited to the foregoing example. For example, the object setting unit 213 can set all predetermined types of subjects detected from the captured image as tracking objects. Predetermined types include humans, etc. Additionally, after the user selects a tracking object, the object setting unit 213 can set a subject located near a tracking object selected by the user in the captured image as a tracking object. Furthermore, the object setting unit 213 can set a subject having characteristics similar to a tracking object selected by the user in the captured image as a tracking object. In this way, compared to the case where the user needs to select all tracking objects that are set as tracking objects, the burden of setting tracking objects is reduced.

[0073] Figure 6 This is a flowchart illustrating the tracking process. The tracking process is the process by which the camera 100 tracks a group of tracking objects, which has multiple tracking objects. In this embodiment, when the camera 100 is set to a mode for tracking the group of tracking objects, the tracking process begins if an image is generated by capturing images using the camera 100.

[0074] The acquisition unit 111 of the camera 100 acquires the captured image generated by the camera (step S301).

[0075] The detection unit 113 of the camera 100 detects subjects of the same type as the tracking object set in the extraction unit 115 from the captured images acquired by the acquisition unit 111 (step S302). If the tracking object set in the extraction unit 115 is a human, the detection unit 113 detects a human as a candidate subject for use as a tracking object from the captured images. In this case, the detection unit 113 identifies the subject by detecting the region of the subject and detecting the features of the subject, thereby generating information for identifying the subject.

[0076] The processing in steps S301 and S302 can be related to Figure 5 The processes in steps S101 and S102 are the same. Additionally, if the processes in steps S301 and S302, as well as the processes in steps S101 and S102, are the same processes for the same captured image, then the processes in steps S301 and S302 can be omitted.

[0077] Extraction unit 115 determines whether multiple tracking objects have been extracted from the captured image (step S303). Extraction unit 115 extracts tracking objects from the captured image and performs the determination in step S303 based on whether the number of extracted tracking objects is more than one.

[0078] If the extraction unit 115 does not extract the tracking object from the captured image, or if the number of tracking objects extracted by the extraction unit 115 is one ("No" in step S303), the tracking process ends. Here, if the extraction unit 115 does not extract the tracking object from the captured image, the control unit 117 does not control the drive unit 109. Alternatively, if the number of tracking objects extracted by the extraction unit 115 is one, the control unit 117 causes the drive unit 109 to perform pan, tilt, and zoom operations to track the extracted single tracking object.

[0079] In addition, if multiple tracking objects are extracted using the extraction unit 115 ("Yes" in step S303), the calculation unit 116 calculates the reference position and size of the tracking object group (step S304).

[0080] Extraction unit 115 performs a determination process (step S305). Although details will be described below, in this determination process, extraction unit 115 determines, from the plurality of tracking objects extracted by extraction unit 115, tracking objects that meet predetermined conditions related to size or length as representative tracking objects among the plurality of tracking objects. The representative tracking object determined in the determination process may be referred to as a representative object below. In addition, in the determination process, extraction unit 115 determines an indicator of the size or length of the representative object.

[0081] Control unit 117 determines whether the index of the size or length of the representative object determined in the determination process is equal to or greater than a threshold value that has been predetermined as a lower limit (step S306). For example, control unit 117 may determine whether the rectangular image 413 generated for the representative object (refer to) Figure 4A The control unit 117 can determine whether the length of the long side of the rectangular image 413 generated for the representative object is greater than or equal to a threshold value used as a lower limit. Additionally, for example, the control unit 117 can determine whether the area enclosed by the rectangular image 413 generated for the representative object is equal to or greater than a threshold value used as a lower limit. Furthermore, for example, the control unit 117 can determine whether the length of the diagonal connecting two vertices in the rectangular image 413 generated for the representative object is equal to or greater than a threshold value used as a lower limit. In other words, the index for the representative object used as the basis for the determination in step S306 only needs to be an index that allows comparison of the size or length of the representative object.

[0082] Additionally, for example, the threshold value as a lower limit is the minimum size or shortest length of the subject that the detection unit 113 can detect from the captured image. Alternatively, for example, the threshold value as a lower limit can be a value obtained by adding a predetermined size or length to the minimum size or shortest length of the subject that the detection unit 113 can detect from the captured image. The predetermined size or length can be any value, but it can be a value determined to be detectable by the detection unit 113. Alternatively, for example, the threshold value as a lower limit can be the minimum size or shortest length of a feature of the subject that the detection unit 113 can detect from the captured image. Alternatively, for example, the threshold value as a lower limit can be a value obtained by adding a predetermined size or length to the minimum size or shortest length of a feature of the subject that the detection unit 113 can detect from the captured image. Alternatively, the threshold value as a lower limit can be a value set by the user's operation on the camera 100 or controller 200. Alternatively, the threshold value as a lower limit can be a value set to ensure the image quality of the tracked object.

[0083] If the index for the representative object is equal to or greater than the lower limit ("Yes" in step S306), the process proceeds to the next step. Control unit 117 calculates the difference between the size calculated by calculation unit 116 for the tracked object group and the value preset by the user as the target size of the tracked object group in the captured image, and calculates the amount of zoom control based on the calculated difference (step S307). In this case, control unit 117 determines whether to zoom in or out of drive unit 109 to reduce the calculated difference. Additionally, control unit 117 determines the zoom speed, increasing it as the calculated difference becomes larger. In this case, zoom is controlled so that the size calculated for the tracked object group approaches the target value.

[0084] Even if the calculated difference is the same, the control unit 117 can vary the zoom speed when the drive unit 109 is magnified and reduced. Specifically, the control unit 117 can set a higher zoom speed when the drive unit 109 is reduced compared to when it is magnified. In this case, the tracking responsiveness of the camera 100 for the dispersion of the tracked object group is improved.

[0085] Furthermore, if the camera 100 tracks and images a group of objects, the movement speed of the tracked objects may be higher than when the camera 100 tracks and images a single object, potentially requiring high-speed zoom. Therefore, the control unit 117 can vary the zoom speed both when the camera 100 tracks and images a group of objects and when it tracks and images a single object. Specifically, if the camera 100 tracks and images a group of objects, the control unit 117 can set a higher zoom speed compared to the zoom speed when the camera 100 tracks and images a single object.

[0086] Furthermore, if the index for the representative object is less than the lower limit ("No" in step S306), the control unit 117 restricts the zoom control of the drive unit 109 (step S308). As a restriction on zoom, for example, the control unit 117 may prohibit the drive unit 109 from zooming. Additionally, as a restriction on zoom, compared to not restricting zoom, the control unit 117 may set a lower zoom speed or a shorter zoom time. Specifically, the zoom restricted in step S308 is reduction. That is, in step S308, the control unit 117 does not necessarily restrict the magnification performed by the drive unit 109. However, in step S308, the control unit 117 may restrict both magnification and reduction using the drive unit 109.

[0087] The control unit 117 compares the reference position calculated by the calculation unit 116 for the tracked object group with the position set by the user as the target position of the tracked object group in the captured image, and calculates the control quantities for pan and pitch based on the comparison result (step S309). In this case, the control unit 117 determines the direction and speed of pan and pitch so that the reference position calculated for the tracked object group approaches the position set as the target. In addition, the control unit 117 determines the speed of pan and pitch so that the speed of pan and pitch increases as the deviation between the reference position calculated for the tracked object group and the position set as the target increases. In this case, pan and pitch are controlled so that the reference position calculated for the tracked object group approaches the target position.

[0088] The control unit 117 operates the drive unit 109 based on step S308 or step S309, and the details determined in step S309 (step S310). More specifically, the control unit 117 provides the drive unit 109 with instructions for operations instructing the details determined in step S308 or step S309. Therefore, the drive unit 109 performs panning, tilting, and zooming operations according to the instructions from the control unit 117.

[0089] If the termination conditions are met after the processing in step S310, the tracking process can be terminated. Termination conditions include receiving an instruction to terminate the tracking process from the camera 100 or controller 200, the date and time reaching a predetermined date and time, and a predetermined time having elapsed since the tracking process began. Alternatively, if the termination conditions are not met, the processing from step S301 can be repeated using the newly generated captured image.

[0090] Figure 7 This shows the process of determining (reference) Figure 6 The flowchart of step S305 in the process.

[0091] Extraction unit 115 identifies the tracking object that meets the minimum condition from the multiple tracking objects extracted in step S303 of the tracking process as the representative object (step S501). The minimum condition is a predetermined condition related to the size or length of the tracking object. In this embodiment, the tracking object with the smallest size and the tracking object with the shortest length in the captured image are identified as the minimum conditions. The index used by extraction unit 115 to identify the size or length of the tracking object that meets the minimum condition is the same index used in step S306 of the tracking process.

[0092] Extraction unit 115 determines whether the index of the size or length of the representative object has changed from the index determined in the previous determination process (step S502).

[0093] If the indicator representing the size or length of the object does not change from the indicator determined in the previous determination process ("No" in step S502), the determination process ends. In this case, the indicator determined in the previous determination process is taken over as the indicator used for the judgment in step S306 of the tracking process.

[0094] Furthermore, the indicator used to represent the size or length of the object can be changed from the indicator determined in the previous determination process ("Yes" in step S502). In this case, the extraction unit 115 will newly determine the indicator for the size or length of the object as the indicator used for the judgment in step S306 of the tracking process (step S503).

[0095] If the current determination process is the first determination process, the extraction unit 115 determines that the index for the size or length of the representative object has changed from the index determined in the previous determination process.

[0096] In addition, information about the index determined by the index used to indicate the judgment in step S306 of the tracking process in the determination process is stored in the storage unit 112 of the camera 100.

[0097] In this way, in this embodiment, if the size of the smallest tracked object in the captured image or the length of the shortest tracked object in the captured image is equal to or greater than a threshold value serving as a lower limit, both zooming in and zooming out of the camera 100 are performed without any restrictions. Furthermore, if the size of the smallest tracked object in the captured image or the length of the shortest tracked object in the captured image is less than the threshold value serving as a lower limit, zooming out of the camera 100 is restricted.

[0098] In this case, compared to a situation where the camera 100 can zoom out regardless of the size or length of the tracked object, the occurrence of tracked objects that become so small in the captured image that the camera 100 cannot detect them is suppressed. For this reason, while tracking each of the multiple detected tracked objects, the camera 100 can zoom out so that they are included in the view of the group of tracked objects.

[0099] Furthermore, this embodiment has described an example of camera control based on the size or length of all tracked objects, but it is not limited to this, and camera control can also be based on the size of a subset of the tracked objects. For example, camera control can be based solely on one or more representative subjects similar to the tracked objects (e.g., tracked objects with higher priority, such as subjects that have been personally authenticated). That is, if the size or length of these tracked objects with relatively higher priority is less than a threshold, i.e., less than a reference value, the reduction of camera 100 is limited. In this case, other tracked objects with relatively lower priority can have a size smaller than the reference value.

[0100] Furthermore, the methods for determining multiple subjects to be tracked are not limited to those described above, and other methods can be used. For example, the primary subject determination process can be based on the detection of specific subjects such as people (faces, heads) or animals, the position or size of subjects detected using personal authentication, or user selection. The selected primary subject can be set as the tracking object, and subjects that exist in its vicinity within a specified distance and are detected by the same specific subject detection process can be grouped together and set as tracking objects.

[0101] (Modified Example)

[0102] Next, variations of the tracking process will be described. The tracking process in this embodiment is not limited to... Figure 6 The tracking process is shown.

[0103] Figure 8 This is a flowchart illustrating another process of tracking processing as a variation. Figure 8 The processing of steps S601 to S604 in the tracking process shown is related to... Figure 6 The steps S301 to S304 in the tracking process shown are the same as those steps.

[0104] The extraction unit 115 performs a determination process (step S605). Although details will be described below, the technique used in the determination process performed in step S605 to determine the representative object using the extraction unit 115 differs from that in... Figure 6 The technique used in the determination process is shown in step S305 of the tracking process.

[0105] Control unit 117 determines whether the index of the size or length of the representative object determined in the determination process is equal to or less than a threshold that has been predetermined as an upper limit (step S606). For example, control unit 117 may determine whether the rectangular image 413 generated for the representative object (refer to) Figure 4A The control unit 117 can determine whether the length of the long side of the rectangular image 413 generated for the representative object is equal to or less than a threshold value, which is an upper limit value. Additionally, for example, the control unit 117 can determine whether the area enclosed by the rectangular image 413 generated for the representative object is equal to or less than a threshold value, which is an upper limit value. Furthermore, for example, the control unit 117 can determine whether the length of the diagonal connecting two vertices in the rectangular image 413 generated for the representative object is equal to or less than a threshold value, which is an upper limit value. In other words, the index for the representative object used as the basis for the determination in step S606 only needs to be an index that allows comparison of the size or length of the representative object.

[0106] Additionally, for example, the threshold value as an upper limit is the maximum size or longest length of the subject that the detection unit 113 can detect from the captured image. Alternatively, for example, the threshold value as an upper limit can be a value obtained by subtracting a predetermined size or length from the maximum size or longest length of the subject that the detection unit 113 can detect from the captured image. The predetermined size or length can be any value, but can be a value determined to allow detection by the detection unit 113. Alternatively, for example, the threshold value as an upper limit can be the maximum size or longest length of a feature of the subject that the detection unit 113 can detect from the captured image. Alternatively, for example, the threshold value as an upper limit can be a value obtained by subtracting a predetermined size or length from the maximum size or longest length of a feature of the subject that the detection unit 113 can detect from the captured image. Alternatively, the threshold value as an upper limit can be a value set by the user's operation on the camera 100 or controller 200. Alternatively, the threshold value as an upper limit can be a value set to ensure the image quality of the tracked object.

[0107] If the index for the representative object is equal to or less than the upper limit value, the process proceeds to step S607. Furthermore, the processing in step S607 is related to... Figure 6 The process in step S307 is the same as the process in the previous step.

[0108] Furthermore, if the index for the representative object exceeds the upper limit, the control unit 117 restricts the zoom control of the drive unit 109 (step S608). As a limitation on zoom, for example, the control unit 117 may prohibit the drive unit 109 from zooming. Additionally, as a limitation on zoom, compared to not limiting zoom, the control unit 117 may set a lower zoom speed or a shorter zoom time. Specifically, the zoom restricted in step S608 is magnification. That is, in step S608, the control unit 117 does not necessarily need to restrict the reduction performed by the drive unit 109. However, in step S608, the control unit 117 may restrict both magnification and reduction using the drive unit 109.

[0109] Furthermore, the processing in steps S609 and S610 is related to... Figure 6 The processing of steps S309 and S310 is the same.

[0110] Figure 9 This shows the determination process as a variation example (see reference). Figure 8 The flowchart of another process (step S605) in the process.

[0111] Extraction unit 115 identifies the tracking object that meets the maximum condition from among the multiple tracking objects extracted in step S303 of the tracking process as the representative object (step S701). The maximum condition is a predetermined condition related to the size or length of the tracking object. Furthermore, in this embodiment, the tracking object with the largest size and the tracking object with the longest length in the captured image are identified as the maximum condition. The index used by extraction unit 115 to identify the tracking object that meets the maximum condition, which is the same index used in step S606 of the tracking process, is the same index used for judgment.

[0112] Extraction unit 115 determines whether the index for the size or length of the representative object has changed from the index determined in the previous determination process (step S702).

[0113] If the indicator for the size or length of the representative object does not change from the indicator determined in the previous determination process ("No" in step S702), the determination process ends. In this case, the indicator determined in the previous determination process is taken over as the indicator used for the determination in step S706 of the tracking process.

[0114] Furthermore, the indicator for the size or length of the representative object can be changed from the indicator determined in the previous determination process ("Yes" in step S702). In this case, the extraction unit 115 will newly determine the indicator for the size or length of the representative object as the indicator used for the judgment in step S606 of the tracking process (step S703).

[0115] If the current determination process is the first determination process, the extraction unit 115 determines that the index for the size or length of the representative object has changed from the index determined in the previous determination process.

[0116] In addition, information about the index determined by the index used to indicate the judgment in step S606 of the tracking process in the determination process is stored in the storage unit 112 of the camera 100.

[0117] In this way, in this embodiment, if the size of the largest tracked object in the captured image or the length of the longest tracked object in the captured image is equal to or less than a threshold value, the zoom control of the camera 100 is freely adjusted for both magnification and reduction. Furthermore, if the size of the largest tracked object in the captured image or the length of the longest tracked object in the captured image is greater than the threshold value, the magnification of the camera 100 is restricted.

[0118] In this case, compared to a situation where the magnification of the camera 100 is not limited regardless of the size or length of the tracked object, the occurrence of tracked objects that grow to a size that the camera 100 cannot detect in the captured image is suppressed. For this reason, while tracking each of the multiple detected tracked objects, the camera 100 can zoom out so that they are included in the field of view of the group of tracked objects.

[0119] <Second Embodiment>

[0120] Next, the camera system 1 of the second embodiment will be described. For the camera system 1 of the second embodiment, a configuration different from that of the camera system 1 of the first embodiment will be described, and descriptions of configurations identical to those of the camera system 1 of the first embodiment will be omitted.

[0121] If the tracking conditions are met, the control unit 117 of the camera 100 in this embodiment controls the zoom using the drive unit 109; if the tracking conditions are not met, the control unit 117 restricts the zoom using the drive unit 109. The tracking conditions are conditions used by the control unit 117 to determine whether the camera 100 can track each tracking object even when the drive unit 109 has zoomed. The tracking conditions will be described in detail below.

[0122] Figure 10 This is an illustrative diagram showing the relationship between the area where the tracked object 402 is located in the captured image 401 and the limitations of magnification using the camera 100.

[0123] exist Figure 10 The captured image 401 shows multiple tracking objects 402 consisting of tracking objects 402a and 402b, and multiple non-tracking objects 403.

[0124] In this embodiment, in the captured image 401, there exists a region identified as a restricted area R. The restricted area R is the region magnified and limited by the camera 100 when the tracked object 402 is located. The restricted area R can be any region. However, in the illustrated example, it is the region indicated by a diagonal line as the circumferential edge of the captured image 401.

[0125] When the tracked object 402 is located within the restricted area R, if the camera 100 zooms in, the tracked object 402 may be cut off from the captured image 401, or the tracked object 402 may not be shown in the captured image 401. Furthermore, when the tracked object 402 is cut off from the captured image 401, or when the tracked object 402 is not shown in the captured image 401, the camera 100 finds it difficult to track the tracked object 402 because the detection unit 113 no longer detects it. Therefore, in this embodiment, the area that the camera 100 may no longer detect when the camera 100 zooms in is determined to be the restricted area R. Moreover, one of the determined tracking conditions is that the area detected by the detection unit 113 as the location of each of the plurality of tracked objects is not included in the restricted area R. In the illustrated example, since a portion of the rectangular image 413 corresponding to the tracked object 402a is included in the restricted area R, the tracking condition is not met.

[0126] Tracking conditions may include excluding any portion of the rectangular image 413 of each tracked object 402 from the restricted region R. Additionally, tracking conditions may include excluding portions of the rectangular image 413 of each tracked object 402 that occupy a predetermined proportion or a larger proportion from the restricted region R. The predetermined proportion can be any proportion, but for example, half.

[0127] Furthermore, the tracking conditions determined by the relationship between the region detected by the position of each of the multiple tracking objects (detection unit 113) and the restricted region R are the tracking conditions applied when the control unit 117 amplifies the drive unit 109. That is, if the control unit 117 shrinks the drive unit 109, it is not necessary to apply the tracking conditions determined by the relationship between the region detected by the position of each of the multiple tracking objects (detection unit 113) and the restricted region R. In other words, the control unit 117 can shrink the drive unit 109 regardless of whether the tracking conditions determined by the relationship between the region detected by the position of the tracking object (detection unit 113) and the restricted region R are met.

[0128] Figures 11A to 11C This is an explanatory diagram illustrating the relationship between the presence or absence of the tracked object 402 detected by the detection unit 113 in the captured image 401 and the limitation on the zoom of the camera 100.

[0129] exist Figure 11A The captured image 401 shown depicts multiple tracking objects 402, consisting of tracking object 402a and tracking object 402b. In this case, it is assumed that the detection unit 113 detects... Figure 11AThe number of tracked objects 402 detected in the captured image 401 shown is "two", including tracked object 402a and tracked object 402b.

[0130] Here, it is assumed that generation Figure 11B The captured image 401 shown is as Figure 11A The next frame of the captured image 401 shown. In Figure 11B In the captured image 401 shown, the tracked object 402b is shown, while the tracked object 402a is not shown (reference). Figure 11A Furthermore, in this embodiment, one of the determined tracking conditions is that the number of people (i.e., the multiple tracking objects extracted by the extraction unit 115) has not decreased.

[0131] exist Figure 11B In the example shown, the number of tracking objects 402 extracted by extraction unit 115 from captured image 401 is "one" (which is the only tracking object 402b). In this case, the tracking condition is not met because the number of people (i.e., the tracking objects extracted by extraction unit 115) has been reduced.

[0132] Additionally, assuming generation Figure 11C The captured image 401 shown is as Figure 11A The next frame of the captured image 401 shown. In Figure 11C In the captured image 401 shown, the tracked object 402b is shown, while the tracked object 402a is not shown (reference). Figure 11A Additionally, in Figure 11C In the captured image 401 shown, a new tracking object 402c is shown. Furthermore, in this embodiment, one of the determined tracking conditions is that the tracking object 402 detected by the extraction unit 115 in the first captured image 401 is also extracted by the extraction unit 115 in the second captured image 401, which was captured at a later time than the first captured image 401.

[0133] exist Figure 11C In the example shown, since the number of tracking objects 402 extracted by extraction unit 115 from captured image 401 is "two," including tracking objects 402b and 402c, the number of people (that is, the number of tracking objects extracted by extraction unit 115) is not reduced. On the other hand, since in Figure 11A The tracked object 402a extracted in the captured image 401 shown is not in Figure 11C The image 401 shown was extracted from it, therefore the tracking conditions are not met.

[0134] Additionally, although not shown, one of the determined tracking conditions is that the size or length of each of the plurality of tracking objects 402 extracted by the extraction unit 115 is equal to or greater than a threshold predetermined as a lower limit value. This threshold, serving as the lower limit value, is used for... Figure 6 The threshold value determined in step S306 of the tracking process shown. Additionally, one of the determined tracking conditions is that an index related to the size or length of each of the plurality of tracking objects 402 extracted by the extraction unit 115 is equal to or less than a threshold value predetermined as an upper limit. This threshold value as an upper limit is used for... Figure 8 The threshold for judgment in step S606 of the tracking process shown.

[0135] In this way, in this embodiment, multiple conditions are determined as tracking conditions. Furthermore, if all of the multiple conditions determined as tracking conditions are met, the control unit 117 controls the zoom operation using the drive unit 109. Conversely, if at least one of the multiple conditions determined as tracking conditions is not met, the control unit 117 restricts a portion of the zoom operation using the drive unit 109.

[0136] Figure 12 This is a flowchart illustrating another process of the tracking process in the second embodiment. Figure 12 The processing steps S1201 to S1203 in the tracking process shown are related to Figure 6 The steps S301 to S303 in the tracking process shown are the same as those steps.

[0137] If no multiple tracking objects are extracted from the captured image ("No" in step S1203), the extraction unit 115 determines whether multiple tracking objects have been extracted from the captured image in the previous tracking process (step S1204). The captured image used to determine whether multiple tracking objects have been extracted in the previous tracking process is a captured image generated by the camera before the captured image used to determine whether multiple tracking objects have been extracted in the current tracking process.

[0138] Even if multiple tracking objects are not extracted from the captured image in the previous tracking process ("No" in step S1204), the tracking process ends. If no tracking object is extracted in the current tracking process, control of the drive unit 109 by the control unit 117 is not performed. Alternatively, if only a single tracking object is extracted in the current tracking process, the control unit 117 causes the drive unit 109 to perform pan, tilt, and zoom operations to track the extracted single tracking object.

[0139] Additionally, if multiple tracking objects are extracted from the captured image in the current tracking process ("Yes" in step S1203), or if multiple tracking objects were extracted from the captured image in a previous tracking process ("Yes" in step S1204), then the process proceeds to step S1205. The processing in step S1205 is related to... Figure 6 The process of step S304 in the tracking process shown is the same as the process in the example.

[0140] Control unit 117 determines whether the tracking conditions are met (step S1206). More specifically, control unit 117 determines whether all of the above conditions have been met as tracking conditions.

[0141] If the tracking condition is met ("Yes" in step S1206), the control unit 117 calculates the amount of zoom control based on the difference between the size calculated by the calculation unit 116 for the group of tracked objects and the target value preset by the user as the size of the group of tracked objects in the captured image. Furthermore, at this time, the control unit 117 calculates the amount of zoom control such that the tracking condition is met even after zoom control by the drive unit 109 (S1207). More specifically, if the positions of each tracked object constituting the group of tracked objects do not change, the control unit 117 calculates the amount of zoom control such that the tracking condition is met even after zooming by the drive unit 109 based on the control amount calculated in step S1207. Furthermore, the case where the positions of each tracked object constituting the group of tracked objects do not change means that the positions of each tracked object constituting the group of tracked objects do not change before and after zooming by the drive unit 109 based on the control amount calculated in step S1207.

[0142] Furthermore, if the tracking conditions are not met ("No" in step S1206), the control unit 117 determines whether the tracking object that does not meet the tracking conditions is a specific tracking object (step S1208). A specific tracking object is... Figure 5 The predetermined tracking object is set in step S107. The specific tracking object can be any tracking object, but for example, it is a tracking object that is determined to be of high importance for tracking. In addition, the specific tracking object can be set by the user's operation on the camera 100 or the controller 200, or it can be set by the detection unit 113 of the camera 100 or the object setting unit 213 of the controller 200.

[0143] If the tracked object that does not meet the tracking conditions is not a specific tracked object ("No" in step S1208), the control unit 117 determines whether the mitigation conditions are met (step S1209). The mitigation conditions are conditions used by the control unit 117 to determine whether the limitation on zooming using the drive unit 109 is mitigated. The mitigation conditions include: a predetermined time has elapsed since zooming using the drive unit 109 began in the state where the tracking conditions are not met. Additionally, the mitigation conditions include: no user operation has been performed on the camera 100 or controller 200 to instruct on the limitation on zooming using the drive unit 109. Furthermore, the mitigation conditions include: the distance from the area detected by the detection unit 113 as the location of the tracked object that does not meet the tracking conditions to the reference position of the tracked object group is equal to or longer than a predetermined distance. Additionally, the mitigation conditions include: the detection unit 113 has detected that the tracked object that does not meet the tracking conditions is not moving. The detection unit 113 can identify the presence or absence of movement of the tracked object by comparing the captured image acquired in the current tracking process with the captured image acquired in a previous tracking process.

[0144] If the mitigation condition is met ("Yes" in step S1209), the process proceeds to step S1210. Furthermore, the processing in step S1210 is related to... Figure 6 The process of step S307 in the tracking process shown is the same as the process described above.

[0145] Additionally, if the tracked object that does not meet the tracking conditions is a specific tracked object ("Yes" in step S1208), or if the mitigation conditions are not met ("No" in step S1209), then the process proceeds to step S1211. Furthermore, the processing in step S1211 is related to... Figure 6 The process of step S308 in the tracking process shown is the same as the process in the example.

[0146] Furthermore, after step S1207, step S1210, or step S1211, the process proceeds to step S1212. Additionally, the processing in steps S1212 and S1213 is related to... Figure 6 The processes in steps S309 and S310 of the tracking process shown are the same.

[0147] As described above, the control unit 117 controls the zoom of the camera unit so that multiple tracked objects are included within the field of view of the camera unit. Furthermore, the control unit 117 controls the zoom so that the size of each tracked object, at least a portion of the multiple tracked objects, falls within a predetermined range. The size of the tracked object includes the size detected by the detection unit 113 as the size or length of the tracked object. Additionally, the predetermined range includes a range equal to or greater than a threshold (lower limit) and equal to or less than a threshold (upper limit).

[0148] In this configuration, compared to zooming without controlling the size of each tracked object, the occurrence of multiple detected tracked objects becoming undetectable after zooming is suppressed. For this reason, the camera unit can be controlled such that while the camera unit tracks each of the multiple detected tracked objects, these multiple tracked objects are included within the field of view of the camera unit.

[0149] Furthermore, the processing of multiple tracking objects to be tracked during the recording process using the camera 100 can also be considered as setting. Additionally, the processing of the camera system 1 detecting the size of the tracking objects in the image captured by the camera 100 can also be considered as size detection. Furthermore, the processing of the camera system 1 controlling the zoom of the camera 100 to include multiple tracking objects within the field of view of the camera unit can also be considered as control. In this control, the zoom is controlled such that the size of each tracking object among at least a portion of the multiple tracking objects falls within a predetermined range.

[0150] Furthermore, the function of the camera system 1 in setting multiple tracking objects to be tracked during the recording process of the camera 100 can also be considered a setting function. Additionally, the function of the camera system 1 in detecting the size of the tracking objects in the image captured by the camera 100 can also be considered a size detection function. Furthermore, the function of the camera system 1 in controlling the zoom of the camera unit to include multiple tracking objects within the field of view of the camera unit can also be considered a control function. This control function controls the zoom so that the size of each tracking object among at least a portion of the multiple tracking objects falls within a predetermined range.

[0151] Images used to detect multiple tracked objects are not limited to captured images.

[0152] For example, if the detection unit 113 has detected a subject in the captured image, an image based on the captured image can be generated by overlaying information related to the detected subject onto the captured image. In this case, the extraction unit 115 can detect multiple tracking objects from the image generated by the detection unit 113. In this way, the image generated by processing the captured image is also included in the image generated based on the image captured by the camera unit.

[0153] Furthermore, if the size of each tracked object in at least a portion of the tracked objects exceeds a threshold value, the control unit 117 limits the zoom magnification. In this case, compared to a configuration that performs zoom magnification regardless of the size of each tracked object, the possibility of multiple tracked objects no longer being detected after magnification is suppressed.

[0154] Furthermore, if the size of each tracked object in at least a portion of the tracked objects falls below a threshold, which is a lower limit, the control unit 117 restricts the zoom reduction. In this case, compared to a configuration that zooms in regardless of the size of each tracked object, the possibility that multiple tracked objects will no longer be detected after zooming in is suppressed.

[0155] Additionally, the detection unit 113 detects a specific subject from the captured image. Furthermore, the extraction unit 115 sets a tracking object from the subject detected by the detection unit 113. In this case, the tracking object can be tracked from the time it is set.

[0156] Furthermore, as described above, the predetermined range is determined based on the size or length of the subject that the detection unit 113 can detect, or by adding or subtracting a predetermined size relative to the size of the subject that the detection unit 113 can detect. In other words, the predetermined range has been determined in relation to the detection capability of the detection unit 113. In this case, if a group of tracked objects is being tracked, scaling of the size of each of the multiple tracked objects that can no longer be detected is suppressed.

[0157] In addition, if the predetermined range is set by the user through operation of the camera 100 or the controller 200, the camera unit can track and capture images of the group of tracked objects while maintaining the size of the tracked object as desired by the user.

[0158] Additionally, if the number of people (i.e., the tracked objects extracted by extraction unit 115) has decreased, or if the tracked objects already extracted by extraction unit 115 are no longer being extracted, then control unit 117 limits zoom. In other words, if detection unit 113 no longer detects any of the tracked objects among the multiple tracked objects, then control unit 117 limits zoom.

[0159] In this case, compared to the configuration where zooming is performed regardless of whether the detection unit 113 no longer detects each tracked object, the situation where multiple tracked objects are no longer detected after zooming is suppressed.

[0160] Additionally, if the size of the smallest tracking object among the multiple tracking objects falls below a threshold that serves as a lower limit, the control unit 117 restricts the zoom reduction. The smallest tracking object among the multiple tracking objects includes the tracking object that meets the minimum condition.

[0161] In this case, compared to a configuration that zooms in regardless of the size of the smallest tracking object among multiple tracking objects, the situation where the smallest tracking object among multiple tracking objects is no longer detected after zooming out is suppressed.

[0162] Additionally, if the size of the largest tracking object among multiple tracking objects exceeds a threshold value, the control unit 117 limits the zoom magnification. The largest tracking object among multiple tracking objects includes the tracking object that meets the maximum condition.

[0163] In this case, compared to a configuration that zooms in regardless of the size of the largest tracking object among multiple tracking objects, the situation where the largest tracking object among multiple tracking objects is no longer detected after zooming is suppressed.

[0164] Additionally, if the positions of at least a portion of the tracked objects are not at predetermined positions, the control unit 117 limits the zoom magnification. Predetermined positions include locations within the area shown in the captured image that are not included in the restricted area R.

[0165] In this case, compared to a configuration that zooms in regardless of the position of each tracked object, it suppresses the situation where multiple tracked objects are no longer detected after zooming.

[0166] Additionally, if the size of each tracked object in at least a portion of the tracked objects does not fall within a predetermined range, the control unit 117 limits zoom. This limitation includes: the control unit 117 prohibiting zoom, the control unit 117 setting a lower zoom speed, or the control unit 117 setting a shorter zoom time.

[0167] In this situation, the recording of each of the multiple tracking objects using camera 100 can no longer continue.

[0168] Furthermore, if the size of each tracked object among at least a portion of the tracked objects does not fall within a predetermined range, the control unit 117 limits zoom. Additionally, even if the size of each tracked object among at least a portion of the tracked objects does not fall within the predetermined range, the control unit 117 does not limit zoom if predetermined conditions are met. These predetermined conditions include mitigation conditions. In this case, if the predetermined conditions are met, the camera 100 can prioritize zooming over tracking of the multiple tracked objects.

[0169] Additionally, the tracked object includes a specific tracked object. Furthermore, if the size of the specific tracked object does not fall within a predetermined range, the control unit 117 limits zooming even if the predetermined conditions are met (see reference). Figure 12 (Steps S1208 and S1211, etc.). In this case, if predetermined conditions are met, the camera 100 can prioritize tracking a specific object over zooming.

[0170] Even in Figure 6 A negative result is obtained in step S306 ("No" in step S306), or even if Figure 8 If a negative result is obtained in step S606 ("No" in step S606), and the mitigation conditions are met, then the control unit 117 does not necessarily need to limit zoom. Furthermore, if in Figure 6 In step S306, a negative result is obtained ("No" in step S306), or if in Figure 8 If a negative result is obtained in step S606 ("No" in step S606), then the detection result of the detection unit for the specific tracked object may not meet the tracking conditions. In this case, the control unit 117 can limit zoom regardless of whether the mitigation conditions are met.

[0171] Furthermore, as described above, multiple conditions have been presented as tracking conditions, but the situation is not limited to this. Regarding tracking conditions, any one or more of the aforementioned conditions only needs to be a tracking condition.

[0172] Furthermore, zoom control or zoom limitation using control unit 117 does not affect non-tracked objects. For example, even if the detection result of detection unit 113 for a non-tracked object does not meet the tracking conditions, if the detection result of detection unit 113 for each of the multiple tracked objects meets the tracking conditions, control unit 117 does not limit zoom using drive unit 109.

[0173] Furthermore, this disclosure describes how the controller 200 sets which subjects will be tracked, but it is not limited thereto. For example, the camera 100 can set which subjects will be tracked. That is, the camera 100 can have the functionality of the object setting unit 213 of the controller 200.

[0174] Furthermore, this disclosure describes the processing performed by the camera 100, such as detecting indicators related to the subject, extracting the tracked object, and whether to restrict zooming using the drive unit 109, but it is not limited thereto. For example, the controller 200 can perform processing such as detecting indicators related to the subject, extracting the tracked object, and whether to restrict zooming using the drive unit 109. That is, the controller 200 can have the functions of the camera 100's detection unit 113, extraction unit 115, calculation unit 116, and control unit 117.

[0175] In addition, this disclosure also includes situations where a software program for implementing the functions of each of the foregoing embodiments is supplied directly from a recording medium or via wired / wireless communication to a system or apparatus having a computer capable of executing the program, and the program is executed.

[0176] Therefore, the program code supplied to and installed in a computer to implement the aforementioned functional processing of this disclosure also implements this disclosure. That is, this disclosure also includes the computer program itself for implementing the functional processing of this disclosure. In this case, as long as it is used as a program, the program can take any form, such as object code, a program executed by an interpreter, or script data supplied to an OS. For example, the recording medium used to supply the program can be a hard disk, a magnetic recording medium such as magnetic tape, an optical / magneto-optical storage medium, or a non-volatile semiconductor memory. In addition, regarding the method of supplying the program, the following method can also be considered: storing the computer program forming this disclosure in a server on a computer network, and connecting client computers downloading and programming the computer program.

[0177] The embodiments of this disclosure can also be implemented by reading and executing computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more functions in the above-described embodiments and / or by a computer of a system or device including one or more circuits (e.g., application-specific integrated circuits (ASICs)) for performing one or more functions in the above-described embodiments and by the following method, wherein the computer of the system or device performs the above-described method by, for example, reading and executing computer-executable instructions from the storage medium to perform one or more functions in the above-described embodiments and / or controlling the one or more circuits to perform one or more functions in the above-described embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessor unit (MPU)) and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. For example, these computer-executable instructions can be provided to a computer from a network or storage medium. The storage medium may include one or more of the following: a hard disk, random access memory (RAM), read-only memory (ROM), the storage unit of a distributed computing system, optical discs (such as CDs, DVDs, or Blu-ray Discs™), flash memory devices, and memory cards.

[0178] Furthermore, the operating system (OS) and other components running in the computer can perform some or all of the actual processing based on the instructions in the program code, and the various functions of the above embodiments can be implemented through this processing. Additionally, the program code read from the storage medium can be written into a memory installed in a function expansion board inserted into the computer or a function expansion unit connected to the computer. Furthermore, based on the instructions in the program code, a CPU or other component installed in the function expansion board or function expansion unit can perform some or all of the actual processing. Even in this case, the various functions of the above embodiments are implemented.

[0179] Other embodiments

[0180] Embodiments of the present invention can also be implemented by providing software (including computer program products of computer programs) that performs the functions of the above embodiments to a system or device via a network or various storage media, and the computer (central processing unit (CPU) or microprocessor unit (MPU) of the system or device) reads and executes the computer program.

[0181] While this disclosure has been described with reference to embodiments, it should be understood that this disclosure is not limited to the disclosed embodiments. The scope of the appended claims should be given the broadest interpretation to cover all such modifications and equivalent structures and functions.

[0182] According to this disclosure, a camera unit can be controlled such that while the camera unit is tracking each of a plurality of detected tracking objects, the plurality of tracking objects are included within the field of view of the camera unit.

[0183] This application claims the benefit of Japanese Patent Application 2024-148135, filed on August 30, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. An image processing apparatus comprising: setting means for setting a plurality of tracking objects to be tracked during imaging by an imaging means; size detecting means for detecting sizes of the tracking objects in an image captured by the imaging means; and control means for controlling zooming of the imaging means so that the plurality of tracking objects are included within a field of view of the imaging means, wherein the control means controls the zooming so that each of at least some of the plurality of tracking objects falls within a predetermined range in size.

2. The image processing apparatus according to claim 1, wherein, in a case where the size of each of at least some of the plurality of tracking objects exceeds a threshold value as an upper limit value, the control means limits magnification of the zooming. wherein 3. The image processing apparatus according to claim 1, wherein, in a case where the size of each of at least some of the plurality of tracking objects falls below a threshold value as a lower limit value, the control means limits reduction of the zooming. wherein 4. The image processing apparatus according to claim 3, wherein the threshold value as the lower limit value is a value set by an operation of a user. wherein 5. The image processing apparatus according to claim 1, further comprising: subject detecting means for detecting a specific subject from the captured image, wherein the control means sets the tracking objects from the subject detected by the subject detecting means.

6. The image processing apparatus according to claim 5, wherein, in a case where each of at least some of the plurality of tracking objects is no longer detected by the subject detecting means, the control means limits the zooming. wherein 7. The image processing apparatus according to claim 1, wherein, in a case where the size of a tracking object having a smallest size among the plurality of tracking objects falls below a threshold value as a lower limit value, the control means limits reduction of the zooming. wherein 8. The image processing apparatus according to claim 1, wherein, in a case where the size of a tracking object having a largest size among the plurality of tracking objects exceeds a threshold value as an upper limit value, the control means limits magnification of the zooming. wherein 9. The image processing apparatus according to claim 1, further comprising: position detecting means for detecting positions of the tracking objects in the captured image, wherein, in a case where the position of each of at least some of the plurality of tracking objects is not a predetermined position, the control means limits magnification of the zooming.

10. The image processing apparatus according to claim 1, wherein, in a case where the size of each of at least some of the plurality of tracking objects does not fall within the range, the control means limits the zooming, and wherein, wherein the limiting includes prohibiting the zooming by the control means, reducing a zooming speed by the control means, or shortening a time of the zooming by the control means.

11. The image processing apparatus according to claim 1, ​ wherein the control means limits the zooming in a case where the size of each of at least some of the plurality of tracking objects does not fall within the range, and wherein the control means does not impose the limitation even in a case where the size of each of at least some of the plurality of tracking objects does not fall within the range, if a predetermined condition is satisfied.

12. The image processing apparatus according to claim 11, wherein, the tracking objects include a specific tracking object, and wherein the control means imposes the limitation in a case where the size of the specific tracking object does not fall within the range, even if the predetermined condition is satisfied.

13. The image processing apparatus according to any one of claims 1 to 12, wherein the zooming is controlled so that a zooming speed in a case where the imaging means is caused to zoom out is higher than a zooming speed in a case where the imaging means is caused to zoom in.

14. An image processing apparatus comprising: setting means for setting a plurality of tracking objects to be tracked during imaging by an imaging means; and control means for controlling zooming of the imaging means so that the plurality of tracking objects are included within a field of view of the imaging means, wherein the control means controls the zooming so that a zooming speed in a case where the imaging means is caused to zoom out is higher than a zooming speed in a case where the imaging means is caused to zoom in.

15. An image processing apparatus comprising: setting means for setting a plurality of tracking objects to be tracked during imaging by an imaging means; size detecting means for detecting sizes of the tracking objects in an image captured by the imaging means; and control means for controlling zooming of the imaging means so that the plurality of tracking objects are included within a field of view of the imaging means, wherein the control means limits zooming out of the zooming in a case where the size of each of at least some of the plurality of tracking objects falls below a threshold value as a lower limit value, and wherein the threshold value as the lower limit value is a value set by an operation of a user.

16. A method for controlling an image processing apparatus, comprising: a setting step of setting a plurality of tracking objects to be tracked during imaging by an imaging means; a size detecting step of detecting sizes of the tracking objects in an image captured by the imaging means; and a control step of controlling zooming of the imaging means so that the plurality of tracking objects are included within a field of view of the imaging means, wherein in the control step, the zooming is controlled so that the size of each of at least some of the plurality of tracking objects falls within a predetermined range.

17. A method for controlling an image processing apparatus, comprising: a setting step of setting a plurality of tracking objects to be tracked during imaging by an imaging means; and a control step of controlling zooming of the imaging means so that the plurality of tracking objects are included within a field of view of the imaging means, ​ ​ ​ In the control step, the zoom is controlled so that the zoom speed is higher in the case where the imaging section is caused to zoom out than in the case where the imaging section is caused to zoom in.

18. A non-transitory recording medium storing a control program of an image processing apparatus, the control program causing a computer to perform the steps of the method according to claim 16 or 17.

19. A computer program product comprising a control program of an image processing apparatus, the control program causing a computer to perform the steps of the method according to claim 16 or 17.

Citation Information

Patent Citations

  • Controller, imaging device, control method of controller, control program of controller, and storage medium

    JP2020112648A

  • Medicine identification support system, medicine identification support method, and program

    JP2024148135A