Camera controller and camera control method

The camera control device addresses subject transfer issues in PTZ cameras by implementing automatic tracking with warning mechanisms, ensuring stable subject tracking through user notification and mode switching.

JP2025164053APending Publication Date: 2025-10-30CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024067780
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-18
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Automatic tracking technologies in PTZ cameras face issues with subject transfer, where the tracked subject switches with another subject, causing disruptions in continuous tracking.

Method used

A camera control device and method that includes video input, automatic tracking, and a warning display mechanism to notify users of potential subject transfers by identifying similar subjects and prompting manual intervention or mode switch.

Benefits of technology

Prevents photography accidents by alerting users to impending subject transfers, allowing them to take corrective actions and maintain stable tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025164053000001_ABST
    Figure 2025164053000001_ABST
Patent Text Reader

Abstract

To provide a camera controller that notifies a user, in an easy-to-understand manner, of a situation where transfer of a tracking target may occur, and urges appropriate measures.SOLUTION: A camera controller has: video input means that receives input of a video obtained by picking up images of a plurality of persons; automatic tracking means that, on the basis of the video input to the video input means, automatically tracks a predetermined person in the video; and warning display instruction means that, when the predetermined person and a second person similar to the predetermined person are present in the video, issues an instruction to display a warning to the display.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a camera control device and a camera control method. [Background technology]

[0002] Generally, cameras known as PTZ cameras, which are capable of adjusting pan, tilt, and zoom, have a technology that automatically tracks a target object detected from a captured image in response to a user request. Such automatic tracking technology automatically controls pan, tilt, and zoom so that the target object is positioned at a desired position within the camera's angle of view. Patent Document 1 discloses a method for calculating the control parameters of a PTZ camera required to display a target object at the center of the screen based on the coordinates of the target object displayed on the screen.

[0003] [Patent Document 1] Patent Publication No. 9-181961 Summary of the Invention [Problem to be solved by the invention]

[0004] Automatic tracking technology has a problem called "subject transfer," in which the tracked subject crosses over with another subject, causing the tracked subject to switch. When a subject transfer occurs, automatic tracking cannot continue normally. Therefore, the present invention aims to notify the user of situations in which a subject transfer may occur. [Means for solving the problem]

[0005] The camera control device of the present invention comprises a video input means for inputting video of multiple people, an automatic tracking means for automatically tracking a specified person in the video based on the video input to the video input means, and a warning display instruction means for instructing a display device to display a warning when the specified person and a second person similar to the specified person are present in the video. [Effects of the Invention]

[0006] According to the present invention, when there is a possibility of a transfer occurring, it is possible to notify the user of the situation. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating a system configuration according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing a block configuration of each device of the first embodiment. [Figure 3] FIG. 2 is a diagram showing an operation flow of the camera 100 of the first embodiment. [Figure 4] FIG. 2 is a diagram showing an operation flow of the camera 100 of the first embodiment. [Figure 5] FIG. 2 is a diagram showing an operation flow of the camera 100 of the first embodiment. [Figure 6] FIG. 4 is a diagram showing an operation flow of the controller 200 of the first embodiment. [Figure 7] 10A and 10B are diagrams showing images displayed on the controller 200 of the first embodiment. [Figure 8] 10A and 10B are diagrams showing a captured image and a display image of the controller 200 in the first embodiment. [Figure 9] 10A and 10B are diagrams showing a captured image and a display image of the controller 200 in the first embodiment. [Figure 10] FIG. 10 is a diagram showing an operation flow of the camera 100 of the second embodiment. [Figure 11] FIG. 10 is a diagram showing an operation flow of the controller 200 of the second embodiment. [Figure 12] FIG. 10 is a diagram showing an operation flow of the controller 200 of the second embodiment. [Figure 13] FIG. 10 is a diagram showing a display image of the controller 200 of the second embodiment. [Figure 14] FIG. 10 is a diagram showing a display image of the controller 200 of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] First Embodiment A camera control device including a camera and a controller according to a first embodiment of the present invention will be described below.

[0009] 1 is a diagram showing the system configuration for implementing the present invention. A camera 100 and a controller 200 are connected via a LAN (Local Area Network) 300, and the devices are configured to be able to communicate with each other.

[0010] The basic operation of each device is as follows: First, camera 100 has an imaging unit and transmits captured images and detection results (described later) to each device via each network or video cable. Furthermore, camera 100 has a driving unit 109 (described later) and has a mechanism that allows pan / tilt operations to change the imaging direction by rotating the device.

[0011] Furthermore, camera 100 detects a subject from the captured image and changes the image capturing direction based on the detection result. Controller 200 acquires images and detection results from camera 100 via LAN 300, and operates camera 100 based on user operations.

[0012] Using this camera control device, the user can select a subject to be tracked via the controller 200, and the camera 100 can automatically track and photograph the selected subject.

[0013] In the first embodiment, a mode in which the camera 100 automatically tracks and captures images is called an automatic tracking mode. On the other hand, a mode in which the user controls the pan, tilt, and zoom of the camera 100 using the controller 200 as a general shooting operation with a PTZ camera is called a normal shooting mode. <Explanation of each device>

[0014] Next, the block configuration of each device will be explained using Figure 2. The left side of Figure 2 shows the configuration of camera 100, which is a video input means. CPU 101, RAM 102, ROM 103, network I / F 105, image processing unit 106, drive I / F 108, and inference unit 110 are interconnected via internal bus 111.

[0015] The CPU 101 controls the entire camera 100. The RAM 102 is a high-speed storage device such as DRAM, into which the OS, various programs, and various data are loaded, and is also used as a work area for the OS and various programs. The ROM 103 is a non-volatile storage device such as flash memory, HDD, SSD, or SD card, and is used as a permanent storage area for the OS, various programs, and various data, as well as a short-term storage area for various data.

[0016] Details of the various programs stored in RAM 102 and ROM 103 of camera 100 will be described later. An image sensor 107 made of a CCD or CMOS is connected to image processing unit 106, and image data acquired from image sensor 107 is converted into a predetermined format, compressed as necessary, and transferred to RAM 102.

[0017] Furthermore, when acquiring an image from the image sensor 107, the image processing unit 106 may perform image quality adjustments such as color correction, exposure control, and sharpness correction on the image, and may also perform cropping to cut out only a predetermined area of ​​the image data. These operations are performed in accordance with instructions from an external communication device such as the controller 200 via the network I / F 105.

[0018] The network I / F 105 is an I / F for connecting to the above-mentioned LAN 300, and is responsible for communication with external devices such as the controller 200 via a communication medium such as Ethernet (registered trademark). Note that communication may also be performed via another I / F such as a serial communication I / F (not shown).

[0019] The inference unit 110 is an inference unit that reads the captured video from the RAM 102 and estimates the position and presence or absence of a predetermined object from the captured video, and is configured with a computing device specialized for image processing and inference processing, such as a so-called GPU (Graphics Processing Unit). While a GPU is generally effective for use in learning processing, equivalent functions may also be realized by a reconfigurable logic circuit such as an FPGA (Field-Programmable Gate Array). Furthermore, the processing of the inference unit 110 may be performed by the CPU 101.

[0020] The drive I / F 108 is a connection unit with the drive unit 109, which will be described later, and is responsible for communication for sending and receiving control signals and the like to the drive unit 109. The drive unit 109 is a rotation mechanism for changing the imaging direction of the camera 100, and is composed of a mechanical drive system, a drive source motor, and the like.

[0021] Based on instructions received from CPU 101 via drive I / F 108, drive unit 109 performs rotations such as pan and tilt operations to adjust the shooting angle of view in the horizontal and vertical directions, and zoom operations to optically change the shooting angle of view.

[0022] Next, a description will be given of the controller 200 on the right side of Fig. 2. A CPU 201, a RAM 202, a ROM 203, a network I / F 204, a display unit 205, and a user input I / F 206 are interconnected via an internal bus 207.

[0023] The CPU 201 controls the entire controller 200. The RAM 202 is a high-speed storage device such as a DRAM, into which the OS, various programs, and various data are loaded, and is also used as a work area for the OS and various programs. The ROM 203 is a non-volatile storage device such as a flash memory, HDD, SSD, or SD card, and is used as a permanent storage area for the OS, various programs, and various data, as well as a short-term storage area for various data.

[0024] The various programs stored in the RAM 202 and ROM 203 of the controller 200 will be described in detail later.

[0025] The network I / F 204 is an I / F for connecting to the above-mentioned LAN 300, and is responsible for communication with an external communication device such as the camera 100 via a communication medium such as Ethernet (registered trademark). For example, this communication may involve sending and receiving control commands to the camera 100 and receiving camera images from the camera 100.

[0026] Display unit 205 is a display unit for displaying images captured by camera 100, detection results, and a setting screen for controller 200. Display unit 205 may be provided integrally with controller 200 or camera 100. Display unit 205 may also be referred to as a display device.

[0027] The user input I / F 206 is an interface for receiving user operations on the controller 200, and examples thereof include buttons, dials, joysticks, touch panels, etc. This concludes the description of each component that constitutes each device. <Explanation of the operation of each device>

[0028] Next, the operation of controlling camera 100 so that it tracks a specific tracking subject among subjects detected from the video captured by camera 100, and the operation of operating camera 100 by controller 200 will be described with reference to FIG.

[0029] 3, 4, and 5 show the operational flow of the camera 100, and Fig. 6 shows the operational flow of the controller 200. First, the operation of the camera 100 will be described with reference to Fig. 3.

[0030] In S101, the CPU 101 determines whether or not the automatic tracking mode is specified. If the result of the determination is that the automatic tracking mode is specified, the CPU 101 starts the processing (steps S102 to S107) in this embodiment, which will be described below.

[0031] When the normal shooting mode is specified, the CPU 101 starts a general shooting operation, which, as described above, is when the user uses the controller 200 to control the pan, tilt, and zoom of the camera 100 to take a picture.

[0032] In S102, the CPU 101 instructs the image processing unit 106 to read the image captured by the image sensor 107 and write it to the RAM 102. The CPU 101 then reads the captured image from the RAM 102, inputs it to the inference unit 110, and instructs the inference unit 110 to execute inference. In response to this instruction, the inference unit 110 writes the inference result to the RAM 102.

[0033] The inference unit 110 reads out a trained model created using a machine learning technique such as deep learning from the ROM 103, receives an image as input data, and outputs position information and an ID, which is identification information, of an object on the image as output data.

[0034] The trained model used in this example is based on data linking a set of images of a person taken from multiple angles with identifiable information about the person, and has been trained to increase the similarity of features between images of the same person.

[0035] If multiple objects are present in the captured footage, inference results will be output for each object, and each object will be assigned a unique ID.

[0036] The CPU 101 stores the inference result for the current frame in the RAM 102, and also stores the inference result for the previous frame as the inference result for the past frame in the RAM 102. In order to improve the processing speed of the inference, the CPU 101 may reduce the size of the captured video, write it to the RAM 102, and input the reduced captured video to the inference unit 110.

[0037] Next, in S103 , the CPU 101 reads out the captured video and the inference result from the RAM 102 and transmits them to the controller 200 via the network I / F 105 .

[0038] Next, in S104, CPU 101 branches the process depending on whether or not an ID indicating the subject to be tracked (hereinafter referred to as a tracking subject ID) has been sent from controller 200 via network I / F 105. If a tracking subject ID has not been sent, the process proceeds to S106. If a tracking subject ID has been sent, the process proceeds to S105.

[0039] In S105, the CPU 101 writes the tracking subject ID sent from the controller 200 into the RAM 102, updates the tracking subject ID, and then the process proceeds to S106.

[0040] In S106, CPU 101 refers to RAM 102 and branches the process depending on whether the ID of the subject to be tracked has been decided. If the ID of the subject to be tracked has been decided, the process proceeds to S107. If the ID of the subject to be tracked has not been decided, the process proceeds to S101.

[0041] The processing of S107 shown in Fig. 4 is a characteristic processing of the present invention. Before explaining the processing of S107, a case where the tracking subject switches due to the tracking subject and another subject intersecting in the captured video will be explained using Fig. 7(A) to Fig. 9(B). Note that in the embodiment, "video" may be replaced with "image" and is an expression that includes both still images and moving images.

[0042] The inference unit 110 assigns the identification information (ID) of the person being tracked to the tracking subject, and assigns the identification information of another person to the other subject. However, if the tracking subject and the other subject cross paths in the captured video, the inference unit 110 may assign the identification information of the person being tracked to the other subject, and the identification information of the other person to the tracking subject. In other words, the identification information may be swapped and assigned. This is a phenomenon known as "possession."

[0043] Using Figures 7(A) to 8(B), we will explain how the tracked subject and another subject intersect in the captured video, and how the subjects end up overlapping each other. Using Figures 8(C) and 8(D), we will explain how "transfer" does not occur after the subjects overlap. Using Figures 9(A) and 9(B), we will explain how "transfer" occurs after the subjects overlap.

[0044] Fig. 7(A) is a diagram showing the inference result when two people (401, 402) are present in the video 400 captured by the camera 100. The people 401 and 402 are each represented using a schematic diagram of the human body. Fig. 7(A) also shows that the people 401 and 402 are planning to move in directions 461 and 462, respectively. Note that the directions 461 and 462 are figures used to supplement the explanation, and are not figures that are actually output as the result of the inference.

[0045] When the captured video 400 is input to the inference unit 110, the inference unit 110 outputs the coordinates indicating the upper left and lower right vertices of a rectangle 410 that includes the entire body of the person as display information indicating the position of the person 401. The inference unit 110 also outputs the coordinates indicating the upper left and lower right vertices of a rectangle 411 as display information indicating the position of the person 402.

[0046] Furthermore, the inference unit 110 assigns ID1 (420 in FIG. 7A) which is identification information given to the corresponding person to a rectangle 410 indicating the position of person 401. Furthermore, the inference unit 110 assigns ID2 (421 in FIG. 7A) which is identification information given to the corresponding person to a rectangle 411 indicating the position of person 402. In other words, the inference unit 110 has a function of identification assigning means that assigns the identification information of the person corresponding to the display information to the display information.

[0047] When the identification information ID1 (420) of the person to be tracked is assigned to the rectangle 410 indicating the position of the person 401, the CPU 101 performs pan and tilt operations to position the person 401 at the center of the angle of view. In this case, the person 401 is originally captured at the center of the angle of view in the captured video 400. However, in Figures 7(A) to 9(B), in order to make it easier to explain the overlapping situation, the angle of view is set so that the movements of multiple subjects can be seen, rather than centering on a specific subject.

[0048] The display information indicating the position of a person output by the inference unit 110 is not limited to the whole body, but may be information indicating the position of the head as shown by 450 and 451 in Fig. 7(B), or information indicating the position of the face. Furthermore, the display information indicating the position is not limited to the coordinates of the upper left vertex and lower right vertex of a rectangle, but may also be information such as the center coordinates, width, and height.

[0049] In addition, in order to detect parts of a person, such as the head or face, rather than detecting a person, the trained model stored in ROM 103 can be changed to a trained model trained based on training data corresponding to the desired output.

[0050] Although it has been stated that object detection such as human body parts is performed using machine learning, the detection method is not limited to this. For example, methods such as the SIFT method, which detects by matching local feature points in an image, or the pattern matching method, which detects by finding the similarity with a pre-registered "pattern (shape)," may also be used.

[0051] Furthermore, the CPU 101 assigns an ID to the current inference result for each frame. At this time, the CPU 101 compares the inference result of a past frame with the inference result of the current frame, so that the same ID is assigned to display information of the same person even if the person moves. The method of assigning the ID will be specifically described with reference to FIG. 7(C).

[0052] 7(C) shows only rectangles that are display information of the detection results for each frame of person 401 and person 402, and omits the schematic diagrams of the human bodies of person 401 and person 402. In FIG. 7(C), the detection results for the current frame are shown as rectangle 410b and rectangle 411b. The detection results for the frame one frame before that are shown as rectangle 410a and rectangle 411a. The detection results for the frame two frames before that are shown as rectangle 410 and rectangle 411.

[0053] The inference unit 110 acquires the feature amount of the image within the rectangle in the current frame and compares it with the feature amount of the image within each rectangle in past frames. If the comparison result is equal to or greater than a predetermined threshold, it determines that the similarity is high and assigns the same ID to the rectangle that is the closest display information. The comparison and determination are performed starting from the most recent past frame and working backward in time. If a determination is possible in the most recent past frame, the ID assignment ends there. However, if a determination cannot be made in that frame because the rectangle to be compared has temporarily disappeared, for example, a comparison is attempted in the previous past frame.

[0054] The feature may be acquired from color information of the image inside the rectangle, which is the display information. Alternatively, the feature may be acquired using a trained model that is trained to output similar feature values ​​for images of a specific person viewed from different angles and output different feature values ​​for images of different people. Furthermore, similarity may be compared and determined by pattern matching.

[0055] Furthermore, a method may be used in which the position of a rectangle in the current frame is predicted using a Kalman filter or the like, and the same ID is assigned to rectangles that are close to the predicted position of the rectangle and whose feature approximation is equal to or greater than a predetermined threshold.

[0056] In the case of Fig. 7(C), the inference unit 110 determines that the image in rectangle 410 from the previous frame has features most similar to those of the image in rectangle 410a, and assigns to rectangle 410a the same ID1 (420) as rectangle 410. Then, the inference unit 110 determines that the image in rectangle 410a from the previous frame has features most similar to those of the image in rectangle 410b, and assigns to rectangle 410b the same ID1 (420) as rectangle 410a.

[0057] 8(A) shows, using a schematic diagram of the human body, a state in which person 401 and person 402 intersect and overlap. Because the two people overlap, the inference unit 110 assigns ID1 (420) to rectangle 410c, which is a single piece of display information, and outputs it.

[0058] 8(B) shows only the rectangles of the detection results for each frame up to the point where person 401 and person 402 overlap in the video, and omits the schematic diagrams of the bodies of person 401 and person 402. Person 401 overlaps in front of person 402, and one rectangle 410c has been detected as a detection result for the current frame. Inference unit 110 determines that rectangle 410b has features most similar to rectangle 410c from the previous frame, and assigns rectangle 410c the same ID1 (420) as rectangle 410b.

[0059] Fig. 8(C) shows, using a schematic diagram of the human body, a state in which person 401 and person 402 have overlapped and then moved in the lower right and upper left directions, respectively, as indicated by moving directions 461 and 462 in Fig. 7(A). Fig. 8(C) also shows a state in which IDs identifying rectangles, which are each piece of display information, have been properly assigned.

[0060] Fig. 8(D) shows only the rectangles of the detection results for each frame up to the state shown in Fig. 8(C). The detection results for the current frame are shown as rectangle 410d and rectangle 411d. The detection result for the frame one frame before is rectangle 410c. The detection results for the frame two frames before are rectangle 410b and rectangle 411b shown in Fig. 7(C).

[0061] In the case of FIG. 8(D), the inference unit 110 determines that rectangle 410d is most similar to rectangle 410c, and assigns the same ID1 (420) to rectangle 410d as rectangle 410c. In this scene, there are no rectangles in the previous frame to search to assign an ID to rectangle 411d. This is because the ID of the only rectangle 410c in the previous frame has already been assigned to rectangle 410d. Therefore, the inference unit 110 searches for display information that is similar to rectangle 411d from the frame two frames earlier. The inference unit 110 determines that rectangle 411d is most similar to rectangle 411b, and assigns the same ID2 (421) to rectangle 411d.

[0062] 9(A) shows a state where an ID identifying a person has been assigned to the wrong display information, a so-called "transfer." Originally, ID2 (421) should have been assigned to the display information displaying person 402, and ID1 (420) should have been assigned to the display information displaying person 401, but the identification information has been swapped.

[0063] Fig. 9(B) shows only the rectangles of the detection results for each frame up to the state in Fig. 9(A) where the identification information has been swapped. The detection results for the current frame are shown as rectangle 410d and rectangle 411d. The detection result for the frame one frame before is rectangle 410c. The detection results for the frame two frames before are rectangles 410b and 411b in Fig. 7(C).

[0064] In the case of Figure 9(B), the inference unit 110 determines that rectangle 411d is most similar to rectangle 410c, and assigns rectangle 411d the same ID1 as rectangle 410c. Ideally, rectangle 410d should be assigned the same ID1 as rectangle 410c, but if the similarity between person 401 and person 402 is high, this type of interchange of identification information will occur. Furthermore, even when position prediction such as the Kalman filter described above is combined, if the direction of travel changes suddenly and significantly, as in traveling direction 461 and traveling direction 462 shown in Figure 7(A), position prediction becomes difficult, and there is a possibility of interchange of identification information.

[0065] In this way, if the identification information ID1 (420) of the person being tracked attached to the display information of the tracking subject is replaced with the identification information of another person attached to the display information of another person, the tracking subject will be swapped. This situation may lead to a shooting accident. <Description of Characteristic Operations of the First Embodiment>

[0066] Returning to Fig. 4, the characteristic processing of the first embodiment will now be described. In order to prevent photography accidents due to vehicle transfer, it is conceivable that the camera 100 detects the possibility of vehicle transfer in advance and notifies the user via the controller 200 to prompt them to take action. This allows the user to take action by closely monitoring the situation in preparation for a vehicle transfer, or by switching to normal shooting mode and operating the camera themselves as a safety measure.

[0067] 4, the CPU 101 determines whether or not there is a subject that has a high similarity to the tracked subject (hereinafter referred to as a similar subject). In this determination method, as explained in FIG. 7, the feature amount of the image within the rectangle of each person is acquired, and if it is equal to or greater than a predetermined threshold, it is determined that the similarity is high.

[0068] As described above, the feature amount may be acquired from color information, spatial feature amount obtained by pattern matching, feature amount obtained by a machine learning model, etc. If the result of this determination indicates that a similar subject is present, the process proceeds to S111. Note that even if a part of the subject is not within the angle of view, it may be possible to determine whether a similar subject is present or not, as long as it is possible to determine whether the similarity is equal to or greater than a predetermined threshold.

[0069] In S111, CPU 101 determines whether or not the similar subject is within a predetermined distance. The determination of whether or not the similar subject is within the predetermined distance is made by determining whether or not the center of the rectangle of the similar subject is inside a circle of radius L1, the center of which is the center of the rectangle of the tracking subject. If the result of this determination is that the similar subject is within the predetermined distance, the process proceeds to S112.

[0070] In S112, the CPU 101 transmits to the controller 200 via the network I / F 105 a warning display instruction to instruct the display of a warning to prompt the user to take measures in preparation for the occurrence of a transfer.

[0071] The warning display instruction is an instruction to display a warning to prompt the user to take measures to prepare for the occurrence of a transfer, or an instruction to hide the warning. The warning display instruction may include information on the ID of the similar subject.

[0072] The controller 200 displays or hides a warning in response to the warning display instruction. If the result of the determination in S110 is that a similar subject does not exist, the process proceeds to S113. If the result of the determination in S111 is that a similar subject exists farther than a predetermined distance, the process proceeds to S113.

[0073] In S113, the CPU 101 transmits a warning display instruction to hide the warning to the controller 200 via the network I / F 105. After that, the process exits from S107, returns to the operation flow of FIG. 3, and the process from S101 is repeated.

[0074] The controller 200 displays a warning when a similar subject is present within a predetermined distance, and hides the warning when the similar subject is present farther away than the predetermined distance. This allows the user to take measures to prepare for the occurrence of a transfer at a more appropriate time. Note that when hiding the warning, no instruction may be given.

[0075] Next, the operation of operating the camera 100 using the controller 200 will be described with reference to FIGS. 6, 9(C), and 9(D).

[0076] 6 shows the operation flow of the controller 200. First, in S121, the CPU 201 branches the process depending on whether or not it has detected a change in shooting mode caused by a user operating a joystick, button, touch panel, or the like via the user input I / F 206. In the initial state, the normal shooting mode is selected as the shooting mode. If a change in shooting mode is detected, the process proceeds to S122. If a change in shooting mode is not detected, the process proceeds to S123.

[0077] As explained above in S101, the shooting modes include an auto-tracking mode and a normal shooting mode. In S122, the CPU 201 writes the shooting mode to the RAM 202 and transmits it to the camera 100 via the network I / F 204.

[0078] In S123, the CPU 201 reads the shooting mode from the RAM 202 and determines the shooting mode. If the result of the determination is that the shooting mode is the auto-tracking mode, the process proceeds to S124. If the shooting mode is the normal shooting mode, an operation for a general shooting operation is started.

[0079] In S124, the CPU 201 receives the captured image and the inference result from the camera 100 via the network I / F 204, and stores them in the RAM 202.

[0080] Next, in S125 , the CPU 201 reads out the captured image and the inference result from the RAM 202 and displays them on the display unit 205 .

[0081] Next, in S126, the CPU 201 branches the process depending on whether or not it has detected a selection of a subject to be tracked by a user operating a joystick, button, touch panel, or the like via the user input I / F 206. If it has detected a selection of a subject to be tracked, it transitions the process to S127. If it has not detected a selection of a subject to be tracked, it transitions the process to S128. In S127, the CPU 201 transmits the ID of the selected subject to be tracked to the camera 100 via the network I / F 204.

[0082] Here, an example of the display on display unit 205 after the user has selected a tracking subject will be described with reference to Fig. 9(C). Display image 470 in Fig. 9(C) is an example of an image displayed on display unit 205 of controller 200.

[0083] 9(C), person 401 and person 402 are detected as people, and rectangles 481 and 482 containing the whole bodies of the people are displayed as display information indicating their respective positions. Person 401 has been selected as the subject to be tracked, and Tracking label 430 is displayed to indicate that person 401 is the subject to be tracked. Note that CPU 101 performs pan / tilt control so as to position person 401, the subject to be tracked, at the center of the angle of view, resulting in an image in which person 401 is captured at the center of the angle of view.

[0084] Now, we return to the description of the operational flow of the controller 200 in Fig. 6. In S128, the CPU 201 branches the process depending on whether or not a warning display instruction has been received from the camera 100 via the network I / F 204.

[0085] The warning display instruction is an instruction to display or not display a warning to prompt the user to take measures in preparation for the occurrence of a transfer, and is information that the camera 100 transmits to the controller 200 in S112 or S113.

[0086] If a warning display instruction is received, the process transitions to S129. If a warning display instruction is not received, the process transitions to S121. At S129, the CPU 201 branches the process depending on whether the content of the warning display instruction is to display a warning or not to display a warning. If a warning is to be displayed, the process transitions to S130. If a warning is not to be displayed, the process transitions to S131. At S130, the CPU 201 displays a warning on the display unit 205 to prompt the user to take action in preparation for the occurrence of a transfer. At S131, the CPU 201 does not display a warning on the display unit 205. If a warning was displayed, the warning is hidden.

[0087] Here, an example of a warning display for urging the user to take measures to prepare for the occurrence of a subject transfer will be described with reference to Fig. 9(D). Fig. 9(D) shows a state in which a person 402 determined to have a high degree of similarity is present within a predetermined distance of a person 401 who is a tracking subject. A Warning label 431 is assigned to the person 402 to indicate that the person is a similar subject.

[0088] As described above, if a similar subject exists near the tracking subject, a warning can be sent to the user to prepare for a possible transfer, allowing the user to take measures such as keeping a close eye on the situation in case a transfer occurs, or switching to normal shooting mode and operating the camera themselves as a safety measure.

[0089] In this embodiment, in S130, the CPU 201 displays a warning using the Warning label 431, but the warning may also be notified to the user by highlighting the rectangle 482 by changing the color or thickness of the lines, for example.

[0090] Furthermore, the warning may be displayed by highlighting the window that displays the display image 470 by changing the frame color, or by displaying a warning window or a warning message in a predetermined message display area.

[0091] In the subject change handling process of S107 in Fig. 4, a determination is made in S111 as to whether a warning display is to be made if a similar subject is present within a predetermined distance from the tracking subject. However, in order to secure more time to prepare for a subject change, the subject change handling process of S107 may be made to be a process that does not determine whether a similar subject is present within the predetermined distance, as shown in Fig. 5. That is, CPU 101, which is a warning display instruction means, may perform a process of transmitting an instruction to display a warning if a similar subject is present (S112), and transmitting an instruction to hide the warning if a similar subject is not present (S113). Note that no instruction may be made if a similar subject is not present.

[0092] Furthermore, the movement direction of the tracking subject and the similar subject may be predicted, and if the trajectories of the movement directions intersect, a determination may be made to display a warning.

[0093] Furthermore, the determination of whether a similar subject exists within the predetermined distance relates to the distance between the tracking subject and the similar subject in the video, not the difference in the distance between the two subjects and the camera. If the difference in distance between the two subjects and the camera is large, the size of the two subjects displayed in the video will differ. Even if two subjects overlap, if the identification information of the two subjects can be determined based on the size information of the subjects, the identification information assigned to the two subjects after the overlap is unlikely to be swapped. Therefore, when determining whether to display or hide a warning, it is possible to include a condition based on the comparison of the subject sizes, such as the ratio of the absolute value of the difference in size between the two subjects to the size of one subject. For example, if the ratio is equal to or less than a predetermined threshold, the warning may be displayed or hidden according to the flow chart in FIG. 4 or FIG. 5, and if the ratio is greater than the threshold, the warning may be hidden. The size of the two subjects may be determined using the length or width of a rectangle indicating the subject's position information, the area of ​​the rectangle, or the like.

[0094] Furthermore, if a similar subject is tracked due to a change of vehicle, that person may be positioned at the center of the angle of view, which may cause the person who should be photographed to move out of the angle of view. Therefore, after transmitting the warning display instruction in S112, CPU 101 may apply processing to automatically stop tracking.

[0095] Furthermore, if a similar subject is not detected or if the similar subject is located outside a predetermined distance, the possibility of a transfer is reduced, so a process of resuming automatic tracking may be applied after transmitting a warning display instruction to hide the warning in S113. Note that if a similar subject is not detected or if the similar subject is located outside a predetermined distance, no instruction may be issued. <Second embodiment>

[0096] In the first embodiment, an example was described in which, when the tracking subject and the similar subject are within a predetermined distance, it is determined that there is a high possibility of a transfer occurring, and the camera 100 transmits a warning to the controller 200 to notify the user.

[0097] In the second embodiment, when the tracking subject and the similar subject cross each other, it is determined that there is a high possibility of a transfer occurring, and the camera 100 transmits a warning display instruction to the controller 200.

[0098] In the second embodiment, as described in the first embodiment, the inference unit 110 also has the function of an identification information assigning means that assigns a person's identification information to display information. In the second embodiment, the inference unit 110 also has an identification information switching accepting means that accepts an instruction to switch the identification information assigned to the display information via an acceptance button or the like displayed on the display unit.

[0099] Furthermore, in order for the user to notice the transfer and quickly return the tracking subject to its original position, the person originally designated as the tracking subject needs to be within the angle of view. Therefore, the camera 100 sends a warning to the controller 200 and immediately starts a zoom operation to widen the angle of view. This prevents the person originally designated as the tracking subject from leaving the angle of view, allowing time for the user to return the tracking subject to its original position. The above is an overview of the characteristic processing of the second embodiment. Note that the second embodiment also shares the system configuration, block configuration, operation flow of FIG. 3, and part of the operation flow of FIG. 6, etc., described in the first embodiment, and therefore redundant explanations will be omitted.

[0100] The following mainly describes the differences from the first embodiment. The processing of S107 in camera 100 described in Figures 4 and 5 is different in the second embodiment, and therefore this processing will be described using Figure 10. The operation flow of controller 200 described in Figure 6 is also different in the second embodiment, and therefore will be described using Figures 11 and 12. The example of the display image on controller 200 described in Figures 9(C) and 9(D) is also different in the second embodiment, and therefore will be described using Figures 13 and 14. <Description of Characteristic Operations of the Second Embodiment>

[0101] First, the processing of S107 of the camera 100 in the second embodiment will be described with reference to Fig. 10. In S201, the CPU 101 branches the processing depending on whether or not it has detected that the display information indicating the positions of the tracking subject and the similar subject has changed into one. If it has detected this, the processing transitions to S202. If it has not detected this, the processing transitions to S206. Here, detection refers to whether or not two people have overlapped, as shown in Fig. 8(A), and the inference unit 110 has detected this as one piece of display information.

[0102] 10, CPU 101 turns on warning mode and writes it to RAM 102. Next, in S203, CPU 101 transmits a warning display instruction to controller 200 via network I / F 105 to display a warning. Next, in S204, CPU 101 starts a zoom operation to change the angle of view to a wider angle. That is, CPU 101, as an angle of view changing unit, changes the angle of view of the camera for shooting.

[0103] This change in the angle of view is done to allow time for the user to return the tracking subject to its original position. Furthermore, even if a possession change occurs, the person who should be tracked can be captured within the angle of view and filmed, which reduces the possibility that the viewer will perceive the footage as an accident.

[0104] The field of view change process to widen the field of view is performed when two people intersect and overlap, and the display information changes to one, regardless of whether a person actually transitions. This field of view change process may be performed by presenting an image that intentionally captures the intersecting scene, and by controlling the field of view to slowly widen with the two people at the center, so that the viewer can accept the image naturally.

[0105] In a switcher system, if the switcher follows the warning display instruction and switches from the main line video to the standby video, consideration for viewers is not required. Therefore, in the process of widening the field of view by the CPU 101 in S204, there is no problem even if the speed at which the field of view changes is increased. The switcher (not shown) may communicate with the switcher via LAN to obtain information on whether the captured video is being used for the main line video or the standby video, and automatically adjust the speed at which the field of view changes as the field of view changes.

[0106] Next, in S205, the CPU 101 transmits, via the network I / F 105, to the controller 200, an instruction to display a switch button that accepts an instruction to switch the identification information added to the display information. In S206, the CPU 101 acquires warning mode information from the RAM 102, and branches the processing depending on whether the warning mode is ON or OFF. If the warning mode is ON, the processing transitions to S207. If the warning mode is OFF, the processing of S107 is omitted, and the operation flow of FIG. 3 is returned to, and the processing from S101 is repeated.

[0107] In S207, the CPU 101 branches the process depending on whether or not a switch button press has been received from the controller 200 via the network I / F 105. If a switch button press has been received, the process proceeds to S209. If a switch button press has not been received, the process proceeds to S208.

[0108] In S208, the CPU 101 branches the process depending on whether a predetermined time has elapsed since the warning mode was turned on. If the predetermined time has elapsed, the process proceeds to S210. If the predetermined time has not elapsed, the process skips S107 and returns to the operation flow of FIG. 3, where the process from S101 is repeated.

[0109] In S209, CPU 101 updates the tracking subject ID in order to switch the current tracking subject to the similar subject with the highest similarity. Next, in S210, CPU 101 transmits an instruction to controller 200 via network I / F 105 to hide the switching button.

[0110] Next, in S211, the CPU 101 returns the widened angle of view to its original state by a zoom operation. Next, in S212, the CPU 101 transmits a warning display instruction to the controller 200 via the network I / F 105 to instruct the controller 200 not to display the warning.

[0111] Next, in S213, the CPU 101 turns off the warning mode and writes the result to the RAM 102. After that, the process exits from S107, returns to the operation flow of FIG. 3, and the process from S101 is repeated.

[0112] If no transfer actually occurs, the CPU 101 does not receive a switch button press from the controller 200. Therefore, in S210 and S212, if a predetermined time has elapsed, the CPU 101 transmits to the controller 200 an instruction to hide the warning display and the switch button. Then, in S211, if a predetermined time has elapsed after changing the angle of view, the CPU 101 returns the angle of view that has been widened to the original angle of view.

[0113] In addition, in S211, when the CPU 101 returns the widened angle of view to its original state, the CPU 101 may perform control to return the viewer to the original state by slowly zooming in on the two people so that the viewer can naturally accept the video.

[0114] Next, the operation of operating the camera 100 with the controller 200 will be described with reference to Figures 11 and 12. Figure 11 shows the operation flow of the controller 200. The processes from S301 to S307 in Figure 11 are the same as those from S121 to S127 in Figure 6, so the description will be omitted. Figure 12 shows the operation flow that follows the operation flow in Figure 11.

[0115] 12, in S308, the CPU 201 branches the process depending on whether or not a warning display instruction has been received from the camera 100 via the network I / F 204. The warning display instruction is an instruction to display a warning to prompt the user to take action in preparation for the occurrence of a transfer, or an instruction to hide the warning, and is information that the camera 100 transmits to the controller 200 in S203 or S212. If a warning display instruction to display a warning has been received, the process transitions to S309. If a warning display instruction to hide the warning has been received, the process transitions to S301.

[0116] In S309, the CPU 201 branches the process depending on whether the warning display instruction is to display or not display a warning. If the warning is to be displayed, the process proceeds to S310. If the warning is to be not displayed, the process proceeds to S311. In S310, the CPU 201 displays a warning on the display unit 205 to prompt the user to take measures to prepare for the occurrence of a transfer.

[0117] Next, in S312, the CPU 201 branches the process depending on whether or not a switching button display instruction has been received from the camera 100 via the network I / F 204. If a switching button display instruction has been received, the process transitions to S313. If a switching button display instruction has not been received, the process transitions to S301. In S313, the CPU 201 branches the process depending on an instruction to display the switching button or an instruction to hide the switching button. If an instruction to display the switching button has been received, the process transitions to S314. If an instruction to hide the switching button has been received, the process transitions to S315. In S314, the CPU 201 displays a switching button on the display unit 205 that allows the tracking subject to immediately return to its original position if a subject transfer occurs.

[0118] The flow up to S310 and S314 and how it is displayed on the display unit 205 of the controller 200 will be described with reference to FIGS. 13(A) and 13(B).

[0119] Display image 500 in Fig. 13(A) is an example of an image displayed on display unit 205. In Fig. 13(A), person 501 and person 502 are detected as people, and rectangles 510 and 511 including the whole bodies of the people are displayed as display information indicating their respective positions. Fig. 13(A) shows a state in which person 501 has been selected as a tracking subject, and a Tracking label 530 has been assigned to person 501. Assume that person 502 has a high similarity to person 501, who is the tracking subject, and is recognized as a similar subject. In display image 500, CPU 101 has performed pan / tilt control to position person 501, who is the tracking subject, at the center of the angle of view, and therefore person 501 is captured at the center of the angle of view.

[0120] 13(B) shows a situation in which person 501 and person 502 intersect and overlap, resulting in a single rectangle 510, which is display information. As described in the first embodiment, when people overlap and display information becomes a single piece, if the similarity between these people is high, there is a possibility that IDs will be swapped and a "transfer" will occur.

[0121] 13B, a warning message 531 is displayed as a result of the processing in S310 to alert the user that an intersection with a similar subject has occurred. This is to prompt the user to take measures to prepare for the possibility of possession.

[0122] The switch button 532 in FIG. 13(B) is displayed as a result of the processing in S314. The switching target of the switch button 532 is limited to only the tracking subject and its similar subject. When the identification information for the subject is switched by pressing the switch button 532, an operation is performed to switch between the similar subject and the tracking subject. If a subject transfer occurs, the person who was originally the similar subject becomes the tracking subject, and the person who was originally the tracking subject becomes the similar subject. By pressing this switch button 532, the user can instantly return to the intended tracking subject.

[0123] Also, in consideration of the case where there are multiple similar subjects, a method may be used in which the cursor is moved to a subject in descending order of similarity each time the switch button 532 is pressed, and the tracking subject is switched by pressing the confirm button.

[0124] Returning to Fig. 12, the description of the operation flow of the controller 200 will continue. In S316, the CPU 201 branches the process depending on whether or not it has detected, via the user input I / F 206, that the switch button has been pressed by the user operating a button, touch panel, or the like. If a press is detected, the process transitions to S317. If a press is not detected, the process transitions to S301. In S317, the CPU 201 transmits, via the network I / F 204, to the camera 100, a notification that the switch button has been pressed.

[0125] The flow up to S317 will be explained using Fig. 13(C) to Fig. 14(A). Fig. 13(C) shows a situation in which the tracking subject has moved from person 501 to person 502. Tracking label 530 has also been assigned to person 502, and the angle of view has also become an image in which person 502 is captured at the center of the angle of view. Furthermore, Fig. 13(C) also shows a situation in which the angle of view has widened compared to Fig. 13(B) due to the processing of S204 of camera 100. Furthermore, Fig. 13(D) shows a situation in which the angle of view has widened even further compared to Fig. 13(C).

[0126] By widening the angle of view in this way, person 501, who should originally be photographed as a tracking subject, can be captured within the angle of view, and time can be saved before returning the tracking subject to its original state using switch button 532. Furthermore, since person 501, who should originally be photographed, can be captured within the angle of view and photography can continue, the possibility of viewers perceiving the footage as an accidental image is reduced.

[0127] FIG. 14A shows a state in which, as a result of the processing in S317 of FIG. 12, the tracking subject has returned to the person 501 who should have been tracked.

[0128] Returning now to Fig. 12, the description of the operational flow of the controller 200 will continue. In S212 of Fig. 10, a warning display instruction to hide the warning is sent from the camera 100 to the controller 200, so if a warning display has already been displayed in S311 of Fig. 12, the CPU 201 hides the warning. Also, in S210 of Fig. 10, a switching button display instruction to hide the switching button is sent to the controller 200, so if the switching button has already been displayed, the CPU 201 hides the switching button in S315.

[0129] 14(B) and 14(C) will be used to explain the image on the display unit 205. Fig. 14(B) shows the image on the display unit 205 when the warning message 531 and the switching button 532 are hidden as a result of the controller 200 executing the processes of S311 and S315, respectively.

[0130] Furthermore, Fig. 14(B) also shows how the angle of view is gradually returning to the original zoom size compared to Fig. 14(A) due to the processing of S211 of the camera 100. Furthermore, Fig. 14(C) shows how the zoom size of the angle of view has been returned to its original state due to the processing of S211 of the camera 100, and the angle of view has returned to the same state as at the time of Fig. 13(B).

[0131] As described above, in the second embodiment, when it is detected that the display information displaying the positions of the tracking subject and the similar subject has changed into one, the camera 100 determines that there is a high possibility of a transfer occurring and sends a warning to the controller 200.

[0132] Furthermore, the controller 200 displays a switching button so that if the subject is transferred, the tracking subject can be quickly returned to its original position. Furthermore, the camera 100 transmits a warning display instruction to the controller 200 and immediately starts a zoom operation to widen the angle of view. This allows time for the user to return the tracking subject to its original position. Note that, as described in the first embodiment, the CPU 101 may apply processing to automatically stop tracking after transmitting the warning display instruction in S112.

[0133] As a result, in preparation for a subject being possessed, the user can take measures such as keeping a close eye on the situation, or as a safety measure, switching to normal shooting mode and operating the camera themselves. And even if a subject is possessed, the user can calmly and easily return the tracked subject to its original state.

[0134] Even while the tracked subject is being returned to its original position, the camera can continue to capture the subject that should be tracked within the field of view, reducing the possibility that viewers will perceive the footage as accidental.

[0135] In S201, the determination as to whether the display information displaying the positions of the tracked subject and the similar subject has changed to one has been explained using Fig. 13(B) in the case where two people overlap when crossing each other. Note that, as a determination method by CPU 101 in S201, it may also be determined that the display information has changed to one when there is an obstruction near the two people and one of the people is hidden behind the obstruction.

[0136] In addition, in S204, the CPU 101 widens the angle of view by zooming in response to the warning. Alternatively, the CPU 101 may read out a zoom value preset in advance in the ROM in response to the warning and perform control to switch to the preset angle of view.

[0137] The present disclosure includes the following configurations. (Configuration 1) a video input means for inputting video of a plurality of people; an automatic tracking means for automatically tracking a predetermined person in the video based on the video input to the video input means; A camera control device characterized by having a warning display instruction means that instructs a display device to display a warning when the specified person and a second person similar to the specified person are present in the image. (Configuration 2) a video input means for inputting video of a plurality of people; an automatic tracking means for automatically tracking a predetermined person in the video based on the video input to the video input means; A camera control device characterized by having a warning display instruction means for instructing a display device to display a warning when display information indicating the position of the specified person and display information indicating the position of a second person similar to the specified person in the video change to one. (Configuration 3) The camera control device according to configuration 1, wherein the warning display instruction means issues an instruction to display the warning when the specified person and the second person are present within a specified distance in the video. (Configuration 4) The method further includes a means for predicting a moving direction of the predetermined person and the second person in the video, The camera control device according to configuration 1, characterized in that the warning display instruction means gives an instruction to display the warning when, based on the result predicted by the means for predicting the movement direction, there is a possibility that the display information indicating the position of the specified person and the display information indicating the position of the second person will change to one in the image. (Configuration 5) The camera control device according to configuration 1, characterized in that the warning display instruction means issues an instruction to display the warning when a value obtained by comparing the size of the specified person with the size of the second person in the video is equal to or smaller than a specified threshold value. (Configuration 6) The camera control device according to configuration 1 or 2, characterized in that whether the predetermined person and the second person in the video are similar or not is determined by a trained model that has been trained to increase the similarity of features for the same person. (Configuration 7) The camera control device according to configuration 1 or 2, characterized in that whether the predetermined person and the second person in the video are similar or not is determined by pattern matching that determines the similarity with a pre-registered pattern. (Configuration 8) The camera control device according to configuration 1 or 2, characterized in that whether the predetermined person and the second person in the video are similar or not is determined from color information of the input video. (Configuration 9) The camera control device described in configuration 2, characterized in that when the display information indicating the position of the specified person and the display information indicating the position of the second person in the video change into one, it is when the specified person and the second person intersect and overlap, causing the display information indicating the positions of the people to become one. (Configuration 10) The camera control device described in configuration 2, characterized in that in the video, the display information indicating the position of the specified person and the display information indicating the position of the second person change to one when there is an obstruction near the specified person and the second person, and one of the people is hidden behind the obstruction, causing the display information indicating the people's positions to become one. (Configuration 11) an identification information assigning means for assigning identification information corresponding to a person to the display information; 3. The camera control device according to configuration 2, further comprising an identification information switching receiving means for receiving an instruction to switch the identification information. (Configuration 12) 12. The camera control device according to claim 11, wherein the identification information switching accepting means accepts an instruction to switch the identification information only between the predetermined person and the second person. (Configuration 13) 3. The camera control device according to configuration 2, further comprising a field angle changing means for changing the field angle of the video when the warning display instruction means issues an instruction to display the warning. (Configuration 14) 14. The camera control device according to claim 13, wherein the angle of view changing means widens the angle of view by a zoom operation. (Configuration 15) 14. The camera control device according to claim 13, wherein the angle of view changing means changes the angle of view to a preset angle of view that is registered in advance. (Configuration 16) 14. The camera control device according to configuration 13, wherein the angle of view changing means automatically adjusts the speed of the angle of view change depending on whether the video is used as a main line video or a standby video. (Configuration 17) 14. The camera control device according to configuration 13, wherein the angle of view changing means changes the angle of view so that the predetermined person and the second person do not deviate from the angle of view. (Configuration 18) 14. The camera control device according to claim 13, wherein the angle-of-view changing means returns the angle of view to its original state when a predetermined time has elapsed since the angle of view was changed. (Configuration 19) 3. The camera control device according to configuration 1 or 2, wherein the automatic tracking means stops the automatic tracking in accordance with an instruction to display the warning. (Configuration 20) 3. The camera control device according to configuration 1 or 2, wherein the instruction to display a warning includes an instruction to display a warning message. (Configuration 21) 3. The camera control device according to configuration 1 or 2, wherein the instruction to display the warning includes an instruction to highlight display information indicating the position of the second person. (Configuration 22) 3. The camera control device according to configuration 1 or 2, wherein the instruction to display the warning includes an instruction to highlight a window displayed on the display unit. (Configuration 23) a video input step of inputting video of a plurality of people; an automatic tracking step of automatically tracking a predetermined person in the video based on the video input step; A camera control method characterized by comprising a warning display instruction step of instructing a display device to display a warning when the specified person and a second person similar to the specified person are present in the video. (Configuration 24) a video input step of inputting video of a plurality of people; an automatic tracking step of automatically tracking a predetermined person in the video based on the video input step; A camera control method characterized by comprising a warning display instruction step of instructing a display device to display a warning when, in the video, display information indicating the position of the specified person and display information indicating the position of a second person similar to the specified person change to one.

[0138] The present invention has been described in detail above based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications are possible based on the spirit of the present invention, and these modifications are not excluded from the scope of the present invention. [Explanation of symbols]

[0139] 100 cameras 200 Controller 300 LAN 101 CPU 106 Image processing section 110 Reasoning Department 201 CPU 205 Display section

Claims

1. a video input means for inputting video of a plurality of people; an automatic tracking means for automatically tracking a predetermined person in the video based on the video input to the video input means; A camera control device characterized by having a warning display instruction means that instructs a display device to display a warning when the specified person and a second person similar to the specified person are present in the image.

2. a video input means for inputting video of a plurality of people; an automatic tracking means for automatically tracking a predetermined person in the video based on the video input to the video input means; A camera control device characterized by having a warning display instruction means for instructing a display device to display a warning when display information indicating the position of the specified person and display information indicating the position of a second person similar to the specified person in the image change to one.

3. 2. The camera control device according to claim 1, wherein the warning display instruction means issues an instruction to display the warning when the predetermined person and the second person are present within a predetermined distance in the video.

4. The method further includes means for predicting a moving direction of the predetermined person and the second person in the video, The camera control device described in claim 1, characterized in that the warning display instruction means instructs to display the warning when, based on the results predicted by the means for predicting the movement direction, there is a possibility that the display information indicating the position of the specified person and the display information indicating the position of the second person in the image will change to one.

5. 2. The camera control device according to claim 1, wherein the warning display instruction means instructs the camera to display the warning when a value obtained by comparing the size of the specified person with the size of the second person in the video is equal to or smaller than a specified threshold value.

6. 3. The camera control device according to claim 1, wherein whether the specified person and the second person in the video are similar is determined using a trained model that has been trained to increase the similarity of features for the same person.

7. 3. The camera control device according to claim 1, wherein whether the predetermined person and the second person in the video are similar or not is determined by pattern matching that determines the degree of similarity with a pre-registered pattern.

8. 3. The camera control device according to claim 1, wherein whether or not the predetermined person and the second person in the video are similar is determined from color information of the input video.

9. The camera control device described in claim 2, characterized in that when the display information indicating the position of the specified person and the display information indicating the position of the second person in the image change into one, it is when the specified person and the second person intersect and overlap, causing the display information indicating the image of the person to become one.

10. The camera control device described in claim 2, characterized in that when the display information indicating the position of the specified person and the display information indicating the position of the second person in the video change to one, it is when there is an obstruction near the specified person and the second person, and one of the people is hidden behind the obstruction, causing the display information indicating the people's positions to become one.

11. an identification information assigning means for assigning identification information corresponding to a person to the display information; 3. The camera control device according to claim 2, further comprising: an identification information switching receiving unit that receives an instruction to switch the identification information.

12. 12. The camera control device according to claim 11, wherein the identification information switching accepting means accepts an instruction to switch the identification information only between the predetermined person and the second person.

13. 3. The camera control device according to claim 2, further comprising a field angle changing means for changing the field angle of the image when the warning display instruction means issues an instruction to display the warning.

14. 14. The camera control device according to claim 13, wherein the angle of view changing means widens the angle of view by a zoom operation.

15. 14. The camera control device according to claim 13, wherein the angle of view changing means changes the angle of view to a preset angle of view that is registered in advance.

16. 14. The camera control device according to claim 13, wherein the angle-of-view changing means automatically adjusts the speed of the angle-of-view change depending on whether the video is used as a main line video or a standby video.

17. 14. The camera control device according to claim 13, wherein the angle-of-view changing means changes the angle of view so that the predetermined person and the second person do not deviate from the angle of view.

18. 14. The camera control device according to claim 13, wherein the angle-of-view changing means returns the angle of view to the original angle when a predetermined time has elapsed since the angle of view was changed.

19. 3. The camera control device according to claim 1, wherein the automatic tracking means stops the automatic tracking in response to an instruction to display the warning.

20. 3. The camera control device according to claim 1, wherein the instruction to display a warning includes an instruction to display a warning message.

21. 3. The camera control device according to claim 1, wherein the instruction to display the warning includes an instruction to highlight the display information indicating the position of the second person.

22. 3. The camera control device according to claim 1, wherein the instruction to display the warning includes an instruction to highlight a window displayed on the display unit.

23. a video input step of inputting video of a plurality of people; an automatic tracking step of automatically tracking a predetermined person in the video based on the video input step; A camera control method characterized by comprising a warning display instruction step of instructing a display device to display a warning when the specified person and a second person similar to the specified person are present in the image.

24. a video input step of inputting video of a plurality of people; an automatic tracking step of automatically tracking a predetermined person in the video based on the video input step; A camera control method characterized by having a warning display instruction process for instructing a display device to display a warning when display information indicating the position of the specified person and display information indicating the position of a second person similar to the specified person in the video change to one.