Image processing device for recognizing object in captured image, image processing system, method for controlling image processing device, and program for image processing device

The image processing apparatus dynamically determines a region of interest based on distortion and position information within the captured image, enhancing the accuracy of recognizing predetermined actions across varying object positions.

WO2025126655A1PCT designated stage expired Publication Date: 2025-06-19CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/036808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-10-16
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing image processing systems struggle to accurately recognize predetermined actions in captured images, especially when the position of the object within the imaging angle of view varies, leading to decreased recognition accuracy due to fixed trimming sizes and resolution reductions.

Method used

An image processing apparatus and system that dynamically determine a region of interest based on distortion information and the position of interest in the captured image, generating a region-of-interest image scaled to a predetermined size for high-precision recognition of actions.

Benefits of technology

Enables high-accuracy recognition of predetermined actions by objects in captured images regardless of their position within the imaging angle of view, improving the reliability of automatic imaging systems for sports and other applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036808_19062025_PF_FP_ABST
    Figure JP2024036808_19062025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To provide an image processing device and the like that are capable of recognizing a prescribed behavior by an object in a captured image with high accuracy regardless of the position on an imaging view angle. [Solution] The present invention is characterized by comprising: a distortion information generation unit that generates distortion information indicating distortion from a known shape, which is the shape of a first object in a captured image which is inputted by an input unit; a calculation unit (step S705) that calculates coordinates of a position of interest that is the position of a second object in the captured image; a determination unit that determines a region of interest on the basis of the distortion information and the coordinates of the position of interest; an image generation unit (step S402) that generates an image signal 601 for a region of interest where an object of interest is present, from the captured image; a magnification / reduction unit (step S404) that magnifies or reduces a region-of-interest image to a prescribed size; and a recognition unit (step S406) that performs prescribed recognition on the basis of the magnified / reduced region-of-interest image. The present invention is also characterized in that the determination unit determines the size, of the region of interest, corresponding to the actual size on the basis of the distortion information and the coordinates of the position of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing system, control method for image processing device, and program for image processing device for recognizing objects in captured images

[0001] The present invention relates to an image processing device, an image processing system, a control method for an image processing device, and a program for recognizing an object in a captured image.

[0002] In recent years, there has been a need for automatic imaging of highlight scenes during sports and other games. For example, in the case of basketball, a highlight scene can be a shot by a player. To automatically capture a shot, one method is to determine whether a shot is being taken based on the positional relationship between the ball and the player and the player's posture from the camera's image data, and then control the camera to capture the player taking the shot.

[0003] A known method for estimating a player's posture from image data is to detect the "player's joint coordinates." In this case, a posture estimation trained model (posture estimation model) is used to detect the "player's joint coordinates." Specifically, by inputting an image of a player into the posture estimation model, the coordinates of the player's joints, such as the top of the head, neck, wrists, elbows, shoulders, and waist, are output, thereby obtaining the "joint coordinates." Taking basketball as an example, when estimating a player's posture, if image data showing the entire basketball court is input into the "posture estimation model," the input resolution is high, resulting in a long processing time that is not suitable for real-time processing. Therefore, a known method is to reduce the input resolution and shorten the processing time by reducing the image data before inputting it into the "posture estimation model."

[0004] However, when imaging data is reduced, the resolution of the player in the imaging data also decreases accordingly, which may make it impossible to detect the player's joint coordinates with high accuracy. In other words, it may be impossible to determine whether the player is shooting or not. Therefore, Patent Document 1 discloses a method for performing predetermined recognition at high speed while increasing the input resolution to a trained model including a posture estimation model. Patent Document 1 discloses a technology for detecting the coordinates of a ball from imaging data and determining the player's actions (behavior) from an image cropped around the ball. Depending on the type of sport, there is a high possibility that the player will perform a predetermined behavior near the ball. Therefore, this technology makes it possible to determine the player's behavior from the coordinates of the ball and the player's joint coordinates.

[0005] Japanese Patent Application Laid-Open No. 2020-054748

[0006] However, when capturing an image capturing the entire basketball court, players farther from the camera are drawn smaller, while players closer to the camera are drawn larger. In the conventional technology disclosed in Patent Document 1, the cropping size is always constant. Therefore, if cropping is performed based on players farther from the camera, players closer to the camera may be cut off. Since the player's joint coordinates cannot be detected in the "cut-off areas," it becomes difficult to accurately estimate the player's posture. On the other hand, if cropping is performed based on players closer to the camera, the resolution of players farther from the camera may be low. Low resolution further reduces the resolution due to the reduction process before inputting the data into the trained model, making it difficult to accurately recognize players. When performing predetermined recognition from data captured over a large area, such as a basketball court, cropping to a fixed size without considering the distance from the camera to the subject (e.g., player) results in insufficient cropping for recognition. This results in a problem of reduced recognition accuracy.

[0007] The present invention provides an image processing device, an image processing system, a control method, and a program that enable a predetermined action of an object in a captured image to be recognized with high accuracy regardless of its position within the captured image field angle.

[0008] The image processing device of the present invention is an image processing device capable of recognizing a specified scene in a captured image, and comprises: an input unit that inputs a captured image of a first object; a distortion information generation unit that generates distortion information indicating distortion of the shape of the first object in the captured image input by the input unit from a known shape; a calculation unit that calculates a focus position in the captured image; a determination unit that determines a focus position based on the distortion information generated by the distortion information generation unit and the focus position calculated by the calculation unit; an image generation unit that generates a focus position image from the captured image based on the focus position determined by the determination unit; a scaling unit that scales the focus position image generated by the image generation unit to a specified size; and a recognition unit that performs specified recognition based on the focus position image scaled by the scaling unit, wherein the determination unit determines a focus position size, which is the size of the focus position corresponding to the actual size, based on the distortion information generated by the distortion information generation unit and the focus position calculated by the calculation unit.

[0009] According to the present invention, it is possible to provide an image processing device, an image processing system, a control method, and a program that can recognize a predetermined action by an object in a captured image with high accuracy regardless of its position within the captured field of view.

[0010] FIG. 1 is a configuration diagram of an image processing system according to the first to third embodiments. FIG. 2 is a hardware configuration diagram of an image processing system according to the first to third embodiments. FIG. 3 is a functional configuration diagram of an image processing system according to the first to third embodiments. FIG. 4 is a flowchart showing processing performed by an overhead camera according to the first to third embodiments. FIG. 5 is a flowchart showing processing performed by an image processing device according to the first to third embodiments. FIG. 6 is a flowchart showing processing performed by a target camera according to the first to third embodiments. FIG. 7 is an explanatory diagram of an overhead image signal according to the first to third embodiments. FIG. 8 is an explanatory diagram of a target area image signal according to the first to third embodiments. FIG. 9 is an explanatory diagram of a reduced image signal according to the first to third embodiments. FIG. 10 is a flowchart showing a target area image generation process according to the first embodiment. FIG. 11 is an explanatory diagram of a target area image generation process according to the first embodiment. FIG. 12 is a flowchart showing a target area image generation process according to the second embodiment. FIG. 13 is an explanatory diagram (side perspective view) of a target area image generation process according to the second embodiment. FIG. 14 is an explanatory diagram (top perspective view) of a target area image generation process according to the second embodiment. FIG. 15 is a flowchart showing a target area image generation process according to the third embodiment. FIG. 16 is an explanatory diagram (perspective view) of a target area image generation process according to the third embodiment. FIG. 17 is an explanatory diagram (plan view) of a target area image generation process according to the third embodiment.

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited to the configurations described in the embodiments. First, a first embodiment of the present invention will be described. In the accompanying drawings, the same or similar components are designated by the same reference numerals, and duplicate explanations will be omitted.

[0012] First, a first embodiment of the present invention will be described. In the first embodiment, a basketball will be used as an example to describe a process of determining whether or not a shooting scene is occurring from image data including an entire basketball court and capturing video including the shooting scene. Note that the use of a basketball as the imaging target in the first embodiment is merely for the purpose of specific description and does not limit the imaging target of the present invention. Furthermore, in the following description, players will be numbered 109 and balls will be numbered 110. When describing multiple objects, only numbers will be used, and when describing an individual object, a number with an additional letter, such as "109a," will be used.

[0013] (FIG. 1: Image processing system configuration diagram) FIG. 1 is a configuration diagram of an image processing system including an image processing device 103 that executes shooting scene judgment processing according to the first embodiment. This image processing system has the Internet 100, a local network 101, an overhead camera 102, an image processing device 103, a target camera 104, a learning server 105, a data collection server 106, and a client terminal 107. This image processing system also includes a position sensor receiving device 115 as appropriate.

[0014] (Internet 100 / Local Network 101) The Internet 100 is a network for communicating required information with the learning server 105, data collection server 106, etc. The Internet 100 is also capable of communicating with electronic devices connected to the local network 101 via the local network 101. The local network 101 is a network for the overhead camera 102, image processing device 103, target camera 104, and client terminal 107 to communicate required information with each other. The local network 101 is also capable of communicating required information with electronic devices connected to the Internet 100 via the Internet 100.

[0015] (Overhead camera 102) The overhead camera 102 is a camera for capturing images of the entire basketball court 108. The overhead camera 102 transmits the captured images to the image processing device 103 via a video cable 113. In this embodiment, the overhead camera 102 is assumed to capture images at 30 frames per second (30 fps), but this does not limit the frame rate. The imaging optical axis 114 of the overhead camera 102 is positioned so as to be parallel to the halfway line of the basketball court 108, and the imaging angle of view is set to an "angle of view θ" that includes the side line in front of the basketball court 108.

[0016] (Image processing device 103 / focused camera 104) The image processing device 103 performs object detection of the ball 110, estimation of the posture of the player 109, determination of whether or not a shooting scene has occurred, etc. on the captured image received from the overhead camera 102, and controls the focused camera 104 when capturing an image of a shooting scene. The focused camera 104 is equipped with a pan-tilt mechanism, a zoom mechanism, etc., and is a camera for capturing an image of a shooting scene in accordance with control from the image processing device 103. In this embodiment, the focused camera 104 is assumed to capture images at 60 frames per second (60 (fps)), but this does not limit the frame rate.

[0017] (Learning Server 105) The learning server 105 is a server for generating "trained models" that are used when the image processing device 103 performs object detection and pose estimation. The learning server 105 has the following two trained models: a trained object detection model (hereinafter referred to as the "object detection model") that detects the ball 110 from the image data captured by the overhead camera 102, and a trained pose estimation model (hereinafter referred to as the "pose estimation model") that estimates the joint coordinates of the player 109 from the attention area image signal 601.

[0018] (Data collection server 106 / client terminal 107 / position sensor receiver 115) The data collection server 106 is a server that collects and stores training data for machine learning required for the learning server 105 to generate a trained model. The client terminal 107 is a terminal that operates electronic devices connected to the Internet 100 and the local network 101 or instructs data transmission and reception. The type of client terminal 107 is not particularly limited and can be realized, for example, by a general-purpose PC. The position sensor receiver 115 reads ball position information from a position sensor attached to the ball 110 and measures the position of the ball 110 on the basketball court 108.

[0019] The above has described an example of the configuration of a system that executes the process of determining a shooting scene according to the first embodiment. This system reads the position of the ball 110, the posture of the player 109, etc. from a basketball game, detects a shooting scene based on the conditions of the ball 110 and the player 109, and automatically captures an image of the shooting scene.

[0020] (Fig. 2: Hardware configuration diagram of image processing system) (Hardware configuration of overhead camera 102) Fig. 2 is a hardware configuration diagram of each device that makes up the system in Fig. 1. The overhead camera 102 has a CPU 201, RAM 202, ROM 203, video engine 204, NIC 205, video I / F 206, image sensor 207, zoom driver 208, and tilt sensor 227. These are connected to a system bus 200 and configured to be able to send and receive required information to and from each other.

[0021] The CPU 201 controls the overhead camera 102. The CPU 201 controls each unit described below and performs control operations such as capturing images using the image sensor 207 and transmitting captured image data via the video I / F 206. The RAM 202 is a volatile memory that allows stored information to be rewritten, and has a work area for a program that describes the control details of the overhead camera 102. The RAM 202 is implemented, for example, by a volatile memory (such as a DRAM) that uses semiconductor elements. The ROM 203 is a non-volatile memory that non-volatilely stores a program that describes the control details of the overhead camera 102 and various parameters. The ROM 203 is implemented, for example, by a flash memory. When the user turns on the overhead camera 102, the CPU 201 reads the program from the ROM 203, expands it into the RAM 202, and executes it, thereby starting control of the overhead camera 102.

[0022] The video engine 204 is an image processing engine that converts the electric charges obtained from the image sensor 207 into image signals. The NIC 205 is a network interface card (NIC) that is used by the overhead camera 102 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series. The video I / F 206 is an interface that connects the overhead camera 102 and the image processing device 103 via a video cable 113.

[0023] The overhead camera 102 transmits imaging data to the image processing device 103 via the video I / F 206. The type of the video I / F 206 is not particularly limited as long as it is an interface that can communicate with the image processing device 103. For example, HDMI (registered trademark) (HIGH-DEFINITION MULTIMEDIA INTERFACE) is an example. The overhead camera 102 may also be a network camera, and may transmit imaging data to the image processing device 103 via the NIC 205.

[0024] The image sensor 207 receives light emitted from the object to be imaged and converts the brightness and color of the light into an electric charge. Examples of this sensor include a photodiode, a CCD (Charge Coupled Device) sensor, and a CMOS (Complementary Metal Oxide Semiconductor) sensor. The zoom driver 208 is a unit for changing the degree of "enlargement / reduction" of the zoom lens provided in the overhead camera 102. The tilt sensor 227 is a sensor for detecting the angle at which the main body of the overhead camera 102 is tilted relative to the horizontal.

[0025] (Hardware Configuration of Image Processing Device 103) The image processing device 103 includes a CPU 210, a RAM 211, a ROM 212, a GPU 213, a video I / F 214, a NIC 215, and an HDD 216. These are connected to a system bus 209 and configured to be able to send and receive required information to and from each other.

[0026] The CPU 210 controls the image processing device 103. The CPU 210 controls each unit described below and performs operations according to imaging data received from the video I / F 206 and communication data received from the NIC 215. The RAM 211 is a volatile memory in which stored information can be rewritten, and has a work area for a program in which the control details of the image processing device 103 are described. The RAM 211 is realized, for example, by a volatile memory (DRAM, etc.) using semiconductor elements. The ROM 212 is a non-volatile memory that non-volatilely stores a program in which the control details of the image processing device 103 are described and various parameters. When the user turns on the image processing device 103, the CPU 210 reads the program from the ROM 212, expands it into the RAM 211, and executes it. This starts control of the image processing device 103. The ROM 212 is realized, for example, by a flash memory or the like.

[0027] The GPU 213 is a unit that performs parallel arithmetic processing of data. When performing a large number of multiply-and-accumulate operations in inference processing, the GPU 213 can execute the processing faster and more efficiently than the CPU 210. The GPU 213 generally uses an LSI called a "Graphics Processing Unit," but it may also be possible to realize functions equivalent to those of an LSI using reconfigurable logic circuit hardware called an FPGA.

[0028] The video I / F 214 is an interface for connecting the overhead camera 102 and the image processing device 103 via the video cable 113. The image processing device 103 receives image data from the overhead camera 102 via the video I / F 214. The video I / F 214 may be any interface capable of communicating with the overhead camera 102, and the type of interface is not particularly limited. For example, a USB (Universal Serial Bus) interface is exemplified. The NIC 215 is a network interface card, and is used by the image processing device 103 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series.

[0029] The HDD 216 is a storage device that stores image data received from the overhead camera 102 and a "trained model" received from the learning server 105. In this embodiment, the HDD 216 is a hard disk drive (HDD) that uses a magnetic storage method, but the type of HDD 216 is not limited to this. For example, the HDD 216 may be an external storage device such as a solid state drive (SSD) that uses semiconductor elements.

[0030] (Hardware Configuration of Attention Camera 104) The attention camera 104 has a CPU 218, RAM 219, ROM 220, a video engine 221, a NIC 222, a video I / F 223, an image sensor 224, a zoom driver 225, and a pan / tilt driver 226. These are connected to a system bus 217 and configured to be able to transmit and receive required information to and from each other.

[0031] The CPU 218 controls the target camera 104. The CPU 218 controls each unit described below, and executes operations such as capturing images using the image sensor 224 and receiving communication data using the NIC 222. The RAM 219 is a rewritable memory, and has a work area for a program describing the control details of the target camera 104. The RAM 219 is realized, for example, by a volatile memory using semiconductor elements. The ROM 220 is a non-volatile memory, and non-volatilely stores a program describing the control details of the target camera 104 and various parameters. When the user turns on the power to the target camera 104, the CPU 218 reads the program from the ROM 220, loads it into the RAM 219, and executes it. This starts control of the target camera 104. The ROM 220 is realized, for example, by a flash memory.

[0032] The video engine 221 is an image processing engine that converts the electric charges obtained from the image sensor 224 into image signals. The NIC 222 is a network interface card that allows the target camera 104 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series. The video I / F 223 is an interface for transmitting image data from the target camera 104. When using this system to capture a basketball game, the image data transmitted from the video I / F 223 becomes the desired highlight scene image. For this reason, the video I / F 206 is used to connect an external storage device (not shown) for storing the highlight scene image. Note that the type of the video I / F 206 is not particularly limited as long as it is an interface that can communicate with the image processing device 103. For example, an HDMI interface is an example.

[0033] The image sensor 224 is a sensor that receives light emitted by the object to be imaged and converts the brightness and color of that light into an electric charge. Examples of this sensor include a photodiode, a CCD sensor, and a CMOS sensor. The zoom driver 225 is a unit that changes the degree of expansion and contraction of the zoom lens provided in the target camera 104. The pan / tilt driver 226 is a unit that changes the degree of rotation of the pan rotation mechanism and tilt rotation mechanism provided in the target camera 104.

[0034] (Hardware configuration of the learning server 105) The learning server 105 has a CPU 229, RAM 230, ROM 231, GPU 232, HDD 233, and NIC 234, which are configured to be able to send and receive required information to and from each other via a system bus 228. The CPU 229 controls the learning server 105. The CPU 229 controls each unit described below and performs operations according to communication data received from the NIC 234. The RAM 230 is a rewritable memory and has a working area for a program describing the control details of the learning server 105. The RAM 230 is realized, for example, by a volatile memory using semiconductor elements. The ROM 231 is a nonvolatile memory and nonvolatilely stores a program describing the control details of the learning server 105, various parameters, etc. When the user turns on the power to the learning server 105, the CPU 229 reads the program from the ROM 231, expands it into the RAM 230, and executes it, thereby starting control of the learning server 105. The ROM 231 is realized by, for example, a flash memory.

[0035] The GPU 232 is a unit that performs parallel arithmetic processing of data, and performs efficient calculations by processing more data in parallel. Therefore, when performing learning multiple times using a learning model such as deep learning, it is effective for the GPU 232 to perform the processing. In this embodiment, the learning server 105 uses the GPU 232 in addition to the CPU 229 for the learning process. Specifically, when executing a learning program including a learning model, the CPU 229 and the GPU 232 cooperate to perform calculations to perform learning. Note that the learning process may be performed by only either the CPU 229 or the GPU 232. The GPU 232 is realized, for example, by an LSI called a "Graphics Processing Unit," but functions equivalent to those of an LSI may also be realized by reconfigurable logic circuit hardware called an FPGA.

[0036] The HDD 233 is a storage device that non-volatilely stores training data received from the data collection server 106, trained models trained from the training data, and the like. In this embodiment, the HDD 233 is a hard disk drive that uses a magnetic storage method, but the type of HDD 233 is not limited to this. For example, the HDD 233 may be an external storage device such as a solid-state drive that uses semiconductor elements. The NIC 234 is a network interface card that is used by the training server 105 to communicate with other devices via the Internet 100. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series.

[0037] (Hardware Configuration of Data Collection Server 106) The data collection server 106 has a CPU 236, RAM 237, ROM 238, GPU 239, HDD 240, and NIC 241, which are configured to be able to send and receive required information to and from each other via a system bus 235. The CPU 236 controls the data collection server 106. The CPU 236 controls each unit described below and performs operations according to communication data received from the NIC 241. The RAM 237 is a rewritable memory, and has a working area for a program in which the control details of the data collection server 106 are written. The RAM 237 is realized, for example, by a volatile memory using a semiconductor element.

[0038] The ROM 238 is a non-volatile memory that stores a program describing the control details of the data collection server 106, various parameters, and the like in a non-volatile manner. When the user turns on the power to the data collection server 106, the CPU 236 reads the program from the ROM 238, expands it into the RAM 237, and executes it, thereby starting control of the data collection server 106. The ROM 238 is realized, for example, by a flash memory. The GPU 239 is a unit that performs parallel arithmetic processing of data. The GPU 239 is generally realized by an LSI called a "Graphics Processing Unit," but functions equivalent to those of an LSI may also be realized by reconfigurable logic circuit hardware called an FPGA.

[0039] The HDD 240 is a storage device that stores teacher data requested by the learning server 105. In this embodiment, the HDD 240 is a hard disk drive that uses a magnetic storage method, but the type of HDD 240 is not particularly limited, and a storage device such as a solid state drive that uses semiconductor elements may be used as the HDD 240. The NIC 241 is a network interface card that is used by the data collection server 106 to communicate with other devices via the Internet 100. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series.

[0040] (Hardware Configuration of Client Terminal 107) The client terminal 107 has a CPU 243, RAM 244, ROM 245, GPU 246, HDD 247, NIC 248, display unit 249, and input unit 250. These are configured to be able to send and receive required information to and from each other via a system bus 242. The CPU 243 is responsible for controlling the client terminal 107. The CPU 243 controls each unit described below, and performs operations in accordance with input data input from the input unit 250 and communication data received from the NIC 241. The RAM 244 is a rewritable memory, and has a working area for a program in which the control contents of the client terminal 107 are written. The RAM 244 is realized, for example, by a volatile memory using a semiconductor element.

[0041] The ROM 245 is a non-volatile memory that stores a program describing the control details of the client terminal 107 and various parameters in a non-volatile manner. When the user turns on the power of the client terminal 107, the CPU 243 reads the program from the ROM 245, deploys it in the RAM 244, and executes it. This starts control of the client terminal 107. The ROM 245 is realized, for example, by a flash memory. The GPU 246 is a unit that performs parallel arithmetic processing of data, and is used to display image data received from the image processing device 103 via the NIC 248 on a display unit 249. The GPU 246 is generally realized by an LSI called a "Graphics Processing Unit," but functions equivalent to those of an LSI may also be realized by reconfigurable logic circuit hardware called an FPGA.

[0042] The HDD 247 is a storage device that nonvolatilely stores image data received from the image processing device 103 via the NIC 248. In this embodiment, the HDD 247 is a hard disk drive that uses a magnetic storage system, but the type of HDD 247 is not particularly limited. Storage devices such as solid-state drives that use semiconductor elements may also be used as the HDD 247. The NIC 248 is a network interface card that allows the client terminal 107 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series. The display unit 249 is a display device that displays image data received from the image processing device 103 via the NIC 248. The display unit 249 is realized, for example, by an LCD (liquid crystal display). The input unit 250 is an input device that accepts operations and instructions from users of the system. The input unit 250 is realized by, for example, a pointing device such as a mouse and an input device such as a keyboard.

[0043] (Fig. 3: Functional configuration diagram of image processing system) Fig. 3 is a functional configuration diagram of software realized by the cooperation of the hardware and programs shown in the hardware configuration diagram of Fig. 2. Note that general-purpose software such as the OS is omitted from this software functional configuration.

[0044] <Functional Configuration of the Overhead Camera 102> The overhead camera 102 has an imaging unit 800, an operation unit 801, and a communication unit 802. These functional units are realized by the CPU 201 loading a program stored in the ROM 203 into the RAM 202 and executing it. The imaging unit 800 acquires imaging data with an angle of view θ that includes the entire basketball court 108 by controlling the video engine 204, image sensor 207, and zoom drive unit 208 by the CPU 201. The operation unit 801 performs operations such as starting and stopping imaging in accordance with instructions received from the client terminal 107 by controlling the NIC 205 by the CPU 201. The communication unit 802 transmits imaging data and camera information to the image processing device 103 by controlling the NIC 205 and video I / F 206 by the CPU 201.

[0045] <Functional Configuration of Image Processing Device 103> The image processing device 103 has a communication unit 803, a processing / calculation unit 804, an inference unit 805, and a data storage unit 806. These functional units are realized by the CPU 210 expanding a program stored in the ROM 212 into the RAM 211 and executing it. The communication unit 803 receives data from the overhead camera 102 or the learning server 105, or transmits data to the target camera 104, by the CPU 210 controlling the video I / F 214 and the NIC 215. The processing / calculation unit 804 performs image processing such as reduction and Hough transform on the imaging data received by the CPU 210 from the overhead camera 102. The processing / calculation unit 804 also cuts out an image from the imaging data in accordance with the inference result of the inference unit 805 (described later). The processing / calculation unit 804 also determines whether a shooting scene is present in the image based on the cut-out image, or calculates a control value for the target camera 104.

[0046] (Inference of Object Detection Model, Inference of Pose Estimation Model) The inference unit 805 performs inference processing of object detection and pose estimation on the "image data" received from the overhead camera 102 using the "trained model" trained by the learning server 105, by the CPU 210 controlling the GPU 213. Note that the "inference of object detection model" takes image data including the ball 110 as input data and the "rectangular coordinates" circumscribing the ball 110 in the "image data" described above as output data. Also, the "inference of posture estimation model" takes image data including the player 109 as input data and the "joint coordinates" of the player 109 in the "image data" described above as output data. The data storage unit 806 stores and manages the trained models received from the learning server 105, by the CPU 210 controlling the HDD 216.

[0047] <Functional Configuration of the Attention Camera 104> The attention camera 104 has an imaging unit 807, a drive control unit 808, and a communication unit 809. These functional units are realized by the CPU 218 loading a program stored in the ROM 220 into the RAM 219 and executing it. The imaging unit 807 captures an image with an angle of view that includes the player 109 and the ball 110 by the CPU 218 controlling the video engine 221, the image sensor 224, and the zoom driving unit 225. The drive control unit 808 controls the amount of rotation of the pan / tilt mechanism and the amount of expansion / contraction of the zoom lens in accordance with instructions received from the image processing device 103 by the CPU 218 controlling the NIC 205 and the video I / F 206. The communication unit 809 receives each control value of the pan / tilt driving unit 226 and the zoom driving unit 225 from the image processing device 103 by the CPU 218 controlling the NIC 205 and the video I / F 206.

[0048] <Functional Configuration of Data Collection Server 106> The data collection server 106 has a data collection / provision unit 813 and a data storage unit 814. These functional units are realized by the CPU 236 loading a program stored in the ROM 238 into the RAM 237 and executing it.

[0049] (Teacher Data for "Object Detection" and "Pose Estimation") The data collection / provision unit 813 has a function of collecting "teacher data" requested by the learning server 105 from the Internet 100 and a function of transmitting the collected "teacher data" to the learning server 105, by the CPU 236 controlling the NIC 241. In this embodiment, an image including the ball 110 and the "position coordinates" of the ball 110 in the image are used as the "teacher data" for object detection. Furthermore, a "learning dataset for pose estimation" consisting of a combination of an image showing the entire body of a person and the position coordinates of each joint of the person shown in the image is used as the "teacher data" for pose estimation. The data storage unit 814 stores and manages the teacher data collected from the Internet 100, by the CPU 236 controlling the HDD 240.

[0050] <Functional Configuration of Client Terminal 107> The client terminal 107 has a display control unit 815, an operation unit 816, and a communication unit 817. These functional units are realized by the CPU 243 loading a program stored in the ROM 245 into the RAM 244 and executing it.

[0051] The display control unit 815 displays image data received from the image processing device 103 on the display unit 249 by the CPU 243 controlling the display unit 249. The operation unit 816 accepts operations and instructions from a user who uses this image processing system by the CPU 243 controlling the input unit 250. The communication unit 817 transmits control commands to devices connected to the Internet 100 and the local network 101 by the CPU 243 controlling the NIC 248.

[0052] <Functional Configuration of Learning Server 105> The learning server 105 has a learning data generation unit 810, a learning unit 811, and a data storage unit 812. These functional units are realized by the CPU 229 expanding a program stored in the ROM 231 into the RAM 230 and executing it.

[0053] The training data generation unit 810 generates training data by performing data augmentation on the "teacher data" received by the CPU 229 from the data collection server 106. In "data augmentation," first, an image transformation process such as affine transformation is performed on the images of the teacher data and the correct answer data corresponding to those images. This generates an image that differs from the original image. Then, information on the correct answer data is also transformed in accordance with the image transformation, increasing the amount of training data.

[0054] The learning unit 811 performs learning of a "learning model for object detection" and a "learning model for pose estimation" by the CPU 229 controlling the GPU 232, and generates trained models for object detection and pose estimation. Specific examples of machine learning algorithms used in the learning models include nearest neighbor methods, naive Bayes methods, decision trees, and support vector machines. Another example of a learning model is deep learning, which generates features and connection weighting coefficients for learning using a neural network. Any of the above-mentioned algorithms can be appropriately selected and applied to this embodiment.

[0055] <"Error Detection Unit" and "Update Unit" Provided in Learning Unit 811> The learning unit 811 may be configured to include an "error detection unit" and an "update unit." The "error detection unit" calculates the error between the training data and output data output from the output layer of the neural network in response to input data input to the input layer. Furthermore, the "error detection unit" may use a loss function to calculate the error between the training data and output data from the neural network. The "update unit" updates the connection weighting coefficients between the nodes of the neural network based on the error calculated by the error detection unit so as to reduce the error. The "update unit" updates the connection weighting coefficients, etc., using, for example, backpropagation. Here, backpropagation is a technique for adjusting the connection weighting coefficients, etc., between the nodes of each neural network so as to reduce the error.

[0056] The data storage unit 812 stores and manages the training data received from the data collection server 106 and the trained model trained by the training unit 811 by the CPU 229 controlling the HDD 233 .

[0057] <Prerequisites for Processing Flow> Next, before describing the specific processing flow of this image processing system, the prerequisite conditions and operations for the processing flow will be described. The processing flow of this image processing system is based on the premise that the "object detection model" and "pose estimation model" trained by the learning server 105 are already stored in the HDD 216 of the image processing device 103. This eliminates the need for the image processing device 103 to request trained models each time the image processing device 103 executes the processing of this embodiment, and allows the image processing device 103 to execute each inference process using the "trained models" stored in the HDD 216.

[0058] Here, a method for preparing a trained model will be described using an "object detection model" as an example. The basic operations of the client terminal 107, the data collection server 106, and the image processing device 103 are the same when preparing a trained model.

[0059] Before starting the processing of this embodiment, the CPU 243 of the client terminal 107 requests the CPU 229 of the learning server 105 to do the following via the NIC 248 of the client terminal 107. That is, the CPU 243 of the client terminal 107 requests the generation of an "object detection model" that detects the ball 110 from the reduced image signal 604 of the overhead image signal 600.

[0060] Next, CPU 229 of learning server 105 requests CPU 236 of data collection server 106 to collect "teacher data" including ball 110 via NIC 234 of learning server 105. Next, CPU 236 of data collection server 106 collects "teacher data" including ball 110 via Internet 100.

[0061] Next, the CPU 236 of the data collection server 106 transmits the collected training data to the learning server. The CPU 229 of the learning server 105 generates training data by performing data augmentation on the training data received from the data collection server 106. Next, the CPU 229 of the learning server 105 controls the GPU 232 of the learning server 105 to generate an "object detection model" based on the generated training data. Next, the CPU 229 of the learning server 105 transmits the generated trained model to the image processing device 103. Then, the CPU 210 of the image processing device 103 stores the "trained model" received from the learning server 105 in the HDD 216.

[0062] The above describes a method for preparing an "object detection model" for detecting the ball 110. Note that a method for preparing a "posture estimation model" involves preparing "image data" including the player 109 and "joint coordinate data" of the player 109 as "teacher data" in the above-described method, and performing learning using a network model for posture estimation.

[0063] (Figs. 4A to 4C: Flowcharts showing the processing procedures of the overhead camera, image processing device, and camera of interest / Figs. 5A to 5C: Explanatory diagrams of the overhead image signal, area of ​​interest image signal, and reduced image signal) Next, specific processing of this system will be described with reference to Figs. 4A to 5C. Note that Figs. 5A, 5B, and 5C are explanatory diagrams that schematically show the content explained in Figs. 4A, 4B, and 4C. Fig. 4A is a flowchart showing processing executed by the overhead camera 102, and shows processing executed by the imaging unit 800, operation unit 801, and communication unit 802 of Fig. 3. Below, processing of the overhead camera 102 will be described with reference to Fig. 4A.

[0064] (FIG. 4A: Processing of Overhead Camera 102) First, CPU 201 reads software for imaging unit 800 from ROM 203, loads it into RAM 202, and executes step S300. Specifically, in step S300, CPU 201 controls video engine 204 and image sensor 207 to generate imaging data (hereinafter referred to as "overhead image signal 600") with an angle of view θ that includes the entire basketball court 108. Note that FIG. 5A is an explanatory diagram of overhead image signal 600. CPU 201 then stores the generated overhead image signal 600 in RAM 202 and proceeds to step S301.

[0065] Next, the CPU 201 reads the software of the communication unit 802 from the ROM 203, loads it into the RAM 202, and executes step S301. Specifically, in step S301, the CPU 201 transmits the overhead image signal 600 stored in the RAM 202 in step S300 to the image processing device 103 via the video I / F 206. Then, the CPU 201 proceeds to step S302.

[0066] Next, the CPU 201 reads out the software of the operation unit 801 from the ROM 203, loads it into the RAM 202, and executes step S302. Specifically, in step S302, the CPU 201 determines whether an instruction to end image capture has been received from an operation on a user interface (UI) (not shown) of the overhead camera 102 or from the client terminal 107 via the NIC 205. If it is determined that an instruction to end image capture has been received (YES), the CPU 201 ends image capture. On the other hand, if it is determined that an instruction to end image capture has not been received (NO), the CPU 201 returns the process to step S300 and repeatedly executes the processes of steps S300 and S301.

[0067] (FIG. 4B: Processing of Image Processing Device 103) FIG. 4B is a flowchart showing the processing of the image processing device 103, and shows the processing executed by the communication unit 803, processing / calculation unit 804, inference unit 805, and data storage unit 806 of FIG. 3. The processing executed by the image processing device 103 will be described below with reference to FIG. 4B. First, the CPU 210 reads software for the communication unit 803 from the ROM 212, loads it into the RAM 211, and executes step S400. Specifically, in step S400, the CPU 210 determines whether the overhead-view image signal 600 (see FIG. 5A ) transmitted from the overhead-view camera 102 in step S301 has been received via the video I / F 214. If it is determined that the overhead-view image signal 600 has been received (YES), the CPU 210 stores the overhead-view image signal 600 in the RAM 211 and proceeds to step S401. On the other hand, if it is determined that the overhead image signal 600 has not been received (NO), the CPU 210 waits for reception of the overhead image signal 600 in step S400.

[0068] Next, the CPU 210 reads the software of the processing / calculation unit 804 from the ROM 212, loads it into the RAM 211, and executes step S401. Specifically, in step S401, the CPU 210 reads the overhead-view image signal 600 stored in the RAM 211 in step S400, and generates a reduced image signal 604 (see FIG. 5C ) by reducing the overhead-view image signal 600. Note that the reduction ratio (hereinafter referred to as the "reduction ratio") is a ratio suitable for input to a trained model used in subsequent inference processing, but may be any ratio specified by the user. The CPU 210 then stores the generated reduced image signal 604 and the reduction ratio in the RAM 211 and proceeds to step S402.

[0069] Next, the CPU 210 reads the software of the inference unit 805 from the ROM 212, loads it into the RAM 211, and executes step S402. Specifically, in step S402, the CPU 210 reads the "object detection model" from the HDD 216 and stores it in the RAM 211. Thereafter, the CPU 210 controls the GPU 213 to read the reduced image signal 604 stored in the RAM 211 in step S401, and inputs the reduced image signal 604 into the "object detection model" stored in the RAM 211 to obtain "ball coordinates 602c." Thereafter, the CPU 210 stores the detected ball coordinates 602c in the RAM 211 and proceeds to step S403. Note that if the ball coordinates cannot be detected, the ball coordinates detected immediately before may be used as the next ball coordinates. Alternatively, if the ball coordinates cannot be detected, the movement trajectory may be predicted from the time series of the ball coordinates detected immediately before, and predicted coordinates calculated from the prediction result may be used as the ball coordinates.

[0070] Next, CPU 210 reads out software for processing / calculation unit 804 from ROM 212, loads it into RAM 211, and executes step S403. Specifically, in step S403, CPU 210 reads out ball coordinates 602c and the reduction ratio stored in RAM 211 in step S402, and uses the reduction ratio to convert ball coordinates 602c in reduced image signal 604 into ball coordinates 602a in overhead image signal 600. CPU 210 then stores the converted ball coordinates 602a in RAM 211. Based on ball coordinates 602a stored in RAM 211, CPU 210 then crops an area including ball 110 and its vicinity from overhead image signal 600 to generate area-of-interest image signal 601 (see FIG. 5B ). Note that in the process of step S403, the process of determining the cropping size is a characteristic process of the present invention, and will be described in detail later with reference to FIGS. 6 and 7 . Thereafter, the CPU 210 stores the generated attention area image signal 601 in the RAM 211 and moves the process to step S404.

[0071] Here, the attention area image signal 601 is an image corresponding to the "attention area size," which is the size of the "attention area." The "attention area" is an area in the captured image that corresponds to the location (basketball court 108) being imaged by the overhead camera 102. The CPU 210 (determination unit) determines the "attention area" based on the distortion information and the attention position (ball coordinates 602). The CPU 210 also determines the "attention area size," which corresponds to the actual size, based on the "distortion information" and the "attention position." As will be made clear later, the "distortion information" is information indicating the distortion of the shape of the first object (basketball court 108) in the captured image captured by the overhead camera 102 from a known shape, and the CPU 210 (distortion information generation unit) generates the distortion information. The attention area image signal 601 includes objects of interest (second and third objects).

[0072] That is, in order to fit the horizontally elongated rectangular basketball court 108 within the angle of view θ of the overhead camera 102, the captured image is distorted into a quadrangular shape (e.g., a trapezoid) as shown in FIGS. 5A and 11A . Therefore, as an example (third embodiment), the CPU 210 generates "distortion information" based on a projection transformation matrix between coordinates indicating the vertices of the first object (basketball court 108) and coordinates indicating the vertices of a known shape of the first object. The CPU 210 then calculates a "position of interest (ball coordinates 602)" in the captured image. Next, the CPU 210 determines a "region of interest" based on the "distortion information" and the calculated "position of interest." Next, the CPU 210 determines a "region of interest size," which is the size of the region of interest. An image corresponding to the "region of interest size" becomes a "region of interest image (region of interest image signal 601)." In this embodiment of the present invention, the "region of interest" specifically refers to an image capture area for capturing an image of a shooting scene, the ball, etc. The CPU 210 can also determine the “area of ​​interest” based on the position information (joint coordinates) of the second object (player 109 ) whose position is estimated in the image captured by the overhead camera 102 .

[0073] After executing step S403, the CPU 210 reads the software of the processing / calculation unit 804 from the ROM 212, loads it into the RAM 211, and executes step S404. Specifically, in step S404, the CPU 210 reads the attention area image signal 601 stored in the RAM 211 in step S403, and scales the attention area image signal 601 to match the input size to the pose estimation model. The scaling factor is a factor appropriate for input to a trained model used in inference processing, but may be any factor specified by the user. The CPU 210 then stores the scaled attention area image signal 601 in the RAM 211 and proceeds to step S405. By scaling (enlarging or reducing) the attention area image signal 601, the resolution of the attention area image signal 601 changes.

[0074] Next, the CPU 210 reads out the software of the inference unit 805 from the ROM 212, loads it into the RAM 211, and executes step S405. Specifically, in step S405, the CPU 210 reads out the "posture estimation model" from the HDD 216 and stores it in the RAM 211. Thereafter, the CPU 210 controls the GPU 213 to read out the attention area image signal 601 stored in the RAM 211 in step S404, and detects the "joint coordinates 603 (see FIG. 5B )" of the player 109a using the "posture estimation model" stored in the RAM 211. Note that if multiple players 109 are shown in the attention area image signal 601, the joint coordinates 603 are detected for each player 109. The CPU 210 stores the detected joint coordinates 603 in the RAM 211 and proceeds to step S406.

[0075] Next, the CPU 210 reads out the software of the processing / calculation unit 804 from the ROM 212, loads it into the RAM 211, and executes step S406. Specifically, in step S406, the CPU 210 determines whether or not the player 109a is shooting, based on the ball coordinates 602a stored in the RAM 211 in step S403 and the joint coordinates 603 stored in the RAM 211 in step S405. Specifically, the CPU 210 converts the coordinate system of the ball coordinates 602a from the coordinate system on the overhead image signal 600 to the coordinate system on the attention area image signal 601, and stores the converted ball coordinates 602b in the RAM 211.

[0076] The CPU 210 then reads out the ball coordinates 602b and joint coordinates 603 stored in the RAM 211 and compares the relative positions of the coordinates. For example, if the CPU 210 determines that the coordinate value of the elbow joint of the player 109a is higher than the coordinate value of the shoulder joint, and that the coordinate value of the wrist joint is higher than the coordinate value of the elbow joint, it determines that the player 109a is in a state where his / her arm is stretched upward. Here, a "higher coordinate value" means a larger coordinate value on the vertical coordinate axis.

[0077] Furthermore, if the coordinate value of the ball coordinates 602b is higher than the coordinate value of the wrist joint of player 109a, the CPU 210 determines that player 109a is shooting. This determination process is performed for each player 109 whose joint coordinates 603 are detected, and the joint coordinates 603 whose center of gravity of the ball coordinates 602b is closest to the coordinate value of the wrist joint of player 109 is selected. In this way, the CPU 210 can determine whether or not a shot is being taken.

[0078] In this embodiment, whether or not a shot has been taken is determined by comparing the ball coordinates with the player's joint coordinates, but this determination method is not limited to this. For example, a "decision tree model" may be prepared that inputs an image of the player 109a in action, such as the attention area image signal 601, and outputs a determination result of whether or not a shot has been taken, and the determination result from this decision tree model may be used to determine whether or not a shot has been taken. In this embodiment, the type of action to be determined is "shooting" because the purpose is to capture the player's shooting scene, and the type of action is not particularly limited. For example, the present invention can be applied to any action that can be identified from the player's posture, such as "passing" or "dribbling." If it is determined that the player 109a is taking a shot (YES), the CPU 210 proceeds to step S407. On the other hand, if it is determined that the player 109a is not taking a shot (NO), the CPU 210 proceeds to step S408.

[0079] Next, based on the processing result of step S406, the CPU 210 reads out the software of the processing / calculation unit 804 from the ROM 212, loads it into the RAM 211, and executes step S407. Specifically, in step S407, the CPU 210 calculates a control value for the pan / tilt driver 226 of the target camera 104 to orient the imaging direction of the target camera 104 toward the shooting scene. Here, the "control value" is a combination of a "control direction" and a "control speed" in this embodiment. For example, the control value for panning is expressed as "a speed of 10 degrees / second to the right," and the control speed of the pan is determined based on the difference between the current pan position and the target pan position. Note that the control value may also be a direct specification of the target pan position. Thus, in step S407, the CPU 210 generates a camera control signal required to orient the target camera 104 toward the shooting scene.

[0080] If the imaging direction and position of the overhead camera 102 differ significantly from the imaging direction and position of the target camera 104, the positions of the overhead camera 102 and the target camera 104 may be calibrated in advance. In this case, it is possible to convert the ball coordinates 602a and the player's joint coordinates 603 as viewed from the overhead camera 102 into the coordinate system of the target camera 104. This makes it possible to apply the present invention even when the imaging direction and position of the overhead camera 102 differ significantly from the imaging direction and position of the target camera 104. An example of a position calibration method is a method in which camera parameters, including information on the relative positions and orientations of the cameras, are calculated by simultaneously capturing images of a calibration pattern, such as a grid pattern, using cameras at two locations. Furthermore, the calculated camera parameters can be used to calculate corresponding points between the imaging coordinates of the overhead camera 102 and the target camera 104 based on epipolar geometry.

[0081] Next, the CPU 210 performs the following calculation based on the ball coordinates 602a stored in the RAM 211 and the attention area size (hereinafter referred to as the size of the attention area) of the attention area image signal 601. That is, the CPU 210 calculates a control value for the zoom driver 225 of the attention camera 104 to fit the ball 110 and the player 109 positioned nearby within the angle of view θ.

[0082] Here, a specific calculation process for the lens drive control values ​​will be described. First, the maximum and minimum values ​​of the focal length of the target camera 104 during image capture are defined. Specifically, before starting the processing of this embodiment, the user measures the focal length T (maximum value) at which the ball 110 and the entire body of the player 109 fit within the angle of view θ when the ball 110 is at the innermost position within the basketball court 108. The user also measures the focal length W (minimum value) at which the ball 110 and the entire body of the player 109 fit within the angle of view θ when the ball 110 is at the foremost position within the basketball court 108. Here, the "foremost position" and "foremost position" refer to the following: if the installation positions of the target camera 104 and the overhead camera 102, whose imaging optical axis is perpendicular to the sideline of the basketball court 108, are close to each other, the image captured by the target camera 104 can be approximated by the overhead image signal 600. In this case, the upper sideline in the overhead-view image signal 600 is the "rearmost position," and the lower sideline in the overhead-view image signal 600 is the "foremost position." Furthermore, if the imaging optical axis of the target camera 104 is installed perpendicular to the end lines of the basketball court 108 as shown in FIG. 1 , the left end line, as viewed from the front of the overhead-view camera 102, is the "rearmost position," and the right end line is the "foremost position." Note that, although the focal lengths T and W are set to focal lengths that allow the ball 110 and the entire body of the player 109 to fit within the angle of view θ, the focal lengths are not necessarily limited to this and may be set to focal lengths desired by the user.

[0083] The measured focal lengths T and W are input to client terminal 107, then transmitted to image processing device 103 and stored in RAM 211. Next, CPU 210 converts the installation position, innermost position, and foremost position of target camera 104 into the coordinate system of bird's-eye image signal 600 based on the result of position calibration. Next, CPU 210 sets the distance from target camera 104 to the innermost position on bird's-eye image signal 600 to "LL1," the distance from target camera 104 to the foremost position to "LL2," and the distance from target camera 104 to ball coordinates 602a to "LL3." CPU 210 can then calculate zoom target position Z, which is associated with the distance from focal length T to focal length W, using the following equation 1:

[0084]

[0085] The control direction of the lens drive is determined based on the current zoom position and the zoom target position Z. For example, if the current zoom position is 30 mm and the target position is 100 mm, the control direction is the "TELE direction." When the zoom drive unit 225 captures a close-up image of a subject, the telephoto direction is referred to as the "TELE direction." The control speed of the lens drive is determined based on the magnitude of the difference between the current zoom position and the zoom target position Z. For example, if the difference DD between the current zoom position and the target zoom position is less than 10 mm, the control speed is set to 0 mm / sec. If the difference DD is 10 mm or greater, the control speed is set to (DD x 10) mm / sec. These numerical values ​​are merely examples and do not limit the control content. For example, the control speed may be set to an upper limit.

[0086] In this embodiment, the zoom control value is calculated based on the installation position of the target camera 104, but the calculation process is not limited to this. For example, a new path for capturing the image captured by the target camera 104 may be provided to the image processing device 103, and object detection and pose estimation may be performed on the image captured by the target camera 104, as with the overhead image signal 600. Then, the positions of the player 109 and the ball 110 within the imaging angle of view θ of the target camera 104 may be estimated, and zoom control may be performed so that the ball 110 and the entire body of the player 109 located nearby fit within the angle of view θ. In this embodiment, the control value of the zoom driving unit 225 is set to a control value necessary to fit the ball 110 and the entire body of the player 109 located nearby within the angle of view θ. However, the control value of the zoom driving unit 225 is not limited to this. For example, when capturing an image of a player in a bust-up composition, the control value may be calculated as described below. That is, the joint coordinates 603 of the player 109 may be converted from the coordinate system on the attention area image signal 601 to the coordinate system on the overhead image signal 600, and the control values ​​may be set as necessary to fit the joint coordinates of the player 109 above the waist joint within the angle of view θ.

[0087] After executing step S407, the CPU 210 stores the control values ​​(hereinafter "camera control signals") of the pan / tilt driver 226 and the zoom driver 225 calculated in step S407 in the RAM 211, and proceeds to step S409. Based on the processing result of step S406, the CPU 210 reads software for the processing / calculation unit 804 from the ROM 212, loads it into the RAM 211, and executes step S408. Specifically, in step S408, the CPU 210 calculates a control value for the pan / tilt driver 226 of the target camera 104 to orient the imaging direction of the target camera 104 toward the ball coordinates 602a, based on the ball coordinates 602a stored in the RAM 211 in step S402. Furthermore, the CPU 210 calculates a control value for the zoom driver 225 of the target camera 104 to fit the ball 110 and a "predetermined area" around it within the angle of view θ, based on the ball coordinates 602a stored in the RAM 211.

[0088] Here, the "predetermined area" may be an area that includes all players whose joint coordinates 603 are detected in the attention area image signal 601. Alternatively, the "predetermined area" may be an area that includes only the player who is closest to the ball coordinates 602a. Note that the calculation process for the control values ​​of the pan / tilt driver 226 and the zoom driver 225 is the same as in step S407.

[0089] Note that in step S408, even if it is determined in step S406 that the scene is not a shooting scene, the imaging direction of the attention camera 104 is directed toward the ball coordinates 602a because the shot will occur starting from the ball position. This minimizes the amount of movement of the attention camera 104 when a shooting scene occurs, thereby reducing the time delay until the attention camera 104 finishes facing the shooting scene. The CPU 210 then stores the camera control signal calculated in step S407 in the RAM 211 and proceeds to step S409. That is, in both steps S407 and S408, the CPU 210 generates a camera control signal to direct the attention camera 104 in a direction according to the recognition result of whether or not the player 109a has taken a shot, based on the ball coordinates 602a and the joint coordinates 603 in step S406. In step S407, a camera control signal is generated to point the target camera 104 toward the shooting scene, and in step S408, a camera control signal is generated to point the target camera 104 toward the ball coordinates, so in either case, a camera control signal is generated to point the target camera 104 in the imaging direction toward the target area.

[0090] Next, the CPU 210 reads the software of the communication unit 803 from the ROM 212, loads it into the RAM 211, and executes step S409. Specifically, in step S409, the CPU 210 transmits the camera control signal stored in the RAM 211 in step S407 to the target camera 104 via the NIC 215.

[0091] (FIG. 4C: Processing of the target camera 104) FIG. 4C is a flowchart showing processing of the target camera 104, and is an explanatory diagram of processing of the imaging unit 807, drive control unit 808, and communication unit 809 in FIG. 3. The processing of the target camera 104 will be described below with reference to FIG. 4C.

[0092] First, the CPU 218 reads out the software of the communication unit 809 from the ROM 220, loads it into the RAM 219, and executes step S500. Specifically, in step S500, the CPU 218 determines whether or not the "camera control signal" transmitted from the image processing device 103 in step S409 has been received via the NIC 222. If it is determined that a camera control signal has been received (YES), the CPU 218 stores the received camera control signal in the RAM 219 and proceeds to step S501. On the other hand, if it is determined that a camera control signal has not been received (NO), the CPU 218 waits for reception in step S500.

[0093] Next, the CPU 218 reads out the software of the drive control unit 808 from the ROM 220, loads it into the RAM 219, and executes step S501. Specifically, in step S501, the CPU 218 controls the pan / tilt drive unit 226 and the zoom drive unit 225 based on the camera control signal stored in the RAM 219 in step S500.

[0094] Next, the CPU 218 reads out software for the imaging unit 807 from the ROM 220, loads it into the RAM 219, and executes step S502. Specifically, in step S502, the CPU 218 controls the video engine 221 and image sensor 224 to generate imaging data with an angle of view θ that includes the player 109 and ball 110 in the imaging direction. That is, in step S502, the imaging unit 807 images a shooting scene of the player 109 in the imaging direction and an area of ​​interest that includes the position of the ball in the imaging direction.

[0095] (FIG. 6: Flowchart showing attention area image generation processing / FIG. 7: Explanatory diagram) Next, the generation processing of the attention area image signal 601, which is a characteristic processing of this system, will be described with reference to FIGS. 6 and 7. Note that FIG. 7 is an explanatory diagram that schematically shows the content explained in FIG. 6. FIG. 6 is a flowchart showing detailed processing of step S403 (subroutine) in the flowchart showing the processing executed by the image processing device 103. The processing of the image processing device 103 will be described below with reference to FIG. 6.

[0096] First, in step S700, the CPU 210 performs a calculation based on a Hough transform on the overhead image signal 600 stored in the RAM 211. Here, the rectangle formed by the resulting straight lines is considered to be the court lines of the basketball court 108. Note that if the overhead image signal 600 contains court lines of a different color than the court lines of the basketball court 108, the color of the court lines of the basketball court 108 can be specified before the calculation. Also, if the floor color inside and outside the basketball court 108 is different, the boundary line of the floor color may be used as the court line. The CPU 210 stores the determined court lines of the basketball court 108 in the RAM 211 and proceeds to step S701. Note that the calculation of the court lines is not limited to this. For example, the court lines may be determined by calculating the boundaries of a rectangle by performing edge extraction processing, such as a SOBEL filter or the CANNY method, on the overhead image signal 600. Thus, in step S700, the CPU 210 performs a Hough transform on the overhead image signal 600 to extract the court lines of the basketball court 108.

[0097] Next, in step S701, the CPU 210 determines a quadrangle formed based on the court lines of the basketball court 108 stored in RAM 211 in step S700, and calculates the coordinate values ​​of the four vertices of the quadrangle (hereinafter referred to as "court coordinates"). As shown in FIG. 7, the four vertices of the quadrangle are designated as "A" for the upper left vertex, "B" for the lower left vertex, "C" for the lower right vertex, and "D" for the upper right vertex. In other words, the quadrangle formed based on the court lines is the quadrangle ABCD. The CPU 210 stores the calculated court coordinates in RAM 211 and proceeds to step S702.

[0098] Next, in step S702, the CPU 210 calculates an "upper base L1" consisting of line segment AD of the rectangle ABCD and a "lower base L2" consisting of line segment BC (see FIG. 7) based on the court coordinates stored in RAM 211 in step S701. Thereafter, the CPU 210 stores the calculated positions and lengths of the upper base L1 and lower base L2 in RAM 211 and proceeds to step S703.

[0099] Next, in step S703, the CPU 210 calculates a straight line that is parallel to the upper base L1 and the lower base L2 and passes through the center of gravity of the ball coordinates 602, based on the ball coordinates 602 stored in RAM 211 and the positions of the upper base L1 and the lower base L2 stored in RAM 211 in step S701. A case in which the upper base L1 and the lower base L2 on the overhead image signal 600 are not parallel will be described in the third embodiment. Next, the CPU 210 calculates a line segment PQ, which is the line segment L3, by defining "P" as the intersection with the line segment AB and "Q" as the intersection with the line segment CD (see FIG. 7 ). The CPU 210 stores the calculated length of the line segment L3 in RAM 211 and proceeds to step S704.

[0100] Next, in step S704, the CPU 210 calculates a "region of interest size," which is the image size of the region of interest image signal 601, from the ratio of the lengths of the line segments based on the length of the lower base L2 stored in RAM 211 in step S702 and the length of the line segment L3 stored in RAM 211 in step S703. Specifically, before starting the processing of this embodiment, the user captures an overhead image signal 600 when the player 109 is positioned on the lower sideline in FIG. 1 and measures the full-body size of the player 109 in the image. The CPU 210 then stores the full-body size input by the user via the client terminal 107 in RAM 211 and generates a "reference region of interest size" based on the full-body size. Here, the "reference region of interest size" is the length of one side of a rectangle forming the "region of interest," and the vertical length is twice the height of the player 109 on the overhead image signal 600, such that the ratio of the horizontal length to the vertical length is 1:1. The "standard attention area size" is not limited to this, and may be any size that allows trimming without cutting out the player taking the shot.

[0101] The CPU 210 then stores the generated "reference attention area size" in the RAM 211. The CPU 210 calculates the attention area size based on the reference attention area size stored in the RAM 211 and the ratio between the lower base L2 and the line segment L3. When the reference attention area size is "SB", the length of the lower base L2 is "L2", and the length of the line segment L3 is "L3", the attention area size S can be calculated using the following formula 2.

[0102]

[0103] Thereafter, the CPU 210 stores the attention area size in the RAM 211 and proceeds to step S705. Through the above process, the attention area size becomes smaller as the distance from the overhead camera 102 to the ball 110 increases. Conversely, the attention area size becomes larger as the distance from the overhead camera 102 to the ball 110 decreases.

[0104] Next, in step S705, the CPU 210 calculates "attention area coordinates 701 (see FIG. 7)" based on the ball coordinates 602 stored in RAM 211 and the attention area size stored in RAM 211 in step S704. Here, the "attention area coordinates 701" are the coordinates of the upper left vertex and the lower right vertex of a rectangular area (or a rectangular area obtained by expanding the rectangular area in all directions by a preset expansion value) centered on the center of gravity of the ball coordinates 602 and one side of which is the attention area size. The CPU 210 stores the calculated attention area coordinates 701 in RAM 211 and proceeds to step S706.

[0105] Next, in step S706, the CPU 210 trims the "area of ​​interest" from the overhead image signal 600 based on the area of ​​interest coordinates 701 stored in the RAM 211 in step S705, and generates an area of ​​interest image signal 601. The CPU 210 then stores the generated area of ​​interest image signal 601 in the RAM 211. The above describes the processing of step S403 in the processing of the image processing device 103. The subsequent processing continues to step S404 (see FIG. 4B ) executed by the image processing device 103.

[0106] The processing executed by the configuration of FIG. 1 in the first embodiment has been described above. As described above, a player (second object) in a captured image can be trimmed to a size based on the position of the ball (third object) on the basketball court 108 (first object). This reduces the possibility of the player being cut off because the trimming size is too small compared to the player's size, or conversely, the subject being crushed when reduced before recognition because the trimming size is too large, resulting in failure to perform the required recognition. The scope of application of the present invention is not limited to the above and can be applied to various scenes, such as sports other than basketball, live music performances, and lectures in auditoriums. For example, the present invention is suitable for sports such as soccer and tennis, where player actions such as shooting, serving, and smashing are recognized near the ball.

[0107] In the first embodiment, an example of performing automatic imaging using imaging data from a single overhead camera 102 has been described, but the present invention is not limited to this. The present invention can also be applied to a system using multiple overhead cameras. In this case, position calibration is performed on the multiple overhead cameras, and the imaging data is integrated based on the calibration results to generate a single imaging data set. The present invention can then be applied to the integrated imaging data set. This makes it easy to apply the present invention to sports and locations with large playing areas, such as soccer and rugby.

[0108] Furthermore, when applying the present invention to scenes in which subjects, such as a live music concert or lecture in an auditorium, perform a predetermined action, such as posing for a song or pointing at a blackboard, it may be difficult to detect court coordinates from the overhead image signal 600. In such cases, the present invention is applicable because the court lines can be calculated by placing target lines surrounding the stage. Alternatively, instead of court lines, markers of a specific color and shape may be attached to the real-world positions that represent the vertices of the rectangular stage, and the rectangle ABCD that forms the stage may be calculated by calculating these markers from the image data captured by the overhead camera 102. In other words, the CPU 210 may be configured to detect markers placed at various locations in the real world that correspond to the vertices of the basketball court 108 object, and acquire the object shape of the basketball court 108 by defining the area with each vertex as a vertex.

[0109] (Attachment of Position Sensor) In the first embodiment, in step S402 of the processing flow of the image processing device 103, the CPU 210 detects the position coordinates of the ball 110 using an object detection model. However, the ball detection process is not limited to this. For example, a position sensor may be attached to the ball 110, and the system configuration of FIG. 1 may be configured with a position sensor receiver 115 that receives a signal from the position sensor. Position calibration may be performed in advance between the coordinate system of the position sensor receiver 115 and the coordinate system of the overhead-view image signal 731 captured by the overhead camera 102. The position sensor receiver 115 receives the position of the position sensor attached to the ball 110. The position sensor receiver 115 transmits the position obtained from the position sensor to the image processing device 103. The image processing device 103 converts the position of the position sensor received from the position sensor receiver 115 into the coordinate system of the overhead-view image signal 600 and uses the converted position of the position sensor as ball coordinates 602. In this manner, a sensor for detecting the position of the ball 110 may be used. In this way, the CPU 210 may acquire position information of a sensor attached to the ball 110, which is a real-world object corresponding to the ball object (third object), to detect the position, and calculate the focus position based on the position information acquired by the sensor.

[0110] Furthermore, in this embodiment, the position coordinates of the ball 110 are detected as a position of interest from the overhead image signal 600 on the assumption that a highlight scene occurs near the ball 110. However, the present invention is not limited to this. For example, the position coordinates of a specific player 109, rather than the ball 110, may be detected as a position of interest. In this case, before starting the processing of this embodiment, image data including the player 109 may be input, and an object detection model may be used that performs inference by outputting the "rectangle coordinates" circumscribing the player 109 in the image data. Specifically, first, in step S402 of FIG. 4B , the CPU 210 of the image processing device 103 controls the GPU 213 to detect the position coordinates of the player 109 from the reduced image signal 604 using the prepared trained model. Next, in step S403 of FIG. 4B , the CPU 210 of the image processing device 103 generates a focus area image signal 601 from the overhead image signal 600 based on the detected position coordinates of the player 109, thereby similarly executing the processing of the present invention. This achieves the effect of capturing highlight scenes of only the focused player.

[0111] Furthermore, in this embodiment, in step S700, which is the processing of the image processing device 103, the CPU 210 calculates the court lines of the basketball court 108 using a Hough transform. However, the method of calculating the court lines is not limited to this. The present invention can be applied to any calculation method that can calculate the court lines. For example, a user interface (UI) that allows a user to specify vertices of the basketball court 108 may be used. Specifically, before starting the processing of this embodiment, the CPU 243 of the client terminal 107 operates the overhead camera 102 via the NIC 248 of the client terminal 107 to receive an overhead image signal 600. Next, the CPU 243 displays the received overhead image signal 600 on the display unit 249 of the client terminal 107. Next, the CPU 243 accepts operations on the input unit 250 of the client terminal. The user operating the client terminal 107 specifies four vertices of the basketball court 108 on the displayed overhead image signal 600. Next, the CPU 243 transmits the four coordinates on the designated overhead image signal 600 to the image processing device 103 via the NIC 248 .

[0112] Next, the present invention can be similarly applied if the CPU 210 determines that the court lines are a rectangle formed by four coordinates on the overhead image signal 600 received from the client terminal 107. This provides the effect of enabling the present invention to be applied even when the court lines cannot be calculated from the overhead image signal 600. Furthermore, the court lines may be detected by estimating a distorted shape on the overhead image signal 600 from the known shape of the basketball court 108, for example. Specifically, the CPU 210 of the image processing device 103 first estimates the distorted shape of the basketball court 108 based on the imaging angle of the overhead camera 102 relative to the basketball court 108. The CPU 210 of the image processing device 103 obtains a distorted rectangle (i.e., the trapezoidal basketball court 108 shown in FIG. 7 ) based on the estimation result.

[0113] Next, the CPU 210 of the image processing device 103 detects the position and shape of an object such that the distorted rectangle is similar to the calculation result obtained by performing edge calculation processing on the overhead image signal 600. The CPU 210 of the image processing device 103 can then similarly apply the present invention by acquiring the court lines based on the position and shape of the detected object. This makes it possible to calculate the court lines of the basketball court 108 and apply the present invention even when the court lines cannot be calculated from the overhead image signal 600. In this way, the "distortion information" is generated by the CPU 210 estimating the distorted shape, performing predetermined detection, etc.

[0114] Next, a second embodiment of the present invention will be described. In the second embodiment, a process for determining the shape of the basketball court 108 using camera information acquired from the overhead camera 102 and generating a region-of-interest image signal 601 will be described. Because the basic configuration of the image processing system is the same as in the first embodiment, a redundant description will be omitted and only the different process of the image processing device 103 will be described. Furthermore, the second embodiment is premised on the premise that the imaging optical axis 114 of the overhead camera 102 is parallel to the half line of the basketball court 108, as shown in FIG. 1 , and that the camera is installed so that the sideline in front of the basketball court 108 is included in the angle of view θ. In other words, the horizontal length of the overhead image signal 600 and the length of the sideline in front of the basketball court 108 in the overhead image signal 600 are the same or approximately the same.

[0115] (FIG. 8: Flowchart showing attention area image generation processing according to the second embodiment / FIG. 9: Explanatory diagram) FIG. 8 is a diagram showing the generation processing of an attention area image signal 601, which is a characteristic process of the image processing device 103 according to the second embodiment, and FIGS. 9A and 9B are explanatory diagrams of the generation processing of FIG. 8. The processing from when the CPU 210 receives the overhead-view image signal 600 from the overhead-view camera 102 and reduces the overhead-view image signal 600 to when the ball position is detected from the reduced image signal 604 is the same as up to step S402 in the processing in the first embodiment (see FIG. 4B). The following description will cover the content of the processing in step S403 in the second embodiment.

[0116] First, in step S900, the CPU 210 of the image processing device 103 requests the CPU 201 of the overhead camera 102 via the NIC 215 to transmit information indicating the focal length of the lens (not shown) of the overhead camera 102 and the tilt of the overhead camera 102. In response to this, the CPU 201 acquires the focal length and "tilt k" of the overhead camera 102 from the zoom drive unit 208 and the tilt sensor 227. Next, the CPU 201 transmits the acquired focal length and "tilt k" to the image processing device 103 via the NIC 205. In response to this, the CPU 210 acquires the focal length and "tilt k" received from the overhead camera 102, stores them in the RAM 211, and proceeds to step S901.

[0117] Next, in step S901, the CPU 210 calculates the angle of view θ of the overhead camera 102 from the focal length stored in RAM 211 in step S900. Next, the CPU 210 calculates the "distance d1" from the overhead camera 102 to the near sideline based on the angle of view θ and the length (SL) of the near sideline of the basketball court 108 in the real world. Specifically, in step S901, when the length of the near sideline of the basketball court 108 in the real world is "SL," the "distance d1" from the overhead camera 102 to the sideline SL is calculated using the following equation 3:

[0118]

[0119] The CPU 210 stores the calculated "distance d1" in RAM 211 and proceeds to step S902. In step S902, the CPU 210 calculates the height h of the overhead camera 102 from the plane on which the basketball court 108 in the real world BR>E is located, based on the "distance d1" stored in RAM 211 and the "tilt k" stored in RAM 211 in step S900. Specifically, when the tilt of the overhead camera 102 with respect to the horizontal direction is "k" and the distance from the overhead camera 102 to the sideline SL is "d1," the CPU 210 calculates the "height h" of the overhead camera 102 using the following equation 4:

[0120]

[0121] The CPU 210 stores the calculated "height h" in RAM 211. The CPU 210 calculates a "distance d2" from the overhead camera 102 to the far sideline based on the "height h" stored in RAM 211, the "distance d1" stored in RAM 211 in step S901, and the length "EL" of the end line of the basketball court 108 in the real world. Specifically, the CPU 210 first determines the distance from the overhead camera 102 to the near sideline SL as "d1," the height of the overhead camera 102 as "h," and the length "EL" of the end line of the basketball court 108 in the real world. The CPU 210 then calculates the "distance d2" from the overhead camera 102 to the far sideline using the following equation 5. The CPU 210 stores the calculated "distance d2" in RAM 211 and proceeds to step S903.

[0122]

[0123] Next, in step S903, the CPU 210 calculates the "distance d3" (described below) based on the "tilt k" and "height h" of the overhead camera 102, as well as the "distance d1" and "distance d2" stored in the RAM 211. That is, the CPU 210 calculates the "distance d3" of the line connecting the sideline SL in front of the basketball court 108 in the real world to the point where the line perpendicularly intersects with the line of "distance d1" and the line of "distance d2." Specifically, the CPU 210 defines the tilt of the overhead camera 102 with respect to the horizontal direction as "k," the height of the overhead camera 102 as "h," and the distance from the overhead camera 102 to the sideline SL in front as "d1." Furthermore, when the distance from the overhead camera 102 to the sideline in the back is defined as "d2," the CPU 210 calculates the "distance d3" from the sideline SL to "distance d2" using the following equation 6: Thereafter, the CPU 210 stores the calculated "distance d3" in the RAM 211 and moves the process to step S904.

[0124]

[0125] Next, in step S904, CPU 210 identifies a trapezoid whose upper base, lower base, and height correspond to "distance d1," "distance d2," and "distance d3," respectively, based on "distance d1," "distance d2," and "d3" stored in RAM 211. Specifically, the distance from overhead camera 102 to the near sideline SL is set to "d1," and the distance from overhead camera 102 to the far sideline is set to "d2." At this time, the upper and lower bases of the trapezoid found from overhead image signal 600 satisfy the following relational expression 7.

[0126]

[0127] Next, the distance from the sideline SL to the line segment at the distance d2 is set to d3. At this time, the height of the trapezoid obtained from the overhead image signal 600 satisfies the following relational expression 8.

[0128]

[0129] Then, the CPU 210 calculates the coordinates of each vertex of the trapezoid in the overhead image signal 600 based on the identified trapezoid (hereinafter referred to as "trapezoid coordinates"). The CPU 210 stores the calculated trapezoid coordinates in the RAM 211. The CPU 210 calculates a line segment L3 based on the ball coordinates 602 stored in the RAM 211 according to the processing of step S703 in FIG. 6 described in the first embodiment and the positions of the upper and lower bases of the trapezoid calculated from the trapezoid coordinates. The CPU 210 stores the calculated length of the line segment L3 in the RAM 211 and proceeds to step S905. As shown in FIG. 7, the line segment L3 is a line segment PQ that is parallel to the upper and lower bases of the identified trapezoid and passes through the center of gravity of the ball's coordinates.

[0130] Next, in step S905, the CPU 210 executes the same process as in step S704 in the first embodiment to calculate the "area of ​​interest size." That is, the CPU 210 calculates the area of ​​interest size based on the ratio of the length of the lower base L2 to the length of the line segment L3. The lower base L2 here is shown in FIG. 7. The CPU 210 stores the calculated area of ​​interest size in the RAM 211 and proceeds to step S906.

[0131] Next, in step S906, the CPU 210 executes the same process as in step S705 in the first embodiment to calculate the attention area coordinates 701. That is, the CPU 210 calculates the attention area coordinates 701 based on the ball coordinates 602 and the attention area size. The CPU 210 stores the calculated attention area coordinates 701 in the RAM 211 and proceeds to step S907.

[0132] Next, in step S907, the CPU 210 executes the same process as in step S706 in the first embodiment to generate the attention area image signal 601. That is, the CPU 210 calculates the attention area image signal 601 based on the attention area coordinates 701. The CPU 210 then stores the generated attention area image signal 601 in the RAM 211. Above, the process of generating the attention area image signal 601, which is a characteristic process of the image processing device 103 in the second embodiment, has been described. The subsequent process follows step S404 in the process of the image processing device 103 in the first embodiment (see FIG. 4B ).

[0133] The above has described the processing of the second embodiment of the present invention in the configuration of the image processing system in Figure 1. In this way, the object shape of the basketball court 108 can be determined based on the camera information acquired from the overhead camera 102. As a result, even if the object shape of the basketball court 108 cannot be detected from the overhead image signal 600 by the processing executed by the image processing device 103, the present invention can be applied without requiring any additional operation by the user.

[0134] Next, a third embodiment of the present invention will be described. In the third embodiment, a process for generating an area-of-interest image signal 601 when the imaging optical axis 114 of the overhead camera 102 is not parallel to the half line of the basketball court 108 will be described. Since the basic system configuration is the same as in the first embodiment, a duplicated explanation will be omitted and only the process of the image processing device 103, which is the difference, will be described.

[0135] (Fig. 10: Flowchart showing attention area image generation processing of the third embodiment / Figs. 11A and 11B: Explanatory diagrams) Fig. 10 is a diagram showing the generation processing of an attention area image signal 601, which is a characteristic process of the image processing device 103 according to the third embodiment. Figs. 11A and 11B are explanatory diagrams of the processing of Fig. 10. The processing from receiving an overhead-view image signal 600 from the overhead-view camera 102 and reducing the overhead-view image signal 600 to detecting the ball position from the reduced image signal 604 is the same as the processing executed by the image processing device 103 according to the first embodiment (see Fig. 4B). In other words, the processing up to step S402 is the same as in the first embodiment. The following explanation will be about the processing content of step S403 in the first embodiment.

[0136] First, in step S1100, the CPU 210 executes the same process as step S700 in the first embodiment, performing a calculation based on the Hough transform on the overhead image signal 600. Here, the rectangle formed by the straight lines obtained as a result of the calculation is regarded as the court lines of the basketball court 108. The CPU 210 stores the calculated court lines of the basketball court 108 in the RAM 211 and proceeds to step S1101.

[0137] Next, in step S1101, the CPU 210 performs the same process as step S701 in the first embodiment to calculate court coordinates. That is, the CPU 210 calculates the coordinates of the four vertices A, B, C, and D of a rectangle formed by the court lines. The four vertices A, B, C, and D are shown in FIG. 11 . Note that in the third embodiment, the imaging optical axis 114 of the overhead camera 102 is not parallel to the half line of the basketball court 108, so the rectangle of the basketball court 108 captured in the overhead image signal 600 has a distorted shape (see FIG. 11A ). Accordingly, the calculated court coordinates are different from the court coordinates in the first embodiment and are the coordinates of each vertex of the distorted rectangle. The CPU 210 stores the calculated court coordinates in the RAM 211 and proceeds to step S1102.

[0138] Next, in step S1102, CPU 210 performs a projective transformation between the coordinate system of overhead image signal 600 and a basketball court coordinate system defined based on the known shape of basketball court 108. Specifically, CPU 210 generates a "projective transformation matrix" between the court coordinates in the coordinate system of overhead image signal 600 stored in RAM 211 in step S1101 and the vertex coordinates of the four vertices of the basketball court in the basketball court coordinate system. CPU 210 stores the calculated projective transformation matrix in RAM 211 and proceeds to step S1103.

[0139] Next, in step S1103, the CPU 210 performs projective transformation of the ball coordinates 602a from the coordinate system on the overhead image signal 600 to the basketball court coordinate system based on the projective transformation matrix stored in the RAM 211 in step S1102. As a result, the ball coordinates become 602d as shown in FIG. 11B. The CPU 210 stores the ball coordinates 602d in the basketball court coordinate system in the RAM 211 and proceeds to step S1104.

[0140] Next, in step S1104, the CPU 210 calculates a line that is parallel to the end line of the basketball court 108 and passes through the center of gravity of the ball coordinates 602d, based on the ball coordinates 602d stored in RAM 211 in step S1103. The CPU 210 calculates the line segment RT as "line segment L4" when the intersection of the calculated line with line segment AD of the rectangle ABCD consisting of the basketball court coordinate system is defined as "R" and the intersection with line segment BC is defined as "T." The CPU 210 stores the length of the calculated "line segment L4" in RAM 211 and proceeds to step S1105.

[0141] Next, in step S1105, the CPU 210 calculates the "area of ​​interest size," which is the image size of the area of ​​interest image signal 601, from the ratio at which the center of gravity of the ball coordinates 602d divides the length of line segment L4 based on the length of line segment L4 stored in RAM 211 in step S1104. Specifically, before starting the processing of this embodiment, the CPU 210 stores the full-body size of the player 109 on the overhead image signal 600 when the player 109 is positioned on the front and back sidelines of the basketball court 108. Then, the CPU 210 generates a "front reference area of ​​interest size" and a "back reference area of ​​interest size" based on the image sizes described above.

[0142] Here, the "reference attention area size" is a size whose vertical size is twice the size of the player 109 on the bird's-eye view image signal 600 and whose horizontal size and vertical size have a ratio of "1:1." The "reference attention area size" may be any size that allows trimming without cutting off the player taking the shot. The CPU 210 stores the generated "reference attention area sizes" for the front and back in the RAM 211. The CPU 210 calculates the line segment from intersection R on the line segment L4 to the center of gravity of the ball coordinates 602d as the "line segment L5." The CPU 210 stores the calculated length of the "line segment L5" in the RAM 211.

[0143] Then, the CPU 210 generates a new region of interest size by enlarging or reducing the reference region of interest size in accordance with the ratio of the length of line segment L5 to the length of line segment L4. Specifically, the CPU 210 calculates a new "region of interest size S" using the following equation 9, where the reference region of interest size for the front is "SF," the reference region of interest size for the back is "SB," the length of line segment L4 is "L4," and the length of line segment L5 is "L5." The CPU 210 stores the length of the new region of interest size in the RAM 211 and proceeds to step S1106.

[0144]

[0145] Next, in step S1106, the CPU 210 calculates attention area coordinates 701 based on the ball coordinates 602d stored in the RAM 211 according to the processing of step S705 in the first embodiment and the attention area size stored in the RAM 211 in step S1105. The CPU 210 stores the calculated attention area coordinates 701 in the RAM 211 and proceeds to step S1107.

[0146] Then, in step S1107, the CPU 210 follows the processing of step S706 in the first embodiment to crop the attention area from the overhead image signal 600 based on the attention area coordinates 701 stored in the RAM 211 in step S1106, thereby generating an attention area image signal 601. The CPU 210 stores the generated attention area image signal 601 in the RAM 211.

[0147] The above has described the process of generating the attention area image signal 601, which is a characteristic process of the image processing device 103 in this embodiment. In this way, the CPU 210 can generate distortion information based on the projective transformation matrix between the coordinates indicating the vertices of the first object (basketball court 108) and the coordinates indicating the vertices of a known shape of the first object. The CPU 210 can then determine the "attention area" based on the distortion information and the calculated ball coordinates 602 (attention position), and determine its size, the "attention area size." An image corresponding to the "attention area size" becomes the "attention area image."

[0148] The subsequent processing continues from step S404 in the processing flow of the image processing device 103 of the first embodiment (see FIG. 4B). The processing of the third embodiment of the present invention has been described above using the configuration of FIG. 1. In this way, the court coordinates in the coordinate system on the overhead image signal 600 can be projectively transformed into a basketball court coordinate system with a known shape. This makes the present invention applicable even when the imaging optical axis 114 of the overhead camera 102 is not parallel to the half line of the basketball court 108.

[0149] The CPU 210 (or the CPU 243) may also include a user interface (UI unit) that acquires an object shape accepted by user operation as the shape of the object (basketball court 108). The CPU 210 may also acquire optical information, such as the optical axis direction and focal length of the lens of the overhead camera 102 that captures the overhead image, and generate distortion information based on the acquired optical information. Shape information for each of a group of objects consisting of objects of multiple types of shapes may also be stored in the HDD 216 or the like. In this case, the CPU 210 may execute the following process. That is, if there is a correlation between any of the multiple types of shape information stored in the HDD 216 or the like and an object in the image captured by the overhead camera 102, the CPU 210 acquires the object of that shape information as an object of the basketball court 108.

[0150] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the present invention. For example, the present invention can be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or recording medium, and having a computer processor in the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more of the functions.

[0151] This application claims priority based on Japanese Patent Application No. 2023-210167, filed on December 13, 2023, the entire contents of which are incorporated herein by reference.

[0152] REFERENCE SIGNS LIST 100 Internet 101 Local network 102 Bird's-eye view camera 103 Image processing device 104 Target camera 105 Learning server 106 Data collection server 107 Client terminal 108 Basketball court 109 Player 110 Ball 113 Video cable 114 Imaging optical axis 115 Position sensor receiving device

Claims

1. An image processing device capable of recognizing a specified scene in a captured image, comprising: an input unit for inputting a captured image of a first object; a distortion information generation unit for generating distortion information indicating distortion of the shape of the first object in the captured image input by the input unit from a known shape; a calculation unit for calculating a position of interest in the captured image; a determination unit for determining a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit; an image generation unit for generating a region of interest image from the captured image based on the region of interest determined by the determination unit; a scaling unit for scaling the region of interest image generated by the image generation unit to a specified size; and a recognition unit for performing specified recognition based on the region of interest image scaled by the scaling unit, wherein the determination unit determines a region of interest size, which is the size of the region of interest corresponding to the actual size, based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit.

2. The image processing device according to claim 1, characterized in that the distortion information generating unit generates distortion information based on one or more of the degree of distortion, size, and shape of the object shape from a known shape, among the information regarding the first object in the captured image.

3. An image processing device as described in claim 1 or 2, further comprising a shape acquisition unit that acquires a shape of the first object in the captured image, and the distortion information generation unit generates distortion information based on the shape of the first object acquired by the shape acquisition unit.

4. The image processing device according to claim 3, wherein the shape acquisition unit further comprises a UI unit that acquires an object shape accepted by a user operation as the shape of the first object.

5. An image processing device as described in claim 3, further comprising a shape memory unit that stores shape information for each of a group of objects consisting of objects of multiple types of shapes, and wherein the shape acquisition unit further acquires the object of the shape information as the shape of the first object when there is a correlation between any of the multiple types of shape information stored by the shape memory unit and an object in the captured image.

6. The image processing device according to claim 3, characterized in that the shape acquisition unit further detects landmarks provided at various locations in the real world corresponding to each vertex of the first object, and acquires the area with each landmark as a vertex as the shape of the first object.

7. The image processing device according to claim 3, characterized in that the distortion information generation unit generates distortion information based on a projection transformation matrix between coordinates indicating vertices of the first object acquired by the shape acquisition unit and coordinates indicating vertices of a known shape of the first object.

8. The image processing device according to claim 1 or 2, characterized in that the determination unit determines the attention area based on position information of a second object whose position is estimated in the captured image.

9. An image processing device as described in claim 1 or 2, further comprising a position information acquisition unit that acquires position information from a sensor that detects a position attached to a real-world object corresponding to a third object, and the calculation unit calculates the focus position based on the position information acquired by the sensor.

10. The image processing device according to claim 1 or 2, characterized in that the recognition unit determines whether or not a second object in the attention area image has performed a predetermined action.

11. The image processing device according to claim 10, characterized in that the recognition unit determines whether the second object has performed a predetermined action based on a position of a third object and a position of the second object.

12. The image processing device described in claim 11, further comprising a first trained model for estimating the position of a second object from the attention area image, and characterized in that the position of the second object is estimated from the attention area image using the first trained model.

13. The image processing device described in claim 11, further comprising a second trained model for detecting the position of a third object from the attention area image, and wherein the position of the third object is detected from the reduced attention area image using the second trained model.

14. The image processing device described in claim 11, characterized in that the position of the second object is the position of a joint of a player competing in a competitive event in a real-world arena corresponding to the first object, and the position of the third object is the position of an object used in the competitive event.

15. An image processing device as described in claim 1 or 2, further comprising an optical information acquisition unit that acquires optical information including an optical axis direction and focal length of a lens of a camera that captures the captured image, and the distortion information generation unit generates distortion information based on the optical information acquired by the optical information acquisition unit.

16. An image processing system having a first camera capable of capturing an image of a first object, a second camera capturing an image in a specified capturing direction, and an image processing device capable of controlling the capturing direction of the second camera, wherein the first camera comprises: a first imaging unit capable of acquiring an image of the first object; and a first transmission unit transmitting the image acquired by the first imaging unit to the image processing device, and the image processing device comprises: a first receiving unit receiving the captured image transmitted by the first transmission unit; a distortion information generating unit generating distortion information indicating distortion of a shape of the first object from a known shape in the captured image received by the first receiving unit; a calculation unit calculating a position of interest in the captured image; a determination unit determining an area of ​​interest based on the distortion information generated by the distortion information generating unit and the position of interest calculated by the calculation unit; an image generation unit generating an area of ​​interest image from the captured image based on the area of ​​interest determined by the determination unit; a scaling unit scaling the area of ​​interest image generated by the image generation unit to a predetermined size; and a recognition unit performing a predetermined recognition from the area of ​​interest image scaled by the scaling unit. an image processing system comprising: a control signal generating unit that generates a control signal for directing the second camera in an imaging direction according to the recognition result of the recognition unit; and a second transmitting unit that transmits the control signal generated by the control signal generating unit to the second camera, wherein the determination unit has a function of determining an attention area size, which is the area size of the attention area corresponding to an actual size, based on distortion information generated by the distortion information generating unit and the attention position calculated by the calculation unit, and the second camera comprises: a second imaging unit capable of imaging; a second receiving unit that receives the control signal transmitted by the second transmitting unit; and a control unit that controls the imaging direction of the second imaging unit based on the control signal received by the second receiving unit.

17. A control method for an image processing device capable of recognizing a predetermined scene in a captured image, comprising: an input step of inputting a captured image of a first object; a distortion information generation step of generating distortion information indicating distortion of the shape of the first object in the captured image input by the input step from a known shape; a calculation step of calculating a position of interest in the captured image; a determination step of determining a region of interest based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step; an image generation step of generating a region of interest image from the captured image based on the region of interest determined by the determination step; a scaling step of enlarging or reducing the region of interest image generated by the image generation step to a predetermined size; and a recognition step of performing a predetermined recognition based on the region of interest image enlarged or reduced by the scaling step, wherein the determination step is a step of determining a region of interest size, which is the size of the region of interest corresponding to the actual size, based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step.

18. A program for causing a computer to execute a control method for an image processing device capable of recognizing a specified scene in a captured image, the control method comprising: an input step of inputting a captured image of a first object; a distortion information generation step of generating distortion information indicating distortion of the shape of the first object in the captured image input by the input step from a known shape; a calculation step of calculating a position of interest in the captured image; a determination step of determining a region of interest based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step; an image generation step of generating a region of interest image from the captured image based on the region of interest determined by the determination step; a scaling step of enlarging or reducing the region of interest image generated by the image generation step to a specified size; and a recognition step of performing a specified recognition based on the region of interest image enlarged or reduced by the scaling step, wherein the determination step is a step of determining a region of interest size, which is the size of the region of interest corresponding to the actual size, based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step.

Citation Information

Patent Citations

  • System and program for estimating position of object

    JP2016099941A

  • Object detection device and object detection method

    JP2019219804A

  • Object detection apparatus, object detection method, object detection program and learning apparatus

    JP2021071757A