Information processing device, image processing system, control method and program
The image processing system addresses the challenge of accurate action recognition in wide-area imaging by generating distortion information and adjusting the region of interest, enhancing recognition accuracy and reducing subject distortion issues.
Patent Information
- Application Number
- JP2023210167
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-25
AI Technical Summary
Existing image processing systems struggle to accurately recognize predetermined actions in captured images, particularly in wide-area imaging scenarios like basketball courts, due to varying resolutions and distortions that affect the detection of joint coordinates, leading to reduced recognition accuracy.
An image processing apparatus and system that generates distortion information from known object shapes, calculates a region of interest, scales it to a predetermined size, and performs recognition, using a determination unit to adjust the region size based on distortion and position, enabling accurate action recognition regardless of the image angle.
Enables high-accuracy recognition of predetermined actions in captured images by adjusting the region of interest based on distortion and position, improving recognition accuracy and reducing the risk of cutting off or crushing subjects during image processing.
Smart Images

Figure 2025094547000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing system, a control method, and a program.
Background Art
[0002] In recent years, in imaging such as sports, there has been a need for automatic imaging of highlight scenes during a game. For example, in the case of basketball, a highlight scene may be a shooting scene of a player. When automatically imaging a shooting scene, for example, it is determined whether a shot is being taken based on the positional relationship between the ball and the player being focused on and the state of the player's posture from the imaging data of the camera, and there is a method of controlling the imaging of the camera for the player who is taking a shot.
[0003] As a method of estimating a player's posture from imaging data, a method of detecting "joint coordinates of a player" is known. At this time, as a method of detecting "joint coordinates of a player", there is a method using a posture estimation learned model (posture estimation model). Specifically, by inputting an image of a player into the posture estimation model, the coordinates of joints such as the top of the player's head, neck, wrists, elbows, shoulders, and waist are output to obtain "joint coordinates". When performing posture estimation of a player using basketball as an example, if imaging data showing the entire basketball court is input into the "posture estimation model", the processing time becomes long because the input resolution is high, and it is not suitable for real-time processing. Therefore, a method is known in which the imaging data is reduced before being input into the "posture estimation model" to lower the input resolution and shorten the processing time.
[0004] However, when the captured data is reduced, the resolution of the player shown in the captured data also decreases according to the reduction, so there is a possibility that the joint coordinates of the player cannot be detected with high accuracy. That is, there is a possibility that it cannot be determined whether the player is shooting or not. Therefore, Patent Document 1 discloses a method of performing predetermined recognition at high speed while increasing the input resolution to a learned model including a pose estimation model. Patent Document 1 discloses a technique of detecting the coordinates of a ball from captured data and determining the action of a player from an image obtained by trimming the vicinity of the ball. Depending on the type of sport, there is a high possibility that the player performs a predetermined action near the ball, so it is possible to determine the action of the player from the coordinates of the ball and the joint coordinates of the player by this technique.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, when imaging is performed so that the entire basketball court is captured, players far from the camera are drawn small, and players close to the camera are drawn large. In the prior art disclosed in Patent Document 1, since the trimming size is always constant, if trimming is performed based on a player far from the camera, there is a possibility that a player close to the camera will be cut off. Since the "cut-off part" cannot detect the joint coordinates of the player, it becomes difficult to accurately estimate the posture of the player. On the other hand, if trimming is performed based on a player close to the camera, the resolution of a player far from the camera may decrease. If the resolution is low, the resolution further decreases due to the reduction process before inputting to the learned model, making it difficult to accurately recognize the player. When performing a predetermined recognition from data obtained by imaging a wide area such as a basketball court, if trimming is performed with a fixed size without considering the distance from the camera to the subject (such as a player), trimming suitable for recognition cannot be performed. As a result, there is a problem that the recognition accuracy decreases.
[0007] An object of the present invention is to provide an image processing apparatus, an image processing system, a control method, and a program that enable a predetermined action by an object in a captured image to be recognized with high accuracy regardless of the position within the captured image angle.
Means for Solving the Problems
[0008] In order to achieve the above object, one aspect of the image processing apparatus of the present invention is an image processing apparatus capable of recognizing a predetermined scene in a captured image, including: an input unit that inputs a captured image of a first object; a distortion information generation unit that generates distortion information indicating distortion from a known shape of the shape of the first object in the captured image input by the input unit; a calculation unit that calculates a position of interest in the captured image; a determination unit that determines a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit; an image generation unit that generates a region-of-interest image from the captured image based on the region of interest determined by the determination unit; a scaling unit that scales the region-of-interest image generated by the image generation unit to a predetermined size; and a recognition unit that performs a predetermined recognition based on the region-of-interest image scaled by the scaling unit. The determination unit is characterized in that it determines the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit.
Advantages of the Invention
[0009] According to the present invention, there is an effect that an image processing apparatus, an image processing system, a control method, and a program can be provided that enable a predetermined action by an object in a captured image to be recognized with high accuracy regardless of the position within the captured image angle.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in the embodiments. First, the first embodiment of the present invention will be described. In the attached drawings, the same or similar configurations are denoted by the same reference numerals, and redundant explanations are omitted.
[0012] <First Embodiment> First, the first embodiment according to the present invention will be described. In the first embodiment, taking basketball as an example, the process of determining whether it is a shooting scene from imaging data including the entire basketball court and imaging a video including the shooting scene will be described. Note that taking a basketball as the imaging object in the first embodiment is only for the purpose of specific explanation and does not limit the imaging object of the present invention. Also, in the following description, player 109 and ball 110 are numbered, and in the description of multiple objects, only numbers are used to represent them, and when identifying and explaining individuals, numbers with added alphabets such as "109a" are attached.
[0013] (Fig. 1: Image Processing System Configuration Diagram) FIG. 1 is a configuration diagram of an image processing system including an image processing apparatus 103 that executes determination processing for a shooting scene according to the first embodiment. This image processing system includes the Internet 100, a local network 101, an overhead camera 102, an image processing apparatus 103, a target camera 104, a learning server 105, a data collection server 106, and a client terminal 107. Further, this image processing system appropriately includes a position sensor receiving device 115.
[0014] (Internet 100: Local network 101) The Internet 100 is a network for communicating required information with the learning server 105, the data collection server 106, etc. Also, the Internet 100 can communicate with electronic devices connected to the local network 101 via the local network 101. The local network 101 is a network for the overhead camera 102, the image processing apparatus 103, the target camera 104, and the client terminal 107 to communicate required information with each other. Also, the local network 101 can communicate required information with electronic devices connected to the Internet 100 via the Internet 100.
[0015] (Overhead camera 102) The overhead camera 102 is a camera for imaging the entire basketball court 108. Also, the overhead camera 102 transmits the captured image to the image processing apparatus 103 via a video cable 113. In the embodiment, it is assumed that the overhead camera 102 captures images at 30 frames per second (30 (fps)), but this does not limit the frame rate. Also, the imaging optical axis 114 of the overhead camera 102 is arranged to be parallel to the half line of the basketball court 108, and the imaging angle of view is assumed to be imaged at an "angle of view θ" in which the sideline in front of the basketball court 108 is included.
[0016] (Image processing apparatus 103: Target camera 104) When the image processing device 103 performs object detection of the ball 110, posture estimation of the player 109, determination of the presence or absence of a shooting scene, etc. on the captured image received from the overhead camera 102 and captures the shooting scene, it controls the attention camera 104. The attention camera 104 includes a pan-tilt mechanism, a zoom mechanism, etc., and is a camera for capturing a shooting scene according to the control from the image processing device 103. In the embodiment, it is assumed that the attention camera 104 captures images at 60 frames per second (60 (fps)), but this does not limit the frame rate.
[0017] (Learning server 105) The learning server 105 is a server for generating a "trained model" used when the image processing device 103 performs object detection and posture estimation. The learning server 105 includes the following two trained models. A trained model for object detection (hereinafter referred to as "object detection model") that detects the ball 110 from the imaging data of the overhead camera 102, and a trained model for posture estimation (hereinafter referred to as "posture estimation model") that estimates the joint coordinates of the player 109 from the attention area image signal 601.
[0018] (Data collection server 106: Client terminal 107: Position sensor receiving device 115) The data collection server 106 is a server that collects and stores the teacher data for machine learning required for the learning server 105 to generate a trained model. The client terminal 107 is a terminal that performs operations on electronic devices connected to the Internet 100 and the local network 101 and instructs data transmission and reception. The type of the client terminal 107 is not particularly limited and can be realized by, for example, a general-purpose PC. The position sensor receiving device 115 reads the position information of the ball from the position sensor attached to the ball 110 and measures the position of the ball 110 on the basketball court 108.
[0019] The configuration example of the system that executes the determination process of the shooting scene according to the first embodiment has been described above. This system reads the position of the ball 110, the posture of the player 109, etc. from a basketball game, detects the shooting scene based on the conditions of the ball 110 and the player 109, and performs automatic imaging of the shooting scene.
[0020] (Fig. 2: Hardware configuration diagram of the image processing system) (Hardware configuration of the overhead camera 102) Fig. 2 is a hardware configuration diagram of each device constituting the system of Fig. 1. The overhead camera 102 includes a CPU 201, a RAM 202, a ROM 203, a video engine 204, a NIC 205, a video I / F 206, an image sensor 207, a zoom drive unit 208, and a tilt sensor 227. These are connected to the system bus 200 and are configured to be able to transmit and receive the required information to and from each other.
[0021] The CPU 201 controls the overhead camera 102. The CPU 201 controls each unit described later and performs control operations such as imaging by the image sensor 207 and transmission of imaging data by the video I / F 206. The RAM 202 is a volatile memory capable of rewriting stored information, and a work area for a program describing the control content of the overhead camera 102 is formed. The RAM 202 is realized by a volatile memory (such as DRAM) using semiconductor elements, for example. The ROM 203 is a non-volatile memory that stores a program describing the control content of the overhead camera 102 and various parameters in a non-volatile manner. The ROM 203 is realized by a flash memory, for example. In response to the user turning on the power of the overhead camera 102, the CPU 201 reads the program from the ROM 203, expands it in the RAM 202, and executes it to start the control of the overhead camera 102.
[0022] The imaging engine 204 is an image processing engine that converts the charge obtained from the image sensor 207 into an image signal. The NIC 205 is a network interface card (NIC), which is used for the overhead camera 102 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized in ETHERNET (registered trademark) or the IEEE802.3 series. The video I / F 206 is an interface for connecting the overhead camera 102 and the image processing device 103 with a video cable 113.
[0023] The overhead camera 102 transmits imaging data to the image processing device 103 via the video I / F 206. Note that the video I / F 206 only needs to be an interface capable of communicating with the image processing device 103, and its type is not particularly limited. For example, HDMI (registered trademark) (HIGH-DEFINITION MULTIMEDIA INTERFACE) is exemplified. Also, the overhead camera 102 may be a network camera, or may transmit imaging data to the image processing device 103 via the NIC 205.
[0024] The image sensor 207 is a sensor that receives light emitted from the imaging target and converts its brightness and color into charge. For example, a photodiode, a CCD (CHARGE COUPLED DEVICE) sensor, a CMOS (COMPLEMENTARY METAL OXIDE SEMICONDUCTOR) sensor, etc. can be mentioned. The zoom drive unit 208 is a unit for changing the degree of "zooming in and out" of the zoom lens provided in the overhead camera 102. The tilt sensor 227 is a sensor that detects the angle at which the main body of the overhead camera 102 tilts with respect to the horizontal direction.
[0025] (Hardware Configuration of Image Processing Device 103) The image processing device 103 includes a CPU 210, a RAM 211, a ROM 212, a GPU 213, a video I / F 214, a NIC 215, and an HDD 216. These are connected to the system bus 209 and are configured to be able to transmit and receive the required information to and from each other.
[0026] The CPU 210 controls the image processing apparatus 103. The CPU 210 controls each unit described below and performs operations according to the imaging data received from the video I / F 206 and the communication data received from the NIC 215. The RAM 211 is a volatile memory capable of rewriting stored information, and forms a work area for a program in which the control content of the image processing apparatus 103 is described. The RAM 211 is realized by, for example, a volatile memory (such as DRAM) using semiconductor elements. The ROM 212 is a non-volatile memory and stores a program in which the control content of the image processing apparatus 103 is described and various parameters in a non-volatile manner. When the user turns on the power of the image processing apparatus 103, the CPU 210 reads the program from the ROM 212, expands it in the RAM 211, and executes it. Thereby, the control of the image processing apparatus 103 is started. The ROM 212 is realized by, for example, a flash memory or the like.
[0027] The GPU 213 is a unit that performs parallel arithmetic processing of data. When performing a large number of multiply-accumulate operations in the inference process, the GPU 213 is faster and more efficient in executing the processing than the CPU 210. Generally, an LSI called "GRAPHICS PROCESSING UNIT" is used for the GPU 213, but a reconfigurable logic circuit hardware called FPGA may be used to realize an equivalent function.
[0028] The video I / F 214 is an interface for connecting the overhead camera 102 and the image processing apparatus 103 with a video cable 113. The image processing apparatus 103 receives imaging data from the overhead camera 102 via the video I / F 214. The video I / F 214 only needs to be an interface capable of communicating with the overhead camera 102, and its type is not particularly limited. For example, a USB (UNIVERSAL SERIAL BUS) interface or the like is exemplified. The NIC 215 is a network interface card and is used for the image processing apparatus 103 to communicate with other devices via the local network 101. For example, communication based on a communication method standardized by ETHERNET (registered trademark) or the IEEE802.3 series is performed.
[0029] The HDD 216 is a storage device that stores the imaging data received from the overhead camera 102 and the "trained model" received from the learning server 105. In this embodiment, the HDD 216 is a hard disk drive (HDD) that uses a magnetic storage method, but the type of the HDD 216 is not limited to this. For example, an external storage device such as a solid state drive (SSD) that uses semiconductor elements may be used as the HDD 216.
[0030] (Hardware Configuration of the Focus Camera 104) The focus camera 104 includes a CPU 218, a RAM 219, a ROM 220, a video engine 221, a NIC 222, a video I / F 223, an image sensor 224, a zoom drive unit 225, and a pan-tilt drive unit 226. These are connected to the system bus 217 and are configured to be able to transmit and receive the required information to and from each other.
[0031] The CPU 218 controls the focus camera 104. The CPU 218 controls each unit described later and executes operations such as imaging by the image sensor 224 and reception of communication data by the NIC 222. The RAM 219 is a memory whose information can be rewritten, and a work area for a program in which the control content of the focus camera 104 is described is formed. The RAM 219 is realized by, for example, a volatile memory that uses semiconductor elements. The ROM 220 is a non-volatile memory that stores a program in which the control content of the focus camera 104 is described and various parameters in a non-volatile manner. In response to the power-on of the focus camera 104 by the user, the CPU 218 reads the program from the ROM 220, expands it in the RAM 219, and executes it. Thereby, the control of the focus camera 104 is started. The ROM 220 is realized by, for example, a flash memory.
[0032] The video engine 221 is an image processing engine that converts the charge obtained from the image sensor 224 into an image signal. The NIC 222 is a network interface card and is used for the target camera 104 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE802.3 series. The video I / F 223 is an interface for transmitting the imaging data of the target camera 104. When the system is used to image a basketball game, the imaging data transmitted from the video I / F 223 becomes the video of the desired highlight scene. Therefore, the video I / F 206 is used to connect an external storage device (not shown) or the like for storing the video of the highlight scene. Note that the video I / F 206 only needs to be an interface capable of communicating with the image processing device 103, and its type is not particularly limited. For example, an HDMI interface is exemplified.
[0033] The image sensor 224 is a sensor that receives the light emitted by the imaging target and converts its brightness and color into charge. For example, a photodiode, a CCD sensor, a CMOS sensor, etc. The zoom drive unit 225 is a unit for changing the degree of zooming of the zoom lens provided in the target camera 104. The pan-tilt drive unit 226 is a unit for changing the degree of turning of the pan turning mechanism and the tilt turning mechanism provided in the target camera 104.
[0034] (Hardware Configuration of the Learning Server 105) The learning server 105 has a CPU 229, a RAM 230, a ROM 231, a GPU 232, an HDD 233, and a NIC 234, which are configured to be able to transmit and receive the necessary information to and from each other via a system bus 228. The CPU 229 controls the learning server 105. The CPU 229 controls each unit described later and performs operations according to communication data received from the NIC 234. The RAM 230 is a memory capable of rewriting information, and a work area for a program in which the control content of the learning server 105 is described is formed. The RAM 230 is realized by a volatile memory using, for example, semiconductor elements. The ROM 231 is a non-volatile memory that non-volatily stores a program in which the control content of the learning server 105 is described, various parameters, and the like. In response to the user turning on the power of the learning server 105, the CPU 229 reads the program from the ROM 231, expands it in the RAM 230, and executes it, thereby starting the control of the learning server 105. The ROM 231 is realized by, for example, a flash memory.
[0035] The GPU 232 is a unit that performs parallel arithmetic processing of data and performs efficient arithmetic by processing more data in parallel. Therefore, when learning is performed multiple times using a learning model such as deep learning, it is effective for the GPU 232 to execute the processing. In the present embodiment, in addition to the CPU 229, the GPU 232 is used for the learning processing of the learning server 105. Specifically, when executing a learning program including a learning model, the CPU 229 and the GPU 232 cooperate to perform arithmetic operations to perform learning. Note that the learning process may be executed by only one of the CPU 229 or the GPU 232. The GPU 232 is realized by, for example, an LSI called "GRAPHICS PROCESSING UNIT", but may realize an equivalent function with a reconfigurable logic circuit hardware called an FPGA.
[0036] The HDD 233 is a storage device that non-volatilely stores teacher data received from the data collection server 106, a learned model learned from the teacher data, and the like. In the present embodiment, the HDD 233 is a hard disk drive using a magnetic storage method, but the type of the HDD 233 is not limited thereto. For example, an external storage device such as a solid state drive using semiconductor elements may be used as the HDD 233. The NIC 234 is a network interface card and is used for the learning server 105 to communicate with other devices via the Internet 100. For example, communication based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series is performed.
[0037] (Hardware Configuration of Data Collection Server 106) The data collection server 106 includes a CPU 236, a RAM 237, a ROM 238, a GPU 239, an HDD 240, and a NIC 241, which are configured to be able to transmit and receive required information to and from each other via a system bus 235. The CPU 236 controls the data collection server 106. The CPU 236 controls each unit described later and performs operations according to communication data received from the NIC 241. The RAM 237 is a memory capable of rewriting information, and a work area for a program in which the control content of the data collection server 106 is described is formed. The RAM 237 is realized by a volatile memory using, for example, semiconductor elements. The ROM 238 is a non-volatile memory and non-volatilely stores a program in which the control content of the data collection server 106 is described, various parameters, and the like. In response to the user turning on the power of the data collection server 106, the CPU 236 reads the program from the ROM 238, expands it in the RAM 237, and executes it, thereby starting the control of the data collection server 106. The ROM 238 is realized by a flash memory, for example. The GPU 239 is a unit that performs parallel arithmetic processing of data. The GPU 239 is generally realized by an LSI called a "GRAPHICS PROCESSING UNIT", but may realize an equivalent function with a reconfigurable logic circuit hardware called an FPGA.
[0038] The HDD 240 is a storage device that stores the teacher data requested from the learning server 105. In this embodiment, the HDD 240 is a hard disk drive that uses a magnetic storage method. However, the type of the HDD 240 is not particularly limited, and a storage device such as a solid state drive that uses semiconductor elements may be used as the HDD 240. The NIC 241 is a network interface card and is used for the data collection server 106 to communicate with other devices via the Internet 100. For example, communication based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series is performed.
[0039] (Hardware Configuration of the Client Terminal 107) The client terminal 107 includes a CPU 243, a RAM 244, a ROM 245, a GPU 246, an HDD 247, a NIC 248, a display unit 249, and an input unit 250. These are configured to be able to transmit and receive necessary information to and from each other via a system bus 242. The CPU 243 controls the client terminal 107. The CPU 243 controls each unit described later and performs operations according to input data input from the input unit 250 or communication data received from the NIC 241. The RAM 244 is a memory capable of rewriting information, and a work area for a program in which the control content of the client terminal 107 is described is formed. The RAM 244 is realized by, for example, a volatile memory using semiconductor elements.
[0040] The ROM 245 is a non-volatile memory that stores a program describing the control content of the client terminal 107 and various parameters in a non-volatile manner. In response to the user turning on the power of the client terminal 107, the CPU 243 reads the program from the ROM 245, expands it in the RAM 244, and executes it. Thereby, the control of the client terminal 107 is started. The ROM 245 is realized by, for example, a flash memory. The GPU 246 is a unit that performs parallel arithmetic processing of data and is used to display the image data received from the image processing device 103 via the NIC 248 on the display unit 249. The GPU 246 is generally realized by an LSI called "GRAPHICS PROCESSING UNIT", but it may also realize an equivalent function with a reconfigurable logic circuit hardware called an FPGA.
[0041] The HDD 247 is a storage device that stores the image data received from the image processing device 103 via the NIC 248 in a non-volatile manner. In this embodiment, the HDD 247 is a hard disk drive using a magnetic storage method, but the type of the HDD 247 is not particularly limited. A storage device such as a solid state drive using semiconductor elements may be used as the HDD 247. The NIC 248 is a network interface card and is used for the client terminal 107 to communicate with other devices via the local network 101. For example, communication is performed based on a communication method standardized by ETHERNET (registered trademark) or the IEEE 802.3 series. The display unit 249 is a display device that displays the image data received from the image processing device 103 via the NIC 248. The display unit 249 is realized by, for example, an LCD (LIQUID CRYSTAL DISPLAY). The input unit 250 is an input device that receives the operations and instructions of the user who uses this system. The input unit 250 is realized by, for example, a pointing device such as a mouse and an input device such as a keyboard.
[0042] (Fig. 3: Functional block diagram of the image processing system) Figure 3 is a functional block diagram of software realized by the cooperation of the hardware and the program shown in the hardware configuration diagram of Figure 2. Note that general-purpose software such as the OS is omitted in the functional configuration of this software.
[0043] <Functional Configuration of Overhead Camera 102> The overhead camera 102 includes an imaging unit 800, an operation unit 801, and a communication unit 802. These functional units are realized by the CPU 201 expanding and executing the program stored in the ROM 203 in the RAM 202. The imaging unit 800 acquires imaging data of an angle of view including the entire basketball court 108 by the CPU 201 controlling the video engine 204, the image sensor 207, and the zoom drive unit 208. The operation unit 801 performs operations such as imaging start and imaging stop according to the instructions received from the client terminal 107 by the CPU 201 controlling the NIC 205. The communication unit 802 transmits imaging data and camera information to the image processing apparatus 103 by the CPU 201 controlling the NIC 205 and the video I / F 206.
[0044] <Functional Configuration of Image Processing Apparatus 103> The image processing apparatus 103 includes a communication unit 803, a processing / arithmetic unit 804, an inference unit 805, and a data storage unit 806. These functional units are realized by the CPU 210 expanding and executing the program stored in the ROM 212 in the RAM 211. The communication unit 803 receives data from the overhead camera 102 or the learning server 105 or transmits data to the attention camera 104 by the CPU 210 controlling the video I / F 214 and the NIC 215. The processing / arithmetic unit 804 performs image processing such as reduction processing and Hough transform on the imaging data received from the overhead camera 102. Further, the processing / arithmetic unit 804 cuts out an image from the imaging data according to the inference result of the inference unit 805 described later, determines whether there is a shooting scene in the cut-out image, or calculates a control value for the attention camera 104.
[0045] (Inference of Object Detection Model, Inference of Pose Estimation Model) The inference unit 805 executes inference processing for object detection and pose estimation on the "imaging data" received from the overhead camera 102 using the "trained model" learned by the learning server 105 by controlling the GPU 213 with the CPU 210. Note that the "inference of the object detection model" inputs image data including the ball 110 and outputs the "rectangular coordinates" circumscribing the ball 110 in the aforementioned "imaging data". Also, the "inference of the pose estimation model" inputs image data including the player 109 and outputs the "joint coordinates" of the player 109 in the aforementioned "imaging data". The data storage unit 806 stores and manages the trained model received from the learning server 105 by controlling the HDD 216 with the CPU 210.
[0046] <Functional configuration of the attention camera 104> The attention camera 104 includes an imaging unit 807, a drive control unit 808, and a communication unit 809. These functional units are realized by the CPU 218 expanding and executing the program stored in the ROM 220 in the RAM 219. The imaging unit 807 captures an image of an angle of view including the player 109 and the ball 110 by the CPU 218 controlling the video engine 221, the image sensor 224, and the zoom drive unit 225. The drive control unit 808 controls the turning amount of the pan-tilt mechanism and the zoom amount of the zoom lens according to the instruction received from the image processing apparatus 103 by the CPU 218 controlling the pan-tilt drive unit 226 and the zoom drive unit 225. The communication unit 809 receives the respective control values of the pan-tilt drive unit 226 and the zoom drive unit 225 from the image processing apparatus 103 by the CPU 218 controlling the NIC 205 and the video I / F 206.
[0047] <Functional configuration of the data collection server 106> The data collection server 106 includes a data collection / providing unit 813 and a data storage unit 814. These functional units are realized by the CPU 236 expanding and executing the program stored in the ROM 238 in the RAM 237.
[0048] (Teacher data for "object detection" and "pose estimation") The data collection / provision unit 813 has a function of collecting the "teacher data" requested from the learning server 105 from the Internet 100 by the CPU 236 controlling the NIC 241, and a function of transmitting the collected "teacher data" to the learning server 105. In this embodiment, as the "teacher data" for object detection, an image including the ball 110 and the "position coordinates" of the ball 110 in the image are used. Also, as the "teacher data" for pose estimation, a "learning data set for pose estimation" consisting of an image showing the whole body of a person and a combination of the position coordinates of each joint of the person shown in the image is used. The data storage unit 814 stores and manages the teacher data collected from the Internet 100 by the CPU 236 controlling the HDD 240.
[0049] <Functional configuration of the client terminal 107> The client terminal 107 includes a display control unit 815, an operation unit 816, and a communication unit 817. These functional units are realized by the CPU 243 expanding and executing the program stored in the ROM 245 in the RAM 244.
[0050] The display control unit 815 displays the image data received from the image processing apparatus 103 on the display unit 249 by the CPU 243 controlling the display unit 249. The operation unit 816 receives the operations and instructions of the user who uses this image processing system by the CPU 243 controlling the input unit 250. The communication unit 817 transmits a control command to the apparatus connected to the Internet 100 and the local network 101 by the CPU 243 controlling the NIC 248.
[0051] <Functional configuration of the learning server 105> The learning server 105 includes a learning data generation unit 810, a learning unit 811, and a data storage unit 812. These functional units are realized by the CPU 229 expanding and executing the program stored in the ROM 231 in the RAM 230.
[0052] The learning data generation unit 810 generates learning data by performing data augmentation on the "teacher data" received by the CPU 229 from the data collection server 106. "Data augmentation" means generating an image different from the original image by performing image conversion processing such as affine transformation on the image of the teacher data and the correct answer data corresponding to the image, and also converting the information of the correct answer data according to the image conversion to increase the number of learning data.
[0053] The learning unit 811 performs learning of each of the "object detection learning model" and the "pose estimation learning model" by the CPU 229 controlling the GPU 232, and generates each learned model of object detection and pose estimation. Specific algorithms of machine learning performed by the learning model include the nearest neighbor method, the naive Bayes method, decision trees, support vector machines, etc. Also, deep learning (deep learning) that generates its own feature amounts and combined weight coefficients for learning using a neural network is also mentioned as a learning model. Appropriate ones among the above algorithms can be selected and applied to this embodiment as appropriate.
[0054] <The "error detection unit" and "update unit" included in the learning unit 811> Note that the learning unit 811 may be configured to include an "error detection unit" and an "update unit". The "error detection unit" obtains the error between the output data output from the output layer of the neural network and the teacher data according to the input data input to the input layer. Further, the error detection unit may calculate the error between the output data from the neural network and the teacher data using a loss function. The "update unit" updates the combined weight coefficients between the nodes of the neural network so that the error becomes smaller based on the error obtained by the error detection unit. The update unit updates the combined weight coefficients and the like using, for example, the error backpropagation method. Here, the error backpropagation method is a method of adjusting the combined weight coefficients between the nodes of each neural network so that the error becomes smaller.
[0055] The data storage unit 812 stores and manages the teacher data received from the data collection server 106 and the learned model learned by the learning unit 811 by the CPU 229 controlling the HDD 233.
[0056] <Premises of the processing flow> Next, before explaining the specific processing flow of this image processing system, the conditions and operations that are the premises of the processing flow will be explained. In the processing flow of this image processing system, it is assumed that the "object detection model" and "pose estimation model" that have been learned by the learning server 105 are already stored in the HDD 216 of the image processing apparatus 103. As a result, the image processing apparatus 103 does not need to request the learned model every time it executes the processing of this embodiment, and can execute each inference process using the "learned model" stored in the HDD 216.
[0057] Here, a method for preparing a learned model will be explained by taking the "object detection model" as an example. The basic operations of the client terminal 107, the data collection server 106, and the image processing apparatus 103 are the same when preparing the learned model.
[0058] Before the start of the processing of this embodiment, the CPU 243 of the client terminal 107 requests the following to the CPU 229 of the learning server 105 via the NIC 248 of the client terminal 107. That is, the CPU 243 of the client terminal 107 requests the generation of an "object detection model" that detects the ball 110 from the reduced image signal 604 of the aerial image signal 600.
[0059] Next, the CPU 229 of the learning server 105 requests the CPU 236 of the data collection server 106 to collect "teacher data" including the ball 110 via the NIC 234 of the learning server 105. Next, the CPU 236 of the data collection server 106 collects the "teacher data" including the ball 110 via the Internet 100.
[0060] Next, the CPU 236 of the data collection server 106 transmits the collected teacher data to the learning server. The CPU 229 of the learning server 105 generates learning data by performing data augmentation on the teacher data received from the data collection server 106. Next, the CPU 229 of the learning server 105 controls the GPU 232 of the learning server 105 to generate an "object detection model" based on the generated learning data. Next, the CPU 229 of the learning server 105 transmits the generated learned model to the image processing device 103. Then, the CPU 210 of the image processing device 103 stores the "learned model" received from the learning server 105 in the HDD 216.
[0061] The method of preparing the "object detection model" for detecting the ball 110 has been described above. The method of preparing the "pose estimation model" may be to perform learning using a network model for pose estimation by preparing "image data" including the player 109 and "joint coordinate data" of the player 109 as "teacher data" in the method described above.
[0062] (FIG. 4: Flowchart showing the processing procedures of the overhead camera, the image processing device, and the attention camera: FIG. 5: Explanatory diagram of the overhead image signal, the attention area image signal, etc.) Next, with reference to FIGS. 4 and 5, the specific processing of this system will be described. Note that FIGS. 5(a), 5(b), and 5(c) are explanatory diagrams schematically showing the contents described in FIGS. 4(a), 4(b), and 4(c). FIG. 4(a) is a flowchart showing the processing executed by the overhead camera 102, and shows the processing executed by the imaging unit 800, the operation unit 801, and the communication unit 802 in FIG. 3. Hereinafter, the processing of the overhead camera 102 will be described with reference to FIG. 4(a).
[0063] (FIG. 4(a): Processing of the overhead camera 102) First, the CPU 201 reads the software of the imaging unit 800 from the ROM 203, expands it in the RAM 202, and executes S300. Specifically, in S300, the CPU 201 controls the video engine 204 and the image sensor 207 to generate imaging data (hereinafter referred to as "overhead image signal 600") of the viewing angle including the entire basketball court 108. Note that FIG. 5(a) is an explanatory diagram of the overhead image signal 600. After that, the CPU 201 stores the generated overhead image signal 600 in the RAM 202 and transfers the process to S301.
[0064] Next, the CPU 201 reads the software of the communication unit 802 from the ROM 203, expands it in the RAM 202, and executes S301. Specifically, in S301, the CPU 201 transmits the overhead image signal 600 stored in the RAM 202 to the image processing device 103 via the video I / F 206. Then, the CPU 201 transfers the process to S302.
[0065] Next, the CPU 201 reads the software of the operation unit 801 from the ROM 203, expands it in the RAM 202, and executes S302. Specifically, in S302, the CPU 201 determines whether there is an operation from the user interface (UI) of the overhead camera 102 (not shown) or an instruction to end imaging from the client terminal 107 via the NIC 205. If it is determined that there is an instruction to end imaging (YES), the CPU 201 ends the imaging. On the other hand, if it is determined that there is no instruction to end imaging (NO), the CPU 201 returns the process to S300 and repeatedly executes the processes of S300 and S301.
[0066] (FIG. 4(b): Processing of the image processing device 103) FIG. 4(b) is a flowchart showing the processing of the image processing apparatus 103, and is a diagram showing the processing executed by the communication unit 803, the processing / arithmetic unit 804, the inference unit 805, and the data storage unit 806 in FIG. 3. Hereinafter, the processing executed by the image processing apparatus 103 will be described with reference to FIG. 4(b). First, the CPU 210 reads the software of the communication unit 803 from the ROM 212, expands it in the RAM 211, and executes S400. Specifically, in S400, the CPU 210 determines whether or not it has received the bird's-eye view image signal 600 (see FIG. 5(a)) from the bird's-eye view camera 102 via the video I / F 214. If it is determined that the bird's-eye view image signal 600 has been received (YES), the CPU 210 stores the bird's-eye view image signal 600 in the RAM 211 and shifts the processing to S401. On the other hand, if it is determined that the bird's-eye view image signal 600 has not been received (NO), the CPU 210 waits for the reception of the bird's-eye view image signal 600 in S400.
[0067] Next, the CPU 210 reads the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S401. Specifically, in S401, the CPU 210 reads the bird's-eye view image signal 600 stored in the RAM 211, and generates a reduced image signal 604 (see FIG. 5(c)) obtained by reducing the bird's-eye view image signal 600. The reduction ratio (hereinafter referred to as "reduction ratio") is a ratio suitable for input to the learned model used in the subsequent inference process, but may be an arbitrary ratio specified by the user. Then, the CPU 210 stores the generated reduced image signal 604 and the reduction ratio in the RAM 211 and shifts the processing to S402.
[0068] Next, the CPU 210 reads out the software of the inference unit 805 from the ROM 212, expands it in the RAM 211, and executes S402. Specifically, in S402, the CPU 210 reads out the "object detection model" from the HDD 216 and stores it in the RAM 211. After that, the CPU 210 controls the GPU 213 to read out the reduced image signal 604 stored in the RAM 211, inputs the reduced image signal 604 into the "object detection model" stored in the RAM 211, and obtains the "ball coordinates 602c". Then, the CPU 210 stores the detected ball coordinates 602c in the RAM 211 and transfers the process to S403. If the ball coordinates cannot be detected, the ball coordinates detected immediately before may be used as the next ball coordinates. Also, a moving trajectory may be predicted from the time series of the ball coordinates detected up to immediately before, and the predicted coordinates calculated from the prediction result may be used as the ball coordinates.
[0069] Next, the CPU 210 reads out the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S403. Specifically, in S403, the CPU 210 reads out the ball coordinates 602c and the reduction ratio stored in the RAM 211, and uses the reduction ratio to convert the ball coordinates 602c on the reduced image signal 604 to the ball coordinates 602a on the bird's-eye view image signal 600. Then, the CPU 210 stores the converted ball coordinates 602a in the RAM 211. After that, based on the ball coordinates 602a stored in the RAM 211, the CPU 210 trims an area including the ball 110 and its vicinity from the bird's-eye view image signal 600 to generate a region-of-interest image signal 601 (see FIG. 5(b)). Note that the process of determining the trimming size in the process of S403 is a characteristic process of the present invention, and will be described in detail with reference to FIGS. 6 and 7 later. Then, the CPU 210 stores the generated region-of-interest image signal 601 in the RAM 211 and transfers the process to S404.
[0070] Here, the target area image signal 601 is an image corresponding to the "target area size", which is the size of the "target area". The "target area" is an area in the captured image corresponding to the imaging target location (basket court 108) of the overhead camera 102. The CPU 210 (determination unit) determines the "target area" based on the distortion information and the target position (ball coordinates 602). Further, the CPU 210 determines the "target area size" corresponding to the actual size based on the "distortion information" and the "target position". Also, as will be clarified in the description of the following embodiments, the "distortion information" is information indicating the distortion from the known shape of the first object (basket court 108) in the captured image captured by the overhead camera 102, and the CPU 210 (distortion information generation unit) generates the distortion information. There are objects to be targeted (second and third objects) in the target area image signal 601.
[0071] That is, in order to fit the horizontally long rectangular basket court 108 within the viewing angle θ of the overhead camera 102, its captured image is distorted into a quadrilateral shape (such as a trapezoid) as shown in FIGS. 5(a) and 11. For this reason, as an example (third embodiment), the CPU 210 generates the "distortion information" based on the projection transformation matrix between the coordinates indicating the vertices of the first object (basket court 108) and the coordinates indicating the vertices of the known shape of the first object. Therefore, the CPU 210 calculates the "target position (ball coordinates 602)" in the captured image, and then determines the "target area" based on the "distortion information" and the calculated "target position", and determines the "target area size", which is its size. The image corresponding to the "target area size" becomes the "target area image (target area image signal 601)". In the embodiments of the present invention, the "target area" is specifically an imaging area for imaging a shooting scene, a ball, etc. Note that the CPU 210 can also determine the "target area" based on the position information (joint coordinates) of the second object (player 109) whose position is estimated in the captured image of the overhead camera 102.
[0072] Next, the CPU 210 reads the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S404. Specifically, in S404, the CPU 210 reads the target area image signal 601 stored in the RAM 211, and scales it according to the input size to the subsequent posture estimation model. Note that the scaling ratio is a ratio suitable for input to the learned model used in the subsequent inference process, but it may be an arbitrary ratio specified by the user. Then, the CPU 210 stores the scaled target area image signal 601 in the RAM 211 and transfers the process to S405. By scaling (enlarging or reducing) the target area image signal 601, the resolution of the target area image signal 601 changes.
[0073] Next, the CPU 210 reads the software of the inference unit 805 from the ROM 212, expands it in the RAM 211, and executes S405. Specifically, in S405, the CPU 210 reads the "posture estimation model" from the HDD 216 and stores it in the RAM 211. Then, the CPU 210 controls the GPU 213 to read the target area image signal 601 stored in the RAM 211 and uses the "posture estimation model" stored in the RAM 211 to detect the "joint coordinates 603 (see FIG. 5(b))" of the player 109a. Note that when multiple players 109 are shown in the target area image signal 601, the joint coordinates 603 are detected for each player 109. The CPU 210 stores the detected joint coordinates 603 in the RAM 211 and transfers the process to S406.
[0074] Next, the CPU 210 reads the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S406. Specifically, in S406, the CPU 210 determines whether the player 109a is shooting based on the ball coordinates 602a and the joint coordinates 603 stored in the RAM 211. Specifically, the CPU 210 converts the coordinate system of the ball coordinates 602a from the coordinate system on the bird's-eye view image signal 600 to the coordinate system on the target area image signal 601, and stores the converted ball coordinates 602b in the RAM 211.
[0075] Then, the CPU 210 reads out the ball coordinates 602b and joint coordinates 603 stored in the RAM 211 and compares the positional relationships of the respective coordinates. For example, when the CPU 210 grasps that the coordinate value of the elbow joint is higher than the coordinate value of the shoulder joint of the player 109a and the coordinate value of the wrist joint is higher than the coordinate value of the elbow joint, it determines that the player 109a is in a state of stretching the arm upward. Here, a higher "coordinate value" means a larger coordinate value on the coordinate axis in the vertical direction.
[0076] Furthermore, when the coordinate value of the ball coordinates 602b is higher than the coordinate value of the wrist joint of the player 109a, the CPU 210 determines that the player 109a is shooting. These determination processes are performed for the number of players 109 for which the joint coordinates 603 are detected, and the joint coordinates 603 with the closest center of gravity of the ball coordinates 602b and the coordinate value of the wrist joint of the player 109 are adopted. As described above, the CPU 210 can determine whether a shot is being made.
[0077] In this embodiment, it is determined whether a shot has been made by comparing the ball coordinates with the joint coordinates of the player, but the determination method is not limited to this. For example, a "decision tree model" that inputs an image of the player 109a in action such as the attention area image signal 601 and outputs a determination result as to whether a shot has been made is prepared, and it may be determined whether a shot has been made using the determination result by the decision tree model. Also, in this embodiment, the type of action to be determined is set to "shoot" for the purpose of imaging the shooting scene of the player, and the type of action (action) is not particularly limited. For example, the present invention can be applied to any action that can be specified from the posture of the player, such as "pass" or "dribble". When it is determined that the player 109a is shooting (YES), the CPU 210 transfers the process to S407. On the other hand, when it is determined that the player 109a is not shooting (NO), the CPU 210 transfers the process to S408.
[0078] Next, based on the processing result of S406, the CPU 210 reads the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S407. Specifically, in S407, the CPU 210 calculates the control value of the pan-tilt drive unit 226 of the target camera 104 for turning the imaging direction of the target camera 104 toward the shoot scene. Here, the "control value" is a combination of "control direction" and "control speed" in the present embodiment. For example, the control value of the pan drive is expressed as "a speed of 10 (degrees / second) in the right direction", etc., and the control speed of the pan is determined based on the difference between the current position and the target position of the pan. Note that the control value may directly specify the target position of the pan. In this way, in S407, the CPU 210 generates a camera control signal necessary for turning the target camera 104 toward the shoot scene.
[0079] When the imaging direction and position of the overhead camera 102 are significantly different from the imaging direction and position of the target camera 104, the positions of the overhead camera 102 and the target camera 104 may be calibrated in advance. In this case, it becomes possible to convert the ball coordinates 602a and the joint coordinates 603 of the player viewed from the overhead camera 102 into the coordinate system of the target camera 104. Thereby, the present invention can be applied even when the imaging direction and position of the overhead camera 102 are significantly different from the imaging direction and position of the target camera 104. Note that as an example of the position calibration method, there is a method of calculating camera parameters including information on the relative position and orientation between cameras by simultaneously imaging a calibration pattern such as a grid pattern with two cameras. Using the calculated camera parameters, corresponding points between the imaging coordinates of the overhead camera 102 and the target camera 104 can be calculated based on epipolar geometry.
[0080] Next, the CPU 210 executes the following calculation based on the ball coordinates 602a stored in the RAM 211 and the size of the target area of the target area image signal 601 (hereinafter, refers to the size of the target area). That is, the CPU 210 calculates the control value of the zoom drive unit 225 of the target camera 104 for enclosing the ball 110 and the player 109 located in the vicinity thereof within the viewing angle θ.
[0081] Here, the specific calculation process of the control value for lens driving will be described. First, the maximum and minimum values of the focal length of the target camera 104 in imaging are defined. Specifically, before the start of the processing of this embodiment, the user measures the focal length T (maximum value) at which the entire bodies of the ball 110 and the player 109 are within the viewing angle θ when the ball 110 is at the deepest position in the basketball court 108. Also, the user measures the focal length W (minimum value) at which the entire bodies of the ball 110 and the player 109 are within the viewing angle θ when the ball 110 is at the frontmost position in the basketball court 108. Here, for the "deepest position" and "frontmost position", when the imaging optical axis of the target camera 104 is installed perpendicular to the sideline of the basketball court 108 and the installation position of the overhead camera 102 is close to that of the target camera 104, the captured image of the target camera 104 can be approximated by the overhead image signal 600. In this case, the upper sideline in the overhead image signal 600 is the "deepest position", and the lower sideline in the overhead image signal 600 is the "frontmost position". Also, as shown in FIG. 1, when the imaging optical axis of the target camera 104 is installed perpendicular to the end line of the basketball court 108, the left end line as viewed from the front side of the overhead camera 102 is the "deepest position", and the right end line is the "frontmost position". Note that the focal lengths T and W are set to the focal lengths at which the entire bodies of the ball 110 and the player 109 are within the viewing angle θ, but the focal length is not necessarily limited to this, and it may be set to the focal length desired by the user.
[0082] After the measured focal lengths T and W are input to the client terminal 107, they are transmitted to the image processing device 103 and stored in the RAM 211. Next, based on the result of position calibration, the CPU 210 converts the installation position of the target camera 104, the deepest position, and the frontmost position into the coordinate system of the overhead image signal 600. Next, the CPU 210 sets the distance from the target camera 104 to the deepest position on the overhead image signal 600 as "LL1", the distance from the target camera 104 to the frontmost position as "LL2", and the distance from the target camera 104 to the ball coordinate 602a as "LL3". Then, the CPU 210 can calculate the target position Z of the zoom associated with the distance from the focal length T to the focal length W by the following equation (1).
[0083]
Number
[0084] The control direction of the lens drive is determined based on the current position of the zoom and the zoom target position Z. For example, when the current position of the zoom is "30 (mm)" and the target position is "100 (mm)", the control direction is set to the "TELE direction". When the subject is to be imaged larger in the zoom drive unit 225, telephoto is referred to as the "TELE direction". Also, the control speed of the lens drive is determined based on the magnitude of the difference between the current position of the zoom and the zoom target position Z. For example, when the difference DD between the current position and the target position of the zoom is less than "10 (mm)", the control speed is set to "0 (mm / sec)", and when the difference DD is "10 mm or more", the control speed is set to "(DD × 10) (mm / sec)". These numerical values are just examples and do not limit the control content. For example, an upper limit value may be set for the control speed.
[0085] Note that in this embodiment, the control value of the zoom is calculated based on the installation position of the target camera 104, but the calculation process is not limited to this. For example, a new path for the image processing device 103 to capture the captured image of the target camera 104 is prepared, and object detection and pose estimation are performed on the captured image of the target camera 104 in the same way as the bird's-eye view image signal 600. As a result, the positions of the player 109 and the ball 110 in the imaging angle θ of the target camera 104 can be estimated, and the zoom may be controlled so that the entire bodies of the ball 110 and the player 109 located in the vicinity thereof are within the imaging angle θ. Also, in this embodiment, the control value of the zoom drive unit 225 is set as the control value required to fit the entire bodies of the ball 110 and the player 109 located in the vicinity thereof within the imaging angle θ. However, the control value of the zoom drive unit 225 is not limited to this. For example, when imaging the player in a bust-up composition, the control value may be obtained as follows. That is, the joint coordinates 603 of the player 109 are converted from the coordinate system on the attention area image signal 601 to the coordinate system on the bird's-eye view image signal 600, and the control value required to fit the joint coordinates above the waist joint of the player 109 within the imaging angle θ may be used.
[0086] Next, the CPU 210 stores the calculated control values of the pan-tilt drive unit 226 and the zoom drive unit 225 (hereinafter referred to as "camera control signals") in the RAM 211 and transfers the process to S409. Based on the processing result of S406, the CPU 210 reads the software of the processing / arithmetic unit 804 from the ROM 212, expands it in the RAM 211, and executes S408. Specifically, in S408, based on the ball coordinates 602a stored in the RAM 211, the CPU 210 calculates the control value of the pan-tilt drive unit 226 of the target camera 104 for directing the imaging direction of the target camera 104 toward the ball coordinates 602a. Further, based on the ball coordinates 602a stored in the RAM 211, the CPU 210 calculates the control value of the zoom drive unit 225 of the target camera 104 for enclosing the ball 110 and the "predetermined area" around it within the angle of view θ.
[0087] Here, the "predetermined area" may be an area including all the players in whom the joint coordinates 603 in the target area image signal 601 are detected. Alternatively, the "predetermined area" may be an area including only the player closest to the ball coordinates 602a. Note that the calculation process of the control values of the pan-tilt drive unit 226 and the zoom drive unit 225 is the same as that in S407.
[0088] In addition, in S408, even when it is determined in the determination of S406 that it is not a shooting scene, the imaging direction of the target camera 104 is directed to the ball coordinate 602a because the shot occurs starting from the ball position. As a result, when a shooting scene occurs, the driving amount of the target camera 104 can be minimized as much as possible, and the time delay until the target camera 104 finishes facing the direction of the shooting scene can be reduced. After that, the CPU 210 stores the calculated camera control signal in the RAM 211 and shifts the process to S409. That is, in both S407 and S408, the CPU 210 generates a camera control signal for directing the target camera 104 in a direction according to the recognition result of whether the player 109a has taken a shot based on the ball coordinate 602a and the joint coordinate 603 in S406. In S407, a camera control signal for directing the target camera 104 to the shooting scene is generated, and in S408, a camera control signal for directing the target camera 104 to the ball coordinate is generated. Therefore, in any case, a camera control signal for directing the imaging direction of the target camera 104 to the target area is generated.
[0089] Next, the CPU 210 reads out the software of the communication unit 803 from the ROM 212, expands it in the RAM 211, and executes S409. Specifically, in S409, the CPU 210 transmits the camera control signal stored in the RAM 211 to the target camera 104 via the NIC 215.
[0090] (Fig. 4(c): Processing of the target camera 104) Fig. 4(c) is a flowchart showing the processing of the target camera 104 and is an explanatory diagram of the processing of the imaging unit 807, the drive control unit 808, and the communication unit 809 in Fig. 3. Hereinafter, the processing of the target camera 104 will be described with reference to Fig. 4(c).
[0091] First, the CPU 218 reads the software of the communication unit 809 from the ROM 220, expands it in the RAM 219, and executes S500. Specifically, in S500, the CPU 218 determines whether it has received a "camera control signal" from the image processing apparatus 103 via the NIC 222. If it is determined that the camera control signal has been received (YES), the CPU 218 stores the received camera control signal in the RAM 219 and transfers the process to S501. On the other hand, if it is determined that the camera control signal has not been received (NO), the CPU 218 waits for reception in S500.
[0092] Next, the CPU 218 reads the software of the drive control unit 808 from the ROM 220, expands it in the RAM 219, and executes S501. Specifically, in S501, the CPU 218 controls the pan-tilt drive unit 226 and the zoom drive unit 225 based on the camera control signal stored in the RAM 219.
[0093] Next, the CPU 218 reads the software of the imaging unit 807 from the ROM 220, expands it in the RAM 219, and executes S502. Specifically, in S502, the CPU 218 controls the video engine 221 and the image sensor 224 to generate imaging data of an angle of view θ including the player 109 and the ball 110 in the imaging direction. That is, in S502, the imaging unit 807 images a region of interest including the shooting scene of the player 109 in the imaging direction and the ball position existing in the imaging direction.
[0094] (Fig. 6: Flowchart showing the generation process of the region of interest image: Fig. 7: Explanatory diagram) Next, with reference to FIGS. 6 and 7, the generation process of the region of interest image signal 601, which is a characteristic process of this system, will be described. Note that FIG. 7 is an explanatory diagram schematically showing the content described in FIG. 6. FIG. 6 is a flowchart showing the detailed process of S403 in the flowchart showing the process executed by the image processing apparatus 103. Hereinafter, the process of the image processing apparatus 103 will be described with reference to FIG. 6.
[0095] First, in S700, the CPU 210 performs an operation based on the Hough transform on the bird's-eye view image signal 600 stored in the RAM 211. Here, the quadrilateral formed by the resulting straight lines is regarded as the court line of the basketball court 108. If there is a court line of a color different from that of the court line of the basketball court 108 in the bird's-eye view image signal 600, it is also possible to perform the operation after specifying the color of the court line of the basketball court 108. Also, if the color of the floor is different inside and outside the basketball court 108, the boundary line of the floor color may be used as the court line. The CPU 210 stores the determined court line of the basketball court 108 in the RAM 211 and transfers the process to S701. Note that the operation process of the court line is not limited to this. For example, the boundary of the quadrilateral may be calculated by performing edge extraction processing such as the SOBEL filter or CANNY method on the bird's-eye view image signal 600 to determine the court line. Eventually, in S700, the CPU 210 performs a Hough transform on the bird's-eye view image signal 600 to extract the court line of the basketball court 108.
[0096] Next, in S701, the CPU 210 obtains the quadrilateral formed based on the court line of the basketball court 108 stored in the RAM 211, and calculates the coordinate values of the four vertices of the quadrilateral (hereinafter referred to as "court coordinates"). Here, as shown in FIG. 7, the four vertices of the quadrilateral are the upper left vertex of the quadrilateral as "A", the lower left vertex as "B", the lower right vertex as "C", and the upper right vertex as "D". That is, the quadrilateral formed based on the court line is quadrilateral ABCD. The CPU 210 stores the calculated court coordinates in the RAM 211 and transfers the process to S702.
[0097] Next, in S702, the CPU 210 calculates the "upper base L1" consisting of the line segment AD and the "lower base L2" consisting of the line segment BC of the quadrilateral ABCD based on the court coordinates stored in the RAM 211 (see FIG. 7). Then, the CPU 210 stores the positions and lengths of the calculated upper base L1 and lower base L2 in the RAM 211 and transfers the process to S703.
[0098] Next, in S703, based on the ball coordinates 602 stored in the RAM 211 and the positions of the upper base L1 and the lower base L2, the CPU 210 calculates a straight line that is parallel to the upper base L1 and the lower base L2 and passes through the center of gravity of the ball coordinates 602. In the case where the upper base L1 and the lower base L2 on the bird's-eye view image signal 600 are not parallel, it will be described in the third embodiment later. Next, on the aforementioned straight line, the CPU 210 calculates a line segment PQ as the "line segment L3", where the intersection with the line segment AB is "P" and the intersection with the line segment CD is "Q" (see FIG. 7). The CPU 210 stores the length of the calculated line segment L3 in the RAM 211 and transfers the process to S704.
[0099] Next, in S704, based on the lower base L2 and the length of the line segment L3 stored in the RAM 211, the CPU 210 calculates the "target area size", which is the image size of the target area image signal 601, from the ratio of the lengths of each line segment. Specifically, before the start of the process of this embodiment, the user captures the bird's-eye view image signal 600 when the player 109 is located on the lower side line in FIG. 1 and measures the full body size of the player 109 in the image. Then, the CPU 210 stores the aforementioned full body size input by the user via the client terminal 107 in the RAM 211 and generates a "reference target area size" based on the full body size. Here, the "reference target area size" is the length of one side of the quadrilateral forming the "target area", the vertical length is twice the height of the player 109 on the bird's-eye view image signal 600, and the horizontal length is such that the ratio of the horizontal length to the vertical length is "1:1". Note that the "reference target area size" is not limited to this, and any size that allows the shooting player to be trimmed without being missed is acceptable.
[0100] Then, the CPU 210 stores the generated "reference target area size" in the RAM 211. The CPU 210 calculates the target area size from the ratio of the reference target area size stored in the RAM 211, the lower base L2, and the line segment L3. When the reference target area size is "SB", the length of the lower base L2 is "L2", and the length of the line segment L3 is "L3", the target area size S is obtained by the following formula 2.
[0101]
Equation
[0102] Thereafter, the CPU 210 stores the size of the attention area in the RAM 211 and transfers the process to S705. By the above process, the size of the attention area becomes smaller as the distance from the overhead camera 102 to the ball 110 increases. Conversely, the size of the attention area becomes larger as the distance from the overhead camera 102 to the ball 110 decreases.
[0103] Next, in S705, the CPU 210 calculates "attention area coordinates 701 (see FIG. 7)" based on the ball coordinates 602 stored in the RAM 211 and the size of the attention area. Here, the "attention area coordinates 701" are the coordinates of the upper left vertex and the lower right vertex of an area with one side of the rectangle being the size of the attention area (or a rectangular area enlarged by a preset enlargement value in all four directions) centered on the center of gravity of the ball coordinates 602. The CPU 210 stores the calculated attention area coordinates 701 in the RAM 211 and transfers the process to S706.
[0104] Next, in S706, the CPU 210 trims the "attention area" from the overhead image signal 600 based on the attention area coordinates 701 stored in the RAM 211 to generate an attention area image signal 601. Thereafter, the CPU 210 stores the generated attention area image signal 601 in the RAM 211. The above describes the process of S403 in the process of the image processing apparatus 103. The subsequent process will follow S404 (see FIG. 4(b)) executed by the image processing apparatus 103.
[0105] The processing in the first embodiment executed according to the configuration of FIG. 1 has been described above. As described above, it is possible to trim the player (second object) in the captured image to a size based on the position of the ball (third object) on the basketball court 108 (first object). Thereby, since the trimming size is smaller than the size of the player, it may be cut off, or conversely, since it is larger, the subject may be crushed by the reduction before recognition, and the possibility that predetermined recognition cannot be performed can be reduced. Note that the application range of the present invention is not limited to the above, and it can be applied to various scenes such as other sports different from basketball, music live, lectures in a lecture hall, etc. For example, it is suitable for sports in which a player recognizes actions such as shooting, serving, and smashing in the vicinity of a ball such as soccer and tennis.
[0106] In the first embodiment, an example of automatic imaging from the imaging data of one overhead camera 102 has been described, but the present invention is not limited to this. The present invention can also be applied to a system using a plurality of overhead cameras. In this case, position calibration is executed for the plurality of overhead cameras, and each imaging data is integrated based on the calibration result to generate one imaging data. Then, the present invention may be applied to the integrated imaging data. Thereby, the present invention can be easily applied to sports and locations where the competition area such as soccer and rugby is wide.
[0107] Furthermore, when the present invention is applied to a scene where a subject such as a music live performance or a lecture in a lecture hall performs a predetermined action such as a singing pose or pointing at a blackboard, it may not be possible to detect the court coordinates from the bird's-eye view image signal 600. In this case, the present invention is applicable because the court line can be calculated by installing a target straight line so as to surround the stage. Alternatively, marks of a specific color and a specific shape may be attached to the real-world positions that are the vertices of a rectangular stage instead of the court line, and the rectangle ABCD forming the stage may be calculated by calculating the marks from the imaging data of the bird's-eye camera 102. That is, the CPU 210 may be configured to detect marks provided at respective locations in the real world corresponding to the vertices of the object of the basketball court 108 and acquire the area having the marks as vertices as the object shape of the basketball court 108.
[0108] (Wearing of position sensor) Also, in the first embodiment, in S402 of the processing flow of the image processing apparatus 103, the CPU 210 detected the position coordinates of the ball 110 using the object detection model, but the detection process of the ball is not limited to this. For example, a position sensor may be attached to the ball 110, and in the system configuration of FIG. 1, a position sensor receiving device 115 that receives the signal of the position sensor may be provided. Position calibration may be performed in advance between the coordinate system of the position sensor receiving device 115 and the coordinate system on the bird's-eye view image signal 731 captured by the bird's-eye camera 102. The position sensor receiving device 115 receives the position of the position sensor attached to the ball 110. The position sensor receiving device 115 transmits the position obtained from the position sensor to the image processing apparatus 103. The image processing apparatus 103 converts the position of the position sensor received from the position sensor receiving device 115 into the coordinate system on the bird's-eye view image signal 600, and uses the converted position of the position sensor as the ball coordinates 602. Thus, a sensor for detecting the position of the ball 110 may be used. Thus, the CPU 210 may acquire the position information of the sensor that detects the position of the ball 110, which is an object in the real world corresponding to the ball object (third object), and calculate the attention position based on the position information acquired by the sensor.
[0109] Also, in the present embodiment, the position coordinates of the ball 110 are detected as the position to be noted from the bird's-eye view image signal 600 on the premise that the highlight scene occurs in the vicinity of the ball 110, but the present invention is not limited to this. For example, the position coordinates of a specific player 109 instead of the ball 110 may be detected as the position to be noted. In this case, before the start of the processing of the present embodiment, an object detection model that inputs image data including the player 109 and performs an inference to output the "rectangular coordinates" circumscribing the player 109 in the above-described image data may be used. Specifically, first, in S402 of FIG. 4(b), the CPU 210 of the image processing apparatus 103 detects the position coordinates of the player 109 from the reduced image signal 604 using the learned model prepared by controlling the GPU 213. Next, in S403 of FIG. 4(b), the CPU 210 of the image processing apparatus 103 generates the attention area image signal 601 from the bird's-eye view image signal 600 based on the detected position coordinates of the player 109, and the processing of the present invention can be similarly executed. As a result, an effect of obtaining a highlight scene only of the player of interest can be obtained.
[0110] Also, in the present embodiment, in S700 which is the processing of the image processing apparatus 103, the CPU 210 calculates the court line of the basketball court 108 using the Hough transform, but the method of calculating the court line is not limited to this. The present invention can be applied as long as it is an operation mode capable of calculating the court line. For example, a user interface (UI) may be used to allow the user to specify the vertices of the basketball court 108. Specifically, before the start of the processing of the present embodiment, the CPU 243 of the client terminal 107 operates the bird's-eye view camera 102 via the NIC 248 of the client terminal 107 and receives the bird's-eye view image signal 600. Next, the CPU 243 displays the received bird's-eye view image signal 600 on the display unit 249 of the client terminal 107. Next, the CPU 243 receives an operation of the input unit 250 of the client terminal. The user operating the client terminal 107 designates the four vertices of the basketball court 108 on the displayed bird's-eye view image signal 600. Next, the CPU 243 transmits the four coordinates on the designated bird's-eye view image signal 600 to the image processing apparatus 103 via the NIC 248.
[0111] Next, if the CPU 210 uses a quadrilateral consisting of four coordinates on the bird's-eye view image signal 600 received from the client terminal 107 as the court line, the present invention can be similarly applied. As a result, even when the court line cannot be calculated from the bird's-eye view image signal 600, an effect of applying the present invention can be obtained. Further, for example, a distorted shape on the bird's-eye view image signal 600 may be estimated from the known shape of the basketball court 108, and the court line may be detected. Specifically, first, the CPU 210 of the image processing apparatus 103 estimates the distorted shape of the basketball court 108 based on the imaging angle of the bird's-eye view camera 102 with respect to the basketball court 108. The CPU 210 of the image processing apparatus 103 acquires a distorted quadrilateral (i.e., the trapezoidal basketball court 108 shown in FIG. 7) based on the estimation result.
[0112] Next, the CPU 210 of the image processing apparatus 103 detects the position and shape of an object in a similar relationship with the operation result obtained by performing edge operation processing on the bird's-eye view image signal 600 and the distorted quadrilateral. Next, if the CPU 210 of the image processing apparatus 103 acquires the court line based on the detected position and shape of the object, the present invention can be similarly applied. As a result, even when the court line cannot be calculated from the bird's-eye view image signal 600, the court line of the basketball court 108 can be calculated and the present invention can be applied. Also, in this way, the "distortion information" is generated when the CPU 210 performs estimation of a distorted shape, predetermined detection, and the like.
[0113] <Second Embodiment> Next, a second embodiment of the present invention will be described. In the second embodiment, a process of obtaining the shape of the basketball court 108 by using the camera information acquired from the overhead camera 102 and generating the attention area image signal 601 will be described. Since the basic configuration of the image processing system is the same as that of the first embodiment, duplicate descriptions will be omitted, and the processing of the different image processing apparatus 103 will be described. Further, in the second embodiment, it is assumed that the imaging optical axis 114 of the overhead camera 102 and the half line of the basketball court 108 are parallel as shown in FIG. 1, and the side line in front of the basketball court 108 fits within the angle of view θ. That is, the length in the horizontal direction of the overhead image signal 600 and the length of the side line in front of the basketball court 108 in the overhead image signal 600 are the same or approximately equal.
[0114] (FIG. 8: Flowchart showing the attention area image generation process of the second embodiment: FIG. 9: Explanatory diagram) FIG. 8 is a diagram showing the generation process of the attention area image signal 601, which is a characteristic process of the image processing apparatus 103 according to the second embodiment, and FIGS. 9(a) and 9 are explanatory diagrams of the generation process of FIG. 8. The process until the CPU 210 of the image processing apparatus 103 receives the overhead image signal 600 from the overhead camera 102 and detects the ball position from the reduced image signal 604 after reducing the overhead image signal 600 is the same as that up to S402 in the process in the first embodiment (see FIG. 4(b)). The following description is the content of the process of the second embodiment in S403.
[0115] First, in S900, the CPU 210 of the image processing apparatus 103 requests the CPU 201 of the overhead camera 102 to transmit information indicating the focal length of the lens (not shown) of the overhead camera 102 and the tilt of the overhead camera 102 via the NIC 215. In response to this, the CPU 201 acquires the focal length of the overhead camera 102 and the "tilt k" from the zoom drive unit 208 and the tilt sensor 227. Next, the CPU 201 transmits the focal length and the "tilt k" acquired via the NIC 205 to the image processing apparatus 103. In response to this, the CPU 210 acquires the focal length and the "tilt k" received from the overhead camera 102, stores them in the RAM 211, and shifts the process to S901.
[0116] Next, in S901, the CPU 210 calculates the viewing angle θ of the overhead camera 102 from the focal length stored in the RAM 211. Next, the CPU 210 calculates the "distance d1" from the overhead camera 102 to the front sideline based on the viewing angle θ and the length (SL) of the front sideline of the basketball court 108 in the real world. Specifically, in S901, when the length of the front sideline of the real-world basketball court 108 is "SL", the "distance d1" from the overhead camera 102 to the sideline SL is obtained by the following Equation 3.
[0117]
Equation
[0118] The CPU 210 stores the calculated "distance d1" in the RAM 211 and transfers the process to S902. In S902, the CPU 210 calculates the height h from the plane where the basketball court 108 exists in the real world to the overhead camera 102 based on the "tilt k" and the "distance d1" stored in the RAM 211. Specifically, when the tilt of the overhead camera 102 with respect to the horizontal direction is "k" and the distance from the overhead camera 102 to the sideline SL is "d1", the "height h" of the overhead camera 102 is obtained by the following Equation 4.
[0119]
Equation
[0120] The CPU 210 stores the calculated "height h" in the RAM 211. Based on the "distance d1" and "height h" stored in the RAM 211 and the length "EL" of the end line of the basketball court 108 in the real world, the CPU 210 calculates the "distance d2" from the overhead camera 102 to the back sideline. Specifically, first, the CPU 210 sets the distance from the overhead camera 102 to the front sideline SL as "d1", the height of the overhead camera 102 as "h", and the length of the end line of the basketball court 108 in the real world as "EL". Then, the CPU 210 obtains the "distance d2" from the overhead camera 102 to the back sideline using the following formula 5. The CPU 210 stores the calculated "distance d2" in the RAM 211 and transfers the process to S903.
[0121]
Number
[0122] Next, in S903, the CPU 210 calculates the following "distance d3" based on the "tilt k" and "height h" of the overhead camera 102, "distance d1", and "distance d2" stored in the RAM 211. That is, it calculates the "distance d3" of the straight line connecting from the front sideline SL of the basketball court 108 in the real world to the intersection point of the straight line perpendicular to the straight line of "distance d1" and the straight line of "distance d2". Specifically, the CPU 210 sets the tilt of the overhead camera 102 in the horizontal direction as "k", the height of the overhead camera 102 as "h", and the distance from the overhead camera 102 to the front sideline SL as "d1". Furthermore, when the distance from the overhead camera 102 to the back sideline is "d2", the "distance d3" from the sideline SL to "distance d2" is obtained using the following formula 6. After that, the CPU 210 stores the calculated "distance d3" in the RAM 211 and transfers the process to S904.
[0123]
Number
[0124] Next, in S904, the CPU 210 specifies a trapezoid whose upper base, lower base, and height respectively correspond to "distance d1", "distance d2", and "distance d3" based on the "distance d1", "distance d2", and "distance d3" stored in the RAM 211. Specifically, let the distance from the overhead camera 102 to the front side line SL be "d1", and the distance from the overhead camera 102 to the rear side line be "d2". At this time, the following relational expression 7 holds for the upper base and the lower base of the trapezoid obtained from the overhead image signal 600.
[0125]
Number
[0126] Subsequently, let the distance from the side line SL to the line segment of "distance d2" be "d3". At this time, the following relational expression 8 holds for the height of the trapezoid obtained from the overhead image signal 600.
[0127]
Number
[0128] Then, the CPU 210 calculates the coordinates of each vertex of the trapezoid (hereinafter referred to as "trapezoid coordinates") in the overhead image signal 600 based on the specified trapezoid. The CPU 210 stores the obtained trapezoid coordinates in the RAM 211. The CPU 210 calculates the line segment L3 based on the ball coordinates 602 stored in the RAM 211 according to the process of S703 in FIG. 6 described in the first embodiment and the positions of the upper base and the lower base of the trapezoid obtained from the trapezoid coordinates. The CPU 210 stores the length of the calculated line segment L3 in the RAM 211 and transfers the process to S905. As shown in FIG. 7, the line segment L3 is parallel to the upper base and the lower base of the specified trapezoid and is the line segment PQ passing through the center of gravity of the ball coordinates.
[0129] Next, in S905, the CPU 210 executes the same process as S704 in the first embodiment to calculate the "attention area size". That is, the CPU 210 calculates the attention area size based on the ratio of the length of the lower base L2 and the line segment L3. The lower base L2 here is shown in FIG. 7. The CPU 210 stores the calculated attention area size in the RAM 211 and transfers the process to S906.
[0130] Next, in S906, the CPU 210 executes the same process as S705 in the first embodiment 1 to calculate the attention area coordinates 701. That is, the CPU 210 calculates the attention area coordinates 701 based on the ball coordinates 602 and the attention area size. The CPU 210 stores the calculated attention area coordinates 701 in the RAM 211 and transfers the process to S907.
[0131] Next, in S907, the CPU 210 executes the same process as S706 in the first embodiment to generate the attention area image signal 601. That is, the CPU 210 calculates the attention area image signal 601 based on the attention area coordinates 701. Then, the CPU 210 stores the generated attention area image signal 601 in the RAM 211. The above is the description of the generation process of the attention area image signal 601, which is the characteristic process of the image processing apparatus 103 in the second embodiment. The subsequent processes will follow S404 in the process of the image processing apparatus 103 in the first embodiment (see FIG. 4(b)).
[0132] The above is the description of the process of the second embodiment of the present invention in the configuration of the image processing system in FIG. 1. In this way, it can be obtained based on the camera information acquired from the overhead camera 102 of the object shape of the basketball court 108. Thereby, even when the object shape of the basketball court 108 cannot be detected by the process executed by the image processing apparatus 103 from the overhead image signal 600, the present invention can be applied without the need for additional operations by the user.
[0133] <Third Embodiment> Next, a third embodiment of the present invention will be described. In the third embodiment, a process of generating a region of interest image signal 601 when the imaging optical axis 114 of the overhead camera 102 is not parallel to the half line of the basketball court 108 will be described. Since the basic system configuration is the same as that of the first embodiment, redundant descriptions will be omitted, and the processing of the image processing apparatus 103, which is the difference, will be described.
[0134] (FIG. 10: Flowchart showing the region of interest image generation process of the third embodiment: FIG. 11: Explanatory diagram) FIG. 10 is a diagram showing a process of generating a region of interest image signal 601, which is a characteristic process of the image processing apparatus 103 according to the third embodiment. FIG. 11 is an explanatory diagram of the process of FIG. 10. The process from receiving the overhead image signal 600 from the overhead camera 102 and reducing the overhead image signal 600 to detecting the ball position from the reduced image signal 604 is the same as the process executed by the image processing apparatus 103 of the first embodiment (see FIG. 4(b)). That is, the process up to S402 is the same as that of the first embodiment. The following description is the content of the process of S403 in the first embodiment.
[0135] First, in S1100, the CPU 210 executes the same process as S700 of the first embodiment and performs an operation on the overhead image signal 600 based on the Hough transform process. Here, as a result of the operation, the quadrilateral formed by the obtained straight line is regarded as the court line of the basketball court 108. The CPU 210 stores the court line of the basketball court 108 obtained by the operation in the RAM 211 and shifts the process to S1101.
[0136] Next, in S1101, the CPU 210 executes the same processing as S701 in the first embodiment to calculate the court coordinates. That is, the CPU 210 calculates the vertex coordinates of the four vertices A, B, C, and D of the quadrilateral formed by the court lines. The four vertices A, B, C, and D are shown in FIG. 11. In the third embodiment, since the imaging optical axis 114 of the overhead camera 102 is not parallel to the half line of the basketball court 108, the quadrilateral of the basketball court 108 shown in the overhead image signal 600 has a distorted shape (see FIG. 11(a)). Accordingly, the calculated court coordinates are different from the court coordinates in the first embodiment and are the vertex coordinates of the distorted quadrilateral. The CPU 210 stores the calculated court coordinates in the RAM 211 and transfers the process to S1102.
[0137] Next, in S1102, the CPU 210 performs a projective transformation between the coordinate system of the overhead image signal 600 and the basketball court coordinate system defined based on the known shape of the basketball court 108. Specifically, the CPU 210 generates a "projective transformation matrix" between the court coordinates on the coordinate system of the overhead image signal 600 stored in the RAM 211 and the vertex coordinates of the four vertices of the basketball court on the basketball court coordinate system. The CPU 210 stores the calculated projective transformation matrix in the RAM 211 and transfers the process to S1103.
[0138] Next, in S1103, the CPU 210 projects and transforms the ball coordinates 602a from the coordinate system on the overhead image signal 600 to the basketball court coordinate system based on the projective transformation matrix stored in the RAM 211. As a result, as shown in FIG. 11(b), the ball coordinates become 602d. The CPU 210 stores the ball coordinates 602d in the basketball court coordinate system in the RAM 211 and transfers the process to S1104.
[0139] Next, in S1104, the CPU 210 calculates a straight line that is parallel to the end line of the basketball court 108 and passes through the centroid of the ball coordinates 602d based on the ball coordinates 602d stored in the RAM 211. On the calculated straight line, when the CPU 210 sets the intersection with the line segment AD of the quadrilateral ABCD in the basketball court coordinate system as "R" and the intersection with the line segment BC as "T", the CPU 210 calculates the line segment RT as the "line segment L4". The CPU 210 stores the length of the calculated "line segment L4" in the RAM 211 and transfers the process to S1105.
[0140] Next, in S1105, the CPU 210 calculates the "attention area size", which is the image size of the attention area image signal 601, from the ratio in which the centroid of the ball coordinates 602d internally divides the length of the line segment L4 based on the length of the line segment L4 stored in the RAM 211. Specifically, before the start of the process of this embodiment, the size of the entire body of the player 109 on the overhead image signal 600 when the player 109 is located on the front and back side lines of the basketball court 108 is memorized. Then, based on the above-mentioned image size, a "front reference attention area size" and a "back reference attention area size" are generated.
[0141] Here, the "reference attention area size" is a size in which the vertical size is twice the size of the player 109 on the overhead image signal 600, and the horizontal size has a ratio of "1:1" to the vertical size. Note that the "reference attention area size" may be any size as long as the shooting player can be trimmed without being missed. The CPU 210 stores the generated front and back "reference attention area sizes" in the RAM 211. The CPU 210 calculates the line segment from the intersection R to the centroid of the ball coordinates 602d on the line segment L4 as the "line segment L5". The CPU 210 stores the length of the calculated "line segment L5" in the RAM 211.
[0142] Then, the CPU 210 expands or contracts the reference attention area size according to the ratio of the length of line segment L5 to the length of line segment L4 to generate a new attention area size. Specifically, when the previous reference attention area size is "SF", the reference attention area size at the back is "SB", the length of line segment L4 is "L4", and the length of line segment L5 is "L5", the new "attention area size S" is obtained by the following formula 9. The CPU 210 stores the length of the new attention area size in the RAM 211 and shifts the process to S1106.
[0143]
Number
[0144] Next, in S1106, the CPU 210 calculates the attention area coordinates 701 based on the ball coordinates 602d and the attention area size stored in the RAM 211 according to the process of S705 in the first embodiment. The CPU 210 stores the calculated attention area coordinates 701 in the RAM 211 and shifts the process to S1107.
[0145] Then, in S1107, the CPU 210 trims the attention area from the bird's-eye view image signal 600 based on the attention area coordinates 701 stored in the RAM 211 according to the process of S706 in the first embodiment to generate an attention area image signal 601. The CPU 210 stores the generated attention area image signal 601 in the RAM 211.
[0146] The generation process of the attention area image signal 601, which is the characteristic process of the image processing apparatus 103 in this embodiment, has been described above. In this way, the CPU 210 can generate distortion information based on the projection transformation matrix between the coordinates indicating the vertices of the first object (basket court 108) and the coordinates indicating the vertices of the known shape of the first object. Then, the CPU 210 can determine the "attention area" based on the distortion information and the calculated ball coordinates 602 (attention position), and can determine the "attention area size" which is its size. The image corresponding to the "attention area size" becomes the "attention area image".
[0147] The subsequent processing will follow S404 in the processing flow of the image processing apparatus 103 of the first embodiment (see Fig. 4(b)). Above, the processing of the third embodiment of the present invention has been described with the configuration of Fig. 1. In this way, the court coordinates in the coordinate system on the bird's-eye view image signal 600 can be projected and transformed into the basket court coordinate system which is a known shape. Thereby, the present invention can be applied even when the imaging optical axis 114 of the bird's-eye view camera 102 is not parallel to the half line of the basket court 108.
[0148] Further, the CPU 210 (or the CPU 243) may be provided with a user interface (UI unit) that acquires the object shape received by the user operation as the shape of the object (basket court 108). Further, the CPU 210 may be configured to acquire optical information such as the optical axis direction and focal length of the lens of the bird's-eye view camera 102 that captures the bird's-eye view image, and generate distortion information based on the acquired optical information. Further, the shape information of each object in the object group composed of objects of a plurality of types may be stored in the HDD 216 or the like. In this case, the CPU 210 may execute the following processing. That is, when there is a correlation between any of the plurality of types of shape information stored in the HDD 216 or the like and the object in the captured image of the bird's-eye view camera 102, the CPU 210 acquires the object of the shape information as the object of the basket court 108.
[0149] <Supplementary Note> The disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) An image processing apparatus capable of recognizing a predetermined scene in a captured image, An input unit that inputs a captured image of a first object, A distortion information generation unit that generates distortion information indicating distortion from a known shape of the shape of the first object in the captured image input by the input unit, A calculation unit that calculates a position of interest in the captured image, A determination unit that determines a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit, An image generation unit that generates a region of interest image from the captured image based on the region of interest determined by the determination unit; A scaling unit that scales the region of interest image generated by the image generation unit to a predetermined size; A recognition unit that performs a predetermined recognition based on the region of interest image scaled by the scaling unit, and is provided with The determination unit determines the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit. An image processing apparatus characterized by that. (Configuration 2) The distortion information generation unit Among the information regarding the first object in the captured image, the image processing apparatus according to Configuration 1, characterized in that distortion information is generated based on one or more of the degree of distortion from the known shape of the object shape, size, and shape. (Configuration 3) The image processing apparatus further includes a shape acquisition unit that acquires the shape of the first object in the captured image, and the distortion information generation unit generates distortion information based on the shape of the first object acquired by the shape acquisition unit. The image processing apparatus according to Configuration 1 or 2, characterized by that. (Configuration 4) The shape acquisition unit further The image processing apparatus according to claim 3, further comprising a UI unit that acquires the object shape received by a user operation as the shape of the first object. (Configuration 5) The image processing apparatus further includes a shape storage unit that stores shape information of each object in an object group composed of objects of a plurality of types of shapes, The shape acquisition unit further When there is a correlation between any of the plurality of types of shape information stored by the shape storage unit and the object in the captured image, the image processing apparatus according to Configuration 3, characterized in that the object of the shape information is acquired as the shape of the first object. (Configuration 6) The shape acquisition unit further The image processing apparatus according to Configuration 3, characterized in that it detects landmarks provided at respective locations in the real world corresponding to respective vertices of the first object, and acquires, as the shape of the first object, a region having each landmark as a vertex. (Configuration 7) The distortion information generation unit generates distortion information based on a projective transformation matrix between coordinates indicating vertices of the first object acquired by the shape acquisition unit and coordinates indicating vertices of a known shape of the first object. The image processing apparatus according to Configuration 3. (Configuration 8) The determination unit determines the region of interest based on position information of a second object whose position is estimated in the captured image. The image processing apparatus according to Configuration 1 or 2. (Configuration 9) Further includes a position information acquisition unit that acquires position information from a sensor that performs position detection mounted on an object in the real world corresponding to a third object. The calculation unit calculates the position of interest based on the position information acquired by the sensor. The image processing apparatus according to Configuration 1 or 2. (Configuration 10) The recognition unit determines whether or not the second object in the region of interest image has performed a predetermined action. The image processing apparatus according to claim 1 or 2. (Configuration 11) The recognition unit determines whether or not the second object has performed a predetermined action based on the position of the third object and the position of the second object. The image processing apparatus according to Configuration 10. (Configuration 12) Further includes a first learned model that estimates the position of the second object from the region of interest image. Estimates the position of the second object from the region of interest image using the first learned model. The image processing apparatus according to claim 11. (Configuration 13) Further includes a second learned model that detects the position of the third object from the region of interest image. The image processing apparatus according to configuration 11, wherein the position of the third object is detected from the reduced region-of-interest image using the second learned model. (Configuration 14) The position of the second object is the position of a joint of a player who conducts a battle competition in an arena in the real world corresponding to the first object. The image processing apparatus according to any one of configurations 11 to 13, wherein the position of the third object is the position of an object used for the battle competition. (Configuration 15) The image processing apparatus further includes an optical information acquisition unit that acquires optical information including the optical axis direction and focal length of a lens of a camera that captures the captured image. The distortion information generation unit generates distortion information based on the optical information acquired by the optical information acquisition unit, the image processing apparatus according to configuration 1 or 2. (System 1) An image processing system including a first camera capable of capturing a captured image of a first object, a second camera that captures an image in a designated imaging direction, and an image processing apparatus capable of controlling the imaging direction of the second camera. The first camera includes a first imaging unit capable of acquiring a captured image of the first object, and a first transmission unit that transmits the captured image acquired by the first imaging unit to the image processing apparatus. The image processing apparatus includes a first reception unit that receives the captured image transmitted by the first transmission unit, a distortion information generation unit that generates distortion information indicating distortion from a known shape of the shape of the first object in the captured image received by the first reception unit, a calculation unit that calculates a position of interest in the captured image, a determination unit that determines a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit, an image generation unit that generates a region-of-interest image from the captured image based on the region of interest determined by the determination unit, a scaling unit that scales the region-of-interest image generated by the image generation unit to a predetermined size. A recognition unit that performs a predetermined recognition on the region-of-interest image enlarged and reduced by the enlargement / reduction unit; A control signal generation unit that generates a control signal for directing the second camera in the imaging direction according to the recognition result of the recognition unit; A second transmission unit that transmits the control signal generated by the control signal generation unit to the second camera. The determination unit has a function of determining a region-of-interest size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit. The second camera A second imaging unit capable of imaging; A second reception unit that receives the control signal transmitted by the second transmission unit; A control unit that controls the imaging direction of the second imaging unit based on the control signal received by the second reception unit. An image processing system characterized by comprising these components. (Method 1) A control method for an image processing apparatus capable of recognizing a predetermined scene in a captured image, comprising: An input step of inputting a captured image of a first object; A distortion information generation step of generating distortion information indicating distortion from a known shape of the shape of the first object in the captured image input in the input step; A calculation step of calculating a position of interest in the captured image; A determination step of determining a region of interest based on the distortion information generated in the distortion information generation step and the position of interest calculated in the calculation step; An image generation step of generating a region-of-interest image from the captured image based on the region of interest determined in the determination step; An enlargement / reduction step of enlarging and reducing the region-of-interest image generated in the image generation step to a predetermined size; A recognition step of performing a predetermined recognition based on the region-of-interest image enlarged and reduced in the enlargement / reduction step. The determination step is a step of determining a target region size, which is the size of the target region corresponding to the actual size, based on the distortion information generated in the distortion information generation step and the target position calculated in the calculation step. A control method for an image processing apparatus, characterized in that. (Program 1) A program that causes a computer to execute a control method for an image processing apparatus capable of recognizing a predetermined scene in a captured image, The control method includes: An input step of inputting a captured image of a first object, A distortion information generation step of generating distortion information indicating distortion from a known shape of the shape of the first object in the captured image input in the input step, A calculation step of calculating a target position in the captured image, A determination step of determining a target region based on the distortion information generated in the distortion information generation step and the target position calculated in the calculation step, An image generation step of generating a target region image from the captured image based on the target region determined in the determination step, A scaling step of scaling the target region image generated in the image generation step to a predetermined size, And a recognition step of performing a predetermined recognition based on the target region image scaled in the scaling step. The determination step is a step of determining a target region size, which is the size of the target region corresponding to the actual size, based on the distortion information generated in the distortion information generation step and the target position calculated in the calculation step. A program, characterized in that.
[0150] The preferred embodiments of the present invention have been described above. However, the present invention is not limited to the above-described embodiments, and various modifications and changes are possible within the scope of the gist thereof. For example, the present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a recording medium, and having a processor of a computer of the system or device read and execute the program. Further, the present invention can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
Explanation of Signs
[0151] 100 Internet 101 Local network 102 Overhead camera 103 Image processing device 104 Focus camera 105 Learning server 106 Data collection server 107 Client terminal 108 Basketball court 109 Player 110 Ball 113 Video cable 114 Imaging optical axis 115 Position sensor receiving device
Claims
1. An image processing apparatus capable of recognizing a predetermined scene in a captured image, comprising: an input unit for inputting a captured image obtained by capturing a first object; a distortion information generation unit for generating distortion information indicating distortion from a known shape of the shape of the first object in the captured image input by the input unit; a calculation unit for calculating a position of interest in the captured image; a determination unit for determining a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit; an image generation unit for generating a region-of-interest image from the captured image based on the region of interest determined by the determination unit; a scaling unit for scaling the region-of-interest image generated by the image generation unit to a predetermined size; a recognition unit for performing a predetermined recognition based on the region-of-interest image scaled by the scaling unit, wherein the determination unit determines a region-of-interest size that is the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit. The image processing apparatus is characterized by this.
2. The distortion information generation unit generates distortion information based on at least one of the degree of distortion, size, and shape of the object shape from the known shape of the information regarding the first object in the captured image. The image processing apparatus according to claim 1 is characterized by this.
3. further comprises a shape acquisition unit for acquiring the shape of the first object in the captured image, wherein the distortion information generation unit generates distortion information based on the shape of the first object acquired by the shape acquisition unit. The image processing apparatus according to claim 1 or 2 is characterized by this.
4. The shape acquisition unit further comprises a UI unit for acquiring, as the shape of the first object, the object shape received by a user operation. The image processing apparatus according to claim 3 is characterized by this.
5. further comprises a shape storage unit for storing the shape information of each object in an object group composed of objects of a plurality of types of shapes, wherein the shape acquisition unit further acquires, as the shape of the first object, the object of the shape information when there is a correlation between any one of the plurality of types of shape information stored by the shape storage unit and the object in the captured image. The image processing apparatus according to claim 3 is characterized by this.
6. The shape acquisition unit further The image processing apparatus according to claim 3, wherein marks provided at respective locations in the real world corresponding to respective vertices of the first object are detected, and a region having the marks as vertices is acquired as the shape of the first object.
7. The distortion information generation unit generates distortion information based on a projective transformation matrix between coordinates indicating vertices of the first object acquired by the shape acquisition unit and coordinates indicating vertices of a known shape of the first object. The image processing apparatus according to claim 3.
8. The determination unit determines the region of interest based on position information of a second object whose position is estimated in the captured image. The image processing apparatus according to claim 1 or 2.
9. The image processing apparatus further includes a position information acquisition unit that acquires position information from a sensor that performs position detection attached to an object in the real world corresponding to a third object. The calculation unit calculates the position of interest based on the position information acquired by the sensor. The image processing apparatus according to claim 1 or 2.
10. The recognition unit determines whether or not the second object in the region of interest image has performed a predetermined action. The image processing apparatus according to claim 1 or 2.
11. The recognition unit determines whether or not the second object has performed a predetermined action based on the position of the third object and the position of the second object. The image processing apparatus according to claim 10.
12. The image processing apparatus further includes a first learned model that estimates the position of a second object from a region of interest image. using the first learned model to estimate the position of the second object from the region of interest image. The image processing apparatus according to claim 11.
13. The image processing apparatus further includes a second learned model that detects the position of a third object from a region of interest image. using the second learned model to detect the position of the third object from the reduced region of interest image. The image processing apparatus according to claim 11.
14. The position of the second object is the position of a joint of a player who performs a combat competition in an arena in the real world corresponding to the first object. The position of the third object is the position of an object used in the combat competition. The image processing apparatus according to any one of claims 11 to 13.
15. The image processing apparatus further includes an optical information acquisition unit that acquires optical information including the optical axis direction and focal length of the lens of the camera that captures the captured image, wherein the distortion information generation unit generates distortion information based on the optical information acquired by the optical information acquisition unit. The image processing apparatus according to claim 1 or 2.
16. An image processing system including a first camera capable of capturing an image of a first object, a second camera that captures an image in a designated imaging direction, and an image processing apparatus capable of controlling the imaging direction of the second camera, wherein the first camera includes a first imaging unit capable of acquiring an image of the first object, and a first transmission unit that transmits the image captured by the first imaging unit to the image processing apparatus, wherein the image processing apparatus includes a first reception unit that receives the image captured transmitted by the first transmission unit, a distortion information generation unit that generates distortion information indicating distortion from a known shape of the shape of the first object in the image captured by the first reception unit, a calculation unit that calculates a position of interest in the captured image, a determination unit that determines a region of interest based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit, an image generation unit that generates a region-of-interest image from the captured image based on the region of interest determined by the determination unit, a scaling unit that scales the region-of-interest image generated by the image generation unit to a predetermined size, a recognition unit that performs predetermined recognition on the region-of-interest image scaled by the scaling unit, a control signal generation unit that generates a control signal for directing the second camera in an imaging direction corresponding to the recognition result of the recognition unit, and a second transmission unit that transmits the control signal generated by the control signal generation unit to the second camera. The determination unit has a function of determining a region-of-interest size that is the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation unit and the position of interest calculated by the calculation unit, wherein the second camera includes a second imaging unit capable of imaging, a second reception unit that receives the control signal transmitted by the second transmission unit, and a control unit that controls the imaging direction of the second imaging unit based on the control signal received by the second reception unit. The image processing system is characterized by this.
17. A control method for an image processing apparatus capable of recognizing a predetermined scene in a captured image, An input step of inputting a captured image of a first object; A distortion information generation step of generating distortion information indicating distortion from a known shape of the shape of the first object in the captured image input by the input step; A calculation step of calculating a position of interest in the captured image; A determination step of determining a region of interest based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step; An image generation step of generating a region-of-interest image from the captured image based on the region of interest determined by the determination step; A scaling step of scaling the region-of-interest image generated by the image generation step to a predetermined size; A recognition step of performing predetermined recognition based on the region-of-interest image scaled by the scaling step, and having, The determination step is a step of determining a region-of-interest size that is the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step. A control method for an image processing apparatus characterized by that.
18. A program for causing a computer to execute a control method for an image processing apparatus capable of recognizing a predetermined scene in a captured image, The control method includes: An input step of inputting a captured image of a first object; A distortion information generation step of generating distortion information indicating distortion from a known shape of the shape of the first object in the captured image input by the input step; A calculation step of calculating a position of interest in the captured image; A determination step of determining a region of interest based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step; An image generation step of generating a region-of-interest image from the captured image based on the region of interest determined by the determination step; A scaling step of scaling the region-of-interest image generated by the image generation step to a predetermined size; A recognition step of performing predetermined recognition based on the region-of-interest image scaled by the scaling step, and having, The determination step is a step of determining a region-of-interest size that is the size of the region of interest corresponding to the actual size based on the distortion information generated by the distortion information generation step and the position of interest calculated by the calculation step. A program characterized by that.
Citation Information
Patent Citations
Play analysis device and play analysis method
JP2020054748A