Image processing device, image processing method, and program

JP7898837B2Active Publication Date: 2026-08-03CANON KK
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2021-09-14
Publication Date
2026-08-03

Smart Images

  • Figure 0007898837000004
    Figure 0007898837000004
  • Figure 0007898837000005
    Figure 0007898837000005
  • Figure 0007898837000006
    Figure 0007898837000006
Patent Text Reader

Abstract

To prevent cut-off of a cut-out image while suppressing the restriction of a cut-out range.SOLUTION: An image processing device according to one aspect comprises: first acquisition means which acquires a reference image; second acquisition means which acquires a peripheral image obtained by imaging a peripheral area on the periphery of an imaging area of the reference image; conversion means which converts the peripheral image so as to extend a field angle in the time of imaging of the reference image; and composition means which generates a composite image obtained by combining a converted image converted from the peripheral image and the reference image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] In a sports scene, many videos are acquired, edited, and distributed by a number of cameras. Also, it has become important to analyze the content of a play from the acquired videos to know the characteristics of an opponent or to contribute to the improvement of the strength of one's own team. However, since these processes require a lot of manpower, when targeting students or youths, for example, the videos may not be effectively utilized. For this reason, there is a need for a technology that acquires videos without using manpower to create distribution videos or extracts information necessary for sports analysis.

[0003] In a scene with fast movement such as sports, it is difficult to control a camera in accordance with the movement of a subject. Therefore, a method is known in which a video of an entire view is acquired at a high resolution, and a necessary scene is cut out from the video to create a video. In that case, in order to acquire a video at a high resolution over a wide range, a lot of memory is required and the processing load becomes heavy. Also, it is wasteful to acquire areas without movement as videos.

[0004] Patent Document 1 discloses a video display device that cuts out a portion where the same subject is imaged from a video and inserts it into a still image when displaying a program with a video and a still image distributed in a plurality of channels.

[0005] Patent Document 2 discloses an image display device that adjusts the alignment of a moving image with respect to a panoramic image based on control points for associating the moving image with the area of the panoramic image by extracting feature points between the panoramic image and the moving image.

[0006] Patent Document 3 discloses an imaging device that displays a viewfinder image and also displays a frame, which is an indicator of the video capture area in the viewfinder image. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2005-229451 [Patent Document 2] Japanese Patent Publication No. 2014-82764 [Patent Document 3] Japanese Patent Publication No. 2003-259161 [Overview of the project] [Problems that the invention aims to solve]

[0008] However, the methods disclosed in Patent Documents 1 and 2 may result in cropping depending on the cropping area. Furthermore, Patent Document 3 does not disclose anything regarding image cropping. The problem that this invention aims to solve is to prevent cropping of the image while suppressing limitations on the cropping range. [Means for solving the problem]

[0009] An image processing apparatus according to one aspect of the present invention includes a moving body. and corresponding to the imaging area Reference image 、 A first acquisition means for acquiring a reference image, which is one of the images that make up the video, and a peripheral image, which is a still image, that does not contain any moving objects and was captured before the reference image. Recorded image area outside A second acquisition means for acquiring a peripheral image corresponding to the surrounding region of the reference image; a conversion means for converting the peripheral image so that the field of view of the reference image at the time of imaging is expanded; and a synthesis means for generating a composite image by combining the converted image obtained from the peripheral image and the reference image. Analysis means for analyzing the reference image and determining the cutting position of the cut-out image to be cut out from the composite image; and cutting means for cutting out the cut-out image from the composite image, including a part of the imaging area and a part of the peripheral area, based on the cutting position.The system comprises the following: the reference image includes a subject other than the moving object that extends beyond the imaging area into the peripheral area; the peripheral image includes other parts of the subject that are different from the part included in the reference image; and the conversion means converts the peripheral image based on a first imaging direction of the reference image and a second imaging direction of the peripheral image, such that in the composite image, the part of the subject included in the reference image and the other part of the subject included in the converted image converted from the peripheral image are continuously connected. [Effects of the Invention]

[0010] According to one aspect of the present invention, it is possible to prevent cropping of the cropped image while suppressing limitations on the cropping range. [Brief explanation of the drawing]

[0011] [Figure 1] A block diagram showing an example configuration of the image processing system according to the first embodiment. [Figure 2A] A block diagram showing an example of the functional configuration of an image processing apparatus according to the first embodiment. [Figure 2B] A block diagram showing an example of the hardware configuration of an image processing device according to the first embodiment. [Figure 3A] A diagram showing an example of a cropping area set in a reference image. [Figure 3B] This figure shows an example of a cropped image extracted from the cropped area in Figure 3A. [Figure 4] A flowchart illustrating the image processing of an image processing apparatus according to the first embodiment. [Figure 5A] A diagram showing an example of the peripheral region around a reference image. [Figure 5B] This figure shows an example of a composite image created by combining a reference image and a transformed image. [Figure 6A] A diagram showing an example of a cropping area set for a composite image. [Figure 6B] This figure shows an example of a cropped image extracted from the cropped area in Figure 6A. [Figure 7]A diagram for explaining the image conversion process of the image processing apparatus according to the first embodiment. [Figure 8A] A block diagram showing a functional configuration example of the image processing apparatus according to the second embodiment. [Figure 8B] A block diagram showing a hardware configuration example of the image processing apparatus according to the second embodiment. [Figure 9] A flowchart showing the image processing of the image processing apparatus according to the second embodiment. [Figure 10A] A diagram showing an example of setting the boundary of the cut-out area according to the second embodiment. [Figure 10B] A diagram showing an example of presenting the position of the reference image during zooming according to the second embodiment. [Figure 11] A diagram showing an example of setting the moving object area according to the second embodiment.

Modes for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the present invention, and not all combinations of the features described in the embodiments are essential for the solution means of the present invention. The configuration of the embodiment can be appropriately modified or changed according to the specifications of the apparatus to which the present invention is applied and various conditions (usage conditions, usage environment, etc.). The technical scope of the present invention is determined by the scope of the claims and is not limited by the following individual embodiments.

[0013] In the following embodiments, we will use sports imaging as an example, but the method is not limited to this and can be applied to imaging various events, concerts, lecture scenes, attractions, parades, acrobatics, dance, theater, and meteors. Furthermore, in the following embodiments, we will use an image processing device that functions as an imaging device (network camera) that can connect to a network and communicate with other devices as an example. However, the method is not limited to this and can also be applied to an image processing device that functions as an imaging device that cannot connect to a network. Furthermore, in the following embodiments, we will explain assuming that the image processing device has an imaging function, but the method is not limited to cases where the image processing device has an imaging function; the imaging function may be implemented by a device other than the image processing device, and the image processing device may acquire the captured images from the other device.

[0014] <First Embodiment> This embodiment describes a method for performing high-speed interpolation when extracting a region of interest from a video by acquiring a still image of the area surrounding the area being captured as a video, correcting the still image to match the video's field of view, and retaining it. By performing this processing, it becomes possible to create an image without any cropping, even in areas with significant distortion at the edges, when extracting a video from an overhead view.

[0015] Figure 1 is a block diagram showing an example configuration of an image processing system according to the first embodiment. In Figure 1, the image processing system 10 comprises an image processing device 100 that functions as an imaging device and a client device 200, and the image processing device 100 and the client device 200 are connected in a state that allows them to communicate with each other via a network 300.

[0016] In this embodiment, we take as an example a device (such as a network camera) that can connect to a network 300 and communicate with other devices. However, it is not essential that the image processing device 100 can connect to the network 300. For example, the image processing device 100 and the client device 200 may be directly connected by a cable. Such a cable could use, for example, HDMI (High-Definition Multimedia Interface) or SDI (Serial Digital Interface).

[0017] Based on user operations, the client device 200 sends a distribution request command to the image processing device 100 to request the distribution of a video (image) stream, and a setting command to set various parameters. The client device 200 can be implemented by installing a predetermined program on a computer such as a personal computer, tablet terminal, or smartphone.

[0018] The image processing device 100 distributes a video stream to the client device 200 in response to a distribution request command and stores various parameters in response to a setting command. The image processing device 100 also acquires a reference image and a peripheral image, transforms the peripheral image so that the field of view of the reference image is expanded, and generates a composite image by combining the transformed image from the peripheral image with the reference image. At this time, the image processing device 100 transforms the peripheral image so that subjects extending beyond the imaging area into the peripheral area can be continuously connected. The peripheral image is an image of the area surrounding the imaging area of ​​the reference image. The reference image may be a video, and the peripheral image may be a still image. This reference image may be an overhead image captured with a wide-angle camera.

[0019] Here, to focus on the edges of the overhead image, if we crop the overhead image focusing on the edges, the entire cropped area will not fit within the wide-angle camera's field of view, resulting in cropped image. On the other hand, if we crop the overhead image in a way that avoids cropping, the range of the cropped image will be limited. Therefore, when the image processing device 100 performs cropping mainly from the edge of an overhead image, it performs cropping from a composite image obtained by combining a converted image converted from the surrounding image and a reference image. This allows the image processing device 100 to suppress limitations on the cropping range that is cropped as the area of ​​interest, while preventing the cropped image from being cut off.

[0020] In this case, the image processing device 100 can transform the peripheral image so that its imaging direction becomes equal to that of the reference image when the imaging direction of the reference image and the imaging direction of the peripheral image are different. As a result, the image processing device 100 can generate a composite image in which distortion of the subject extending beyond the imaging area into the peripheral area is suppressed, and the subject extending beyond the imaging area into the peripheral area is continuously connected.

[0021] For example, the image processing device 100 acquires a still image of the peripheral area surrounding the imaging area of ​​a moving image using PTZ (pan, tilt, zoom) without moving the installation position of the same camera, and converts it into a planar image of the imaging position of the moving image and stores it. When using a method called trapezoidal transformation, which converts the cropped image into a rectangle by specifying four vertices, a cropped image without any cropping can be easily obtained by converting the still image of the peripheral area into a planar image of the imaging position of the moving image.

[0022] Figure 2A is a block diagram showing a functional configuration example of an image processing apparatus according to the first embodiment. Of the functional blocks shown in Figure 2A, those implemented by software have programs stored in memory such as ROM (Read Only Memory) to provide the functionality of each functional block. These programs are then read into RAM (Random Access Memory) and executed by the CPU (Central Processing Unit). For functions implemented by hardware, for example, a dedicated circuit can be automatically generated on the FPGA from the program to implement the functionality of each functional block using a designated compiler. FPGA stands for Field Programmable Gate Array. Alternatively, a gate array circuit can be formed in a similar manner to an FPGA and implemented as hardware. Alternatively, it can be implemented using an ASIC (Application Specific Integrated Circuit). Note that the configuration of the functional blocks shown in Figure 2A is just one example; multiple functional blocks may constitute a single functional block, or any functional block may be divided into blocks that perform multiple functions.

[0023] In Figure 2A, the image processing device 100 includes a peripheral image acquisition unit 211, an image conversion unit 212, a reference image acquisition unit 213, an image analysis unit 214, an image synthesis unit 215, and an image cropping unit 216.

[0024] The reference image acquisition unit 213 acquires a reference image. The reference image may be a video. This reference image may also be an overhead image captured by a wide-angle camera. For example, the reference image acquisition unit 213 may acquire a video from the imaging unit 221 in Figure 2B or an external device (not shown) and use one frame included in the video as the reference image. The reference image acquisition unit 213 also generates a reference image using various parameters (various settings) acquired from the storage unit 222 in Figure 2B.

[0025] The peripheral image acquisition unit 211 acquires peripheral images. Peripheral images are images of the peripheral region surrounding the imaging area of ​​the reference image. For example, the peripheral image acquisition unit 211 acquires still images from the imaging unit 221 in Figure 2B or from an external device (not shown) and generates peripheral images corresponding to each frame. The peripheral image acquisition unit 211 also generates peripheral images using various parameters (various settings) acquired from the storage unit 222 in Figure 2B.

[0026] The image conversion unit 212 converts the peripheral image based on the image generated by the peripheral image acquisition unit 211 and various parameters used when acquiring that image (such as camera orientation and zoom value). At this time, the image conversion unit 212 converts the peripheral image so that the field of view of the reference image is expanded. For example, the image conversion unit 212 can convert the peripheral image so that subjects extending beyond the imaging area of ​​the reference image into the peripheral area can be continuously connected. Here, the image conversion unit 212 can convert the peripheral image so that the imaging direction of the peripheral image becomes equal to the imaging direction of the reference image when the imaging direction of the reference image and the imaging direction of the peripheral image are different. For example, the image conversion unit 212 converts the image to one taken by a camera placed at the same position but with different pan, tilt, and zoom values. The image conversion unit 212 stores the converted peripheral image generated by the conversion in the storage unit 222.

[0027] The image analysis unit 214 analyzes the reference image and determines the extraction position of the extracted image to be cut out from the composite image. For example, the image analysis unit 214 performs object detection processing on the reference image acquired by the reference image acquisition unit 213 and extracts the extraction region to be cut out from the image. The object detection method may be, for example, a machine learning method. In this case, the image analysis unit 214 can detect objects and determine the extraction position by generating a classifier that has learned the characteristics of the object to be detected and applying it to the image data. The image analysis unit 214 stores the reference image acquired from the reference image acquisition unit 213 and information regarding the extraction region extracted from that reference image in the storage unit 222.

[0028] The image synthesis unit 215 synthesizes the converted peripheral image stored in the storage unit 222 with the reference image acquired by the reference image acquisition unit 213 to generate a composite image. The generated composite image is sent to the image extraction unit 216.

[0029] The image extraction unit 216 creates a pseudo-PTZ image (hereinafter referred to as a pseudo-PTZ image) from the composite image generated by the image synthesis unit 215, based on the information about the extracted region held in the storage unit 222.

[0030] Figure 2B is a block diagram showing an example of the hardware configuration of an image processing device according to the first embodiment. In Figure 2B, the image processing device 100 includes an imaging unit 221, a storage unit 222, a control unit 223, a communication unit 224, and an accelerator unit 225.

[0031] The imaging unit 221 converts light formed on the light-receiving surface of the image sensor through the lens into an electric charge for each pixel to acquire a moving image. The imaging unit 221 includes, for example, a zoom lens, a focus lens, an image stabilization lens, an aperture, a shutter, an optical low-pass filter, an IR (Infrared Rays) cut filter, a color filter, and an image sensor. The image sensor may be, for example, an image sensor such as a CMOS (Complementary Metal Oxide Semiconductor) or a CCD (Charge Coupled Device).

[0032] The memory unit 222 stores programs for the image processing device 100 to perform various operations. The memory unit 222 can also store data (commands and image data) and various parameters acquired from external devices such as the client device 200 via the communication unit 224. For example, the memory unit 222 stores parameters related to camera orientation and magnification, such as pan, tilt, and zoom, for moving images acquired by the imaging unit 221, as well as parameters related to camera settings, such as camera white balance and exposure. The memory unit 222 may also store parameters related to image data, including the frame rate and image data size (resolution). Furthermore, the memory unit 222 can provide a work area used by the control unit 223 when performing various processes. In addition, the memory unit 222 can function as a frame memory or buffer memory.

[0033] The storage unit 222 may consist of both ROM and RAM, or it may consist of either ROM or RAM. Furthermore, the storage unit 222 may use storage media such as an SSD (Solid State Drive), flexible disk, hard disk, optical disk, magneto-optical disk, CD-ROM, CD-R, magnetic tape, non-volatile memory card, or DVD.

[0034] The control unit 223 controls the entire image processing device 100 by executing a program stored in the storage unit 222. The control unit 223 may also control the entire image processing device 100 in cooperation with the program stored in the storage unit 222 and the OS (Operating System).

[0035] The control unit 223 may be a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit). Furthermore, the control unit 223 may also include a processor such as a DSP (Digital Signal Processor) or an ASIC (Application Specific Integrated Circuit).

[0036] The communication unit 224 transmits and receives wired or wireless signals in order to communicate with the client device 200 via the network 300 shown in Figure 1.

[0037] The accelerator unit 225 is a processing unit added to the camera primarily for performing high-performance processing using machine learning such as Deep Learning. For example, the accelerator unit 225 may be configured to handle the processing of the image analysis unit 214. The accelerator unit 225 may include a CPU, GPU, FPGA, and memory unit.

[0038] In this embodiment, the case where image analysis processing is performed using an image processing device 100 including an accelerator unit 225 is described. Regarding the analysis processing, it may also be performed using an externally connected accelerator unit such as a USB (Universal Serial Bus). Alternatively, video may be input directly from the camera via HDMI (High-Definition Multimedia Interface) or SDI (Serial Digital Interface) to a dedicated device with a GPU or FPGA for analysis. Furthermore, the video may be streamed and then saved to a PC (Personal Computer) for processing, or the video recorded on an SD card attached to the camera may be processed on a PC that is not connected to the camera.

[0039] The functional configuration of the image processing device 100 shown in Figure 2A may be implemented using the hardware shown in Figure 2B, or it may be implemented using software.

[0040] The operation of the image processing device 100 shown in Figure 2A will be explained below using the example of acquiring an overall overview image of a sports competition and creating a composite image from which moving scenes can be seamlessly extracted.

[0041] Figure 3A shows an example of an overhead image used as a reference image. In Figure 3A, the reference image frame 30 includes a scene in which multiple basketball players 310 and a basketball court 320 are captured from an overhead perspective. Here, the position of the area to be extracted when creating a cropped image focusing on the area where the players 310 gather is set as the cropped area 330.

[0042] Figure 3B shows an example of a cropped image extracted from the cropped area in Figure 3A. In Figure 3B, it is assumed that a cropped image frame 31 is generated from the cropping region 330 set in the reference image frame 30 of Figure 3A. The cropped image frame 31 can be generated by converting the cropping region 330, which is enclosed by the four vertices specified on the reference image frame 30, into a rectangle. In this case, for the parts of the cropping region 330 that extend beyond the reference image frame 30, pixel data cannot be obtained, so in the cropped image frame 31, these become cropped regions 340 and 341 with a pixel value of 0.

[0043] To prevent the occurrence of cropped areas 340 and 341, the image processing device 100 acquires a peripheral image capturing the area surrounding the reference image frame 30, and stores a converted peripheral image obtained by image transformation so that it is continuously connected to the reference image frame 30. Then, the image processing device 100 creates a composite image by interpolating the reference image frame 30 with the converted peripheral image, and performs cropping from this composite image. This prevents cropped areas 340 and 341 from occurring in the cropped image even if the cropped area 330 extends beyond the reference image frame 30.

[0044] Figure 4 is a flowchart showing the image processing of the image processing apparatus according to the first embodiment. Figure 5A shows an example of a peripheral region around a reference image, and Figure 5B shows an example of a composite image obtained by combining the reference image and the converted image. Figure 6A shows an example of a cropping region set in the composite image, and Figure 6B shows an example of a cropped image obtained by cropping from the cropping region in Figure 6A.

[0045] Each step in Figure 4 is realized by the control unit 223 reading and executing the program stored in the memory unit 222 in Figure 2B. Furthermore, at least a portion of the flowchart shown in Figure 4 may be implemented in hardware. In the case of hardware implementation, for example, a dedicated circuit can be automatically generated on the FPGA from the program required to implement each step by using a predetermined compiler. Alternatively, a Gate Array circuit may be formed in a similar manner to the FPGA, and the implementation may also be carried out using an ASIC. In this case, each block in the flowchart shown in Figure 4 can be considered a hardware block. Note that multiple blocks may be combined to form a single hardware block, or a single block may be composed of multiple hardware blocks.

[0046] The processing shown in Figure 4 is achieved when the control unit 223 in Figure 2B executes the control program stored in the memory unit 222, performing calculations, processing, and control of each hardware component. In addition, the accelerator unit 225 also performs processing to accelerate the machine learning-based processing.

[0047] In step S41 of Figure 4, the peripheral image acquisition unit 211 and the reference image acquisition unit 213 acquire the settings necessary to generate image data. For example, the peripheral image acquisition unit 211 and the reference image acquisition unit 213 acquire parameters related to the imaging unit 221, which captures images, from the storage unit 222. The parameters related to the imaging unit 221 include the camera's installation position, orientation, field of view, lens focal length, frame rate, and image data size (resolution). For example, the image data size can be 1920 × 1080 pixels and the frame rate can be 30 fps.

[0048] Next, in step S42, the peripheral image acquisition unit 211 instructs the system to image the area surrounding the reference position indicated by the various settings acquired in step S41, and acquires image data corresponding to the peripheral area. The peripheral area is, for example, the area 50A to 50D surrounding the imaging area of ​​the reference image frame 30 in Figure 3A, as shown in Figure 5A. The peripheral area is used for interpolation of the image captured at the reference position, and mainly consists of areas that do not contain important moving elements. Therefore, for example, when a basketball game is being played, as shown in Figure 3A, the movement of the players and the ball may be captured at the reference position, while the peripheral area may be imaged when there are no players on the court before the game starts.

[0049] However, the timing of imaging the peripheral area is not limited to before the start of the match; it can be arbitrarily selected. For example, imaging may be performed during breaks between matches to update the latest image. In particular, imaging may be performed during breaks only when changes occur in lighting or other factors due to the passage of time. Furthermore, the timing of imaging the peripheral area may be after changes in the parameters of the imaging unit 221 that captures the peripheral image. In addition, if the camera is installed in a fixed position, the peripheral image acquisition unit 211 may store peripheral images captured under multiple conditions according to camera parameters such as exposure, and select the peripheral image captured under the imaging conditions closest to the reference image.

[0050] There are several ways to acquire image data corresponding to the surrounding area. For example, the camera or pan / tilt head can be rotated to change the camera's orientation left / right or up / down and take images, or the zoom can be changed to capture a wider area than the reference position. The surrounding image acquisition unit 211 generates one or more image frames corresponding to the surrounding area thus captured and stores them in the storage unit 222.

[0051] Next, in step S43, the image conversion unit 212 generates a converted peripheral image that can be continuously connected to the reference image frame from the image frame of the peripheral region acquired in step S42. Here, the first imaging direction of the reference image and the second imaging direction of the peripheral image are different. In this case, the image conversion unit 212 may map the planar coordinates of the peripheral image captured from the second imaging direction onto a spherical coordinate system, and then map the planar coordinates of the peripheral image mapped onto the spherical coordinate system onto the planar coordinates of the reference image captured from the first imaging direction.

[0052] Next, in step S44, the reference image acquisition unit 213 instructs the acquisition of a reference image based on the various settings acquired in step S41 and generates a reference image frame 30.

[0053] Next, in step S45, the image analysis unit 214 performs object detection processing on the reference image frame acquired in step S44 to detect the target object. For example, in the reference image frame 30 of Figure 3A, the image analysis unit 214 can set the detected targets to be the player 310 and the ball.

[0054] For object detection using image analysis, machine learning, particularly Deep Learning-based methods, can be used to achieve high accuracy and real-time processing. Specifically, examples of object detection methods using image analysis include YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector). SSD is one method for detecting individual objects in an image containing multiple objects.

[0055] To build a classifier that detects player 310 and the ball using an SSD, images containing human bodies are collected from multiple videos and prepared as training data. Specifically, regions containing human bodies and balls are extracted from the videos, and a file is created that records the coordinates of the center position of each region and its size. The classifier that detects human bodies and balls is then trained using this prepared training data.

[0056] When a human body and a ball are detected using the generated classifier, information indicating the location and size of the detected area is obtained. The location information of the area is expressed as coordinates with the top left of the frame as the origin, and the center position of the detected area and the size of its rectangle (width and height) are shown as a ratio to the size of the image. The location and size of the detected area obtained in this way are obtained as a list, as multiple people may be detected within the frame.

[0057] Next, the image analysis unit 214 determines the cropping region 330 based on the acquired detection results of the player 310 and the ball. Specifically, it determines the centroids of the detected positions of the player 310 and the ball, and sets these centroids as the center of the cropping region 330. There are several methods for determining the center of the cropping region. For example, the image analysis unit 214 may calculate the centroids by treating the player 310 and the ball as having the same weight, or it may calculate the centroid of the player 310 only.

[0058] Next, the image analysis unit 214 calculates the four vertices of the cropped region 330. The cropped region 330 is considered as the imaging area when a second camera (a camera with a smaller field of view, placed in the same position as the overhead camera) is pointed at the center of the cropped region 330 calculated above. At this time, a pseudo-PTZ image is generated by cropping, in which PTZ is performed by the second camera.

[0059] Next, in step S46, the image synthesis unit 215 synthesizes the converted peripheral image generated in step S43 with the reference image frame acquired in step S44 to generate a pseudo-ultra-wide-angle composite image 52. The converted peripheral image generated in step S43 is an image that has already been converted so that it can be continuously connected with the reference image frame. Therefore, the image synthesis unit 215 can embed the reference image acquired in step S44 into the converted peripheral image to generate the composite image 52. Furthermore, the image synthesis unit 215 determines parameters to convert the coordinate values ​​on the reference image to the coordinates of the composite image. These parameters are stored as translations in the x and z directions according to the size of the converted peripheral image. Then, the image synthesis unit 215 stores the generated composite image 52 in the storage unit 222.

[0060] For example, as shown in Figure 5B, the image synthesis unit 215 generates a composite image 52 by interpolating the converted surrounding images 51A to 51D acquired in step S43 around the reference image frame 30.

[0061] Next, in step S47, the image cropping unit 216 performs a trapezoidal transformation on the composite image 52 created in step S46 using the positions of the four vertices of the cropped region 330 acquired in step S45, and creates a cropped image frame 33 as shown in Figure 6B. Here, the image cropping unit 216 transforms the coordinates of the four vertices of the cropped region 330 using the translation parameters acquired in step 43, and converts them to coordinates on the composite image 52.

[0062] Next, in step S48, the control unit 223 determines whether or not there is image data to be cropped. If there is image data to be cropped (Yes in step S48), the control unit 223 returns to step S44 and continues processing the next image data. If there is no image data to be cropped (No in step S48), the control unit 223 terminates the process.

[0063] Figure 7 is a diagram illustrating the image conversion process of the image processing apparatus according to the first embodiment. In Figure 7(A), during the image transformation process, spherical coordinates are set for the reference image frame 30 captured by the overhead camera, with the camera position O as the origin. If the center of the reference image frame 30 is R, the area above the center R is the z-coordinate, and the area to the right is the x-coordinate, then the pixel positions on the reference image frame 30 can be mapped one-to-one onto the spherical coordinate system as seen from the camera position O. The coordinates Q(x, y, z) on the reference image frame 30 correspond to the coordinates P(xp, yp, zp) on the sphere 60 with radius r and the camera position O as the origin.

[0064] Here, the spherical coordinates (r, θ, φ) are defined as shown in Figure 7(B). In this case, the coordinates Q(x, y, z) on the reference image frame 30 can be expressed in spherical coordinates as given by the following equation (1).

[0065]

number

[0066] The coordinates P(xp,yp,zp) on a sphere 60 of radius r with camera position O as the origin can be given by the following equation (2).

[0067]

number

[0068] Next, in order to capture the peripheral region of the reference image frame 30, the orientation of the camera 101 is changed as shown in Figure 7(C). At this time, the center of the captured peripheral image frame 51 moves to R', and point T on the peripheral image frame 51 is associated with point S on the sphere 60. In this way, by mapping the information of the peripheral image frame 51, which captures the peripheral region of the reference image frame 30, onto the sphere 60, equation (1) can be used again to associate it with an extended region of the reference image frame 30 with center R.

[0069] The image conversion unit 212 generates a converted peripheral image that can be continuously connected to the reference image frame 30 by image conversion from a peripheral image frame 51 that captures the area surrounding the reference image frame 30. Furthermore, the image conversion unit 212 calculates parameters for converting coordinate values ​​on the reference image frame 30. Specifically, it stores the size of the converted peripheral image (the number of pixels added to the left and above the reference image) and stores the pixel position as a value for translation when specifying the top left of the image as the reference. In this way, the image conversion unit 212 stores the generated converted peripheral image and the translation parameters in the storage unit 222.

[0070] As shown in Figure 7(D), the image cropping unit 216 generates a pseudo-PTZ image by cropping. Specifically, the image cropping unit 216 converts the center coordinates of the cropping region 330 obtained on the reference image frame 30 into a point U(θc,φc) on a spherical coordinate system by the image transformation shown in step S43 of Figure 4. Then, with point U as the center, the image cropping unit 216 obtains the four vertices (F1, F2, F3, F4) of the cropping region 61 on the sphere 60, as shown in equation (3) below, with a horizontal field of view of 2Δθ and a vertical field of view of 2Δφ for the size to be cropped.

[0071]

number

[0072] Then, the image extraction unit 216 converts the four vertices (F1, F2, F3, F4) of the extraction region 61 acquired on the spherical surface 60 into coordinates on the reference image frame 30 using equation (1) again, and acquires them as the four vertices of the extraction region 330. The image extraction unit 216 stores the coordinates of the acquired four vertices of the extraction region 330 in the storage unit 222.

[0073] As described above, according to the first embodiment described above, when the image processing device 100 creates an extracted image frame (pseudo-PTZ image) from an image (reference image frame) to be captured as a video, it pre-images the area surrounding the reference image frame. The image processing device 100 then holds the pre-imaged area as a converted peripheral image, which is image-converted to connect continuously with the reference image frame, and composites it around the reference image frame to create a composite image 52. By creating an extracted image frame 33 from the composite image 52, the image processing device 100 can prevent the occurrence of the cropped areas 340 and 341 in Figure 3B without changing the cropped area 330 in Figure 3A.

[0074] <Second Embodiment> In the first embodiment, when creating a pseudo-PTZ image, a method was described in which a converted peripheral image, which has been image-transformed to be continuously connected to the reference image frame 30, is used to interpolate the cropped areas 340 and 341 shown in Figure 3B. At this time, it is preferable to clearly show the user how much of the surrounding area needs to be imaged. Also, when the converted surrounding image is acquired as a still image, the interpolated area becomes a past still image. In Figure 3A, the gymnasium wall at the top of the reference image frame 30 does not look out of place as a still image, but the lower left and right areas include the basketball court 320, and in some cases it is better not to replace them with still images.

[0075] In the second embodiment, after acquiring the converted peripheral image, the reference image and the converted peripheral image are displayed, and parameters such as the range of video capture are set. Specifically, the image processing device 100 assists in setting parameters related to the generation of a pseudo-PTZ image by showing changes in the pan-tilt range due to the interpolation region and changes in the reference image region when zoomed in. In the second embodiment, the same reference numerals are used for parts similar to those in the first embodiment, and detailed descriptions are omitted.

[0076] Figure 8A is a block diagram showing a functional configuration example of an image processing apparatus according to the second embodiment. In Figure 8A, the image processing device 101 includes a set value acquisition unit 716 instead of the image extraction unit 216 of the image processing device 100 in Figure 2A. Otherwise, the image processing device 101 can be configured in the same way as the image processing device 100.

[0077] The setting value acquisition unit 716 acquires the settings displayed on the composite image display screen. The settings are the boundaries of the cropped images extracted from the composite image, or the position of the reference image based on zoom, pan, or tilt. The settings may also be the range of the motion region of the composite image, or the processing applied to the motion region. The processing applied to the motion region is an enhancement process for the transformed image included in the cropped images extracted from the composite image. It is also possible to set how to combine video and still images. For example, the setting value acquisition unit 716 displays the composite image created by the image synthesis unit 215 on the GUI unit 726 in Figure 8B. Then, based on user input, the setting value acquisition unit 716 acquires the setting parameters necessary to create a pseudo-PTZ image.

[0078] Figure 8B is a block diagram showing an example of the hardware configuration of an image processing device according to the second embodiment. In Figure 8B, the image processing device 101 is the same as the image processing device 100 in Figure 2B, but with the addition of a GUI (Graphical User Interface) unit 726. Otherwise, the image processing device 101 can be configured in the same way as the image processing device 100. The GUI unit 726 displays a GUI that displays composite images and interactively acquires user input.

[0079] Figure 9 is a flowchart showing the image processing of the image processing apparatus according to the second embodiment. The process shown in Figure 9 is achieved when the control unit 223 in Figure 8B executes the control program stored in the storage unit 222, performing calculations, processing, and control of each piece of hardware. Note that the processes other than steps S85 and S86 are the same as those in Figure 4, and therefore their explanation is omitted.

[0080] In step S85, the image synthesis unit 215 (Figure 8A) synthesizes the converted peripheral image generated in step S43 with the reference image frame acquired in step S44 to generate a pseudo-ultra-wide-angle composite image 52. The control unit 223 then presents the composite image created by the image synthesis unit 215 to the user through the GUI unit 726.

[0081] Next, in step S86, the setting value acquisition unit 716 acquires the user's setting values ​​for the composite image and obtains the setting parameters necessary for creating a pseudo-PTZ image.

[0082] Figure 10A shows an example of setting the boundary of the cut-out region according to the second embodiment. In Figure 10A, for example, the left and right boundaries 910 of the range that the cropping area 330 can take can be set as setting parameters on the composite image 53. Although not shown in the figure, the upper and lower boundaries of the range that the cropping area 330 can take can also be set on the composite image 53. The composite image 53 may be a scene before the start of the game, when there are no players 310 etc. on the basketball court 320. In this case, the reference image frame 32 of Figure 5A can be used in the composite image 53 instead of the reference image frame 30 of the composite image 52 of Figure 5B.

[0083] Figure 10B shows an example of the presentation of the reference image position during zooming according to the second embodiment. In Figure 10B, the setting parameter can indicate, for example, the position 920 when the field of view of the reference image is changed by zooming. The setting parameter can also indicate the position when the reference image is moved by panning or tilting.

[0084] Figure 11 shows an example of setting the motion region according to the second embodiment. In Figure 11, a motion region 321 where the presence of a moving object is expected is set on the composite image 53, and the setting value acquisition unit 716 can acquire this setting. The composite image 53 includes an image of the basketball court 320 captured by an overhead camera, but there are areas at the lower left and right of the basketball court 320 that are not included in the field of view. If such areas are replaced with a converted surrounding image generated from past still images, the player will disappear if he enters that area. To prevent this situation, when replacing the motion region 321 with a converted surrounding image, the area can be highlighted by coloring or other means to alert the user. Alternatively, the system can perform motion tracking of the player and change the processing depending on whether it is predicted that the player is in an area where the player appears to disappear, or whether it is not predicted. The system allows the user to make such choices regarding changes in processing.

[0085] As described above, according to the second embodiment described above, the image processing device 100 presents the setting parameters for creating a pseudo-PTZ image from a reference image and a converted peripheral image on a GUI in a way that is easy for the user to specify. As a result, the image processing device 100 can create a pseudo-PTZ image in the form desired by the user.

[0086] (Other embodiments) The present invention can also be realized by a process in which one or more processors read and execute a program that implements one or more of the functions of the above-described embodiments. The program may be supplied to a system or device having a processor via a network or storage medium. Furthermore, the present invention can also be realized by a circuit (e.g., an ASIC) that implements one or more of the functions of the above-described embodiments. [Explanation of symbols]

[0087] 10 image processing systems, 100 image processing devices, 200 client devices, 300 network

Claims

1. A first acquisition means for acquiring a reference image that includes a moving object, corresponds to the imaging area, and is one of the images that constitute a video; A second acquisition means for acquiring a peripheral image that is a still image, which does not contain any moving objects and was captured before the reference image, and which corresponds to a peripheral area outside the imaging area. A conversion means for converting the surrounding image so that the field of view of the reference image is expanded when it is captured, A synthesis means for generating a composite image by combining the converted image obtained from the surrounding image and the reference image, Analysis means for analyzing the aforementioned reference image and determining the cropping position of the cropped image extracted from the composite image, Based on the aforementioned cutting position, a cutting means for cutting out the cut image from the composite image, which includes a part of the imaging area and a part of the surrounding area, Equipped with, The aforementioned reference image includes a subject other than the moving object, and a portion of the subject that extends beyond the imaging area into the peripheral area. The aforementioned peripheral image includes other parts of the subject that are different from the part included in the reference image. The conversion means is characterized by converting the peripheral image based on a first imaging direction of the reference image and a second imaging direction of the peripheral image, such that in the composite image, a part of the subject included in the reference image and another part of the subject included in the converted image converted from the peripheral image are continuously connected.

2. The image processing apparatus according to claim 1, characterized in that the conversion means converts the peripheral image so that the second imaging direction of the peripheral image becomes equal to the first imaging direction of the reference image when the first imaging direction of the reference image and the second imaging direction of the peripheral image are different.

3. The image processing apparatus according to claim 2, characterized in that the conversion means maps the planar coordinates of the peripheral image captured from the second imaging direction onto spherical coordinates, and maps the planar coordinates of the peripheral image mapped onto spherical coordinates onto the planar coordinates of the reference image captured from the first imaging direction.

4. The image processing apparatus according to any one of claims 1 to 3, characterized in that the surrounding image is an image taken before the start of a sports match, during a break in the sports, or after a change in the parameters of the imaging means that takes the surrounding image.

5. The image processing apparatus according to any one of claims 1 to 4, characterized in that the second acquisition means acquires a peripheral image captured under imaging conditions closest to the imaging conditions of the reference image from peripheral images captured under multiple imaging conditions.

6. A first acquisition step involves acquiring a reference image that includes a moving object, corresponds to the imaging area, and is one of the images that constitute the video. A second acquisition step involves acquiring a peripheral image that is a still image, which does not contain any moving objects and was captured before the reference image, and which corresponds to a peripheral area outside the imaging area. A conversion step of transforming the surrounding image so that the field of view of the reference image at the time of acquisition is expanded, A synthesis step of generating a composite image by combining the converted image converted from the surrounding image and the reference image, A determination step involves analyzing the aforementioned reference image and determining the position of the cropped image extracted from the composite image, A cropping step in which, based on the cropping position, a cropped image is extracted from the composite image, including a portion of the imaging region and a portion of the peripheral region, Equipped with, The aforementioned reference image includes a subject other than the moving object, and a portion of the subject that extends beyond the imaging area into the peripheral area. The aforementioned peripheral image includes other parts of the subject that are different from the part included in the reference image. An image processing method characterized in that, in the conversion step, the peripheral image is converted based on a first imaging direction of the reference image and a second imaging direction of the peripheral image, such that in the composite image, a part of the subject included in the reference image and another part of the subject included in the converted image converted from the peripheral image are continuously connected.

7. A program for operating a computer as an image processing device according to any one of claims 1 to 5.